An integrated chain data flow conversion method for ore prospecting prediction based on data lake technology

By integrating geological data through data lake technology and employing federated learning and homomorphic encryption, the problems of high data privacy risks and lagging model updates in geological exploration have been solved. This has enabled secure data collaboration and real-time economic optimization across mining areas, generating efficient exploration plans.

CN120634059BActive Publication Date: 2025-11-07CHINA GEOLOGICAL SURVEY NATURAL RESOURCES COMPREHENSIVE SURVEY COMMAND CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511127486.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-07
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing technologies in geological exploration suffer from high data privacy risks, lagging model updates, and a disconnect from market response. In particular, when collaborating on data across mining areas, the original datasets need to be uploaded centrally, which increases the risk of leakage of sensitive exploration data. Furthermore, the long model update cycle cannot meet the needs of dynamic exploration.

Method used

Data lake technology is used to integrate geophysical, geochemical and remote sensing data. Intelligent prediction models are deployed locally in the mining area through federated learning. Homomorphic encryption algorithm is used to transmit gradient parameters. Combined with geologically constrained Kalman filter algorithm, mineralization probability distribution map and economic evaluation report are constructed, and priority list of exploration target areas and borehole layout plan are generated.

Benefits of technology

It enables secure data collaboration across mining areas and real-time economic parameter optimization decisions, reduces data privacy risks, shortens model update cycles, improves market response efficiency, and generates efficient exploration plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634059B_ABST
    Figure CN120634059B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on data lake technology's prospecting prediction integrated chain data flow conversion method, it is related to geological exploration field, including, collection geophysical data, geochemical data and remote sensing data, collate into original exploration data;Metadata is added in original exploration data, is transmitted to mining local data lake node, forms original exploration data set;Federal learning agent is deployed in mining local data lake node, and the gradient parameter of mining intelligent prediction model is obtained, and the gradient parameter of mining intelligent prediction model is obtained through intelligent analysis algorithm;The gradient parameter of mining intelligent prediction model is encrypted using homomorphic encryption algorithm and uploaded to central aggregation server.The application realizes the safe cooperation of cross-mining area data not out of domain through federal learning and homomorphic encryption, and integrates real-time economic parameter optimization decision, solves the core problems of high data privacy risk, model update lag and market response disconnection in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of geological exploration, and in particular to a prospecting prediction integrated chain data flow method based on data lake technology. BACKGROUND

[0002] The data processing technology in the field of geological exploration has made significant progress, especially the combination of distributed data storage and intelligent analysis technology provides a new solution for prospecting prediction. Generally, a traditional data warehouse architecture is adopted, the geophysical, geochemical and remote sensing data are integrated through an ETL process, and a resource potential evaluation is carried out using a machine learning algorithm. A multi-scale mineral prediction system is used to jointly invert regional gravity, aeromagnetic data and geochemical sampling results to generate a two-dimensional favorable mineralization zoning map, which to some extent solves the problem of data islands and provides data support for mineral exploration decision-making.

[0003] The existing technology still has limitations in processing full-chain data collaboration and privacy protection. The federated learning framework is not fully integrated into the geological exploration scenario, resulting in the need to upload raw data in the cross-mining area data collaboration, and the use of centralized modeling method leads to an increased risk of sensitive exploration data leakage. Moreover, the model update cycle is as long as two weeks, which cannot meet the dynamic exploration demand. The traditional technology lacks a real-time response mechanism for economic parameters, making the target area optimization result out of touch with market supply and demand. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a prospecting prediction integrated chain data flow method based on data lake technology to solve the core problems of high data privacy risk, model update lag and market response disconnection in the prior art.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides a prospecting prediction integrated chain data flow method based on data lake technology, which comprises collecting geophysical data, geochemical data and remote sensing data, and arranging them into raw exploration data; adding metadata to the raw exploration data and transmitting it to a local data lake node in the mining area to form a raw exploration data set;

[0008] A federated learning agent is deployed at the local data lake node in the mining area to obtain a mining area intelligent prediction model, and a gradient parameter of the mining area intelligent prediction model is obtained through an intelligent analysis algorithm;

[0009] The gradient parameter of the mining area intelligent prediction model is encrypted using a homomorphic encryption algorithm and uploaded to a central aggregation server to generate a global prediction model through weighted aggregation;

[0010] Create an integrated three-dimensional geological model, monitor equipment state data and economic parameters in the data lake service layer, construct a two-way mapping relationship between the physical data lake and the digital twin through a Kalman filtering algorithm with geological constraints, obtain a mineralization probability distribution map and an economic evaluation report;

[0011] Based on the mineralization probability distribution map and the economic evaluation report, a multi-objective planning algorithm is used to generate a prospecting target area priority list and a drilling layout scheme.

[0012] As a preferred scheme of the integrated chain data flow method for ore prospecting prediction based on the data lake technology, the geophysical data, geochemical data and remote sensing data are collected and sorted into original exploration data, including the following steps,

[0013] The original magnetic anomaly time series data are obtained by using a proton magnetometer for wiring measurement and correction, and the corrected original magnetic anomaly time series data are used to generate geophysical data by Kriging interpolation;

[0014] The stream sediment samples are collected to obtain original stream samples, the original stream samples are input into an Olympus Vanta XRF instrument, element content data are obtained after correction, and the element content data are output to an IDW algorithm to obtain geochemical data;

[0015] Through the Sentinel-2 L2A level image, the multispectral data after radiation correction, and the PCA transformation formula, the principal component composite image is obtained, the principal component composite image is input into the Crosta algorithm to obtain remote sensing data;

[0016] The geophysical data, geochemical data and remote sensing data are spatially registered to obtain the original exploration data.

[0017] As a preferred scheme of the integrated chain data flow method for ore prospecting prediction based on the data lake technology, metadata is added to the original exploration data and transmitted to the local data lake node of the mining area to form an original exploration data set, including the following steps,

[0018] The original exploration data is checked by adding metadata tags according to the geological information metadata standard to obtain complete original exploration data;

[0019] The complete original exploration data is transmitted to the local data lake node of the mining area through special network encryption to form an original exploration data set.

[0020] As a preferred scheme of the integrated chain data flow transfer method for ore prospecting prediction based on the data lake technology, wherein: a federal learning agent is deployed at a local data lake node of a mining area to obtain a mining area intelligent prediction model, and a gradient parameter of the mining area intelligent prediction model is obtained through an intelligent analysis algorithm, including the following steps,

[0021] A Docker container is installed at the data lake node of the mining area, a federal learning image is loaded, a standardized running environment is obtained, and normalized exploration data is output;

[0022] The normalized exploration data is read from the data lake HDFS directory to obtain a dimension-unified training tensor, a 3-layer fully connected network is constructed using a GeLU activation function, and a mining area intelligent prediction model is obtained;

[0023] The mining area intelligent prediction model is trained using a weighted cross-entropy loss function to obtain a trained mining area intelligent prediction model and a gradient parameter of the mining area intelligent prediction model.

[0024] As a preferred scheme of the integrated chain data flow transfer method for ore prospecting prediction based on the data lake technology, wherein: the gradient parameter of the mining area intelligent prediction model is encrypted using a homomorphic encryption algorithm and uploaded to a central aggregation server, and a global prediction model is generated through weighted aggregation, including the following steps,

[0025] The gradient parameter of the mining area intelligent prediction model is normalized by the maximum absolute value to obtain a standardized gradient parameter, and an encryption public key and a private key are generated based on the central aggregation server;

[0026] The standardized gradient parameter is encrypted using the public key to obtain an encrypted standardized gradient parameter, an ECDSA-SHA256 signature is added to the encrypted standardized gradient parameter, and the encrypted standardized gradient parameter with the signature is uploaded to the central aggregation server to obtain the encrypted standardized gradient parameter with the signature, and the encrypted standardized gradient parameter with the signature is decrypted using a threshold to obtain a global gradient matrix;

[0027] Based on the global gradient matrix, a global prediction model is generated through weighted aggregation.

[0028] As a preferred scheme of the integrated chain data flow transfer method for ore prospecting prediction based on the data lake technology, wherein: an integrated three-dimensional geological model, monitoring equipment state data and economic parameters are created at a data lake service layer, including the following steps,

[0029] A database connection pool is established at the data lake service layer, a GOCAD format file is read from the directory of the data lake, vertex and facet data are parsed, a structured three-dimensional grid object is generated, non-manifold edge detection is performed on the structured three-dimensional grid object, and an integrated three-dimensional geological model is obtained;

[0030] The Kafka consumer subscribes to the topic based on the database connection pool configuration, receives the device state data in real time, and obtains the abnormal score using a sliding window algorithm.

[0031] The economic parameter data is extracted from the LME API, energy directory and mine site labor cost table, and the exchange rate is converted to obtain standardized economic parameters.

[0032] As a preferred scheme of the ore prospecting prediction integrated chain data flow conversion method based on the data lake technology, the two-way mapping relationship between the physical data lake and the digital twin is constructed by the Kalman filtering algorithm with geological constraints, and the mineralization probability distribution map and the economic evaluation report are obtained, including the following steps,

[0033] The geological rules are quantified into constraint matrices to obtain a set of digital constraint parameters, a state space model is constructed, and the initialized filter parameters are obtained;

[0034] The drilling data in the data lake is read to obtain a standardized observation vector, and the Kalman filtering iteration with geological constraints is performed to obtain a three-dimensional space state vector;

[0035] The three-dimensional space state vector is mapped to the three-dimensional geological model to output the updated digital twin grid;

[0036] Based on the updated digital twin grid, the Kriging interpolation is performed to generate the mineralization probability distribution map;

[0037] Based on the LME and the mining cost, the economic evaluation report is obtained.

[0038] As a preferred scheme of the ore prospecting prediction integrated chain data flow conversion method based on the data lake technology, the multi-objective programming algorithm is used based on the mineralization probability distribution map and the economic evaluation report to generate the exploration target area priority list and the drilling layout scheme, including the following steps,

[0039] Based on the mineralization probability map and the economic evaluation report, the objective function is established to obtain a standardized multi-objective programming problem, and the NSGA-II algorithm is executed to obtain a target area coordinate list;

[0040] The A* algorithm is used for drilling path planning to obtain a drilling parameter table of azimuth angle and inclination angle;

[0041] Based on the LME API, the return on investment is calculated to obtain the drilling layout scheme.

[0042] In a second aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, wherein the computer program is executed by the processor to implement any step of the ore prospecting prediction integrated chain data flow conversion method based on the data lake technology according to the first aspect of the present application.

[0043] In a third aspect, the present application provides a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any step of the method for integrated chain data flow of ore-prospecting prediction based on data lake technology according to the first aspect of the present application.

[0044] The present application has the beneficial effects that: through federated learning and homomorphic encryption, safe cooperation of cross-mine data without domain outflow is realized, and real-time economic parameter optimization decision is integrated, thus solving the core problems of high data privacy risk, model update lag and market response disconnection in the prior art; geophysical, geochemical and remote sensing data are collected, stored in the local data lake node of the mine area after standardization processing, then a federated learning agent is deployed, gradient parameters are transmitted by homomorphic encryption, safe and efficient global model aggregation is realized, three-dimensional geological models, real-time monitoring data and economic parameters are integrated in the data lake service layer, a digital twin is constructed by a Kalman filtering algorithm with geological constraints, a mineralization probability distribution map and an economic evaluation report are generated, and finally an optimal exploration scheme is output based on a multi-objective planning algorithm. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Fig. 1 The flowchart of the method for integrated chain data flow of ore-prospecting prediction based on data lake technology.

[0047] Fig. 2 The schematic diagram of the intelligent prediction model training of the mine area.

[0048] Fig. 3 The schematic diagram of the target area deployment.

[0049] Fig. 4 The schematic diagram of the original exploration data collection. DETAILED DESCRIPTION

[0050] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.

[0051] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0052] Second, the term "one embodiment" or "an embodiment" as used herein means an implementation that can include one or more features, structures, or characteristics, but that does not mean that all of the features, structures, or characteristics are required in one implementation. The appearance of the phrase "in one embodiment" or "in an embodiment" in various places in the specification is not meant to be interpreted as an indication that all of the features, structures, or characteristics that are referred to in that embodiment are required in one or more implementations.

[0053] Referring to Figs. 1-4 For one embodiment of the present application, the embodiment provides a data lake technology-based integrated chain data flow conversion method for ore prediction, comprising the following steps:

[0054] S1, collect geophysical data, geochemical data and remote sensing data, and organize them into original exploration data.

[0055] S1.1 uses a proton magnetometer to perform wiring measurement to obtain original magnetic anomaly time series data and performs correction, and generates geophysical data by using Kriging interpolation on the corrected original magnetic anomaly time series data.

[0056] Specifically, the expression is,

[0057] ;

[0058] Wherein, is the corrected magnetic anomaly value, is the original observation value of the magnetometer, is the diurnal value recorded by the base station magnetometer at the observation time , is the time index, is the diurnal value recorded by the base station magnetometer at the reference time , and is the reference time.

[0059] S1.2 collects stream sediment samples to obtain original stream samples, inputs the original stream samples into an Olympus Vanta XRF instrument, performs correction to obtain element content data, and outputs the element content data to an IDW algorithm to obtain geochemical data.

[0060] Specifically, the expression is,

[0061] ;

[0062] Wherein, is the corrected element content data, is the original reading directly measured by the Olympus Vanta XRF instrument.

[0063] ; ​​

[0064] wherein, is the element content prediction value at the grid point, is the measured element content of the th sampling point, is the measured element content of the th sampling point, is the horizontal distance from the grid point to the sampling point

[0065] It should be noted that after collecting the stream sediment sample, the stream original sample is obtained, the stream original sample is input into the Olympus Vanta XRF instrument for measurement, the original element content reading is obtained, the original element content reading is corrected by using the correction formula, and the corrected element content data is obtained; the corrected element content data is input into the inverse distance weighted algorithm, the element content prediction value at the grid point is calculated, and the geochemical data is obtained through the inverse distance weighted algorithm.

[0066] S1.3, through the Sentinel-2 L2A level image, the multi-spectral data after radiation correction, through the PCA transformation formula, the principal component composite image is obtained, and the principal component composite image is input into the Crosta algorithm to obtain the remote sensing data.

[0067] Specifically, the expression is,

[0068] ;

[0069] wherein, is the principal component composite image, is the red band reflectivity, is the green band reflectivity, is the blue band reflectivity.

[0070] S1.4, the geophysical data, the geochemical data and the remote sensing data are spatially registered to obtain the original exploration data.

[0071] Further, the geophysical data, the geochemical data and the remote sensing data are uniformly converted to the CGCS2000 coordinate system, so as to ensure that all data adopt the same projection parameters; the gridded magnetic anomaly map in the geophysical data is resampled so as to keep consistent with the grid resolution of the geochemical data; the principal component composite image in the remote sensing data is geometrically corrected to eliminate the influence of terrain displacement; the sampling point data of the geochemical data is interpolated to the same grid node as the geophysical data by using the nearest neighbor interpolation method; the geophysical data, the geochemical data and the remote sensing data are aligned in the same geographical coordinate range through spatial superposition analysis, and the boundary coincidence degree of each data layer is checked; the registered geophysical data, the geochemical data and the remote sensing data are subjected to data format standardization processing to generate the original exploration data with unified spatial reference.

[0072] S2. Add metadata to the original exploration data and transmit to the local data lake node of the mining area to form the original exploration data set.

[0073] S2.1, add metadata tags to the original exploration data according to the geological information metadata standard, and obtain the complete original exploration data.

[0074] Further, according to the format requirements of the geological information metadata standard, metadata tags are added to the original exploration data, including data source (such as proton magnetometer, Olympus Vanta XRF instrument, Sentinel-2 L2A level image), collection time (UTC timestamp), coordinate range (longitude and latitude boundary under CGCS2000 coordinate system), data precision (such as magnetic anomaly value precision ±1nT) and processing method (such as Kriging interpolation, inverse distance weighted algorithm, PCA transformation) and other core metadata fields; the original exploration data after adding metadata tags is checked for integrity, and whether the required fields are missing is checked; the range of numerical parameters in the metadata tags is checked (such as longitude range 0°-180°, latitude range -90°-90°), and the format of the text parameters in the metadata tags is checked. After passing the check, the original exploration data containing complete metadata tags is generated.

[0075] S2.2, the complete original exploration data is transmitted to the local data lake node of the mining area through the special network encryption to form the original exploration data set.

[0076] Further, the complete original exploration data is transmitted through the 5G special network to establish a secure transmission channel, and the original exploration data containing metadata tags is encrypted using the AES-256 encryption algorithm; the MD5 checksum of the original exploration data is calculated before transmission and attached to the data packet header; the encrypted original exploration data is transmitted to the local data lake node of the mining area in blocks through the breakpoint resume protocol; after receiving the data block, the local data lake node of the mining area first verifies the MD5 checksum to ensure data integrity, and then decrypts the data using the corresponding key; the decrypted original exploration data is stored in the HDFS distributed file system of the local data lake node of the mining area according to the data type (geophysical data, geochemical data, remote sensing data) and is classified into the corresponding directory to form a structured original exploration data set.

[0077] S3, deploy a federal learning agent on the local data lake node of the mining area to obtain an intelligent prediction model of the mining area, and obtain the gradient parameters of the intelligent prediction model of the mining area through intelligent analysis algorithm.

[0078] S3.1, install Docker container on the local data lake node of the mining area, load federal learning image, obtain standardized running environment and output normalized exploration data.

[0079] Further, Ubuntu 20.04 LTS operating unit environment is deployed at the mining area data lake node, Docker engine is installed through the apt-get command, a pre-configured federated learning image (containing PyTorch1.9.0 framework and FATE 1.7 components) is pulled from the central mirror warehouse, and an isolated container running environment is created; the HDFS storage volume of the mining area local data lake node is mounted inside the container, and a read-write channel with the original exploration data set is established; the container network parameters are configured to ensure safe communication with the central aggregation server; after starting the container, a data preprocessing script is automatically executed to normalize the original exploration data according to the feature dimension, and normalized exploration data with a numerical range in the [0, 1] interval is output; the distribution characteristics of the normalized exploration data are verified to ensure the comparability of each feature dimension.

[0080] S3.2, read the normalized exploration data from the data lake HDFS directory to obtain a dimension-unified training tensor, use the GeLU activation function to build a 3-layer fully connected network, and obtain the mining area intelligent prediction model.

[0081] Further, the normalized exploration data is read from the / normalized_data / partition of the data lake HDFS directory, the data format is a Parquet file, and the normalized feature values of geophysical data, geochemical data and remote sensing data are included; the data dimension consistency is checked to ensure that all feature fields are complete and have no missing values; the read normalized exploration data is converted into a Float32 type tensor, and the tensor shape is [sample number x 128], where 128 corresponds to the total sum of the feature dimensions of geophysical data, geochemical data and remote sensing data;

[0082] A 3-layer fully connected network is built using the PyTorch framework, and the network structure is: input layer (128 dimensions), hidden layer (64 dimensions, using GeLU activation function), output layer (1 dimension, Sigmoid activation); the Xavier uniform initialization method is used to set the network weight parameters; the weighted cross-entropy loss function is used for training, the Adam optimizer is used for parameter updating, and the initial learning rate is set to 0.001; after 50 rounds of iterative training, the mining area intelligent prediction model is output.

[0083] S3.3, use the weighted cross-entropy loss function to train the mining area intelligent prediction model, and obtain the trained mining area intelligent prediction model and the gradient parameters of the mining area intelligent prediction model.

[0084] Further, the intelligent prediction model of the mining area is trained by using a weighted cross-entropy loss function, a gradient of each parameter in the intelligent prediction model of the mining area is calculated through a back propagation algorithm to obtain a complete gradient tensor, and the loss function values of the training set and the verification set are monitored in real time during the training process. When the loss of the verification set does not decrease for 5 consecutive rounds, the training is terminated. Finally, the trained intelligent prediction model of the mining area (and the gradient parameter matrix of the intelligent prediction model of the mining area under the current batch of data, the matrix dimension is consistent with the number of parameters of the intelligent prediction model of the mining area, and the Adam optimizer is used to update the parameters during the training process, and the initial learning rate is set to 0.001, and the batch size is 64.

[0085] S4, the gradient parameters of the intelligent prediction model of the mining area are encrypted by using a homomorphic encryption algorithm and uploaded to a central aggregation server, and a global prediction model is generated by weighted aggregation, including the following steps,

[0086] S4.1, the gradient parameters of the intelligent prediction model of the mining area are normalized by using the maximum absolute value, to obtain standardized gradient parameters, and the central aggregation server generates an encryption public key and a private key.

[0087] Further, the gradient parameter matrix obtained by training the intelligent prediction model of the mining area is subjected to a maximum absolute value normalization operation, which represents the maximum value of the absolute values of all elements in the gradient parameter matrix. Through the normalization processing, all gradient parameters are linearly transformed to the interval [-1, 1] to obtain a standardized gradient parameter matrix. The central aggregation server generates a 2048-bit Paillier homomorphic encryption key pair by using the RSA algorithm, and outputs a public key and a private key.

[0088] S4.2, the public key is used to encrypt the standardized gradient parameters to obtain encrypted standardized gradient parameters, and the encrypted standardized gradient parameters are added with ECDSA-SHA256 signatures and uploaded to the central aggregation server to obtain signed encrypted standardized gradient parameters, and the signed encrypted standardized gradient parameters are decrypted by using a threshold to obtain a global gradient matrix.

[0089] Specifically, the expression is,

[0090] ;

[0091] Wherein, is the global gradient matrix, is the aggregation weight of the i-th independent mining area node, is the Paillier homomorphic decryption function, is the standardized gradient parameter encrypted by the i-th independent mining area node, is the independent mining area node.

[0092] ​​S4.3, generating a global prediction model by weighted aggregation based on the global gradient matrix.

[0093] Further, the local gradient matrix is used to reflect the update of each parameter of the local model during the training process, and the global gradient matrix is used to update the global prediction model. The update process is based on the gradient descent method, and the parameter update amount in the global gradient matrix is applied to the global prediction model, so as to adjust the parameter configuration of the global prediction model. The generation of the global prediction model depends on the parameter update direction and amplitude provided by the global gradient matrix. By converting the information in the global gradient matrix into the adjustment of the model parameters, a global prediction model reflecting the overall training process is finally formed.

[0094] S5, creating an integrated three-dimensional geological model, monitoring equipment state data and economic parameters in the data lake service layer.

[0095] S5.1, establishing a database connection pool in the data lake service layer, reading GOCAD format files from the directory of the data lake, parsing vertex and face data, generating structured three-dimensional mesh objects, and performing non-manifold edge detection on the structured three-dimensional mesh objects to obtain an integrated three-dimensional geological model.

[0096] Further, the database connection pool of the PostgreSQL and HDFS hybrid storage architecture is configured in the data lake service layer, the maximum number of connections is set to 50, and the idle timeout is set to 300 seconds. GOCAD format files are read from the / 3d_models / partition of the data lake through the JDBC interface, vertex coordinates are obtained by parsing the VERTEX chapter, and triangular facet indexes are obtained by parsing the TFACE chapter. The vertex coordinates are stored as a floating-point matrix, and the triangular facet indexes are stored as an integer matrix. Structured three-dimensional mesh objects are generated by combining, topological integrity checks are performed on the three-dimensional mesh objects, and all edges are checked to see if they are shared by exactly two facets. Non-manifold edges and isolated vertices are marked, and the detected topological defects are repaired by the Poisson reconstruction algorithm to generate an integrated three-dimensional geological model that meets the Watertight requirement.

[0097] S5.2, configuring a Kafka consumer to subscribe to a topic based on the database connection pool, receiving equipment state data in real time, and using a sliding window algorithm to obtain an anomaly score.

[0098] Further, a dedicated consumer group is configured to subscribe to the equipment state topic, and JSON format drilling parameters (rotational speed, torque) are received in real time. The message body strictly includes the device ID, UTC timestamp and parameter key-value pair. A 60-second time window is used to calculate the anomaly score, and an alarm is triggered when the standard deviation rotational speed > 3.0σ or torque > 2.5σ.

[0099] Specifically, the expression is,

[0100] ;

[0101] wherein, is the abnormal score, is the time of the measured value of the device parameter, is the arithmetic mean of the parameter within the sliding window, is the standard deviation of the parameter within the sliding window, is the smoothing factor.

[0102] S5.3, extract economic parameter data from LME API, energy directory and mine area labor cost table, perform exchange rate conversion to obtain standardized economic parameters.

[0103] Further, by LME API, real-time metal prices (such as copper price USD / ton) are obtained, power / diesel guiding price is read from national energy directory, and human cost table is exported from mine area PostgreSQL database. The converted price data is combined with the local cost of the mine area to generate standardized economic parameters.

[0104] S6, by the Kalman filtering algorithm with geological constraints, the bidirectional mapping relationship of the physical data lake-digital twin is constructed, and the mineralization probability distribution map and the economic evaluation report are obtained.

[0105] S6.1, the geological rules are quantified into constraint matrices to obtain a set of digital constraint parameters, a state space model is constructed, and the initialized filter parameters are obtained.

[0106] Further, the geological expert rules (such as fault occurrence and stratigraphic contact relationship) are quantified into constraint matrices and constraint vectors. The constraint matrix represents the relationship between density change and fault zone. Based on the vertex coordinates and physical parameters of the three-dimensional geological model, a 5-dimensional state vector is constructed, the Kalman filter parameters are initialized, including the state covariance matrix and the process noise matrix, and finally a state space model with geological constraints is formed.

[0107] S6.2, read the drilling data in the data lake to obtain a standardized observation vector, and perform Kalman filtering iteration with geological constraints to obtain a three-dimensional space state vector.

[0108] Specifically, the expression is,

[0109] ;

[0110] wherein, is the three-dimensional space state vector at time , is the prior three-dimensional space state prediction at time , is the three-dimensional space state prediction at time Lower Kalman gain matrix, actual observation value at time , observation matrix, geological constraint matrix, geological constraint vector.

[0111] S6.3, map the three-dimensional space state vector to the three-dimensional geological model, and output the updated digital twin grid.

[0112] Further, the three-dimensional space state vector output by the Kalman filter is mapped to the three-dimensional geological model grid vertex, wherein the coordinate components are corrected by bilinear interpolation to the spatial position of the corresponding vertex, and the physical property parameters are written in the vertex attribute table in the form of attribute fields. Non-manifold edge detection is performed on the updated grid, and abnormal patches with holes or self-intersection are marked and repaired. Through Laplace smoothing algorithm, the grid topology is optimized, and finally a digital twin grid meeting the Watertight requirement is generated, which is stored as a GOCAD format file containing the updated vertex coordinate matrix, triangular patch index matrix and physical property parameter label. The digital twin grid and the exploration data in the physical data lake maintain consistency in spatial reference system and attribute fields, with a planar positioning error controlled within ±0.5m and a vertical error controlled within ±0.2m.

[0113] S6.4, perform Kriging interpolation based on the updated digital twin grid to generate a mineralization probability distribution map.

[0114] Further, the vertex coordinates and corresponding mineralization indicator values are extracted from the digital twin grid, a three-dimensional space index is established to speed up the search of adjacent points, a Kriging matrix is constructed, and point Kriging estimation is performed on the 50m×50m×20m grid nodes to generate a mineralization probability distribution map.

[0115] S6.5, obtain an economic evaluation report based on LME and mining cost.

[0116] Further, real-time metal prices are obtained through LME API, and mining costs are extracted from the mining area database to obtain an economic evaluation report.

[0117] Specifically, the expression is,

[0118] ;

[0119] wherein, NPV is the net present value, is the ore sales revenue in the year, is the time value of money, is the year number, is the initial investment amount, for the first year of production.

[0120] S7, based on the mineralization probability distribution map and the economic evaluation report, a multi-objective programming algorithm is used to generate a prospecting target area priority list and a drilling layout scheme.

[0121] S7.1, based on the mineralization probability map and the economic evaluation report, a target function is established to obtain a standardized multi-objective programming problem, and the NSGA-II algorithm is executed to obtain a target area coordinate list.

[0122] Further, based on the mineralization probability map and the economic evaluation report, a double-objective optimization function is constructed: maximizing the total mineralization probability and minimizing the total production cost, and the NSGA-II algorithm is used to execute optimization, a population is generated through simulated binary crossover and polynomial mutation, and a Pareto optimal solution set is output after non-dominated sorting and congestion, and the final target area coordinate list contains the central position in the CGCS2000 coordinate system and the mineralization probability, and the target area coordinate list is obtained.

[0123] S7.2, use A* algorithm for drilling path planning, and get the drilling parameter table of azimuth and inclination.

[0124] Further, based on the digital twin grid and the target area coordinate list, the A* algorithm is used to search for the optimal drilling path in three-dimensional space, the node cost is obtained through the heuristic function, the azimuth and inclination are obtained according to the path node coordinates, and the path weight is adjusted considering the rock strength and fault avoidance constraints; finally, a drilling parameter table containing drilling ID, azimuth, inclination and design depth is generated, and a prospecting target area priority list is obtained.

[0125] S7.3, based on LME API, calculate the rate of return on investment, and get the drilling layout scheme.

[0126] Specifically, the expression is,

[0127] ;

[0128] wherein, the rate of return on investment, the amount of metal, the price, and Y is the drilling cost.

[0129] The embodiment also provides a computer device suitable for the case of the integrated chain data flow conversion method for ore prospecting prediction based on data lake technology, comprising: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the integrated chain data flow conversion method for ore prospecting prediction based on data lake technology proposed in the above embodiment.

[0130] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, an operator network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0131] The embodiment also provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the method for realizing integrated chained data flow conversion of ore-finding prediction based on a data lake technology as described in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0132] In summary, the present application solves the core problems of high data privacy risk, model update lag and market response disconnection in the prior art by realizing cross-mining area data non-domain safe cooperation through federated learning and homomorphic encryption, and integrating real-time economic parameter optimization decision, collects geophysical, geochemical and remote sensing data, stores them to the local data lake node of the mining area after standardization processing, then deploys the federated learning agent, transmits the gradient parameters by homomorphic encryption, realizes safe and efficient global model aggregation, integrates the three-dimensional geological model, real-time monitoring data and economic parameters in the data lake service layer, constructs the digital twin body through the Kalman filtering algorithm with geological constraints, generates the mineralization probability distribution map and economic evaluation report, and finally outputs the optimal exploration scheme based on the multi-objective planning algorithm.

[0133] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A chain-like data flow method for mineral exploration prediction based on data lake technology, characterized in that: The method comprises the following steps of: collecting geophysical data, geochemical data and remote sensing data to form original exploration data; adding metadata to the original exploration data and transmitting the original exploration data to a local data lake node of a mine area to form an original exploration data set; A federal learning agent is deployed on the local data lake node of the mine area to obtain a mine area intelligent prediction model, and a gradient parameter of the mine area intelligent prediction model is obtained through an intelligent analysis algorithm; A Docker container is installed on the data lake node, a federal learning image is loaded, a standardized running environment is obtained, and normalized exploration data is output; The normalized exploration data is read from the data lake HDFS directory to obtain a dimension-unified training tensor, a 3-layer fully connected network is constructed using a GeLU activation function, and a mine area intelligent prediction model is obtained; The mine area intelligent prediction model is trained using a weighted cross-entropy loss function to obtain a trained mine area intelligent prediction model and a gradient parameter of the mine area intelligent prediction model; The gradient parameter of the mine area intelligent prediction model is encrypted using a homomorphic encryption algorithm and uploaded to a central aggregation server to generate a global prediction model through weighted aggregation; The gradient parameter of the mine area intelligent prediction model is normalized by the maximum absolute value to obtain a standardized gradient parameter, and an encryption public key and a private key are generated based on the central aggregation server; The standardized gradient parameter is encrypted using the public key to obtain an encrypted standardized gradient parameter, an ECDSA-SHA256 signature is added to the encrypted standardized gradient parameter, and the encrypted standardized gradient parameter with the signature is uploaded to the central aggregation server to obtain a signed encrypted standardized gradient parameter, and the signed encrypted standardized gradient parameter is decrypted using a threshold to obtain a global gradient matrix; Based on the global gradient matrix, a global prediction model is generated through weighted aggregation; An integrated three-dimensional geological model, monitoring equipment state data and economic parameters are created in the data lake service layer, a bidirectional mapping relationship between the physical data lake and the digital twin is constructed through a geologically constrained Kalman filtering algorithm, and a mineralization probability distribution map and an economic evaluation report are obtained; A bidirectional mapping relationship between the physical data lake and the digital twin is constructed through a geologically constrained Kalman filtering algorithm to obtain a mineralization probability distribution map and an economic evaluation report, including the following steps, Quantize the geological rules into a constraint matrix to obtain a digital constraint parameter set, construct a state space model, and obtain initialized filter parameters; Read the drilling data in the data lake to obtain a standardized observation vector, perform geologically constrained Kalman filtering iteration, and obtain a three-dimensional space state vector; Map the three-dimensional space state vector to the three-dimensional geological model to output an updated digital twin grid; Based on the updated digital twin grid, perform Kriging interpolation to generate a mineralization probability distribution map; Based on the LME and the mining cost, an economic evaluation report is obtained; Based on the mineralization probability distribution map and the economic evaluation report, a multi-objective planning algorithm is used to generate a target area priority list and a drilling layout scheme; Based on the mineralization probability map and the economic evaluation report, a target function is established to obtain a standardized multi-objective planning problem, and an NSGA-II algorithm is executed to obtain a target area coordinate list; An A* algorithm is used for drilling path planning to obtain a drilling parameter table of azimuth angle and inclination angle; Based on the LME API, the rate of return on investment is calculated to obtain the drilling arrangement scheme.

2. The integrated chain data flow conversion method for ore-finding prediction based on data lake technology according to claim 1, characterized in that: Collect geophysical data, geochemical data and remote sensing data, and organize them into raw exploration data, The method comprises the following steps, Use the proton magnetometer to perform wiring measurement to obtain raw magnetic anomaly time series data, and perform correction; and use the Kriging interpolation to generate geophysical data from the corrected raw magnetic anomaly time series data. Collect stream sediment samples to obtain raw stream samples, input the raw stream samples into an Olympus Vanta XRF instrument, perform correction to obtain element content data, and output the element content data to an IDW algorithm to obtain geochemical data. Through the Sentinel-2 L2A level image, the radiation-corrected multi-spectral data, and the PCA transformation formula, a principal component composite image is obtained; and the principal component composite image is input into a Crosta algorithm to obtain remote sensing data. The geophysical data, the geochemical data and the remote sensing data are spatially registered to obtain raw exploration data.

3. The integrated chain data flow conversion method for ore-finding prediction based on data lake technology according to claim 2, characterized in that: In the raw exploration data, metadata is added, and the raw exploration data is transmitted to a local data lake node of a mining area to form a raw exploration data set, including the following steps, The raw exploration data is checked by adding metadata tags according to the geological information metadata standard to obtain complete raw exploration data. The complete raw exploration data is transmitted to the local data lake node of the mining area through encryption of a special network to form the raw exploration data set.

4. The integrated chain data flow conversion method for ore-finding prediction based on data lake technology according to claim 3, characterized in that: In the data lake service layer, an integrated three-dimensional geological model, monitoring equipment state data and economic parameters are created, including the following steps, In the data lake service layer, a database connection pool is established, GOCAD format files are read from the directory of the data lake, vertex and face data are parsed, structured three-dimensional mesh objects are generated, non-manifold edge detection is performed on the structured three-dimensional mesh objects, and an integrated three-dimensional geological model is obtained; Based on the database connection pool, a Kafka consumer is configured to subscribe to a topic, real-time reception of equipment state data is performed, and a sliding window algorithm is used to obtain an abnormal score; Economic parameter data are extracted from the LME API, an energy directory and a mining area labor cost table, exchange rate conversion is performed, and standardized economic parameters are obtained.

5. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that: The processor executes the computer program to realize the steps of the data lake technology-based integrated chain data flow conversion method for ore prediction according to any one of claims 1-4.

6. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to realize the steps of the data lake technology-based integrated chain data flow conversion method for ore prediction according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method and system for predicting prospecting target area based on geological three-dimensional modeling

    CN119850863A

  • Intelligent planning method for geological mineral exploration analysis model

    CN120124945A