A geological survey optimization method and system based on big data analysis

Through big data analysis and deep learning technology, the geological survey paths and cost allocation are optimized, and the shortcomings of path planning and cost control in traditional geological surveys are solved, and more efficient and accurate geological surveys are achieved.

CN119514887BActive Publication Date: 2025-05-16SHANDONG INST OF GEOLOGICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510072495.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Traditional geological surveys have shortcomings in path planning and cost control, resulting in increased survey costs and wasted time, and it is difficult to reasonably allocate cost weights in different regions.

Method used

The geological survey optimization method based on big data analysis is adopted, and the multi-dimensional geological data is obtained, key geological features are extracted using convolutional neural networks, survey paths and key areas are determined, and cost weights are allocated based on topographic complexity, historical survey data and traffic accessibility. The cost matrix is ​​optimized using a simulated annealing algorithm, and finally the sampling point layout is optimized through the genetic algorithm.

Benefits of technology

It improves the efficiency and accuracy of geological surveys, reduces survey costs, optimizes resource allocation, and enhances the pertinence and effectiveness of survey plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119514887B_ABST
    Figure CN119514887B_ABST
Patent Text Reader

Abstract

The present invention provides a geological survey optimization method and system based on big data analysis, which relates to the field of data processing technology. The method includes: assigning different cost weights to different regions; constructing an initial cost matrix based on geological characteristics, constantly fine-tuning node costs and calculating evaluation function values, gradually cooling down until the final temperature is reached, and outputting an optimized inter-node cost matrix; by initializing distance and access arrays, constantly selecting unvisited nodes and updating the distances corresponding to their adjacent nodes until all nodes are accessed, and finally outputting the path corresponding to the starting point to the target area; establishing a mineral resource evaluation model based on multi-source geological data, and establishing and optimizing a resource prediction model by calculating indicator contribution, screening key areas, evaluating survey priorities, and selecting sampling points, and finally predicting the distribution of geological resources in the entire study area. The present invention can improve the efficiency and accuracy of geological surveys.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a geological survey optimization method and system based on big data analysis. Background Art

[0002] By integrating multi-dimensional geological data, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data, a more comprehensive and detailed geological information model can be constructed. These data not only cover the morphological characteristics of the surface, but also deeply reveal the underground structure and resource distribution patterns.

[0003] However, traditional geological surveys have some shortcomings in path planning and cost control. For example, due to the lack of comprehensive and real-time data support, the selection of survey paths is not optimized, which easily leads to increased survey costs and waste of time. At the same time, the geological conditions and accessibility of different regions vary greatly. How to reasonably allocate cost weights to different regions is also a problem that traditional methods are difficult to solve. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a geological survey optimization method and system based on big data analysis, which can improve the efficiency and accuracy of geological surveys.

[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0006] In a first aspect, a geological survey optimization method based on big data analysis comprises:

[0007] Obtain multi-dimensional data related to geological surveys, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data;

[0008] Preprocessing the multi-dimensional data to obtain preprocessed multi-dimensional data;

[0009] Use convolutional neural networks to automatically extract key geological features from pre-processed multi-dimensional data, and train the model to learn the mapping relationship from data to labels, and finally output geological structural features;

[0010] Determine the starting point of geological survey and the target area to be investigated based on the geological structure characteristics;

[0011] Different cost weights are assigned to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; an initial cost matrix is ​​constructed based on geological characteristics, and the node costs are continuously fine-tuned and the evaluation function values ​​are calculated, gradually cooling down until the final temperature is reached, and the optimized node cost matrix is ​​output;

[0012] By initializing the distance and visit arrays, we continuously select unvisited nodes and update the distances of their adjacent nodes until all nodes are visited, and finally output the path from the starting point to the target area.

[0013] Establish a mineral resource evaluation model based on multi-source geological data, calculate the contribution of indicators, screen key areas, evaluate survey priorities, select sampling points, establish and optimize resource prediction models, and ultimately predict the distribution of geological resources in the entire study area;

[0014] By dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to optimize the survey plan, the corresponding sampling point layout is finally obtained;

[0015] The arrangement of sampling points is optimized to obtain an optimized geological survey plan.

[0016] Furthermore, a convolutional neural network is used to automatically extract key geological features from the preprocessed multi-dimensional data, and the model is trained to learn the mapping relationship from data to labels, and finally outputs geological structural features, including:

[0017] Use the convolutional neural network in the deep learning model to automatically extract key features from the pre-processed multi-dimensional data, including rock formation trends, fault distribution patterns, and element content anomalies;

[0018] The deep learning model is trained using the preset data set and labels. During the training process, the weights and biases of the deep learning model are adjusted so that the deep learning model learns the mapping relationship from input data to output labels to obtain a trained deep learning model.

[0019] After the training is completed, the preprocessed multi-dimensional data is input into the trained deep learning model, and the trained deep learning model will output the corresponding geological structure characteristics.

[0020] Furthermore, the cost weight calculation formula is:

[0021] ;

[0022] in, Indicates from the area To area The cost weight of The weight coefficient representing the complexity of the terrain; Indicates area To area The elevation difference between Indicates the maximum elevation difference in the study area; Indicates area To area The slope difference between represents the maximum slope difference within the study area; Represents the weight coefficient of geological disaster risk; Indicates from the area To area The number of geological disaster points between Indicates the total number of geological hazard points in the study area; Indicates the number of historical activities of geological hazard sites; represents the total number of years in the observation period, which is used to calculate the historical activity frequency; A scaling factor representing the frequency of historical activity; Indicates the area of ​​environmental damage caused by geological disasters; represents the total area of ​​the study area; A weight indicating the severity of environmental damage; A scaling factor indicating the degree of damage caused by environmental destruction; The weight coefficient representing traffic accessibility; Indicates from the area To area travel time; Represents the maximum travel time within the study area.

[0023] Furthermore, an initial cost matrix is ​​constructed based on geological characteristics, the node costs are constantly fine-tuned and the evaluation function values ​​are calculated, the temperature is gradually reduced until the final temperature is reached, and the optimized node cost matrix is ​​output, including:

[0024] Based on the characteristics of the geological survey area, the actual cost weights between each node are determined, and the cost weights are used to construct a The initial cost matrix ,in, is the number of nodes, the initial cost matrix Elements in Represents a slave node To Node Initial cost; define the parameters of simulated annealing, including the initial temperature , Final temperature , cooling rate and maximum number of iterations ;

[0025] Randomly select two different node pairs in the cost matrix and ; Fine-tune the cost between these two node pairs to generate new candidate solutions ;

[0026] Compute a new candidate solution using the same merit function as the current solution The evaluation function value of ;

[0027] Calculate the difference in evaluation function values ,in, Represents the evaluation function value of the current solution;

[0028] Calculate the probability of accepting a new solution based on the acceptance probability of simulated annealing ; Generate a random number in the range [0, 1) ;if , then accept the new solution and set the current solution and ; Reduce the temperature according to the cooling rate; if the temperature drops below the final temperature, the iteration stops, and when it stops, the current solution is output As the optimized cost matrix, the optimized cost matrix is ​​a two-dimensional array, in which each element represents the optimized cost between two nodes in the geological survey area.

[0029] Furthermore, the calculation formula of the evaluation function is:

[0030] ;

[0031] in, is the number of nodes; is the current cost matrix, representing the node and nodes direct costs between is a distance coefficient; Representation Node and nodes The geographical distance between is the height coefficient; Represents a slave node and nodes height changes; Represents the evaluation function value calculated based on the current cost matrix.

[0032] Furthermore, by initializing the distance and visit arrays, unvisited nodes are continuously selected and the distances corresponding to their adjacent nodes are updated until all nodes are visited, and finally the path corresponding to the starting point to the target area is output, including:

[0033] Set the starting point to node B;

[0034] Create an array dist[] of length R to store the shortest distance from the starting point B to each node; initially, set dist[B] to 0, indicating that the distance from the starting point to itself is 0, and the remaining elements are set to infinity, indicating that the initial distance from the starting point to the remaining nodes is unknown;

[0035] Create a Boolean array visited[] of length R to mark whether each node has been visited. Initially, all elements are set to false, indicating that no node has been visited.

[0036] Select an unvisited node u from the array dist[] and mark node u as visited, that is, set visited[u] to true;

[0037] Traverse each node v adjacent to node u, and check whether the distance from the starting point through node u to node v is less than the currently recorded dist[v], that is, dist[u] + C[u][v], where C[u][v] represents the cost from node u to node v; if so, update dist[v] to make it equal to the distance from node u to node v until all nodes have been visited;

[0038] According to the updated array dist[], the constructed target path from the starting point to the target area is output.

[0039] Furthermore, a mineral resource evaluation model is established based on multi-source geological data. By calculating the contribution of indicators, screening key areas, evaluating survey priorities, and selecting sampling points, a resource prediction model is established and optimized, and finally the distribution of geological resources in the entire study area is predicted, including:

[0040] Establish a mineral resource evaluation model based on geological, geophysical and geochemical multi-source data;

[0041] Select the corresponding geological factors as evaluation indicators, calculate the contribution of each evaluation indicator to resource distribution, and derive the probability of resource distribution based on the contribution. The evaluation indicators include lithology, structure and alteration.

[0042] According to the probability of resource distribution, select the corresponding probability areas as the key investigation objects;

[0043] According to the geological structure characteristics, lithological conditions and mineralization laws, the key survey objects are evaluated to obtain the evaluation results; according to the evaluation results, the survey priority of each area is determined;

[0044] In the high-probability area, the sampling point locations are preliminarily selected based on the geological conditions and topographic features;

[0045] Establish resource prediction model based on geological survey data, resource distribution probability and sampling point data;

[0046] Verify the resource prediction model through the data of known mining points to obtain the verification results; optimize the resource prediction model according to the verification results to obtain the optimized resource prediction model;

[0047] The optimized resource prediction model is used to predict the resource distribution of the entire study area to obtain the distribution of geological resources.

[0048] Furthermore, by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to optimize the survey plan, the corresponding sampling point arrangement is finally obtained, including:

[0049] Determine the survey objectives, which include maximizing resource discovery and minimizing cost or time; divide the survey area into grids and use binary coding, where 1 indicates that the grid is selected for survey and 0 indicates that it is not selected;

[0050] Multiple initial survey plans are randomly generated, each plan represents a sampling point arrangement method, and the survey plan is an individual in the population;

[0051] According to the survey objectives and constraints, a fitness function is constructed to evaluate the quality of the survey plan;

[0052] Evaluate the quality of each individual in the population according to the fitness function and select the corresponding individuals to enter the next generation; randomly select two individuals for crossover operation to generate a new survey plan;

[0053] The individuals are mutated with a certain probability, that is, the survey status of the grid is randomly changed, and the selection, crossover and mutation operations are repeated until the preset number of iterations is reached to obtain the optimal individual, and the optimal individual is decoded into a specific survey plan, including the location and number of sampling points.

[0054] Furthermore, the calculation formula of the fitness function is:

[0055] ;

[0056] in, and γ represent weight coefficients; Indicates The predicted resource volume in each grid; represents a binary decision variable, and for each grid, The value is 1 or 0, indicating whether to select the grid for investigation. , then it means grids are selected for investigation; if , it means that the grid is not investigated; Indicates the total number of grids in the survey area, indicating how many grids the entire survey area is divided into; represents the unit cost, the average cost incurred by sampling on each selected grid when performing a geological survey; It represents the total survey time, which means the time required to complete the entire geological survey program.

[0057] In the second aspect, a geological survey optimization system based on big data analysis includes:

[0058] The acquisition module is used to acquire multi-dimensional data related to geological surveys, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data; and pre-process the multi-dimensional data to obtain pre-processed multi-dimensional data;

[0059] The analysis module is used to automatically extract key geological features from pre-processed multi-dimensional data using a convolutional neural network, and through training, the model learns the mapping relationship from data to labels, and finally outputs geological structural features;

[0060] The determination module is used to determine the starting point of the geological survey and the target area to be investigated based on the geological structure characteristics;

[0061] The allocation module is used to assign different cost weights to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; construct an initial cost matrix based on geological characteristics, continuously fine-tune the node cost and calculate the evaluation function value, gradually cool down until the final temperature is reached, and output the optimized node cost matrix; by initializing the distance and access array, continuously select unvisited nodes and update the corresponding distances of their adjacent nodes until all nodes are visited, and finally output the path from the starting point to the target area; establish a mineral resource evaluation model based on multi-source geological data, calculate the indicator contribution, screen key areas, evaluate the survey priority, select sampling points, establish and optimize the resource prediction model, and finally predict the distribution of geological resources in the entire study area;

[0062] The calculation module is used to optimize the survey plan by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to finally obtain the corresponding sampling point layout; the sampling point layout is optimized to obtain the optimized geological survey plan.

[0063] According to a third aspect, a computing device includes:

[0064] one or more processors;

[0065] The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described.

[0066] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described is implemented.

[0067] The above scheme of the present invention includes at least the following beneficial effects.

[0068] Preprocessing multi-dimensional data can not only clean and correct errors and anomalies in the data, but also improve the quality and availability of the data, ensuring the accuracy of subsequent analysis. Deep mining of preprocessed data can more accurately identify geological structural features. By comprehensively considering factors such as terrain complexity, historical survey difficulties and risk points, and traffic accessibility, reasonable cost weights are assigned to different regions, and a cost matrix is ​​constructed, which effectively reduces survey costs and improves work efficiency.

[0069] By combining geological structural characteristics, resource distribution probability and target path, we can accurately delineate key areas for geological surveys, optimize the layout of sampling points, and predict the distribution of geological resources, thereby greatly improving the pertinence and effectiveness of geological surveys. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a flow chart of a geological survey optimization method based on big data analysis provided by an embodiment of the present invention.

[0071] Figure 2 It is a schematic diagram of a geological survey optimization system based on big data analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0073] like Figure 1 As shown, an embodiment of the present invention proposes a geological survey optimization method based on big data analysis, and the method comprises the following steps:

[0074] Step 11, obtaining multi-dimensional data related to geological surveys, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data;

[0075] Step 12, preprocessing the multi-dimensional data to obtain preprocessed multi-dimensional data;

[0076] Step 13, using a convolutional neural network to automatically extract key geological features from the preprocessed multi-dimensional data, and through training, the model learns the mapping relationship from data to labels, and finally outputs geological structural features;

[0077] Step 14, determining the starting point of the geological survey and the target area to be investigated based on the geological structure characteristics;

[0078] Step 15: assign different cost weights to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; construct an initial cost matrix based on geological characteristics, continuously fine-tune the node costs and calculate the evaluation function value, gradually cool down until the final temperature is reached, and output the optimized node cost matrix;

[0079] Step 16, by initializing the distance and visit array, continuously selecting unvisited nodes and updating the distances corresponding to their adjacent nodes until all nodes are visited, and finally outputting the path corresponding to the starting point to the target area;

[0080] Step 17, establish a mineral resource evaluation model based on multi-source geological data, calculate the indicator contribution, screen key areas, evaluate the survey priority, select sampling points, establish and optimize the resource prediction model, and finally predict the geological resource distribution of the entire study area;

[0081] Step 18, by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm to perform selection, crossover, and mutation operations to optimize the survey plan, and finally obtain the corresponding sampling point arrangement;

[0082] Step 19, optimizing the arrangement of sampling points to obtain an optimized geological survey plan.

[0083] In the embodiment of the present invention, step 11, comprehensively collect multi-dimensional data to ensure the integrity and diversity of information, provide rich data sources for subsequent analysis, and improve the accuracy and comprehensiveness of geological surveys. Step 12, through data preprocessing, cleaning and standardization of data, eliminate outliers and noise, and improve data quality. Step 13, the deep learning algorithm can deeply explore the complex relationship between data, more accurately identify geological structural features, improve the automation and intelligence level of geological feature identification, and reduce manual intervention and subjective judgment. Step 14, based on the results of deep learning analysis, scientifically determine the starting point and key areas of the survey, make the survey work more focused and efficient, and avoid waste of resources and ineffective labor. Step 15, comprehensively consider various factors to allocate cost weights to different regions, and construct a cost matrix, which is helpful to more reasonably evaluate the survey cost, optimize resource allocation, and reduce the overall survey cost. Step 16, based on the cost matrix, path planning can ensure the economy and feasibility of the survey path, improve survey efficiency, and reduce unnecessary distance and time consumption. Step 17, comprehensively determine key areas and optimize the layout of sampling points based on various information, which can more accurately predict the distribution of geological resources and improve the pertinence and effectiveness of geological surveys. Step 18, formulating a geological survey plan based on comprehensive analysis can ensure the scientificity and systematicness of the survey work and improve the reliability of the survey results. Step 19, optimizing the plan can further improve the rationality and efficiency of the survey plan, ensure that the survey work can proceed smoothly and achieve the expected goals, while reducing potential risks and costs.

[0084] The above step 11 specifically includes the following steps:

[0085] Collect basic geological data such as stratigraphic structure, rock distribution, fault information, etc. through field surveys and geological mapping. Use geophysical exploration techniques (such as seismic exploration, gravity exploration, magnetic exploration, etc.) to obtain the distribution characteristics of underground physical fields, and then infer the underground geological structure. By collecting and analyzing geochemical samples such as rocks, soils, and water samples, understand the distribution and migration patterns of elements, and provide chemical information for geological surveys. With the help of satellite remote sensing, aerial remote sensing and other technical means, obtain surface information such as landforms, vegetation coverage, surface rock types, etc. on a large scale and quickly. Sort and summarize the information and data of previous geological surveys, including geological maps, survey reports, exploration data, etc., to provide basic information and reference for new geological surveys. Integrate the collected multi-dimensional data to form a unified data set. This includes data format conversion, coordinate system unification, data quality assessment, etc., to ensure the accuracy and reliability of subsequent analysis.

[0086] The above step 12 specifically includes the following steps:

[0087] Check and delete duplicate records in the data set to avoid repeated calculations during analysis; for missing values ​​in the data set, fill in (such as using mean, median, etc.), interpolate, or delete them according to the nature of the data and analysis requirements; identify and process outliers in the data, which may be caused by measurement errors, data entry errors, etc., and need to be reasonably corrected or eliminated; convert data of different dimensions (such as length, weight, time, etc.) so that they are in the same dimension to facilitate subsequent data analysis and comparison; scale the data so that it falls into a small specific interval (such as 0-1 or -1-1) to eliminate dimensional differences and order of magnitude differences between the data and improve the convergence speed and accuracy of the algorithm. For categorical data or qualitative data, necessary coding conversion is performed, such as converting text information such as geological types and rock names into digital codes to facilitate computer processing and analysis; for high-dimensional data, principal component analysis (PCA), factor analysis and other methods are used to reduce dimensionality to reduce data redundancy and complexity and improve analysis efficiency; according to analysis requirements, the data is divided into training sets, validation sets and test sets to facilitate model training, validation and evaluation; when the data volume is large, random sampling, stratified sampling and other methods are used to sample the data to reduce the amount of computational work in data processing and analysis.

[0088] In an embodiment of the present invention, the above step 13, using a convolutional neural network to automatically extract key geological features from the pre-processed multi-dimensional data, and through training, the model learns the mapping relationship between data and labels, and finally outputs geological structure features, may include:

[0089] Step 131, using the convolutional neural network in the deep learning model to automatically extract key features from the pre-processed multi-dimensional data, including rock formation trends, fault distribution patterns, and element content anomalies, specifically including:

[0090] Select a convolutional neural network (CNN) model that can handle multi-dimensional images or data matrices. Design the architecture of the CNN, including convolutional layers, pooling layers, activation functions, etc., based on the characteristics of geological data. Format the preprocessed multi-dimensional data (such as geological images, geophysical data matrices, etc.) into an input form that the CNN model can accept. Further normalize or standardize the data as needed. The convolutional layers of the CNN model automatically extract local features in the input data, such as the edges and textures of the rock formations. As the data propagates forward in the network, the convolutional layers gradually extract higher-level features, such as the direction of the rock formations and the distribution pattern of faults. The pooling layers reduce the dimensionality of the feature maps while retaining important features to enhance the robustness of the model. The last few layers of the CNN model output highly abstract feature maps that capture the key geological structural information in the input data. These feature maps are flattened into one-dimensional vectors as feature inputs for subsequent classification or regression tasks.

[0091] Step 132, using a preset data set and label to train the deep learning model. During the training process, by adjusting the weights and biases of the deep learning model, the deep learning model learns the mapping relationship from input data to output labels to obtain a trained deep learning model, specifically including:

[0092] Prepare a labeled dataset containing preprocessed multi-dimensional geological data and corresponding geological structure feature labels; divide the dataset into training set, validation set and test set for model training, tuning and evaluation; initialize the weights and bias parameters of the CNN model using random initialization or pre-trained models; train the CNN model using the training set, calculate the output of the model through forward propagation, calculate the loss function between the model output and the true label (such as cross entropy loss, mean square error, etc.), and calculate the gradient of the loss function to the model parameters through the back propagation algorithm; use an optimization algorithm (such as gradient descent) to update the weights and bias parameters of the model to minimize the loss function, and repeat the above process until the preset number of training rounds (epochs) is reached; evaluate the performance of the model, such as accuracy, recall, etc., on the validation set; adjust the model's hyperparameters (such as learning rate, batch size, etc.) or network structure based on the verification results to optimize the model performance.

[0093] Step 133, after the training is completed, the pre-processed multi-dimensional data is input into the trained deep learning model, and the trained deep learning model will output the corresponding geological structure features, specifically including:

[0094] Load the CNN model saved after training, including its weights and bias parameters; prepare the geological data to be analyzed, and ensure that it has undergone the same preprocessing steps; format the data into an input format acceptable to the model; input the preprocessed multi-dimensional data into the trained CNN model; the trained CNN model automatically extracts the key geological structural features of the data through forward propagation; the output layer of the trained CNN model will give the corresponding geological structural feature prediction results, such as the classification of rock formation trends and the probability of fault distribution; parse the output of the model and convert it into interpretable geological structural feature information.

[0095] In an embodiment of the present invention, through the convolutional neural network in step 131, key geological structural features, such as rock formation trends, fault distribution patterns, and element content anomalies, can be automatically extracted from complex multi-dimensional data. This automated feature extraction method not only greatly improves the efficiency of feature extraction, but also reduces manual participation and subjective judgment, thereby improving the accuracy and objectivity of feature extraction. In step 132, the deep learning model is trained to learn the mapping relationship from input data to output labels. The deep learning model has strong learning and representation capabilities, and can capture the complex relationships and nonlinear features between data, so as to more accurately identify and predict geological structural features. When processing new multi-dimensional data, the trained deep learning model (step 133) can accurately output the corresponding geological structural features based on the learned mapping relationship. This prediction method based on big data and deep learning has higher prediction accuracy and generalization ability than traditional methods, and provides more reliable analysis results for geological surveys. Accurate analysis of geological structural features provides strong support for geological survey decisions. The analysis results obtained through deep learning algorithms can help investigators more scientifically determine the focus of the investigation, optimize the layout of sampling points, and more accurately predict the distribution of geological resources, thereby improving the overall efficiency of geological surveys and resource utilization.

[0096] The above step 14, determining the starting point of the geological survey and the target area to be investigated based on the geological structure characteristics, may include:

[0097] Detailed analysis of geological structural features extracted by deep learning models, including interpretation of key information such as rock formation orientation, fault distribution pattern, and element content anomalies; evaluation of the geological importance of these features and their potential impact on the survey objectives. For example, certain rock formation orientations may indicate the presence of ore bodies, while fault distribution patterns may be related to groundwater flow or geological hazard risks.

[0098] Based on the results of feature analysis, one or more representative locations are selected as the starting points of the geological survey. These starting points are located in areas with significant geological structural features, rich information and easy access. The selection of starting points should take into account the convenience, safety and cost-effectiveness of the survey, as well as the expansibility of subsequent survey work. According to the distribution and importance of geological structural features, target areas that need to be investigated are delineated. These areas can be areas with complex rock strata, dense faults, and obvious element anomalies. When delineating target areas, multiple factors such as geological background, survey purpose, resource potential, and environmental risks should be considered comprehensively. For the determined starting points and target areas, a detailed geological survey plan is formulated, which includes the planning of survey routes, the arrangement of sampling points, the configuration of required equipment and personnel, etc. The survey plan should ensure that key geological data can be collected efficiently while ensuring the safety and feasibility of the survey work. Before actually conducting a geological survey, one or more field surveys are carried out to verify the rationality of the selection of the starting point and target area. During the field verification process, pay attention to observing and recording the geological phenomena on site, and compare and verify them with the previous analysis results to ensure the accuracy and effectiveness of the survey work. Through the above implementation process, step 14 can ensure that the geological survey work starts from representative and important locations and focus on in-depth investigations in key areas, thereby improving the efficiency and quality of the geological survey.

[0099] In the embodiment of the present invention, the calculation formula of the cost weight is:

[0100] ;

[0101] in, Indicates from the area To area The cost weight of The weight coefficient representing the complexity of the terrain; Indicates area To area The elevation difference between Indicates the maximum elevation difference in the study area; Indicates area To area The slope difference between represents the maximum slope difference within the study area; Represents the weight coefficient of geological disaster risk; Indicates from the area To area The number of geological disaster points between Indicates the total number of geological hazard points in the study area; Indicates the number of historical activities of geological hazard sites; represents the total number of years in the observation period, which is used to calculate the historical activity frequency; A scaling factor representing the frequency of historical activity; Indicates the area of ​​environmental damage caused by geological disasters; represents the total area of ​​the study area; A weight indicating the severity of environmental damage; A scaling factor indicating the degree of damage caused by environmental destruction; The weight coefficient representing traffic accessibility; Indicates from the area To area travel time; Represents the maximum travel time within the study area.

[0102] In the embodiment of the present invention, by integrating detailed terrain data such as elevation difference and slope difference, as well as the historical activity frequency and environmental damage degree of geological disaster points, the formula can more accurately reflect the actual cost differences between regions. This refined processing method helps to reduce estimation errors and improve the accuracy of cost calculation. Multiple weight coefficients and scaling coefficients are introduced in the formula, such as , , , , and , these coefficients can be adjusted according to the characteristics and needs of the specific study area. This design makes the formula more flexible and adaptable in dealing with different scenarios. The risk of geological hazards is incorporated into the calculation formula. By considering factors such as the number of geological hazard points, the number of historical activities and the degree of environmental damage, the formula can better reflect the impact of potential risks on costs. This helps to more comprehensively assess risks in the planning and decision-making process, so as to take more effective response measures. To area The formula can evaluate the impact of transportation accessibility on cost by taking into account the travel time and the maximum travel time in the study area.

[0103] In the embodiment of the present invention, an initial cost matrix is ​​constructed based on geological characteristics, the node costs are continuously fine-tuned and the evaluation function values ​​are calculated, the temperature is gradually reduced until the final temperature is reached, and the optimized node cost matrix is ​​output, including:

[0104] Based on the characteristics of the geological survey area, the actual cost weights between each node are determined, and the cost weights are used to construct a The initial cost matrix ,in, is the number of nodes, the initial cost matrix Elements in Represents a slave node To Node Initial cost; define the parameters of simulated annealing, including the initial temperature , Final temperature , cooling rate and maximum number of iterations ;

[0105] Randomly select two different node pairs in the cost matrix and ; fine-tune the cost between these two node pairs (increase or decrease a small amount) to generate a new candidate solution ;

[0106] Compute a new candidate solution using the same merit function as the current solution The evaluation function value of ;

[0107] Calculate the difference in evaluation function values ,in, Represents the evaluation function value of the current solution;

[0108] Calculate the probability of accepting a new solution based on the acceptance probability of simulated annealing ; Generate a random number in the range [0, 1) ;if , then accept the new solution and set the current solution and ; Reduce the temperature according to the cooling rate; if the temperature drops below the final temperature, the iteration stops, and when it stops, the current solution is output As the optimized cost matrix, the optimized cost matrix is ​​a two-dimensional array, in which each element represents the optimized cost between two nodes in the geological survey area.

[0109] In an embodiment of the present invention, by comprehensively considering the characteristics of the geological survey area, the actual cost weights between each node are determined, thereby constructing a more accurate cost matrix. This matrix not only reflects the physical distance between nodes, but also takes into account multiple factors such as terrain complexity and geological disaster risks, so that the cost calculation is closer to the actual situation. By optimizing the cost matrix through the simulated annealing algorithm, a more reasonable cost configuration between nodes can be found, thereby guiding the reasonable allocation of resources in the geological survey process. This helps to improve work efficiency and reduce unnecessary waste of resources. The simulated annealing algorithm is a heuristic search algorithm that can find an approximate optimal solution in a larger solution space. This means that in the face of a complex and changeable geological survey environment, the method can provide more decision-making options and enhance the flexibility and adaptability of decision-making. The simulated annealing algorithm can avoid falling into a local optimal solution to a certain extent by introducing randomness and probability acceptance criteria, thereby making it more likely to find a global optimal solution. At the same time, the algorithm has good convergence and stability, and can obtain more satisfactory results in a shorter time.

[0110] In the embodiment of the present invention, the calculation formula of the evaluation function is:

[0111] ;

[0112] in, is the number of nodes; is the current cost matrix, representing the node and nodes direct costs between is a distance coefficient; Representation Node and nodes The geographical distance between is the height coefficient; Represents a slave node and nodes height changes; Represents the evaluation function value calculated based on the current cost matrix.

[0113] In the embodiment of the present invention, the evaluation function comprehensively considers the direct cost , Geographical distance and height changes Multiple dimensions such as , can more comprehensively reflect the actual cost situation between nodes. This helps to find a more reasonable solution in the optimization process and improve the accuracy of cost calculation. and height coefficient , the evaluation function can be flexibly adjusted according to different geological survey needs and environmental characteristics. For example, in some cases, geographical distance may be more important, while in other cases, height changes may dominate. The setting of these coefficients makes the evaluation function more adaptable and flexible. The evaluation function provides a clear optimization goal for the simulated annealing algorithm. During the algorithm iteration process, by continuously calculating and comparing the values ​​of the evaluation function, the algorithm can be guided to search in a better direction, thereby finding a better cost matrix configuration.

[0114] In the embodiment of the present invention, step 16, by initializing the distance and access array, continuously selecting unvisited nodes and updating the distances corresponding to their adjacent nodes until all nodes are visited, and finally outputting the path corresponding to the starting point to the target area, may include:

[0115] Step 161, setting the starting point to node B;

[0116] Step 162, create an array dist[] of length R, which is used to store the shortest distance from the starting point B to each node; initially, dist[B] is set to 0, indicating that the distance from the starting point to itself is 0, and the remaining elements are set to infinity, indicating that the initial distance from the starting point to the remaining nodes is unknown;

[0117] Step 163, create a Boolean array visited[] of length R, which is used to mark whether each node has been visited; initially, all elements are set to false, indicating that no node has been visited;

[0118] Step 164, select a node u that has not been visited from the array dist[], and mark the node u as visited, that is, set visited[u] to true;

[0119] Step 165, traverse each node v adjacent to node u, and check whether the distance from the starting point through node u to node v is less than the currently recorded dist[v], that is, dist[u] + C[u][v], where C[u][v] represents the cost from node u to node v; if so, update dist[v] to make it equal to the distance from node u to node v, until all nodes have been visited;

[0120] Step 166, output the constructed target path from the starting point to the target area according to the updated array dist[].

[0121] In an embodiment of the present invention, by calculating the shortest distance from the starting point to each node, it is possible to ensure that the planned path is the most cost-effective. This is particularly important for geological survey tasks with limited resources, because it can help reduce unnecessary cost consumption, such as time, manpower, and materials. The arrays dist[] and visited[] are used to store intermediate results and visit status, avoiding repeated calculations and invalid traversals, thereby improving the efficiency of path planning. The choice of this data structure enables the algorithm to quickly locate unvisited nodes and update the shortest distance in a timely manner. The method does not rely on a specific starting point or target area, and can be flexibly applied to different combinations of starting points and end points. This means that in practical applications, the path planning strategy can be quickly adjusted as needed. The planned target path can be intuitively displayed to decision makers to help them better understand task requirements and resource allocation. The path planning method is not only applicable to static cost matrices, but can also be extended to dynamically changing cost environments according to actual needs. For example, when certain elements in the cost matrix change due to environmental changes, the method can still effectively replan the path.

[0122] In the embodiment of the present invention, step 17, based on multi-source geological data, establishes a mineral resource evaluation model, calculates the index contribution, screens key areas, evaluates the survey priority, selects sampling points, establishes and optimizes the resource prediction model, and finally predicts the geological resource distribution of the entire study area, including:

[0123] Step 171, based on geological, geophysical and geochemical multi-source data, establish a mineral resource evaluation model, specifically including:

[0124] Obtain geological maps, geological profiles, lithology distribution maps, etc. in the study area from geological survey institutions or public databases. Organize the distribution data of mineral deposits (points), including location, mineral type, reserves and other information. Collect mineral geological information such as mineralization type and deposit genesis. Obtain raw measurement data such as gravity exploration and magnetic exploration, which are usually saved in a specific format (such as GRAV, MAG). Ensure the integrity of the data and the accuracy of the measurement parameters. Collect rock and soil samples in the study area and conduct laboratory chemical analysis, such as spectral analysis, X-ray fluorescence analysis, etc., to obtain element content data.

[0125] Perform quality checks on all types of collected data, remove duplicate, erroneous or abnormal data points, and fill or exclude missing data based on data characteristics and context; convert geophysical data from raw formats to formats supported by GIS platforms, such as Shapefile (.shp) or GeoJSON. Standardize geochemical data to eliminate systematic errors in measurement methods and instruments, using Z-score standardization or other appropriate methods.

[0126] Import the cleaned and preprocessed data into GIS software (such as ArcGIS, QGIS, etc.). In the GIS environment, overlay the geological, geophysical and geochemical data layers to create a comprehensive data set; establish an attribute table for each data layer, record metadata (such as data source, collection date, etc.), and add quantitative information such as element content, physical property parameters, etc. to the attribute table.

[0127] Analyze historical mineral resource exploration data, including the location, scale, grade, etc. of ore deposits, and identify geological structural features (such as faults, folds), lithology types (such as intrusive rocks, sedimentary rocks), and geophysical (such as high gravity, magnetic anomalies) and geochemical anomalies (such as element enrichment areas) that are closely related to the distribution and enrichment of mineral resources. Determine which geological, geophysical and geochemical factors are most critical to the evaluation of mineral resource potential. Geological factors include geological structure, lithology, ore deposit type and mineralization; geophysical factors include gravity anomalies, magnetic anomalies and geochemical factors. Geochemical factors include element content and distribution. For example, abnormally high values ​​of elements such as tin, copper, and zinc may indicate corresponding ore deposits. Geochemical factors include element combinations and ratios. The combination relationship and ratio between elements can provide information about the mineralization environment, mineralization process and ore type. For example, the combination of elements such as Rb, Ba, and Sn is of great significance in some ore deposits. Geochemical factors include isotopic characteristics. The distribution of stable isotopes and radioactive isotopes can reveal the source of mineralization and the age of mineralization. Isotope data help distinguish different mineralization events and mineralization environments.

[0128] Based on data analysis and expert opinions, a list of key evaluation factors is compiled, each factor is briefly described, and its role in mineral resource evaluation is determined. For each key evaluation factor, specific quantitative standards are formulated based on its positive or negative impact on the potential of mineral resources. For example, for a certain lithology, it can be divided into three levels: high, medium, and low according to its mineralization, and the corresponding numerical values ​​are assigned to each level. According to the quantitative standards, each evaluation factor in the entire study area is assigned a value, which may involve converting the spatial data into raster or vector format and assigning a value to each cell or point.

[0129] Select a convolutional neural network (CNN) and design the architecture of the neural network, including the input layer, hidden layer, and output layer. Determine the number of neurons in the input layer to match the number of evaluation factors. Set the number of layers and neurons in the hidden layer and find the best configuration through experiments. Arrange the assigned evaluation factor data into a format acceptable to the neural network and perform data standardization or normalization to eliminate the dimensional differences between different factors.

[0130] Based on historical data and expert advice, set initial values ​​for the weights, thresholds, etc. of the neural network model. Random initialization and heuristic initialization methods can be used to determine them. Use historical data sets to train the neural network model, and adjust the weights and thresholds through the back-propagation algorithm. Divide the data set into a training set and a validation set to monitor the performance of the model on unknown data. Adjust the parameters of the neural network (such as learning rate, hidden layer structure, etc.) through multiple experiments and comparative analysis. Use cross-validation, grid search and other techniques to optimize model performance. Evaluate the model performance under different parameter configurations based on performance indicators (such as accuracy, recall, etc.) on the validation set, and select the best performing model as the final mineral resource evaluation model.

[0131] Step 172, select corresponding geological factors as evaluation indicators, calculate the contribution of each evaluation indicator to resource distribution, and obtain the probability of resource distribution according to the contribution, wherein the evaluation indicators include lithology, structure and alteration, specifically including:

[0132] Select evaluation indicators:

[0133] Lithology: Different lithologies have a significant impact on the formation and distribution of mineral resources. For example, certain lithologies may be more conducive to the deposition or formation of minerals. Therefore, lithology is used as an evaluation indicator.

[0134] Structure: Geological structures (such as faults, folds, etc.) play a key role in the distribution of mineral resources. Tectonic activity may lead to the migration, enrichment or formation of specific mineral deposits.

[0135] Alteration: Rock alteration (such as silicification, sericitization, etc.) is often closely related to mineralization. Alteration zones may indicate areas of hydrothermal activity and thus potential mineralization areas.

[0136] According to the comprehensive data set established in step 171, this data set should already contain the superposition information of multiple source data such as geology, geophysics and geochemistry. For lithology data, classify the lithology in the study area according to the lithology classification scheme (such as sedimentary rock, igneous rock, metamorphic rock, etc.), and assign a unique code to each lithology type. For structural data, identify and extract major structural elements such as faults and folds. Record the location (latitude and longitude or coordinates), scale (length, width, etc.) and activity (such as active faults, inactive faults, etc.) information of these structural elements. For alteration data, identify different types of alteration areas (such as silicification areas, sericitization areas, etc.) through image interpretation or field surveys, and record their intensity and spatial distribution. Ensure that the extracted data format is suitable for subsequent statistical analysis. It may be necessary to convert the data into a table form, where each row represents an observation point or area, and each column represents a different geological factor (lithology, structure, alteration, etc.).

[0137] Analyze the correlation between lithology, structure and alteration data and known mineral deposit data using statistical software (e.g., SPSS, R, etc.). This can be achieved by calculating correlation coefficients (e.g., Pearson correlation coefficient). Analyze the frequency or average grade of mineral deposits in different lithology types to determine the contribution of lithology to the distribution of mineral resources. For example, if deposits occur frequently and have high grades in a certain lithology, then this lithology contributes more to the mineral resources. Similarly, analyze the correlation between structural elements (e.g., fault distance, fold axis) and the location or size of the deposit. For example, if deposits tend to occur within a certain distance from the fault, then the distance from the fault can be considered an important contributing factor. For alteration data, evaluate the correlation between different types of alteration and the intensity of mineralization. For example, if a certain alteration type is closely associated with high-intensity mineralization, then this alteration contributes more to the mineral resources. Based on the results of the correlation analysis, assign a contribution value to each of the geological factors such as lithology, structure and alteration. This value can be a relative weight to indicate the relative importance of the factor to the distribution of mineral resources.

[0138] Using the weighted overlay method, the contribution of geological factors such as lithology, structure and alteration is superimposed, which can be achieved by weighted summing the contribution layers of each factor. In the weighted overlay process, ensure that the weight of each factor matches its contribution to the distribution of mineral resources. Based on the results of weighted overlay, a comprehensive resource distribution probability data is generated. In the resource distribution probability data, the value of each pixel or area represents the probability of the existence of mineral resources there. Use numerical ranges to represent different probability levels. For example, high probability areas can be represented by red or a higher numerical range, while low probability areas can be represented by blue or a lower numerical range.

[0139] Step 173, according to the probability of resource distribution, select the corresponding probability area as the key investigation object, specifically including: setting a certain probability threshold to divide different probability level areas, and these thresholds can be determined based on previous experience. For example, the area with a probability value greater than 0.8 can be set as a high probability area, the area with a probability value between 0.5 and 0.8 can be set as a medium probability area, and the area with a probability value less than 0.5 can be set as a low probability area. Using GIS software, the probability layer of resource distribution is loaded into the workspace. Use "reclassification" or similar tools to divide the probability layer according to the set probability threshold to generate a new probability level area layer. In the new layer, different probability level areas should be distinguished by obvious colors or patterns to facilitate subsequent analysis. In the divided probability level area layer, high probability areas can be identified by visual inspection or using selection tools. The query function of GIS can be further used to screen out areas with specific attributes (such as probability values ​​greater than 0.8) to ensure accurate identification of high probability targets. For the identified high probability areas, a comprehensive analysis is conducted in combination with geological background information, known mineral point distribution and other relevant data. Evaluate the mineralization conditions, deposit types, and potential resource volume of these areas to determine their mineralization potential and exploration value. Based on the evaluation results of high-probability areas, give priority to those areas with the highest mineralization potential and exploration value as key survey targets. Develop a detailed survey plan, including survey objectives, methods, schedules, budgets, and staffing. In accordance with the formulated survey plan, organize a professional survey team to conduct field surveys in the selected key areas.

[0140] Step 174, based on the geological structure characteristics, lithological conditions and mineralization laws, the key survey objects are evaluated to obtain evaluation results; based on the evaluation results, the survey priority of each area is determined, including:

[0141] Collect geological structure data of key survey objects, including detailed information such as faults and folds, organize lithology data, record the distribution and characteristics of different lithologies, and collect known mineral deposit data and mineralization law information. Clean the collected data, remove outliers and missing data, and standardize the data; use statistical software to calculate the Pearson correlation coefficient between geological structure characteristics, lithology conditions and known mineral deposits. The Pearson correlation coefficient reflects the degree of correlation between various factors. Convert the Pearson correlation coefficient into a score. For example, the absolute value or square value of the correlation coefficient can be used and multiplied by an appropriate weight to reflect the importance of each factor to the distribution of minerals.

[0142] Clustering algorithms (such as K-means) are used to cluster the key survey objects based on geological structure, lithology and other characteristics. Based on the clustering results, the mineral potential of the group to which each object belongs is evaluated. For example, if a group contains multiple known deposits, the objects in this group may receive a higher score. A cluster analysis score is assigned to each object based on the group's mineral potential and the object's position in the group (such as centrality). Principal component analysis is applied to extract the principal components representing the main geological factors, and the standardized data of each object is projected onto the principal components to obtain the scores of each principal component, which reflect the performance of the object on the main geological factors. According to the variance contribution rate (or explanation) of each principal component, the PCA scores are weighted and summed to obtain the overall PCA score for each object.

[0143] Integrate the correlation analysis scores, cluster analysis scores, and principal component analysis scores, which can be achieved through weighted averaging. Set weights, which should be based on historical data to reflect the importance of each analysis method in mineral prediction. Through the above steps, a comprehensive score can be calculated for each key survey object. Sort the key survey objects according to the comprehensive score. Objects with high scores have higher survey priorities. Different score thresholds can be set to divide the priority levels, such as high, medium, and low. According to the priority ranking, formulate a detailed geological survey and exploration plan, and give priority to the survey work in high-priority areas.

[0144] Step 175, in the high probability area, based on the geological conditions and topographical features, the sampling point locations are preliminarily selected, specifically including: in the high probability area, the geological conditions are analyzed in detail, including the characteristics of the strata, structure, magmatic rocks, etc.; the topographical features of the study area, such as mountains, rivers, lakes, etc., which may affect the accessibility and safety of the sampling points; the locations of the sampling points are preliminarily selected by comprehensively considering the geological conditions and topographical features, as well as the actual exploration needs and resource distribution characteristics. These locations should be able to representatively reflect the geological characteristics and resource potential of the area.

[0145] Step 176, based on the geological survey data, resource distribution probability and sampling point data, establish a resource prediction model; verify the resource prediction model through the data of known mineral points to obtain a verification result; optimize the resource prediction model according to the verification result to obtain an optimized resource prediction model, specifically including:

[0146] Collect all relevant geological survey data (such as strata, structure, lithology, etc.), resource distribution probability data (data obtained through geological exploration, geophysical exploration, etc.) and sampling point data (including location, resource grade, etc.); clean the collected data, remove duplicate, erroneous or invalid data, and ensure the accuracy and reliability of the data; convert all data into a unified format; merge the cleaned and formatted geological survey data, resource distribution probability data and sampling point data to form a complete data set containing multi-dimensional feature information.

[0147] Analyze the integrated data set to understand the data type (such as continuous variables, categorical variables, etc.) and distribution characteristics. Determine the specific goal of the prediction, whether it is to predict the specific reserves, grade distribution, or mineability of resources. Based on the data type and prediction goal, select the random forest prediction model; according to the requirements of the selected model, further preprocess the data, such as normalization, standardization, or feature selection.

[0148] Use feature importance evaluation methods (such as feature importance scoring for tree-based models) to screen features that have the greatest impact on the prediction target, remove redundant or irrelevant features, and reduce model complexity. Divide the integrated dataset into training and test sets at a ratio of 70%-30%, and use a random number generator to ensure the randomness and representativeness of the data partition. Use the training set data to train the random forest model. Adjust model parameters, such as the number of trees, maximum depth, etc., to optimize model performance.

[0149] Use the test set data to evaluate the trained model, calculate evaluation indicators such as prediction accuracy, recall rate, F1 score, etc., use cross-validation technology to more accurately evaluate model performance, and adjust the parameters of the random forest model according to the evaluation results, such as increasing or decreasing the number of trees, adjusting the splitting criterion, etc. Repeat the training and evaluation process until satisfactory prediction performance is achieved to obtain an optimized resource prediction model.

[0150] Step 177, using the optimized resource prediction model to predict resource distribution in the entire study area to obtain the distribution of geological resources, specifically includes:

[0151] Collect geological data of the entire study area, including stratigraphy, structure, lithology, geophysical and geochemical exploration data, etc. Preprocess the collected geological data, including data cleaning, format conversion and necessary spatial analysis, to ensure data quality and consistency.

[0152] Load the resource prediction model optimized in step 176; provide the geological data of the entire study area as input to the resource prediction model; run the optimized resource prediction model to predict the resource distribution of the entire study area to estimate the existence probability, reserves and grade of different types of resources.

[0153] In an embodiment of the present invention, by integrating multi-source data such as geology, geophysics and geochemistry, and comprehensive analysis based on geological structure characteristics and resource distribution probability, the distribution of mineral resources can be predicted more accurately. This helps to reduce the risk of blind exploration and improve the success rate of mineral exploration. By screening out high-probability areas as key survey objects and determining the survey priority of each area based on the evaluation results, geological surveys can be made more focused and efficient. This helps to maximize the acquisition of valuable geological information under limited time and resource conditions. In high-probability areas, the locations of sampling points are preliminarily selected based on geological conditions and topographic features to ensure the representativeness and effectiveness of the sampling points. The resource prediction model is verified and optimized through data from known mineral points, and the prediction performance of the model can be continuously improved. This helps to enhance the model's predictive ability for unknown areas and improve the accuracy and confidence of geological resource exploration.

[0154] In the embodiment of the present invention, step 18, by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm to perform selection, crossover, and mutation operations to optimize the survey plan, and finally obtaining the corresponding sampling point arrangement, includes:

[0155] Step 181, determine the survey objectives, which include maximizing resource discovery and minimizing cost or time; divide the survey area into grids and use binary coding, where 1 means selecting the grid for survey and 0 means not selecting, specifically including: determine the main objectives of the geological survey, which may include maximizing resource discovery and minimizing survey cost or time. Use a geographic information system (GIS) to divide the survey area into uniform grids, each grid represents a potential survey unit, and assign a unique identifier to each grid based on the division of the grid; use binary coding to represent the survey status of each grid. Among them, "1" means selecting the grid for survey and "0" means not selecting; construct a binary string whose length is equal to the total number of grids to represent a complete survey plan.

[0156] Step 182, randomly generate multiple initial survey plans, each plan represents a sampling point arrangement method, and the survey plan is an individual in the population, specifically including: setting an appropriate population size, that is, the number of initial survey plans, according to the complexity of the problem and computing resources; using a random number generator to generate a binary string for each individual (i.e., survey plan). The length of this string is the same as the number of grids divided in step 181; ensure that the generated binary string has sufficient diversity to cover different survey strategies.

[0157] Step 183, constructing a fitness function for evaluating the quality of the survey plan according to the survey objectives and constraints;

[0158] Step 184, evaluate the quality of each individual in the population according to the fitness function, and select the corresponding individuals to enter the next generation; randomly select two individuals for crossover operation to generate a new survey plan, specifically including: using the fitness function constructed in step 183 to evaluate the quality of each individual in the population, which involves calculating the fitness value corresponding to each survey plan; according to the fitness value, adopt a certain selection strategy (such as roulette selection, tournament selection, etc.) to select a part of individuals to enter the next generation, these selected individuals usually have higher fitness values, that is, more likely to meet the survey objectives; randomly select two selected individuals as parents, and perform crossover (or recombination) operation. This usually involves cutting the binary strings of the two parents at a random position and exchanging the cut parts to generate new offspring individuals. This process helps to explore new areas in the solution space and may produce better survey plans.

[0159] Step 185, perform mutation operation on individuals with a certain probability, that is, randomly change the survey state of the grid, repeat the selection, crossover and mutation operations until the preset number of iterations is reached to obtain the optimal individual, and decode the optimal individual into a specific survey plan, including the location and number of sampling points, specifically including:

[0160] Mutate the individuals in the population with a certain probability. This usually involves randomly changing certain bits in the individual binary string, that is, changing from "1" to "0" or from "0" to "1". The mutation operation helps to increase the diversity of the population and may help the algorithm jump out of the local optimal solution; repeat the selection, crossover and mutation operations until the preset number of iterations is reached, and the optimal solution is approached through continuous iteration; after the iteration, the individual with the highest fitness value is selected as the optimal individual. The binary string of this optimal individual is decoded into a specific survey plan, including the location and number of sampling points, which can be achieved by identifying the grid position corresponding to the "1" in the binary string as the sampling point.

[0161] In an embodiment of the present invention, through the construction and optimization process of the fitness function, the survey plan can focus on the area where geological resources are most likely to be enriched, thereby increasing the amount of resource discovery, which helps to improve the efficiency and success rate of geological exploration. The optimization process not only considers the amount of resource discovery, but also takes into account the cost and time of the survey. By reasonably selecting the location and number of sampling points, unnecessary exploration activities can be avoided, and the waste of manpower, material resources and financial resources can be reduced, thereby reducing costs and shortening survey time. The method can adaptively generate survey plans according to different survey objectives and constraints. Whether it is pursuing the maximization of resource discovery or the minimization of cost and time, it can be achieved by adjusting the fitness function. The genetic algorithm has a global search capability and can find an approximate optimal solution in a complex solution space. This means that the generated survey plan can achieve the best coverage of resources in the entire survey area and avoid falling into a local optimum. The method is based on binary coding and genetic algorithms, and is relatively simple to implement and easy to understand. At the same time, by adjusting the parameters of the algorithm (such as crossover rate, mutation rate, number of iterations, etc.), the effect of the survey plan can be further optimized.

[0162] In the embodiment of the present invention, the calculation formula of the fitness function is:

[0163] ;

[0164] in, and γ represent weight coefficients; Indicates The predicted resource volume in each grid; represents a binary decision variable, and for each grid, The value is 1 or 0, indicating whether to select the grid for investigation. , then it means grids are selected for investigation; if , it means that the grid is not investigated; Indicates the total number of grids in the survey area, indicating how many grids the entire survey area is divided into; represents the unit cost, the average cost incurred by sampling on each selected grid when performing a geological survey; It represents the total survey time, which means the time required to complete the entire geological survey program.

[0165] In the embodiment of the present invention, the fitness function allows decision makers to flexibly adjust the priority between resource discovery, cost and time according to actual needs through different weight coefficients. By quantifying multiple goals into a single fitness value, the function provides an objective and comparable numerical basis for selecting the optimal survey plan. This helps to reduce the influence of subjective judgment on the decision-making process. With decision variables The product of encourages the selection of resource-rich grids for investigation, thereby improving the efficiency of resource discovery. At the same time, by considering the unit cost c, the function helps to reduce the overall investigation cost while ensuring the amount of resource discovery. By taking the total investigation time T into consideration, the fitness function helps to optimize the execution time of the investigation plan.

[0166] The above step 19 may include:

[0167] Collect all data related to geological surveys, including geological maps, remote sensing images, historical exploration data, etc. Analyze the collected data to identify key geological features and potential resource distribution patterns. Evaluate the preliminary generated geological survey plan and collect opinions and feedback from all parties, especially on resource discovery potential, cost-effectiveness and feasibility. Based on the feedback collected, adjust the weight coefficients in the fitness function to better reflect project needs and priorities; use the adjusted parameters and grid divisions to rerun the genetic algorithm and find the optimal or approximate optimal solution that meets the new constraints. Compare the optimized geological survey plan with the preliminary plan to evaluate the degree of improvement in resource discovery, cost and time. If the optimized plan shows significant advantages, it will be selected as the final geological survey plan.

[0168] like Figure 2 As shown, an embodiment of the present invention further provides a geological survey optimization system based on big data analysis, comprising:

[0169] The acquisition module is used to acquire multi-dimensional data related to geological surveys, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data; and pre-process the multi-dimensional data to obtain pre-processed multi-dimensional data;

[0170] The analysis module is used to automatically extract key geological features from pre-processed multi-dimensional data using a convolutional neural network, and through training, the model learns the mapping relationship from data to labels, and finally outputs geological structural features;

[0171] The determination module is used to determine the starting point of the geological survey and the target area to be investigated based on the geological structure characteristics;

[0172] The allocation module is used to assign different cost weights to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; construct an initial cost matrix based on geological characteristics, continuously fine-tune the node cost and calculate the evaluation function value, gradually cool down until the final temperature is reached, and output the optimized node cost matrix; by initializing the distance and access array, continuously select unvisited nodes and update the corresponding distances of their adjacent nodes until all nodes are visited, and finally output the path from the starting point to the target area; establish a mineral resource evaluation model based on multi-source geological data, calculate the indicator contribution, screen key areas, evaluate the survey priority, select sampling points, establish and optimize the resource prediction model, and finally predict the distribution of geological resources in the entire study area;

[0173] The calculation module is used to optimize the survey plan by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to finally obtain the corresponding sampling point layout; the sampling point layout is optimized to obtain the optimized geological survey plan.

[0174] The embodiment of the present invention further provides a computing device, comprising: a processor, a memory storing a computer program, wherein when the computer program is executed by the processor, the method described above is executed. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.

[0175] The embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when executed on a computer, enable the computer to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.

[0176] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A geological survey optimization method based on big data analysis, characterized in that: The method comprises: Use convolutional neural networks to automatically extract key geological features from pre-processed multi-dimensional data, and train the model to learn the mapping relationship from data to labels, and finally output geological structural features; Determine the starting point of geological survey and the target area to be investigated based on the geological structure characteristics; Different cost weights are assigned to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; an initial cost matrix is ​​constructed based on geological characteristics, and the node costs are continuously fine-tuned and the evaluation function values ​​are calculated. The temperature is gradually lowered until the final temperature is reached, and the optimized node cost matrix is ​​output; the cost weight calculation formula is: ; in, represents the cost weight from area b to area g; a represents the weight coefficient of terrain complexity; represents the elevation difference between area b and area g; H represents the maximum elevation difference in the study area; represents the slope difference between area b and area g; S represents the maximum slope difference in the study area; Represents the weight coefficient of geological disaster risk; Indicates the number of geological disaster points from area b to area g; Indicates the total number of geological hazard points in the study area; represents the number of historical activities of the geological disaster site; T represents the total number of years in the observation period, which is used to calculate the historical activity frequency; f represents the scaling factor of the historical activity frequency; D represents the area of ​​environmental damage caused by geological disasters; A represents the total area of ​​the study area; w represents the weight of the severity of environmental damage; d represents the scaling factor of the degree of environmental damage; c represents the weight coefficient of traffic accessibility; t represents the travel time from area b to area g; represents the maximum travel time within the study area; By initializing the distance and visit arrays, we continuously select unvisited nodes and update the distances of their adjacent nodes until all nodes are visited, and finally output the path from the starting point to the target area. Establish a mineral resource evaluation model based on multi-source geological data, calculate the contribution of indicators, screen key areas, evaluate survey priorities, select sampling points, establish and optimize resource prediction models, and ultimately predict the distribution of geological resources in the entire study area; By dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to optimize the survey plan, the corresponding sampling point layout is finally obtained; Optimize the arrangement of sampling points to obtain an optimized geological survey plan; construct an initial cost matrix based on geological characteristics, continuously fine-tune the point cost and calculate the evaluation function value, gradually cool down until the final temperature is reached, and output the optimized node cost matrix, including: Based on the characteristics of the geological survey area, the actual cost weights between each node are determined, and the cost weights are used to construct a The initial cost matrix , where N is the number of nodes and the initial cost matrix Elements in represents the initial cost from node i to node j; defines the parameters of simulated annealing, including the initial temperature , Final temperature , cooling rate and maximum number of iterations ; Randomly select two different node pairs in the cost matrix ; Fine-tune the cost between these two node pairs to generate new candidate solutions ; Compute a new candidate solution using the same merit function as the current solution The evaluation function value of ; Calculate the difference in evaluation function values ,in, Represents the evaluation function value of the current solution; Calculate the probability of accepting a new solution based on the acceptance probability of simulated annealing ; Generate a random number r in the range [0, 1); if , then accept the new solution and set the current solution ; Reduce the temperature according to the cooling rate; if the temperature drops below the final temperature, the iteration stops, and when it stops, the current solution is output As the optimized cost matrix, the optimized cost matrix is ​​a two-dimensional array, in which each element represents the optimized cost between two nodes in the geological survey area; the calculation formula of the evaluation function is: ; Where N is the number of nodes; is the current cost matrix, representing the direct cost between node i and node j; is a distance coefficient; represents the geographical distance between node i and node j; is the height coefficient; represents the height change from node i to node j; Represents the evaluation function value calculated based on the current cost matrix.

2. A geological survey optimization method based on big data analysis according to claim 1, characterized in that: Convolutional neural networks are used to automatically extract key geological features from preprocessed multi-dimensional data, and the model is trained to learn the mapping relationship from data to labels, and finally output geological structural features, including: Use the convolutional neural network in the deep learning model to automatically extract key features from the pre-processed multi-dimensional data, including rock formation trends, fault distribution patterns, and element content anomalies; The deep learning model is trained using the preset data set and labels. During the training process, the weights and biases of the deep learning model are adjusted so that the deep learning model learns the mapping relationship from input data to output labels to obtain a trained deep learning model. After the training is completed, the preprocessed multi-dimensional data is input into the trained deep learning model, and the trained deep learning model will output the corresponding geological structure characteristics.

3. A geological survey optimization method based on big data analysis according to claim 2, characterized in that: By initializing the distance and visit array, continuously selecting unvisited nodes and updating the distances corresponding to their adjacent nodes until all nodes are visited, and finally outputting the path corresponding to the starting point to the target area, including: setting the starting point to node B; Create an array dist[] of length R to store the shortest distance from the starting point B to each node; initially, set dist[B] to 0, indicating that the distance from the starting point to itself is 0, and the remaining elements are set to infinity, indicating that the initial distance from the starting point to the remaining nodes is unknown; Create a Boolean array visited[] of length R to mark whether each node has been visited. Initially, all elements are set to false, indicating that no node has been visited. Select an unvisited node u from the array dist[] and mark node u as visited, that is, set visited[u] to true; Traverse each node v adjacent to node u, and check whether the distance from the starting point through node u to node v is less than the currently recorded dist[v], that is, dist[u] + C[u][v], where C[u][v] represents the cost from node u to node v; if so, update dist[v] to make it equal to the distance from node u to node v until all nodes have been visited; According to the updated array dist[], the constructed target path from the starting point to the target area is output.

4. A geological survey optimization method based on big data analysis according to claim 3, characterized in that: A mineral resource evaluation model is established based on multi-source geological data. By calculating the contribution of indicators, screening key areas, evaluating survey priorities, and selecting sampling points, a resource prediction model is established and optimized, and the distribution of geological resources in the entire study area is finally predicted, including: Establish a mineral resource evaluation model based on geological, geophysical and geochemical multi-source data; Select the corresponding geological factors as evaluation indicators, calculate the contribution of each evaluation indicator to resource distribution, and derive the probability of resource distribution based on the contribution. The evaluation indicators include lithology, structure and alteration. According to the probability of resource distribution, select the corresponding probability areas as the key investigation objects; According to the geological structure characteristics, lithological conditions and mineralization laws, the key survey objects are evaluated to obtain the evaluation results; according to the evaluation results, the survey priority of each area is determined; In the high-probability area, the sampling point locations are preliminarily selected based on the geological conditions and topographic features; Establish resource prediction model based on geological survey data, resource distribution probability and sampling point data; Verify the resource prediction model through the data of known mining points to obtain the verification results; optimize the resource prediction model according to the verification results to obtain the optimized resource prediction model; The optimized resource prediction model is used to predict the resource distribution of the entire study area to obtain the distribution of geological resources.

5. A geological survey optimization method based on big data analysis according to claim 4, characterized in that: By dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to optimize the survey plan, the corresponding sampling point layout is finally obtained, including: Determine the survey objectives, which include maximizing resource discovery and minimizing cost or time; divide the survey area into grids and use binary coding, where 1 indicates that the grid is selected for survey and 0 indicates that it is not selected; Multiple initial survey plans are randomly generated, each plan represents a sampling point arrangement method, and the survey plan is an individual in the population; According to the survey objectives and constraints, a fitness function is constructed to evaluate the quality of the survey plan; Evaluate the quality of each individual in the population according to the fitness function and select the corresponding individuals to enter the next generation; randomly select two individuals for crossover operation to generate a new survey plan; The individuals are mutated with a certain probability, that is, the survey status of the grid is randomly changed, and the selection, crossover and mutation operations are repeated until the preset number of iterations is reached to obtain the optimal individual, and the optimal individual is decoded into a specific survey plan, including the location and number of sampling points.

6. A geological survey optimization method based on big data analysis according to claim 5, characterized in that: The calculation formula of fitness function is: ; in, represents the weight coefficient; represents the predicted resource volume in the i-th grid; represents a binary decision variable, and for each grid, The value is 1 or 0, indicating whether to select the grid for investigation. , it means that the i-th grid is selected for investigation; if , it means that the grid is not investigated; Indicates the total number of grids in the survey area, indicating how many grids the entire survey area is divided into; represents unit cost, which is the average cost incurred for sampling on each selected grid when performing a geological survey; T represents the total survey time, which is the time required to complete the entire geological survey program.

7. A geological survey optimization system based on big data analysis, characterized in that: Applied to the method according to any one of claims 1 to 6, comprising: The acquisition module is used to acquire multi-dimensional data related to geological surveys, including geological structure data, geophysical data, geochemical data, remote sensing data, and historical geological survey data; and pre-process the multi-dimensional data to obtain pre-processed multi-dimensional data; The analysis module is used to automatically extract key geological features from pre-processed multi-dimensional data using a convolutional neural network, and through training, the model learns the mapping relationship from data to labels, and finally outputs geological structural features; The determination module is used to determine the starting point of the geological survey and the target area to be investigated based on the geological structure characteristics; The allocation module is used to assign different cost weights to different regions based on terrain complexity, difficulties and risk points in historical survey data, and traffic accessibility; construct an initial cost matrix based on geological characteristics, continuously fine-tune the node cost and calculate the evaluation function value, gradually cool down until the final temperature is reached, and output the optimized node cost matrix; by initializing the distance and access array, continuously select unvisited nodes and update the corresponding distances of their adjacent nodes until all nodes are visited, and finally output the path from the starting point to the target area; establish a mineral resource evaluation model based on multi-source geological data, calculate the indicator contribution, screen key areas, evaluate the survey priority, select sampling points, establish and optimize the resource prediction model, and finally predict the distribution of geological resources in the entire study area; The calculation module is used to optimize the survey plan by dividing the grid, encoding, randomly generating the initial plan, and using the genetic algorithm for selection, crossover, and mutation operations to finally obtain the corresponding sampling point layout; the sampling point layout is optimized to obtain the optimized geological survey plan.

Citation Information

Patent Citations

  • Geochemical exploration scheme screening optimization method based on improved ant colony algorithm

    CN118096425A

  • Hydraulic ring geological survey system and method based on GPS positioning

    CN118428094A