Training method of geological unit classification model, geological unit classification method and device

By integrating multiple data features, a geological unit classification model trained through machine learning solves the problems of imprecise classification of lunar geological units and significant influence from human factors in existing technologies, achieving efficient and intelligent identification and lithological mapping of lunar geological units.

CN114580524BActive Publication Date: 2025-12-30NAT ASTRONOMICAL OBSERVATORIES CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210195346.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-12-30
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

Existing methods for classifying lunar geological units are limited by low data resolution and insufficient coverage, resulting in imprecise classification and significant human influence, making it difficult to meet the needs of lunar chronology research.

Method used

A geological unit classification model was trained using machine learning algorithms. By integrating multiple data features, such as longitude, latitude, grayscale, elevation, slope, and TiO2 abundance, and based on high-resolution lunar imagery and spectral inversion data, spatial overlay operations and feature extraction were performed to construct an efficient geological unit classification model.

Benefits of technology

It achieves efficient and intelligent classification of lunar geological units, reduces human interference, improves classification accuracy and efficiency, and supports the division of geological units and lithological mapping across the entire lunar area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580524B_ABST
    Figure CN114580524B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method of a geological unit classification model. The training method of the geological unit classification model comprises: selecting a target area according to a lunar topographic map framing rule, and determining image data of the target area; generating pixel grid vector data based on the image data, and calculating the longitude and latitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data; performing spatial superposition operation based on the data point vector data to determine at least one data feature; determining sample input data of a to-be-trained geological unit classification model based on the at least one data feature; and training the to-be-trained geological unit classification model by using a machine learning algorithm according to the sample input data. The present disclosure also provides a geological unit classification method, device, equipment and storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of image technology, can be applied to the field of artificial intelligence technology, and more particularly to a geological unit classification model training method, a geological unit classification method, an apparatus, an electronic device, and a storage medium. BACKGROUND

[0002] In the study of lunar geochronology, due to the limited conditions of lunar field survey and lunar sample acquisition, the impact crater size-frequency distribution method (CSFD) is widely used and recognized as an effective method for estimating the geological age of the lunar surface in the field of lunar and planetary science. The application of the CSFD method for geological age estimation first requires the division of geological units, and then the statistics of impact craters in the same geological unit.

[0003] Currently available USGS 1:5 million geological maps and Hiesigner geological unit classification and division cannot meet the requirements of lunar geochronology research using the CSFD method to estimate the geological age of the lunar surface. On the one hand, the USGS 1:5 million geological map was completed in the 1970s, mainly using image data obtained from the Lunar Orbiter mission. Limited by the resolution and type of data available at the time, as well as the low spatial resolution of the geological map scale, it is difficult to clearly express and synthesize the details of the lunar surface topography, making it impossible to achieve more detailed and accurate classification and division of lunar surface geological units. On the other hand, Hiesigner's division of geological units is limited to the lunar mare region, with limited coverage and cannot support full-moon range dating research. The classification of geological units often requires researchers to make subjective judgments about the study area. However, due to the difficulty of identifying lunar geological unit categories, lithology identification cannot rely solely on one type of data, and multiple types of data are needed to confirm. Due to the different sensitivity of the human eye in distinguishing gray-scale images, the classification of geological units often depends on the practical experience of researchers, and the results produced by different researchers vary greatly. The classification results of geological units are greatly influenced by human factors and have low efficiency. SUMMARY

[0004] In view of the above problems, the present disclosure provides a geological unit classification model training method, a geological unit classification method, an apparatus, a device, a medium, and a program product.

[0005] According to a first aspect of the present disclosure, a geological unit classification model training method is provided, comprising: in response to receiving a control instruction, analyzing the control instruction to obtain an analysis result corresponding to the control instruction; wherein the control instruction is transmitted through a 5G network; sending the analysis result to a drone to make the drone execute the control instruction and return data information; and receiving the data information.

[0006] According to embodiments of this disclosure, the step of performing spatial overlay operations based on the data point vector data to determine at least one data feature includes one or more of the following operations: extracting geological unit classification features based on the sample data point vector data through spatial overlay operations; extracting one or more of longitude features, latitude features, and grayscale features based on the sample data point vector data through spatial overlay operations; extracting one or more of elevation features, slope features, and undulation features based on the sample data point vector data through spatial overlay operations; extracting TiO2 features based on the sample data point vector data through spatial overlay operations; and extracting one or more of FeO, pyroxene, olivine, plagioclase, pyroxene, microscopic metallic iron, and optical maturity features based on the sample data point vector data through spatial overlay operations.

[0007] According to embodiments of this disclosure, the step of generating pixel grid vector data based on the image data and calculating the latitude and longitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data includes: taking points at preset distance intervals based on the image data to generate pixel grid vector data with spatial positions; and calculating the latitude and longitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data.

[0008] According to an embodiment of this disclosure, determining the sample input data for the geological unit classification model to be trained based on the at least one data feature includes: calculating the importance of the at least one data feature using an embedded method to obtain at least one importance calculation result; and determining the sample input data for the geological unit classification model to be trained based on the at least one importance calculation result.

[0009] According to embodiments of this disclosure, the method further includes: using a label encoding method to encode the geological unit classification features to obtain a digital code for the geological unit classification features; and using the digital code for the geological unit classification features as sample input data for a geological unit classification model to be trained.

[0010] According to a second aspect of this disclosure, a training method for a geological unit classification model is provided, comprising: a geological unit classification method, comprising: inputting a target image into a geological unit classification model to obtain a geological unit classification result corresponding to the target image; wherein the geological unit classification model is trained according to the method provided in this disclosure.

[0011] A third aspect of this disclosure provides a training apparatus for a geological unit classification model, comprising: a first data processing module for selecting a target region according to lunar topographic map grading rules and determining image data of the target region; a second data processing module for generating pixel grid vector data based on the image data and calculating the latitude and longitude coordinates of the center position of each pixel grid according to the pixel grid vector data to obtain data point vector data; a third data processing module for performing spatial overlay operations based on the data point vector data to determine at least one data feature; a fourth data processing module for determining sample input data for a geological unit classification model to be trained based on the at least one data feature; and a training module for training the geological unit classification model to be trained using a machine learning algorithm based on the sample input data.

[0012] A fourth aspect of this disclosure provides a geological unit classification apparatus, comprising: an acquisition module for inputting a target image into a geological unit classification model to obtain a geological unit classification result corresponding to the target image; wherein the geological unit classification model is trained based on the apparatus provided in this disclosure.

[0013] A fifth aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the methods disclosed above.

[0014] A sixth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods disclosed above. Attached Figure Description

[0015] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0016] Figure 1 A flowchart illustrating a training method for a geological unit classification model according to an embodiment of the present disclosure is shown schematically.

[0017] Figure 2 A flowchart illustrating a geological unit classification method according to an embodiment of the present disclosure is shown schematically.

[0018] Figure 3 A schematic diagram illustrating the structure of a training apparatus for a geological unit classification model according to an embodiment of the present disclosure is shown.

[0019] Figure 4 A schematic diagram illustrating the structure of a geological unit classification device according to an embodiment of the present disclosure; and

[0020] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing a training method and / or a geological unit classification method for a geological unit classification model, according to embodiments of the present disclosure. Detailed Implementation

[0021] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0025] The embodiments of this disclosure provide a training method and apparatus for a geological unit classification model. The method involves selecting a target region based on lunar topographic map grading rules and determining image data for that region; generating pixel grid vector data based on the image data and calculating the latitude and longitude coordinates of the center position of each pixel grid to obtain data point vector data; performing spatial overlay operations on the data point vector data to determine at least one data feature; determining sample input data for the geological unit classification model to be trained based on the at least one data feature; and training the geological unit classification model to be trained using a machine learning algorithm based on the sample input data.

[0026] pass Figure 1The training method of the geological unit classification model of the disclosed embodiments is described in detail.

[0027] Figure 1 A flowchart illustrating a training method for a geological unit classification model according to an embodiment of the present disclosure is shown.

[0028] like Figure 1 As shown, this embodiment includes operations S101 to S105, and the training method based on the geological unit classification model can be executed by a server.

[0029] In operation S101, the target area is selected according to the lunar topographic map grading rules, and the image data of the target area is determined.

[0030] Understandably, the study area can be selected according to the lunar topographic map grading rules, and the base map of the selected study area can be used as the image data for the target area. In order to facilitate flexible data organization and avoid the problem of processing large amounts of global data, this embodiment can divide the global lunar data according to the grading rules of the national standard "Division and Numbering of Lunar Basic Scale Topographic Maps" (GB / T32521-2016).

[0031] For example, imagery data can be divided according to the sheet division rules of the national standard "Division and Numbering of Lunar Basic Scale Topographic Maps" (GB / T 32521-2016): For instance, the polar regions between 84° and 90° north and south latitude are divided into two separate sheets; within the 84° north and south latitude range, starting from the equator, it is divided into multiple sub-projection zones with a latitude difference of 14°. Each sub-zone, from high latitude to the equator, uses 45°, 30°, 24°, 20°, and 18° as the sheet division longitude differences, respectively, dividing the entire lunar region within 84° north and south latitude into 186 sheet regions. Following this method, the entire lunar imagery data is divided into 188 sheet regions. This allows for subsequent selection of study areas based on the sheet division rules to conduct regional or full-lunar geological unit classification studies.

[0032] In operation S220, pixel grid vector data is generated based on image data, and the latitude and longitude coordinates of the center position of each pixel grid are calculated based on the pixel grid vector data to obtain data point vector data.

[0033] This is understandable, such as generating pixel grid vector data by sampling points based on image data; and calculating the latitude and longitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data.

[0034] In operation S230, spatial overlay operation is performed based on data point vector data to determine at least one data feature.

[0035] Understandably, other data can be acquired first, such as USGS full-moon geological map data, or Chang'e-2 lunar image DOM data; spatial overlay operations can be performed on the data point vector data using this other data; at least one data feature can be determined, such as geological unit classification features, or latitude, longitude and grayscale features.

[0036] In operation S240, sample input data for the geological unit classification model to be trained is determined based on at least one data feature.

[0037] Understandably, multi-feature variables (i.e., at least one data feature) are used as sample input data for the geological unit classification model to be trained.

[0038] For example, feature optimization based on at least one data feature can select data features with higher importance and use the selected dataset as sample input data.

[0039] In operation of S250, a geological unit classification model is trained using machine learning algorithms based on the sample input data.

[0040] Understandably, this embodiment uses machine learning algorithms to train the geological unit classification model to be trained. Machine learning algorithms can be KNeighbors, DecisionTree, RandomForest, XGBoost, Bagging, etc.

[0041] For example, based on the sample input data, it is divided into a training set, a validation set, and a test set. The training set and validation set are used to train the geological unit classification model to be trained. The model is continuously iterated to validate and improve it. Finally, the test set is used to perform classification prediction to obtain the geological unit classification result.

[0042] This embodiment constructs different classification models for geological units to be trained by selecting a variety of machine learning algorithms. First, the classification model is trained using training and validation sets. The model's performance is improved through continuous iteration, and the optimal classification model is selected. The optimal classification model is then applied to the test set for classification prediction, and the model's performance is evaluated through the test results.

[0043] In this embodiment, the process of training a geological unit classification model using a machine learning algorithm can include evaluating the classification results, such as assessing the model's performance through test results. For example, metrics such as confusion matrix (G), accuracy (Ac), macro-average precision (Pr), macro-average recall (Re), and macro-average F1 score (F1) can be used to evaluate the effectiveness and results of geological unit classification. Accuracy is the percentage of correctly classified samples; macro-average precision represents the proportion of correctly predicted positive samples; macro-average recall represents the proportion of correctly predicted positive samples; and macro-average F1 score is the harmonic mean of macro-average precision and macro-average recall.

[0044] In this embodiment, the process of training a geological unit classification model using a machine learning algorithm may include optimizing the classification results: optimizing the classification results from two aspects, namely adjusting the parameters of the geological unit classification model to be trained and selecting feature variables, to achieve better classification results until the classification accuracy reaches the expected value.

[0045] The training method for the geological unit classification model provided in this embodiment can train the geological unit classification model based on machine learning and data fusion with multiple feature variables. The trained geological unit classification model can achieve efficient and intelligent classification and identification of lunar surface geological units. It can overcome the limitations of manual identification, such as insensitivity to lunar surface grayscale images, interference from human factors, and low efficiency. It can also be used for lithological mapping of lunar surface geological units. By learning and training on the classification information of geological units in known study areas, it can effectively classify geological units in unknown areas, thereby providing effective support for lithological mapping of lunar surface geological units.

[0046] Spatial overlay operations are performed on sample point vector data to determine at least one data feature, including one or more of the following operations: extracting geological unit classification features based on sample point vector data; extracting one or more of longitude, latitude, and grayscale features based on sample point vector data; extracting one or more of elevation, slope, and undulation features based on sample point vector data; extracting TiO2 features based on sample point vector data; and extracting one or more of FeO, pyroxene, olivine, plagioclase, pyroxene, submicroscopic metallic iron, and optical maturity features based on sample point vector data.

[0047] For example, geological unit classification features can be extracted by spatially overlaying point vector data with USGS full-month geological map data. This can be achieved by performing intersection operations on the point vector data and the geological map vector surface data according to their positions, assigning the attributes of the vector surface containing the intersection points to the sample vector points, thus forming target classification sample points. It is understandable that geological unit classification is a fundamental component of geological units in the USGS full-month geological map, with a total of 49 geological unit classifications across the entire month.

[0048] For example, features such as longitude, latitude, and grayscale can be extracted by spatial overlay operations on the pixel raster grid in the data point vector data and the Chang'e-2 lunar image DOM data. For instance, the corresponding raster attribute value can be extracted from the raster row and column numbers of the corresponding position in the image data and assigned to the sample vector point. It can be understood that the longitude and latitude feature variables represent the spatial location information of the pixel center point, while the grayscale feature variable represents the grayscale value of the image pixel.

[0049] For example, features such as elevation, slope, and undulation can be extracted by spatial overlay operations on the pixel raster grid in the data point vector data and the Chang'e-2 DEM data, and then assigned to the sample vector points. It can be understood that the elevation feature variable refers to the elevation value corresponding to the location of the pixel center point; the slope feature variable refers to the rate of elevation change from one pixel to another in the terrain data; and the undulation feature variable refers to the difference between the maximum and minimum elevation values ​​of all pixels within an eight-neighborhood centered on that pixel.

[0050] For example, TiO2 characteristic variables can be extracted by spatially overlaying data point vector data with data obtained from the Wide-Angle Camera (WAC) of the Lunar Reconnaissance Orbiter (LROC) system. Understandably, TiO2 characteristic variables are TiO2 abundance data retrieved from the raw data obtained from the WAC of the LOC system. Similarly, characteristic variables such as FeO, clinopyroxene, olivine, plagioclase, pyroxene, submicroscopic metallic iron, and optical maturity can be extracted by spatially overlaying data point vector data with inverted data from the Kaguya lunar probe's multi-band imager. Understandably, the characteristic variables such as FeO, pyroxene, olivine, plagioclase, pyroxene, submicroscopic metallic iron, and optical maturity were obtained using multispectral image data of the lunar surface acquired by the Luna Multi-band Imager (MI) at five wavelength positions in the ultraviolet-visible band (UVVIS; 415, 750, 900, 950, 1001 nm) and four wavelength positions in the near-infrared band (NIR; 1000, 1050, 1100, 1250 nm). The FeO content, the content of the four common minerals (pyroxene, olivine, plagioclase, pyroxene), the abundance of submicroscopic metallic iron (SMFe), and the optical maturity (OMAT) data were calculated through inversion, covering an area close to the entire moon.

[0051] The training method for the geological unit classification model provided in this embodiment can achieve automated classification and division of lunar surface geological units based on multi-dimensional characteristic variables such as lunar surface morphology, mineral composition and element content by fusing high-resolution lunar images, topographic and spectral inversion data from the entire moon. It is not limited by the limitations of low data resolution, single data features and insufficient data coverage.

[0052] The process involves generating pixel grid vector data based on image data, and calculating the latitude and longitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data. This includes: taking points at preset distance intervals based on image data to generate pixel grid vector data with spatial location; and calculating the latitude and longitude coordinates of the center position of each pixel grid based on the pixel grid vector data to obtain data point vector data.

[0053] Understandably, the preset distance range can be 500m, 1km, 3km, etc.; for example, points are taken at preset distance ranges (such as 500m, 1km, 3km, 5km, etc.) to generate pixel grid vector data with spatial location, and the latitude and longitude coordinates of the center position of each pixel grid are calculated to generate sample data point vector files.

[0054] The method for determining the sample input data of a geological unit classification model to be trained based on at least one data feature includes: calculating the importance of at least one data feature using an embedded method to obtain at least one importance calculation result; and determining the sample input data of the geological unit classification model to be trained based on at least one importance calculation result.

[0055] Understandably, an embedded approach is used to calculate the importance of data features and select features based on their importance influencing factors. The chosen embedded method can include machine learning algorithms such as Decision Tree, Random Forest, and XGBoost. Based on the calculation results of the machine learning algorithm, a predetermined number of the most important features are selected, such as the four highest-ranking features.

[0056] The training method for the geological unit classification model also includes: using a label encoding method to encode and convert the geological unit classification features to obtain digital codes for the geological unit classification features; and using the digital codes for the geological unit classification features as sample input data for the geological unit classification model to be trained.

[0057] Understandably, feature transformations are performed on the classification characteristics of geological units, such as using label encoding methods to encode and convert character-based geological unit classifications into numerical codes.

[0058] Furthermore, the digital codes of the geological unit classification features will be used as sample input data for the geological unit classification model to be trained, and machine learning algorithms will be used to train the geological unit classification model to be trained.

[0059] For example, based on the feature importance measurement results, four feature variables with higher influence factors can be selected to form a data set X. The encoded and converted geological units can then be classified to form a prediction target set Y. X and Y are the final sample datasets for geological unit classification, and there is a one-to-one correspondence between elements in X and Y. X and Y are then divided into training, validation, and test sets. Specifically: First, the data of X and Y are randomly divided into training and test sets according to a certain ratio (e.g., 70% and 30%, or 80% and 20%). Second, the training set is then randomly divided again into training and validation sets according to the same ratio. The classification model is trained using the training and validation sets, and the model is continuously iterated and improved. Finally, the test set is used to perform classification predictions to obtain the classification results.

[0060] pass Figure 2 The geological unit classification method of the disclosed embodiments is described in detail.

[0061] Figure 2 A flowchart illustrating a geological unit classification method according to an embodiment of the present disclosure is shown schematically.

[0062] like Figure 2 As shown, this embodiment includes operation S201, in which the geological unit classification method can be executed by a server.

[0063] In operation S201, the target image is input into the geological unit classification model to obtain the geological unit classification result corresponding to the target image.

[0064] For example, the geological unit classification model is trained according to the method provided in this disclosure.

[0065] For example, the geological unit classification model is trained according to the method 100 provided in this disclosure.

[0066] The geological unit classification method provided in this embodiment can efficiently and quickly obtain the geological unit classification results corresponding to the target image, avoiding interference from human factors.

[0067] To better understand this disclosure, the following embodiments further illustrate the content of this disclosure, but this disclosure is not limited to the following embodiments.

[0068] For example, to facilitate flexible data organization and avoid the difficulties in processing large volumes of global data, this disclosure divides the lunar global data according to the map sheet division rules of the national standard "Lunar Basic Scale Topographic Map Sheet Division and Numbering" (GB / T 32521-2016). The image data is divided according to the map sheet division rules of the national standard "Lunar Basic Scale Topographic Map Sheet Division and Numbering" (GB / T 32521-2016): the polar regions between 84° and 90° north and south latitude are divided into two separate map sheets; within the 84° north and south latitude range, starting from the equator, it is divided into multiple sub-projection zones with a latitude difference of 14°. Each sub-zone, from high latitude to the equator, selects 45°, 30°, 24°, 20°, and 18° as the longitude difference for the map sheet division, dividing the entire lunar area within 84° north and south latitude into 186 map sheets. Using this method, the entire lunar image data is divided into 188 map sheets. Subsequently, the research area can be selected according to the map division rules to carry out regional or full-month geological unit classification research.

[0069] For example, using the 7-meter high-resolution image data from Chang'e-2 as the base map, points are selected at intervals of a custom size range (such as 500m, 1km, 3km, 5km, etc.) to generate pixel grid vector data with spatial location. The latitude and longitude coordinates of the center position of each pixel grid are calculated to generate a sample data point vector file, and the longitude, latitude, and grayscale attribute values ​​of the pixels are extracted. The specific steps are as follows:

[0070] 1) Based on the custom size range (e.g., 500m, 1km, 3km, 5km, etc.), use ArcMap and Fishnet tools to generate pixel grid vector data at corresponding intervals within the full-month image data range (latitude -180 to 180 degrees, longitude -90 to 90 degrees).

[0071] 2) Based on the pixel grid vector data with latitude and longitude coordinates generated in 1), extract the minimum (Longtitudelow_left, Latitudelow_left) and maximum (Longtitudehigh_right, Latitudehigh_right) corner coordinates of each cell.

[0072] 3) The formula for calculating the latitude and longitude coordinates of the center point of each pixel is:

[0073] Long=(Longtitudelow_left+Longtitudehigh_right) / 2

[0074] Lat=(Latitudelow_left+Latitudehigh_right) / 2

[0075] This refers to the latitude and longitude coordinates corresponding to the required pixel grid, and the ArcMap feature to point tool is used to generate a pixel grid vector point file.

[0076] 4) Based on the pixel grid vector point file generated in 3), the corresponding pixel can be found on the Chang'e-2 7m resolution image data using the ArcMap Identify tool, and the grayscale value of the pixel can be obtained.

[0077] For example, based on the sample point vector data generated above, spatial intersection operations are performed according to the pixel positions using geological map data as the base map, and the geological unit classification attributes of the vector surface where the intersection points are located are assigned to the pixel vector points.

[0078] The specific steps for obtaining the geological unit classification are as follows:

[0079] Step 1: Based on the custom size range (e.g., 500m, 1km, 3km, 5km, etc.), use ArcMap and Fishnet tools to generate pixel grid vector data at corresponding intervals within the full-month image data range (latitude -180 to 180 degrees, longitude -90 to 90 degrees).

[0080] Step 2: Based on the pixel grid vector data with latitude and longitude coordinates generated in Step 1, extract the minimum (Longtitudelow_left, Latitudelow_left) and maximum (Longtitudehigh_right, Latitudehigh_right) corner coordinates of each cell.

[0081] Step 3: The formula for calculating the latitude and longitude coordinates of the center point of each pixel is as follows:

[0082] Long=(Longtitudelow_left+Longtitudehigh_right) / 2

[0083] Lat=(Latitudelow_left+Latitudehigh_right) / 2

[0084] This refers to the latitude and longitude coordinates corresponding to the required pixel grid. At the same time, the ArcMap feature to point tool is used to generate a pixel grid vector point file (i.e., sample points).

[0085] Step 4: Based on the pixel grid vector point file generated in Step 3 and the USGS full-month geological map vector surface data, use the ArcMap intersect tool to perform spatial intersection operation. The result is also a vector point file. Moreover, the geological classification data that comes with the USGS full-month geological map vector surface data will be automatically assigned to this vector point file, thus obtaining the geological classification data corresponding to each pixel grid vector point (i.e., sample point).

[0086] For example, using terrain data as a base map, spatial overlay operations are performed by superimposing latitude and longitude grids and pixel grids, and the elevation, slope, and undulation values ​​of each pixel are extracted from the raster row and column numbers corresponding to each pixel in the terrain data.

[0087] The elevation of a pixel represents its elevation value, which is directly obtained from the corresponding row and column in the terrain data based on the pixel's geographic coordinates. The formula is as follows:

[0088]

[0089] In equation (1), lon is the longitude of the pixel, lat is the latitude of the pixel, oriLon is the longitude of the top left corner of the image, oriLat is the latitude of the top left corner of the image, dx and dy are the horizontal and vertical resolutions of the image, round is the rounding function, and pixel is the function to extract the pixel value based on the number of rows and columns.

[0090] Slope is the rate of change of elevation from one pixel to another in terrain data. Let p be the slope. Let represent the partial derivatives in the x and y directions, respectively. Then:

[0091]

[0092] The elevation relief of a pixel is the difference between the maximum and minimum elevation values ​​of all pixels within an eight-square radius centered on that pixel. The expression for calculating the elevation relief is:

[0093] Δh=h max -h min Equation (3);

[0094] In equation (3), h max h represents the highest pixel elevation value within the eight neighboring regions. min Δh represents the lowest pixel elevation value within its eight-neighborhood, and is the elevation difference within the eight-neighborhood of a pixel. This method uses an eight-neighborhood of 3*3 size.

[0095] For example, using TiO2 abundance map data as a base map, a spatial join operation is performed by overlaying latitude and longitude grids and pixel grids. The TiO2 abundance attribute value of each pixel is then extracted from the raster row and column numbers corresponding to each pixel position in the TiO2 abundance map. The formula is as follows:

[0096]

[0097] In equation (4), lon is the longitude of the pixel, lat is the latitude of the pixel, oriLon is the longitude of the top left corner of the image, oriLat is the latitude of the top left corner of the image, dx and dy are the horizontal and vertical resolutions of the image, round is the rounding function, and pixel is the function to extract the pixel value based on the number of rows and columns.

[0098] For example, olivine mineral content inversion data obtained by the Kaguya lunar probe's multi-band imager is spatially superimposed with generated sample point vector data to extract the percentage of olivine content. The formula is as follows:

[0099]

[0100] In equation (5), lon is the longitude of the pixel, lat is the latitude of the pixel, oriLon is the longitude of the top left corner of the image, oriLat is the latitude of the top left corner of the image, dx and dy are the horizontal and vertical resolutions of the image, round is the rounding function, and pixel is the function to extract the pixel value based on the number of rows and columns.

[0101] The method provided in this embodiment may include the following steps: Step 1, data acquisition; Step 2, feature variable extraction; Step 3, feature transformation; Step 4, feature importance measurement; Step 5, dataset construction; Step 6, data segmentation; Step 7, classification model construction and prediction; Step 8, classification result evaluation; Step 9, classification result optimization. Specifically:

[0102] Step 1: Data acquisition. Select the study area according to the full-month zoning rules. Based on the image data of the study area, perform raster grid segmentation according to the custom interval distance. Each small grid represents an image pixel point, which constitutes the sample data points.

[0103] Step 2: Multi-feature variable extraction. From the USGS full-moon geological map, Chang'e-2 image data, Chang'e-2 topographic data, WAC TiO2 abundance map data, and Kaguya multi-band imager data, pixel segmentation was performed to extract 15 feature variables, including geological unit classification, longitude, latitude, grayscale, elevation, slope, relief, TiO2 abundance, FeO content, pyroxene content, olivine content, plagioclase content, pyroxene content, submicroscopic metallic iron content, and optical maturity. The geological unit classification features were grouped into a separate target set Y0, and the remaining 14 features were combined into an initial feature set X0.

[0104] Step 3: Feature transformation. The tag encoding method is used to transform the feature encoding of set Y0, converting all geological unit classifications from character type to numeric type encoding.

[0105] Step 4: Feature Importance Measurement. Three machine learning algorithms—Decision Tree, Random Forest, and XGBoost—are used to score the importance of data features. This involves estimating the contribution of each feature to each tree in the machine learning algorithm, averaging the contributions, and ranking the features based on their importance. Feature selection is then made according to this ranking. In this invention, the Gini index is used as the evaluation metric for measuring feature importance in these three machine learning algorithms.

[0106] The Gini index is represented by GI, and the feature importance score is represented by VIM. For the 14 features X1, X2, X3, ..., X14, the Gini index score for each feature Xj is calculated. That is, the average impurity of node splitting for the j-th feature across all decision trees in the machine learning algorithm.

[0107] The formula for calculating the Gini index is:

[0108]

[0109] Where K represents the number of categories, p mk This represents the proportion of category k in node m.

[0110] Feature X j The importance of node m, i.e., the change in the Gini index before and after node m branches, is as follows:

[0111]

[0112] Among them, GI l and GI r These represent the Gini indices of the two new nodes after the branching.

[0113] If feature X j If the nodes appearing in decision tree i are set M, then X j The importance of the i-th tree is:

[0114]

[0115] Assuming there are n trees in total in RF, then:

[0116]

[0117] Finally, all the obtained importance scores are normalized, and feature selection is performed based on the ranking.

[0118]

[0119] Step 5: Dataset construction. Based on the feature importance measurement results, four feature variables with higher influence factors are selected to form a data set X. The geological units after encoding and conversion are classified to form a prediction target set Y. X and Y are the final sample datasets for geological unit classification, and there is a one-to-one correspondence between the elements in X and Y.

[0120] Step 6: Data Splitting. X and Y are split into training, validation, and test sets. Specifically: First, the data of X and Y are randomly split into training and test sets at a ratio of 70% and 30%, respectively. Second, the training set is then randomly split again into training and validation sets at the same ratio. The classification model is trained using the training and validation sets, and the model is validated and improved through continuous iteration. Finally, the test set is used to perform classification predictions and obtain the classification results.

[0121] Step 7: Classification Model Construction and Prediction. Machine learning classification algorithms are applied to construct a model for classifying and predicting geological units. Suitable machine learning algorithms include DecisionTree, RandomForest, XGBoost, and Bagging. Through training, iteration, and validation on the training and validation sets, all four classifier models demonstrated excellent classification performance on the test set. The XGBoost algorithm showed the highest classification accuracy and is therefore the preferred choice. XGBoost is an ensemble learning algorithm based on GBDT (Gradient Boosting Decision Tree). Its main algorithm flow is as follows: input is the training set samples I = {(x1, y1), (x2, y2), ..., (xm, ym)}, maximum number of iterations T, loss function L, regularization coefficients λ and γ, and output is the strong learner f(x). For iterations t = 1, 2, ..., T:

[0122] 1) Calculate the loss function L for the i-th sample (i-1, 2, ..., m) in the current round based on f t-1 (x i The first derivative of g) ti Second derivative h ti Calculate the first derivative of all samples and and second derivative and

[0123] 2) Attempt to split the decision tree based on the current node, with a default score of 0. G and H are the sum of the first and second derivatives of the nodes to be split.

[0124] For feature indices k = 1, 2...K:

[0125] a) GL = 0, HL = 0

[0126] b.1) Arrange the samples in ascending order of feature k, and take out the i-th sample in turn. Calculate the sum of the first and second derivatives of the left and right subtrees after the current sample is placed in the left subtree:

[0127] GL = GL + gti, GR = G - GL

[0128] HL = HL + hti, HR = H - HL

[0129] b.2) Try to update the maximum score:

[0130]

[0131] 3) Split subtrees based on the partitioning features and feature values ​​corresponding to the maximum score.

[0132] 4) If the maximum score is 0, the current decision tree is complete. Calculate w for all leaf regions. tj We obtain the weak learner h t (x), update the strong learner f t (x), proceed to the next round of weak learner iteration. If the maximum score is not 0, go to step 2) to continue trying to split the decision tree.

[0133] Step 8: Evaluation of Classification Results. The performance of the model is evaluated through test results. This invention uses metrics such as confusion matrix (G), accuracy (Ac), macro-average precision (Pr), macro-average recall (Re), and macro-average F1 score (F1) to evaluate the effectiveness and results of geological unit classification. Accuracy is the percentage of correctly classified samples; macro-average precision represents the proportion of correctly predicted positive samples; macro-average recall represents the proportion of correctly predicted positive samples; and macro-average F1 score is the harmonic mean of macro-average precision and macro-average recall.

[0134] Assume confusion matrix G:

[0135]

[0136] Where k is the number of geological unit classification categories. The formulas for calculating precision (Ac), macro-average precision (Pr), macro-average recall (Re), and macro-average F1 score (F1) in the confusion matrix G are as follows:

[0137]

[0138]

[0139]

[0140]

[0141] Where gaa represents the number of correctly predicted Class A geological units; gab represents the number of Class A geological units predicted as Class B.

[0142] Step 9: Optimize Classification Results. Optimize the classification results by adjusting the classification algorithm model parameters and feature variable selection to achieve better classification performance until the classification accuracy reaches the expected value. The classification algorithm model parameters can be automatically optimized using Bayesian optimization algorithms or manually tuned. For the XGBoost classification algorithm, the adjustable parameters mainly include: learning_rate, n_estimators, max_depth, min_child_weight, gamma, subsample, colsample_bytree, objective, num_class, and seed. Optimizing feature variable selection involves appropriately increasing or decreasing feature variables to create a new sample dataset with different feature variables. Based on this new sample dataset, re-classify and predict to achieve better classification results. Generally, feature variable selection is more effective in optimizing classification results and can achieve significant results.

[0143] Figure 3 A schematic block diagram of a training apparatus for a geological unit classification model according to an embodiment of the present disclosure is shown.

[0144] like Figure 3 As shown, the training device 300 for the geological unit classification model in this embodiment includes a first data module 310, a second data processing module 320, a third data processing module 330, a fourth data processing module 340, and a training module 350.

[0145] The first data processing module 310 is used to select a target area according to the lunar topographic map grading rules and determine the image data of the target area; the second data processing module 320 is used to generate pixel grid vector data based on the image data and calculate the latitude and longitude coordinates of the center position of each pixel grid according to the pixel grid vector data to obtain data point vector data; the third data processing module 330 is used to perform spatial overlay operation based on the data point vector data to determine at least one data feature; the fourth data processing module 340 is used to determine the sample input data of the geological unit classification model to be trained based on the at least one data feature; and the training module 350 is used to train the geological unit classification model to be trained using a machine learning algorithm according to the sample input data.

[0146] In some embodiments, the third data processing module includes one or more of the following operations: extracting geological unit classification features by performing spatial overlay operations on the sample data point vector data; extracting one or more of longitude features, latitude features, and grayscale features by performing spatial overlay operations on the sample data point vector data; extracting one or more of elevation features, slope features, and undulation features by performing spatial overlay operations on the sample data point vector data; extracting TiO2 features by performing spatial overlay operations on the sample data point vector data; and extracting one or more of FeO, pyroxene, olivine, plagioclase, pyroxene, submicroscopic metallic iron, and optical maturity features by performing spatial overlay operations on the sample data point vector data.

[0147] In some embodiments, the second data processing module is configured to: take points at preset distance intervals based on the image data to generate pixel grid vector data with spatial location; and calculate the latitude and longitude coordinates of the center position of each pixel grid according to the pixel grid vector data to obtain data point vector data.

[0148] In some embodiments, the fourth data processing module is configured to: calculate the importance of the at least one data feature using an embedded method to obtain at least one importance calculation result; and determine the sample input data for the geological unit classification model to be trained based on the at least one importance calculation result.

[0149] In some embodiments, the apparatus further includes: an encoding conversion module, configured to perform encoding conversion on the geological unit classification features using a label encoding method to obtain digital codes of the geological unit classification features; and to use the digital codes of the geological unit classification features as sample input data for a geological unit classification model to be trained.

[0150] According to embodiments of this disclosure, any plurality of modules among the first data processing module 310, the second data processing module 320, the third data processing module 330, the fourth data processing module 340, and the training module 350 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first data processing module 310, the second data processing module 320, the third data processing module 330, the fourth data processing module 340, and the training module 350 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the first data processing module 310, the second data processing module 320, the third data processing module 330, the fourth data processing module 340, and the training module 350 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0151] Figure 4 A schematic block diagram of a geological unit classification device according to an embodiment of the present disclosure is shown.

[0152] like Figure 4 As shown, the geological unit classification device 400 of this embodiment includes an acquisition module 410.

[0153] The module 410 is used to input the target image into the geological unit classification model to obtain the geological unit classification result corresponding to the target image.

[0154] For example, the geological unit classification model is trained based on the apparatus provided in this disclosure.

[0155] According to embodiments of this disclosure, any plurality of modules in obtaining module 410 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of obtaining module 410 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of obtaining module 410 may be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0156] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing a training method and / or a geological unit classification method for a geological unit classification model, according to embodiments of the present disclosure.

[0157] like Figure 5 As shown, an electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0158] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0159] According to embodiments of this disclosure, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0160] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0161] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503 described above.

[0162] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the item recommendation method provided in the embodiments of this disclosure.

[0163] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0164] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0165] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0166] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0168] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0169] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A method for training a geological unit classification model, comprising: selecting a target area according to a lunar topographic map framing rule, and determining image data of the target area; generating pixel grid vector data based on the image data, and calculating longitude and latitude coordinates of each pixel grid center position based on the pixel grid vector data to obtain data point vector data; performing spatial overlay operation based on the data point vector data to determine data characteristics, including the following operations: performing spatial overlay operation on the data point vector data using full-moon geological map data to obtain geological unit classification characteristics, performing spatial overlay operation on the data point vector data and lunar image DOM data to extract longitude characteristics, latitude characteristics and gray scale characteristics, performing spatial overlay operation on the data point vector data and lunar DEM data to extract elevation characteristics, slope characteristics and relief characteristics, performing spatial overlay operation on the data point vector data and data taken by a wide-angle camera of a lunar orbiter system to extract TiO2 characteristics, performing spatial overlay operation on the data point vector data and inversion data of a multi-band imager of a lunar probe to extract FeO characteristics, diopside characteristics, olivine characteristics, plagioclase characteristics, pyroxene characteristics, sub-microscopic metallic iron characteristics and optical maturity characteristics; wherein the spatial overlay operation represents assigning attribute values of each data characteristic to corresponding data point vector data in space; based on the geological unit classification characteristics, the longitude characteristics, the latitude characteristics, the elevation characteristics, the slope characteristics, the relief characteristics, the TiO2 characteristics, the FeO characteristics, the diopside characteristics, the olivine characteristics, the plagioclase characteristics, the pyroxene characteristics, the sub-microscopic metallic iron characteristics and the optical maturity characteristics, determining sample input data of a geological unit classification model to be trained; and training the geological unit classification model to be trained using a machine learning algorithm based on the sample input data to obtain a geological unit classification model for predicting a geological unit classification result of the moon.

2. The method of claim 1, wherein, The generating pixel grid vector data based on the image data, and calculating longitude and latitude coordinates of each pixel grid center position based on the pixel grid vector data to obtain data point vector data, comprises: taking points at intervals of a preset distance range based on the image data to generate pixel grid vector data with spatial positions; and calculating longitude and latitude coordinates of each pixel grid center position based on the pixel grid vector data to obtain data point vector data.

3. The method of claim 1, wherein, The determining sample input data of a geological unit classification model to be trained, comprises: calculating importance of at least one data characteristic using an embedded method to obtain at least one importance calculation result; and determining sample input data of a geological unit classification model to be trained based on the at least one importance calculation result.

4. The method of claim 1, further comprising: performing encoding conversion on the geological unit classification characteristics using a label encoding method to obtain digital encoding of the geological unit classification characteristics; and using the digital encoding of the geological unit classification characteristics as sample input data of a geological unit classification model to be trained. ​ 5. A method for classifying a geological unit, comprising: inputting a target image into a geological unit classification model to obtain a geological unit classification result corresponding to the target image; wherein the geological unit classification model is trained according to the method of any one of claims 1 to 4.

6. A device for training a geological unit classification model, comprising: a first data processing module configured to select a target region according to a sheeting rule of a lunar topographic map and determine image data of the target region; a second data processing module configured to generate pixel grid vector data based on the image data and calculate longitude and latitude coordinates of a center position of each pixel grid based on the pixel grid vector data to obtain data point vector data; a third data processing module configured to perform spatial superposition operation based on the data point vector data to determine data features, including the following operations: performing spatial superposition operation on the data point vector data using full-moon geological map data to obtain geological unit classification features, performing spatial superposition operation on the data point vector data and lunar image DOM data to extract longitude features, latitude features and gray scale features, performing spatial superposition operation on the data point vector data and lunar DEM data to extract elevation features, slope features and relief features, performing spatial superposition operation on the data point vector data and data captured by a wide-angle camera of a lunar orbiter system to extract TiO2 features, performing spatial superposition operation on the data point vector data and inversion data of a multi-band imager of a lunar probe to extract FeO features, diopside features, olivine features, plagioclase features, pyroxene features, sub-microscopic metallic iron features and optical maturity features, wherein the spatial superposition operation represents assigning attribute values of each data feature to corresponding data point vector data in space; a fourth data processing module configured to determine sample input data of a geological unit classification model to be trained based on the geological unit classification features, the longitude features, the latitude features and the gray scale features, the elevation features, the slope features, the relief features, the TiO2 features, the FeO features, the diopside features, the olivine features, the plagioclase features, the pyroxene features, the sub-microscopic metallic iron features and the optical maturity features; and a training module configured to train the geological unit classification model to be trained using a machine learning algorithm based on the sample input data to obtain a geological unit classification model for predicting a geological unit classification result of the moon.

7. A device for classifying a geological unit, comprising: an obtaining module configured to input a target image into a geological unit classification model to obtain a geological unit classification result corresponding to the target image; wherein the geological unit classification model is trained according to the device of claim 6.

8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 5. ​ 9. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any of claims 1-5.

Citation Information

Patent Citations

  • Geological disaster multi-disaster comprehensive risk evaluation method based on random forest

    CN111582386A

  • Training method and device of electric power decomposition curve prediction model and storage medium

    CN113947201A