Multi-factor driven analysis method and system for dynamic distribution and size of lunar rocks

By classifying lunar boulders by size and combining multi-factor analysis, and employing machine learning and interpretable artificial intelligence methods, the problem of multi-factor driving the distribution pattern of lunar boulders was solved, achieving a refined and accurate interpretation of the causes of boulder distribution.

CN122049549APending Publication Date: 2026-05-15INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
Filing Date
2026-01-12
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies cannot actively explain the multi-factor driving mechanisms of the distribution of boulders on the lunar surface, and lack the ability to systematically model the synergistic effects of multiple driving factors, resulting in inaccurate interpretation of the boulder distribution pattern.

Method used

A two-group partitioning strategy based on size thresholds was adopted to fit exponential and power-law functions to boulders larger than and smaller than 10 meters, respectively. A machine learning regression model was constructed by combining multiple environmental factors, and an interpretable artificial intelligence method was used to analyze the factor contribution. A multidimensional factor analysis framework was constructed to reveal the causes of the distribution of boulders.

Benefits of technology

It enables refined, systematic, and interpretable multi-factor driven active attribution analysis of lunar boulder distribution, improving the analytical accuracy and engineering application value of lunar surface evolution processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049549A_ABST
    Figure CN122049549A_ABST
Patent Text Reader

Abstract

The invention provides a multi-factor driving analysis method and system for dynamic distribution and size of lunar rocks, and belongs to the field of planet geology remote sensing analysis and machine learning modeling. The method comprises the following steps: acquiring giant stone distribution data of a target area, and dividing giant stones into a first size group and a second size group according to a preset size threshold value; simultaneously acquiring a plurality of environmental factor data related to the giant stone distribution; function fitting is carried out on the giant stone size-frequency distribution of the first size group and the second size group, and dominant physical mechanism assumptions of different size giant stone distribution are established based on fitting results; independent machine learning regression models are constructed for the groups of different sizes; performing factor contribution degree analysis on the output of the machine learning regression model by using an interpretable artificial intelligence method; in combination with geological time, material attributes and topographic features, a multi-dimensional factor analysis framework is constructed, and complete deduction from data statistics to physical mechanism hypothesis, active attribution and mechanism verification is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of planetary geological remote sensing analysis and machine learning modeling technology, and in particular to a multi-factor driven analysis method and system for the distribution and size dynamics of lunar boulders. Background Technology

[0002] Lunar boulders are crucial records of the impact history and geological evolution of the inner solar system. The Apollo Basin, the largest impact basin within the South Pole-Aitken Basin, with its thin crust and complex geological activity (such as multiple periods of volcanic activity and impact events), is an ideal region for studying lunar crust-mantle interactions. Accurately interpreting the origins of boulder distribution is of great significance for engineering applications such as the study of extraterrestrial surface processes, deep space exploration mission planning, and landing site safety assessments.

[0003] Currently, researchers typically employ statistical modeling methods to describe the spatial distribution characteristics of boulders on the lunar surface, such as fitting the size-frequency relationship of boulders using power-law distributions or exponential functions. While these models can depict the trend of boulder quantity varying with size, they are essentially empirical descriptions, their physical meaning relying on post-hoc interpretation. They lack active, explicit physical driving mechanisms, meaning that existing methods cannot proactively explain the fundamental causes of differences in boulder distribution across different regions and lack proactive attribution capabilities. Furthermore, existing research in mechanism attribution is often limited to single-factor analysis (such as impact alone or thermal stress alone), lacking the ability to systematically model the synergistic effects of multiple driving factors. This can easily lead to attribution biases and hinder accurate interpretation of boulder distribution patterns.

[0004] Therefore, there is an urgent need to provide an analytical method for the multi-factor driving mechanism of lunar boulder distribution, which can quantify the contributions of various geological processes and have the ability to attribute mechanisms, so as to improve the analytical accuracy and engineering application value of lunar surface evolution processes. Summary of the Invention

[0005] The purpose of this application is to provide a multi-factor driven analysis method and system for the distribution and size dynamics of lunar boulders, in order to solve or alleviate the problems existing in the prior art.

[0006] To achieve the above objectives, this application provides the following technical solution: This application provides a multi-factor driven analysis method for the distribution and size dynamics of lunar boulders, including: Acquire data on the distribution of boulders in the target area, and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, acquire data on multiple environmental factors related to the distribution of boulders. Function fitting was performed on the size-frequency distribution of boulders in the first size group and the second size group, respectively. The first size group was fitted with an exponential function and the second size group was fitted with a power-law function. Based on the fitting results, the dominant physical mechanism hypothesis of the distribution of boulders of different sizes was established. Independent machine learning regression models are constructed for the first size group and the second size group respectively, using the multiple environmental factors as inputs, and outputting the predicted value of boulder density in each grid cell. The factor contribution analysis of the output of the machine learning regression model was performed using interpretable artificial intelligence methods to reveal the differential impact of various environmental factors on the distribution of boulders of different sizes. By combining geological time, material properties, and topographic features, an analytical framework is constructed to analyze the synergistic effects of different environmental factors on temporal evolution, material properties, and spatial geomorphology from a geological perspective, thereby achieving a comprehensive attribution of the causes of megalithic distribution.

[0007] Preferably, the preset size threshold is 10 meters, the first size group consists of boulders with a diameter greater than 10 meters, and the second size group consists of boulders with a diameter less than or equal to 10 meters.

[0008] Preferably, the environmental factors include topographic factors, impact characteristic factors, geological unit factors, and thermal radiation factors; The terrain factors include digital elevation model (DEM), slope, aspect, and topographic relief. The impact characteristic factor is the impact crater density; The geological unit factors are categorical variables based on geological genesis and chronological division; The thermal radiation factors include the minimum temperature, the maximum temperature, and visual brightness.

[0009] Preferably, the machine learning regression model is an extreme gradient boosting model.

[0010] Preferably, the interpretable artificial intelligence method is the SHAP method or the LIME method, used to generate a global feature importance ranking, local factor contribution values, and factor dependency graphs to reveal the differentiated impact of various environmental factors on the distribution of boulders of different sizes.

[0011] Preferably, by combining geological time, material properties, and topographic features, an analytical framework is constructed to analyze the synergistic effects of different environmental factors on temporal evolution, material properties, and spatial geomorphology, including: The differences in megalith density in geological units from different geological periods were analyzed to determine the impact of the time dimension on megalith replenishment. Analyze the differences in megalith retention rates in geological units with different rock types to determine the impact of material properties on weathering preservation; Analyze the spatial correlation between topographic features and the spatial aggregation of megaliths to determine the control effect of topography on the migration and deposition of megaliths.

[0012] Preferably, after acquiring the distribution data of boulders in the target area and dividing the boulders into a first size group and a second size group according to a preset size threshold; and simultaneously acquiring multiple environmental factor data related to the boulder distribution, the method further includes: For each particle size group, the study area was divided into regular grid cells. The number of boulders in each grid was counted and the density was calculated. At the same time, environmental factors were aggregated into the same grid cells and standardized.

[0013] Preferably, the method is applied to the analysis of the geological evolution process of the surface of celestial bodies without an atmosphere, or to the safety assessment of the landing area and the prediction of the distribution of surface obstacles in deep space exploration missions.

[0014] This embodiment provides a multi-factor driven analysis system for the distribution and size dynamics of lunar boulders. The system is used to execute the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders described above, including: The data acquisition unit is used to acquire the distribution data of boulders in the target area and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, it acquires multiple environmental factor data related to the distribution of boulders. The fitting unit is used to perform function fitting on the size-frequency distribution of boulders in the first size group and the second size group, respectively. The first size group is fitted with an exponential function, and the second size group is fitted with a power-law function. Based on the fitting results, the dominant physical mechanism hypothesis of boulders of different sizes is established. The model building unit is used to build independent machine learning regression models for the first size group and the second size group respectively, taking the multiple environmental factors as inputs and outputting the predicted value of the boulder density in each grid cell. The contribution analysis unit is used to perform factor contribution analysis on the machine learning regression model using interpretable artificial intelligence methods, quantify the relative importance of each environmental factor to the prediction of boulder density, and identify the direction of their influence. Mechanism attribution unit is used to combine geological time, material properties and topographic features to construct an analytical framework for the combined effects of multiple factors, analyze the synergistic mechanism of different environmental factors in temporal evolution, material properties and spatial geomorphology, and achieve a comprehensive attribution of the causes of megalithic distribution.

[0015] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders described in any of the above embodiments.

[0016] Beneficial effects: The solution provided in this embodiment employs a two-group (first size group and second size group) partitioning and differentiated modeling strategy based on size thresholds. By dividing the boulder into size groups, different functions are used for fitting, and a machine learning model is constructed. Combined with environmental factor analysis, the causes of boulder distribution are ultimately attributed comprehensively, achieving a shift from statistical model distribution pattern description to a comprehensive attribution of active, size-specific, and differentiated physical mechanisms. Through quantitative impact analysis of multiple environmental factors, combined with a multidimensional factor analysis framework, multi-factor driven analysis is integrated. Furthermore, this method alleviates the shortcomings of existing technologies, such as the ambiguity of the physical mechanisms of boulder distribution at certain sizes and the lack of field verification, by introducing interpretable artificial intelligence (XAI) methods and the multidimensional factor analysis framework for interactive verification, thus improving the scientific rigor and rationality of the explanation of boulder distribution mechanisms. In summary, this solution provides a new analytical paradigm for the analysis and research of driving factors in boulder distribution and size dynamics. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the structure of a computer device provided according to some embodiments of this application.

[0018] Figure 2 Histogram statistics of the diameter (D) of the boulder provided according to some embodiments of this application.

[0019] Figure 3 This is a schematic diagram showing the fitting results of exponential and power-law functions for different size groups.

[0020] Figure 4 The XGBoost-SHAP framework is used for the driving factor analysis of large boulders (diameter > 10 meters). Among them, (a) is the global feature importance ranking, also known as the feature average SHAP value histogram, which reflects the average influence strength of each input variable on the model output. (b) is the local factor contribution value, also known as the SHAP value distribution plot, which shows the direction and magnitude of the influence of each variable on the predicted value under different value ranges.

[0021] Figure 5 The XGBoost-SHAP framework is used for the driving factor analysis of small boulders (diameter < 10 meters). (a) is the global feature importance ranking, also known as the feature average SHAP value histogram, which reflects the average influence strength of each input variable on the model output. (b) is the local factor contribution value, also known as the SHAP value distribution plot, which shows the direction and magnitude of the influence of each variable on the predicted value under different value ranges. Detailed Implementation

[0022] This application proposes a multi-factor driven analysis method and system for the distribution and size dynamics of lunar boulders. It comprehensively considers multiple driving mechanisms such as geological unit type, impact crater characteristics, and topographic parameters. Through data acquisition, differential fitting with dual-size threshold grouping to correlate physical mechanisms, factor contribution analysis based on XAI to achieve proactive preliminary attribution, and a multi-dimensional factor analysis framework to reveal synergistic mechanisms, it constructs a multi-factor driven analysis system with both scientific explanatory and predictive capabilities. This effectively solves the core bottlenecks of existing technologies that emphasize description over attribution and single factors over synergistic mechanisms. It achieves refined, mechanistic, and interpretable multi-factor driven proactive, automatic, and accurate attribution analysis of the causes of boulder distribution in the Apollo Basin of the Moon, providing a powerful new research tool for revealing the impact history of the inner solar system, lunar crust-mantle interactions, and surface evolution processes.

[0023] The embodiments of this application will now be described with reference to the accompanying drawings.

[0024] The embodiments of this application can be applied to Figure 1 The computer devices (also known as electronic devices) shown may include, but are not limited to, mobile terminals such as mobile phones, tablets, handheld computers, and personal digital assistants (PDAs), smart home devices such as smart TVs and smart cameras, wearable devices such as smart bracelets, smartwatches, and smart glasses, or other desktop, laptop, notebook, ultra-mobile personal computer (UMPC), netbook, and smart screen computer devices.

[0025] like Figure 1 As shown, the electronic device 200 may include one or more of the following components: a processor 201, a memory 203, a communication interface 202, and a communication bus 204. The memory 203 can be connected to the processor 201 via the bus 204. The bus can transfer data between the processor 201 and the memory 203. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0026] Processor 201 may include one or more processing cores. Processor 201 can connect to various parts within the electronic device 200 using various interfaces and lines. It performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 203, and by calling data stored in memory 203. For example, processor 201 may include an application processor (AP), a modem processor, a CPU, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), and / or a neural network processing unit (NPU). The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; the NPU implements artificial intelligence (AI) functions; and the modem handles wireless communication. Different processing units can be independent devices or integrated into one or more processors. For example, the multiple processing units shown above are all integrated into a single SoC, or the AP is a separate semiconductor chip, while other processing units are integrated into a single SoC. This application does not limit this to any particular type.

[0027] The memory 203 may include random access memory (RAM), read-only memory (ROM), or non-transitory computer-readable storage medium. The memory 203 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 203 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function, such as a multi-factor driven analysis method for the dynamic distribution and size of lunar boulders; the data storage area may store data created based on the use of the electronic device 200, such as input data for numerical solutions.

[0028] In addition, those skilled in the art will understand that the structure of the electronic device 200 shown in the above figures does not constitute a limitation on the electronic device 200. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device 200 may also include components such as a microphone, speaker, radio frequency circuit, sensor, audio circuit, power supply, and Bluetooth module, which will not be described in detail here.

[0029] This embodiment provides a multi-factor driven analysis method for the distribution and size dynamics of lunar boulders. It aims to reveal the physical mechanisms of the generation, migration, and preservation of boulders of different sizes under complex geological backgrounds, influenced by the synergistic effects of multiple environmental factors. This allows for attribution of the causes of the spatial distribution pattern and size of boulders, providing technical support for lunar geological evolution research and deep space exploration mission planning. The method includes: Step S1: Obtain the distribution data of boulders in the target area, and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, obtain the data of multiple environmental factors related to the distribution of boulders.

[0030] In this step, the target area for analysis is first determined. The target area is preferably the Apollo Basin and its surrounding typical geological units within the South Pole-Aitken Basin (SPA) on the far side of the Moon. This area has a complex impact history, multiple phases of volcanic infill, and significant differences in thermodynamic environment, making it an ideal region for studying the multi-factor driving mechanism of megaliths.

[0031] The distribution data of the megaliths is obtained as follows: high-resolution remote sensing images (hereinafter referred to as image data) are acquired first, and the location and size of the megaliths are extracted based on the image data to obtain the distribution data of the megaliths. Among them, the image data can be, for example, images from the narrow-angle camera (NAC) carried by NASA's Lunar Reconnaissance Orbiter (LRO), with a spatial resolution of 0.5–2.0 meters per pixel, covering the entire target area.

[0032] Preferably, the megalith distribution data can also utilize existing research findings, such as the global megalith dataset publicly released by scholars Bickel et al. in 2020, which contains 136,610 megalith points with diameters ranging from 3 to 25 meters and a spatial accuracy of no more than 2 meters. This dataset records the location and size of the megaliths.

[0033] In this step, after obtaining the distribution data of the boulders, a preset size threshold is introduced to group the boulders in order to distinguish the preference of different formation mechanisms for the size of boulders.

[0034] Preferably, the diameter (or simply size) of the boulder can be determined based on statistical results, for example, by setting the median diameter as a size threshold, and then grouping the boulders.

[0035] In one example Figure 2 This shows the histogram statistics of the diameter (D) of the boulders in the target area. For example... Figure 2 As shown, the histogram displays the frequency distribution of the megalith diameter (D), and the median megalith diameter (Dmedium). 中位数 The value is 10.68 meters, indicating that the center of the boulder size distribution is concentrated around 10 meters.

[0036] Based on the above statistical results, preferably, the preset size threshold is 10 meters. Correspondingly, the first size group is boulders with a diameter greater than 10 meters, and the second size group is boulders with a diameter less than or equal to 10 meters. That is, the boulders are divided into the >10-meter group (the object of the exponential function fitting) and the <10-meter group (the object of the power law function fitting) according to their diameter, and two datasets of particle size groups are constructed respectively.

[0037] In this embodiment, the boulder distribution data includes the spatial distribution and size information of the boulders, providing a data foundation for subsequent modeling.

[0038] In this embodiment, multiple environmental factors related to the distribution of boulders include four dimensions: topographic factors, impact characteristic factors, geological unit factors, and thermal radiation factors. Among them, the topographic factors further include: digital elevation model (DEM), slope, aspect, and topographic relief; the impact characteristic factor is the density of impact craters; the geological unit factor is a classification variable based on geological genesis and age division; and the thermal radiation factor includes minimum temperature, maximum temperature, and visual brightness.

[0039] In this embodiment, an environmental factor system is constructed to lay the data foundation for multi-factor driven analysis. When acquiring megalith distribution data, 10 environmental factors closely related to the megalith formation, migration, and preservation processes are simultaneously collected to construct the environmental factor system, as detailed below: (1) Topographic factors (4 items): Digital Elevation Model (DEM): Originating from SLDEM2015 or LOLA global DEM products, with an original resolution of 100m; it should be noted that terrain factors including digital elevation model (DEM) refer to terrain factors including elevation values ​​(grayscale values ​​of DEM data). Slope: Generated by calculating the gradient from the DEM, in degrees (°). Aspect: Generated based on DEM, indicating the direction of maximum surface descent; Roughness: Generated based on DEM.

[0040] Relief degree: Based on DEM data, the standard deviation of elevation is calculated within a 3×3 window to reflect the degree of local fragmentation.

[0041] It should be noted that the specific extraction methods for topographic factors such as slope, aspect, and topographic relief based on DEM data can refer to existing technologies, and will not be elaborated here.

[0042] (2) Impact characteristic factors (1 item): (3) Crater Density, derived from meteorite crater density data, is obtained by performing kernel density analysis on impact craters with a diameter ≥ 1 km based on the International Lunar Impact Crater Database. This index can reflect the cumulative impact flux in the region.

[0043] (4) Geological unit factors (1 item): Geologic Unit: Based on the latest lunar geological map classification system proposed by Guo et al. (2024), the study area is divided into 14 geological units, such as the edge of the Apollo Basin (Ar_Apollo), the SPA plain (SPbsi), and the Copernican impact crater material (Cppmian). For specific geological unit classifications, please refer to the literature provided by Guo et al., which will not be elaborated here.

[0044] (5) Thermal radiation and optical property factors (3 items): Visual brightness (Albedo), minimum temperature (MiniTemperature), maximum temperature (MaxTemperature): Using LRO Diviner data with a resolution of 1km, the diurnal temperature range (ΔT) was calculated using the minimum and maximum temperatures.

[0045] After acquiring the distribution data of boulders and multiple environmental factor data in the target area, the method provided in this embodiment further includes a systematic data preprocessing step to achieve structured alignment of spatial data, standardization of variables, and improved modeling compatibility, ensuring that data from different sources and at different scales can be analyzed together within a unified spatial framework. Specifically, for each particle size group, the study area is divided into regular grid cells, the number of boulders in each grid is counted and the density is calculated, and environmental factors are aggregated into the same grid cells and standardized.

[0046] This step includes two parts: grid statistics and feature engineering.

[0047] Grid statistics are used to transform point-distributed megalithic data into continuous field variables that can be used for regional modeling. This embodiment employs a regular geographic grid division method, that is, for each grain size group, the target study area (such as the Apollo Basin and its surrounding geological units) is divided into regular grid cells, and the number of megaliths within each grid is counted and the density is calculated. For example, the study area can be divided into 1°×1° grids, which corresponds to a spatial range of approximately 30 km × 30 km in the mid-latitude regions of the Moon (considering lunar curvature correction). This ensures sufficient sample density while effectively reducing local noise interference, making it suitable for large-scale geological process analysis.

[0048] For each grain size group (i.e., the first size group D > 10 meters, the second size group D < 10 meters), perform the following statistics to determine the boulder density for the corresponding grain size within each grid: (1) Extract all boulders belonging to that size group that fall within each grid cell; (2) Count the total number of boulders in each grid; (3) Calculate the area of ​​the grid; (4) Calculate the density of the boulder: , In the formula, Represents a grid The density of the boulders , They are grids The total number and area of ​​the megaliths inside.

[0049] After the above steps, all grid cells form a two-dimensional spatial matrix, with each cell corresponding to a boulder density value. This realizes the conversion from discrete point data of boulder distribution to continuous field data, which facilitates spatial matching and regression modeling with rasterized environmental factors. At the same time, the statistical method of grouping by particle size preserves the independent evolutionary information of boulders at different scales, providing a basis for subsequent differential attribution analysis.

[0050] Feature engineering is used to ensure that environmental factor data and boulder density data are comparable in both spatial and numerical dimensions. In this embodiment, feature engineering includes two steps: spatial aggregation and numerical standardization. Spatial aggregation was used to uniformly resample and aggregate the previously acquired environmental factor data (which had different original resolutions) into a 1°×1° grid system identical to the boulder density data. The specific operations were as follows: For continuous variables (such as DEM, diurnal temperature range ΔT, visual brightness Albedo, etc.), bilinear interpolation or regional averaging was used to aggregate high-resolution data (such as 100m topographic data, 1km Diviner data) into the average value of each 1°×1° grid; for categorical variables (such as geological unit type), majority resampling was used, taking the geological unit with the highest frequency within each grid as the representative type of that grid; for impact crater density, since a continuous raster layer (1km resolution) had already been generated during the kernel density analysis stage, it could be directly aggregated into the average value of the 1°×1° grid.

[0051] The specific steps for numerical standardization are as follows: Z-score standardization (also known as standard score transformation) is used to standardize continuous variables.

[0052] For categorical variables (such as geological unit types), the one-hot encoding method is used for numerical processing. The specific steps are as follows: (1) Convert the original 14 geological units (such as Ar_Apollo, SPbsi, Cppmian, etc.) into 14 binary variables (0 / 1); (2) Each 1°×1° grid is marked as 1 at the location of its corresponding geological unit, and 0 for the rest; (3) The encoded variables together with other continuous variables constitute the model input feature vector.

[0053] After the above feature engineering process, all environmental factors are transformed into a feature matrix with consistent dimensions. Each row is a 1°×1° grid cell, and each column is a standardized predictor variable (a total of 10 environmental factors, 14 geological unit dummy variables, and a total of 24-dimensional feature space), which is aligned with the boulder density response variable.

[0054] Step S2: Perform function fitting on the megalith size-frequency distribution of the first size group and the second size group respectively, wherein the first size group is fitted with an exponential function and the second size group is fitted with a power-law function. Based on the fitting results, establish the dominant physical mechanism hypothesis for the distribution of megaliths of different sizes.

[0055] In this embodiment, after completing the grouping and spatial preprocessing of the megalith data by size, megalith distribution modeling is performed. That is, exponential and power-law functions are used to fit the abundance curve of the stones to achieve size-frequency distribution modeling of megaliths of different grain sizes. The statistical regularity is revealed by fitting mathematical functions, and hypotheses about the dominant generation physical mechanism are proposed based on this.

[0056] Specifically, such as Figure 3 As shown, for the first and second size groups, nonlinear least squares fitting is performed using exponential and power-law functions, respectively, and corresponding geological process hypotheses are established based on the goodness of fit and physical background.

[0057] I. First size group: Exponential function fitting of large megaliths (>10 m).

[0058] For large boulders with a diameter greater than 10 meters (belonging to the first size group), their quantity decreases rapidly with increasing size; therefore, this embodiment uses an exponential decay function for fitting. , In the formula, The diameter of the boulder The abundance of megaliths per unit area. This is the proportionality coefficient. Attenuation rate parameter.

[0059] In this embodiment, all areas within the Apollo Basin... For boulders larger than 10 meters, least squares fitting was performed to obtain the optimal parameters: , The model has a high goodness of fit (coefficient of determination). >0.87), indicating that the exponential function can effectively characterize the distribution trend of large boulders. Based on this fitting result, a hypothesis is established that the physical mechanism of surface exfoliation dominated by thermal stress is determined, namely, the grain size... The 10-meter-tall boulder was mainly formed by surface spalling caused by long-term thermal fatigue.

[0060] II. Second size group: Power law function fitting for small-sized boulders (<10 meters).

[0061] For small boulders with a diameter of less than 10 meters (belonging to the second size group), their number increases sharply as the size decreases. Therefore, this embodiment uses a power-law function for fitting: , In the formula, This is a normalization constant; This is the power-law exponent.

[0062] For all diameters within the Apollo Basin For boulders less than 10 meters in height, least squares fitting was performed to obtain the optimal parameters: , The above fitting results show that the power-law function has a diameter The distribution of boulders smaller than 10 meters has good statistical explanatory power, and the power-law exponent... =1.3487 is within the typical impact fragmentation theory prediction range (1.0–1.8), indicating that the region has experienced a cascade fragmentation process triggered by multiple high-energy impact events.

[0063] In this embodiment, different applicable functions are used for stones of different sizes, reflecting the different distribution characteristics of different stones and implying their different fragmentation processes. In other words, different statistical models should be applied to boulders of different size ranges, which essentially reflects the differences in the underlying dominant physical processes. Based on this, the following size grouping and physical mechanism hypothesis system is established: For the first size group (D > 10 meters), an exponential function is applied, and the dominant physical mechanism is inferred to be lamellar spalling caused by thermal stress fatigue; for the second size group (D < 10 meters), a power-law function is applied, and the dominant physical mechanism is inferred to be a cascading effect caused by impact fragmentation.

[0064] Furthermore, the rationality of setting the size threshold was verified iteratively based on the fitting residual analysis: if the sum of squared residuals of the exponential and power-law statistical models changed significantly around 10 meters, it indicated that the statistical behavior had undergone a fundamental change, which showed that setting the size threshold of 10 meters was reasonable.

[0065] The steps described above in this embodiment break through the limitations of traditional single-model description of global distribution, realize the megalith size division and the assumption of dominant physical mechanism, and provide a prior constraint framework for multi-factor attribution analysis in subsequent steps.

[0066] Step S3: Construct independent machine learning regression models for the first size group and the second size group respectively, using multiple environmental factors as inputs, and outputting the predicted value of boulder density in each grid cell.

[0067] Step S4: Use interpretable artificial intelligence methods to perform factor contribution analysis on the output of the machine learning regression model to reveal the differentiated impact of various environmental factors on the distribution of boulders of different sizes, so as to make a preliminary attribution to the hypothesis of the dominant physical mechanism.

[0068] To further reveal the dominant driving factors and their intensity in the distribution of megaliths of different sizes, this embodiment proposes a particle size modeling strategy based on size grouping and statistical modeling. Specifically, independent machine learning regression models are constructed for the first size group (>10 meters) and the second size group (<10 meters). This method can avoid signal confusion between different causal mechanisms and improve the accuracy and physical consistency of attribution analysis.

[0069] This embodiment uses the extreme gradient boosting (XGBoost) model as the core algorithm of the machine learning regression model, and combines it with the SHAP (Shapley Additive exPlanations) interpretability analysis framework to construct a dual-group XGBoost-SHAP model system, thereby quantifying the contribution of driving factors to the density of large and small boulders.

[0070] Specifically, based on the physical mechanism assumptions of step S2, different groups ( The first size group >10 meters and The second size (<10 meters) may be dominated by different physical processes. This embodiment does not use a single global model, but instead trains independent XGBoost regression models for the two particle size groups in step S3: Model 1: Predicts the density of megaliths larger than 10 meters in size; Model 2: Predict the density of small boulders <10 meters in size.

[0071] The two models share the same input variable system, but capture the unique response relationships of different size groups through independent training (with different training samples), thereby achieving a refined analysis of the differences in driving mechanisms under different sizes.

[0072] The two models share the same input variable system, meaning that the input features of the two models are consistent and are both based on multiple environmental factors obtained in step S1, including: visual brightness, impact crater density, DEM, slope, aspect, topographic relief, geological units, minimum temperature, and maximum temperature. All of the above variables have been spatially aggregated and numerically standardized on 1°×1° grid cells to ensure data consistency and modeling stability.

[0073] The training process for each model is as follows: First, the dataset is divided: all 1°×1° grid samples belonging to this group within the target area are randomly divided into training and test sets in a ratio of 8:2, with 80% used for model training and 20% used for independent performance validation to ensure that the evaluation results are free from overfitting bias.

[0074] Secondly, hyperparameter optimization: 10-fold cross-validation combined with grid search is used to search for the optimal hyperparameter combination on the training set. The optimization objective is to minimize the prediction error (MAE). The final optimal parameters are as follows: The maximum depth of the tree (max_depth) = 5 (to control model complexity and prevent overfitting); Learning rate = 0.01 (to improve convergence stability); Bag score = 0.75 (introducing randomness to enhance generalization ability).

[0075] The model was trained on the training set using optimized parameters, and the predicted boulder density for each grid cell was output. The prediction performance of the two models was evaluated on an independent test set. The main metrics are as follows: Mean Absolute Error (MAE) = 0.107, Root Mean Square Error (RMSE) = 0.1467. The above results show that the model has high prediction accuracy and robustness.

[0076] In step S4, after the model training is completed, an interpretable artificial intelligence method is used to analyze the differences in driving factors and quantify the marginal contribution of each environmental factor to the prediction results. The interpretable artificial intelligence method can be either the SHAP method or the LIME method; the SHAP method will be used as an example below for specific explanation.

[0077] For both models, the SHAP method was used to generate global feature importance rankings, local factor contribution values, and factor dependency graphs. Interpretability analysis was performed on the model outputs to quantify the marginal contribution of each environmental factor to the megalith density prediction results. The SHAP value distributions of the two models were analyzed and compared (e.g., ...). Figure 4 , Figure 5 As shown in the figure, we can see that: (a) For large boulders with a diameter greater than 10 meters: (1) Crater Density: The average |SHAP|≈0.031 is the strongest driving factor. Areas with high crater density (such as areas with dense young craters) correspond to positive SHAP values, indicating that recent impact events replenish large boulders through ejecta, verifying the impact replenishment mechanism. (2) Visual Brightness: The average |SHAP|≈0.029 is the second highest. The distribution of SHAP values ​​shows a significant negative correlation—high brightness areas (such as high albedo surfaces) inhibit the formation of boulders, while low brightness areas (such as dark basalt) promote their existence, clearly revealing their inhibitory effect. (3) Digital Elevation Model (DEM, i.e., altitude): The average |SHAP|≈0.025. High-altitude areas located at the edge of basins or ancient highlands have thinner crusts and more frequent tectonic activity, which is conducive to exposing unweathered bedrock, thereby increasing the density of boulders. (4) Geologic Unit: Average |SHAP|≈0.018. (5) Mini Temperature and Max Temperature: Reflect the effects of low and high temperatures on thermal stress, respectively. (6) Relief Degree, Aspect, Slope, and Roughness: The influence is relatively small, and the SHAP value distribution is narrow, indicating that their influence is relatively stable and weak. This further verifies that the migration ability of large boulders (greater than 10 meters) is weak and is mainly controlled by the in-situ fracturing process.

[0078] (II) For small boulders with a diameter of less than 10 meters: (1) Visual brightness: Average |SHAP|≈0.041, the strongest driving factor. The SHAP distribution shows a significant negative correlation - high brightness areas (such as high albedo surfaces) inhibit the formation of boulders, while low brightness areas (such as dark basalt) promote their existence. This effect is stronger than that of the >10m group, indicating that small fragments are more sensitive to thermophysical conditions. (2) Geologic unit: Average |SHAP|≈0.028, ranking second. (3) Crater density: Average |SHAP|≈0.028. (4) Digital elevation model (DEM, altitude): Average |SHAP|≈0.022. (5) Maximum temperature and minimum temperature. (6) Relief Degree, Aspect, Roughness, and Slope: The impact is relatively small, but slightly higher than that of the >10-meter megalithic group, indicating that small and medium-sized boulders are more likely to be transported and accumulated by gravity, and the topographic gradient has a certain regulating effect on their spatial distribution.

[0079] Through comparative analysis, the above process quantitatively reveals the dominant physical mechanisms of the two groups of megaliths: the large-sized megalith group (>10 meters) is based on impact replenishment, superimposed with thermal stress exfoliation and topographic background regulation, forming a composite driving mode; while the small-sized megalith group (<10 meters) is dominated by impact-volcanic activity, with geological units and impact flux being the decisive factors.

[0080] Furthermore, it is possible to compare and analyze the differences in the effects of the same environmental factor on different particle size groups, such as how low temperature only promotes the fragmentation of large boulders, thus enabling visualization and attribution of differences in cross-scale driving paths.

[0081] Step S5 involves constructing an analytical framework that combines geological time, material properties, and topographic features to analyze the synergistic effects of different environmental factors on temporal evolution, material properties, and spatial geomorphology from a geological perspective, thereby achieving a comprehensive attribution of the causes of megalithic distribution.

[0082] Based on the completion of particle size modeling and attribution analysis of driving factors, this embodiment further proposes a multi-dimensional factor analysis framework. Starting from three dimensions—geological time evolution, material property differences, and topographic spatial control—it systematically explains the synergistic mechanism of different environmental factors on the formation, migration, and preservation of megaliths in long-term geological processes. This framework breaks through the limitations of single variable or static models and realizes a comprehensive attribution of the dynamic, material, and geomorphological distribution patterns of megaliths.

[0083] This framework, with the core logic of "impact period-material-topography interaction," reveals the underlying geological factors behind the spatial distribution and abundance differences of boulders of different grain sizes. Specifically, it includes: analyzing the differences in boulder density within geological units of different geological periods to determine the impact of time on boulder replenishment; analyzing the differences in boulder retention rates within geological units of different rock types to determine the impact of material properties on weathering and preservation; and analyzing the spatial correlation between topographic features and the spatial aggregation of boulders to determine the control role of topography in boulder migration and deposition.

[0084] In the temporal dimension, by comparing the megalith density of different geological periods (Copernican, Eratosonic, etc.), the results show that the megalith concentration is higher in younger impact units (such as the Copernican). In the material dimension, by analyzing the correlation between rock type (anthracite, ultramafic rock) and megalith preservation, the results show that younger ejecta have a higher megalith retention rate due to their strong resistance to weathering. In the topographic dimension, through spatial overlay analysis, it is revealed that megaliths are concentrated along the edges of impact craters and steep slopes, and the topographic relief is positively correlated with the megalith density.

[0085] The method provided in this embodiment can be applied to the analysis of geological evolution processes on the surface of atmospheric bodies, or to the safety assessment of landing areas and the prediction of surface obstacle distribution in deep space exploration missions. It can also provide a tool for assessing the risk of boulders in landing areas for missions such as Chang'e 6 (e.g., avoiding areas with high density of boulders). A dynamic equilibrium model of "impact-thermal stress-geological units" is established, the dynamic of which is reflected in the fact that new impacts will produce new boulders of different sizes, and that large boulders will break down and weather into smaller boulders due to heat. This model can be extended to the study of surface processes on other atmospheric bodies (such as Mercury and asteroids).

[0086] In summary, the method provided in this embodiment uses exponential functions to fit the first size group (characterizing surface exfoliation dominated by thermal stress) and power-law functions to fit the second size group (corresponding to the cascade effect of impact fragmentation). Based on the fitting results, a hypothesis of the dominant physical mechanism is established, achieving a preliminary mapping from data fitting to the physical mechanism. Machine learning regression models are trained for boulders of different size groups, and the XAI method is introduced for factor contribution analysis to quantify the contribution of each environmental factor to the density prediction of boulders of different sizes. This achieves an active and preliminary attribution of the distribution of boulders in different size groups, while simultaneously providing data-driven verification of the dominant physical mechanism hypothesis from a quantitative perspective. Furthermore, by integrating geological time evolution, material properties, and topographic features, a three-dimensional factor analysis framework (i.e., a multi-dimensional factor analysis framework) is constructed to explain the synergistic mechanism of multiple factors from a geological perspective, supporting the geological mechanism of the preliminary attribution using the Explainable Artificial Intelligence (XAI) method, and further verifying the rationality of the attribution. Through a triple strategy of subscale modeling, explainable machine learning, and the three-dimensional factor analysis framework, a complete inference chain from data to physical mechanism explanation and attribution is realized.

[0087] The method provided in this embodiment is essentially an analytical method for multi-factor driven mechanisms based on particle size. Implementing this method demonstrates that: 1. Thermal stress dominates the fragmentation of large boulders: the lowest temperature is a key factor in the distribution of boulders >10m in diameter. Diurnal temperature variations induce surface tensile stress, leading to layered spalling (thermal wave penetration depth ~1m, affecting only surface fragmentation). 2. Impact controls the replenishment of small boulders: crater density is strongly positively correlated with the density of small boulders (<10m in diameter), indicating that young impact units directly replenish the boulder source through ejecta. 3. Spatiotemporal constraints of geological units: the boulder density at the edges of young impact craters is 3–5 times higher than that of ancient volcanic plains because the ejecta has not been weathered / buried over a long period; volcanic units have the lowest boulder density due to lava flow coverage.

[0088] Based on the same inventive concept, this embodiment provides a multi-factor driven analysis system for the distribution and size dynamics of lunar boulders. This system is used to execute the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders provided in any of the above embodiments, including: The data acquisition unit is used to acquire the distribution data of boulders in the target area and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, it acquires multiple environmental factor data related to the distribution of boulders. The fitting unit is used to perform function fitting on the size-frequency distribution of boulders in the first size group and the second size group, respectively. The first size group is fitted with an exponential function, and the second size group is fitted with a power-law function. Based on the fitting results, the dominant physical mechanism hypothesis of boulders of different sizes is established. The model building unit is used to build independent machine learning regression models for the first size group and the second size group respectively, taking the multiple environmental factors as inputs and outputting the predicted value of the boulder density in each grid cell. The contribution analysis unit is used to perform factor contribution analysis on the machine learning regression model using interpretable artificial intelligence methods, quantify the relative importance of each environmental factor to the prediction of boulder density, and identify the direction of their influence. Mechanism attribution unit is used to combine geological time, material properties and topographic features to construct an analytical framework for the combined effects of multiple factors, analyze the synergistic mechanism of different environmental factors in temporal evolution, material properties and spatial geomorphology, and achieve a comprehensive attribution of the causes of megalithic distribution.

[0089] The multi-factor driven analysis system for the distribution and size dynamics of lunar boulders provided in this embodiment can realize the steps and processes of the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders provided in any of the above embodiments, and achieve the same technical effect, which will not be described in detail here.

[0090] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A multi-factor driven analysis method for the distribution and size dynamics of lunar boulders, characterized in that, include: Acquire data on the distribution of boulders in the target area, and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, acquire data on multiple environmental factors related to the distribution of boulders. Function fitting was performed on the size-frequency distribution of boulders in the first size group and the second size group, respectively. The first size group was fitted with an exponential function and the second size group was fitted with a power-law function. Based on the fitting results, the dominant physical mechanism hypothesis of the distribution of boulders of different sizes was established. Independent machine learning regression models are constructed for the first size group and the second size group respectively, using the multiple environmental factors as inputs, and outputting the predicted value of boulder density in each grid cell. The factor contribution analysis of the output of the machine learning regression model was performed using interpretable artificial intelligence methods to reveal the differential impact of various environmental factors on the distribution of boulders of different sizes. By combining geological time, material properties, and topographic features, an analytical framework is constructed to analyze the synergistic effects of different environmental factors on temporal evolution, material properties, and spatial geomorphology from a geological perspective, thereby achieving a comprehensive attribution of the causes of megalithic distribution.

2. The method according to claim 1, characterized in that, The preset size threshold is 10 meters. The first size group consists of boulders with a diameter greater than 10 meters, and the second size group consists of boulders with a diameter less than or equal to 10 meters.

3. The method according to claim 1, characterized in that, The environmental factors include topographic factors, impact characteristic factors, geological unit factors, and thermal radiation factors; The terrain factors include digital elevation model (DEM), slope, aspect, and topographic relief. The impact characteristic factor is the impact crater density; The geological unit factors are categorical variables based on geological genesis and chronological division; The thermal radiation factors include the minimum temperature, the maximum temperature, and visual brightness.

4. The method according to claim 1, characterized in that, The machine learning regression model is an extreme gradient boosting model.

5. The method according to claim 1, characterized in that, The interpretable artificial intelligence method is either the SHAP method or the LIME method, used to generate global feature importance ranking, local factor contribution values, and factor dependency graphs to reveal the differentiated impact of various environmental factors on the distribution of boulders of different sizes.

6. The method according to claim 1, characterized in that, By combining geological time, material properties, and topographic features, an analytical framework is constructed to analyze the synergistic effects of different environmental factors on temporal evolution, material properties, and spatial geomorphology, including: The differences in megalith density in geological units during different geological periods were analyzed to determine the impact of the time dimension on megalith replenishment. Analyze the differences in megalith retention rates in geological units with different rock types to determine the impact of material properties on weathering preservation; Analyze the spatial correlation between topographic features and the spatial aggregation of megaliths to determine the control effect of topography on the migration and deposition of megaliths.

7. The method according to claim 1, characterized in that, After acquiring data on the distribution of boulders in the target area and dividing the boulders into a first size group and a second size group according to a preset size threshold, and simultaneously acquiring data on multiple environmental factors related to the boulder distribution, the method further includes: For each particle size group, the study area was divided into regular grid cells. The number of boulders in each grid was counted and the density was calculated. At the same time, environmental factors were aggregated into the same grid cells and standardized.

8. The method according to claim 1, characterized in that, The method can be applied to the analysis of geological evolution processes on the surface of atmospheric celestial bodies, or to the assessment of landing area safety and prediction of surface obstacle distribution in deep space exploration missions.

9. A multi-factor driven analysis system for the distribution and size dynamics of lunar boulders, the system being used to execute the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders as described in any one of claims 1 to 8, characterized in that, include: The data acquisition unit is used to acquire the distribution data of boulders in the target area and divide the boulders into a first size group and a second size group according to a preset size threshold; at the same time, it acquires multiple environmental factor data related to the distribution of boulders. The fitting unit is used to perform function fitting on the size-frequency distribution of boulders in the first size group and the second size group, respectively. The first size group is fitted with an exponential function, and the second size group is fitted with a power-law function. Based on the fitting results, the dominant physical mechanism hypothesis of boulders of different sizes is established. The model building unit is used to build independent machine learning regression models for the first size group and the second size group respectively, taking the multiple environmental factors as inputs and outputting the predicted value of the boulder density in each grid cell. The contribution analysis unit is used to perform factor contribution analysis on the machine learning regression model using interpretable artificial intelligence methods, quantify the relative importance of each environmental factor to the prediction of boulder density, and identify the direction of their influence. Mechanism attribution unit is used to combine geological time, material properties and topographic features to construct an analytical framework that combines multiple factors to analyze the synergistic mechanism of different environmental factors in temporal evolution, material properties and spatial geomorphology, and to achieve a comprehensive attribution of the causes of megalithic distribution.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the multi-factor driven analysis method for the distribution and size dynamics of lunar boulders as described in any one of claims 1 to 8.