A machine learning based geological sample composition identification and anomaly detection system

By integrating X-ray fluorescence spectroscopy and short-wave infrared reflectance spectroscopy through a machine learning system, the problem of inconsistent measurement results in geological sample composition identification and anomaly detection was solved, achieving efficient and reliable composition identification and anomaly detection.

CN121068531BActive Publication Date: 2026-02-10CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511605748.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-10
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

In existing technologies, the measurement results of short-wave infrared reflectance spectroscopy and X-ray fluorescence energy spectroscopy of geological samples lack a unified metrological scale, are easily affected by particle size and surface condition, leading to deviations in composition determination and making it difficult to meet the needs of rapid and reliable composition identification and anomaly detection.

Method used

A machine learning-based geological sample composition identification and anomaly detection system is adopted. The system acquires X-ray fluorescence energy spectrum and short-wave infrared reflectance spectrum through the spectral acquisition module. Combined with atomic number metrology module, parameterized particle size extraction module, clay ratio fitting module, endmember weight calculation module and cross-mode inconsistency module, the data is processed and optimized. Finally, the composition is determined by the composition demixing module.

Benefits of technology

It effectively eliminates systematic interference from particle size, coating, and equipment drift in the determination, improves the accuracy and stability of component identification, ensures the consistency and comparability of results, reduces the number of retesting rounds and resource consumption, and improves the accuracy and timeliness of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121068531B_ABST
    Figure CN121068531B_ABST
Patent Text Reader

Abstract

The application discloses a kind of geological sample component identification and abnormal detection system based on machine learning, it is related to geological sample component identification and detection technical field, first by the ratio of Compton and Rayleigh scattering intensity calculated by ray fluorescence energy spectrum, obtain effective atomic number and its uncertainty;Again, the continuum of short-wave infrared spectrum is removed, the depth of aluminum hydroxyl absorption band and band shape asymmetry are extracted, and the slope is used to represent the granularity, and the monotone spline is trained to obtain the clay proportion estimation;According to the end member set and atomic number table, the theoretical atomic number and uncertainty are obtained, and the measurement side is integrated into a unified scale to obtain the cross-modal inconsistency index.The state variable and the index are used as the target to adaptively optimize the collection and retest, and the consistency constraint is executed to solve the mixture and abnormality determination.The application can carry out high-precision identification and abnormality determination under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of geological sample composition identification and detection, and particularly relates to a geological sample composition identification and anomaly detection system based on machine learning. BACKGROUND

[0002] In the scenes of mine development, exploration sampling, geological scientific research and laboratory rapid screening, short-wave infrared reflectance spectrum and ray fluorescence spectrum need to be obtained at the same measuring point to determine the rock and mineral composition and mark the anomaly point. The depth and shape of the aluminum hydroxyl absorption band in the short-wave infrared reflectance spectrum will change due to the influence of factors such as particle size difference, surface dust or film, water content change, lighting and geometric conditions. The Compton and Rayleigh scattering intensities in the ray fluorescence spectrum will also fluctuate with the change of the average atomic number and matrix effect. This disturbance will be transmitted along the detection chain, so that the two sensing channels give inconsistent results for the same sample, thereby causing deviation in composition identification and anomaly judgment.

[0003] The existing technology mainly adopts single-channel analysis or simple parallel comparison. For example, only the depth threshold of short-wave infrared in the fixed window is used for judgment, or only the intensity ratio of ray fluorescence is used for empirical conversion and manual comparison with the former. Or directly linear unmixing without correcting the overall slope change caused by particle size and coating, lacking self-correction of sampling geometry and equipment drift. More importantly, the existing method generally lacks the step of normalizing the measurement and model errors into the same metrological framework, and cannot consider the measurement uncertainty of the effective atomic number and the uncertainty of the theoretical atomic number derived from the end-member ratio. Therefore, the difference between the results of the two ends lacks a comparable scale, and can only rely on fixed thresholds or empirical rules, making it difficult to adaptively give reliable thresholds with batches and working conditions, and difficult to meet the needs of rapid and reliable judgment on site. SUMMARY

[0004] The purpose of the present application is to solve the problem of lack of unified metrological scale of measurement results of ray fluorescence spectrum channel and short-wave infrared reflectance spectrum channel in the prior art, which is prone to composition determination deviation caused by particle size and surface state interference. A geological sample composition identification and anomaly detection system based on machine learning is proposed.

[0005] To solve the problems existing in the prior art, the present application adopts the following technical scheme:

[0006] A geological sample composition identification and anomaly detection system based on machine learning, comprising:

[0007] A spectrum acquisition module for acquiring ray fluorescence spectrum and short-wave infrared reflectance spectrum of a geological sample;

[0008] An atomic number metrology module is configured to calculate an integral intensity ratio of a Compton scattering peak and a Rayleigh scattering peak based on the ray fluorescence spectrum, and obtain a measured value and a measurement uncertainty according to the integral intensity ratio;

[0009] A band parameter extraction module is configured to remove a continuum from the short-wave infrared reflectance spectrum, to obtain a depth of the hydroxyl aluminum bond absorption band and a band shape asymmetry, and to obtain a particle size proxy according to a slope of the short-wave infrared reflectance spectrum;

[0010] A clay proportion fitting module is configured to perform a monotonic spline regression fitting on the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the particle size proxy, to obtain a clay proportion estimate value;

[0011] An endmember weight calculation module is configured to construct an endmember mixing weight according to the clay proportion estimate value and the particle size proxy, and to obtain a prediction value and a prediction uncertainty through an endmember atomic number table;

[0012] A cross-modal inconsistency module is configured to calculate a synthetic uncertainty based on the measured value and the measurement uncertainty, and to calculate a cross-modal inconsistency index according to the synthetic uncertainty, the prediction value, and the prediction uncertainty;

[0013] An optimization module is configured to design a collection and fusion scheme based on the particle size proxy, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the synthetic uncertainty, and to obtain an optimized cross-modal inconsistency index by taking the cross-modal inconsistency index as an optimization target;

[0014] A component unmixing module is configured to perform component unmixing and anomaly judgment on the short-wave infrared reflectance spectrum based on the optimized cross-modal inconsistency index.

[0015] Preferably, the atomic number metrology module includes the following steps:

[0016] Baseline correction and peak shape fitting are performed on the ray fluorescence spectrum to obtain an integral intensity of the Compton scattering peak and an integral intensity of the Rayleigh scattering peak;

[0017] A monotonic calibration mapping relationship is established according to a standard sample with a known average atomic number;

[0018] The ratio of the integral intensity of the Compton scattering peak to the integral intensity of the Rayleigh scattering peak is converted through the monotonic calibration mapping relationship to obtain a measured value of the effective atomic number;

[0019] Error propagation is performed on the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak according to a covariance of the peak shape fitting to obtain a measurement uncertainty of the effective atomic number.

[0020] Preferably, a continuum removal process is performed on the short-wave infrared reflectance spectrum to obtain the depth and band asymmetry of the hydroxyl aluminum bond absorption band. Based on the slope of the short-wave infrared reflectance spectrum, the particle size distribution is obtained, including:

[0021] Discrete approximations of the first and second derivatives of the reflectance sequence of short-wave infrared reflectance spectra as a function of wavelength are obtained by adjacent differences.

[0022] The local minimum point where the first derivative turns from negative to positive and the second derivative is positive is taken as the candidate point for the center of the hydroxyl-containing aluminum bond absorption band.

[0023] A processing window is generated based on the candidate points at the center of the absorption band containing hydroxyl aluminum bonds;

[0024] Within the processing window, the extreme reflectance points at both ends of the processing window are connected by polynomial fitting to form a continuum baseline;

[0025] Divide the original reflectance within the processing window by the baseline reflectance of the corresponding band to obtain the continuum-removed spectrum;

[0026] The depth of the hydroxyl-containing aluminum bond absorption band is obtained by the difference in reflectance between the continuum baseline and the continuum removed spectrum at the center of the processing window.

[0027] Calculate the integral area from the left boundary to the center of the processing window and the integral area from the right boundary to the center of the window, and obtain the band shape asymmetry of the hydroxyl aluminum bond absorption band based on the ratio of the two areas.

[0028] Within the processing window, the slope of reflectivity with respect to wavelength is calculated using the least squares method, and the slope is used as a granular surrogate quantity.

[0029] Preferably, monotonic spline regression fitting is performed on the depth, asymmetry of the band shape, and particle size surrogate of the hydroxyl aluminum bond absorption band to obtain an estimated value of the clay proportion, including:

[0030] The actual clay ratio of each sample in the standard sample set was determined by experiments, and the short-wave infrared reflectance spectra of the samples were subjected to continuum removal and spectral slope calculation to obtain the depth, band asymmetry and particle size proxy of the hydroxyl aluminum bond absorption band.

[0031] Abnormal standard samples were removed and normalized for the depth, band asymmetry, and particle size proxy of the hydroxyl aluminum bond absorption band of each standard sample in the standard sample set;

[0032] A monotonic spline mapping model is constructed with the actual clay proportion of the standard sample set as the dependent variable, the depth and band asymmetry of the hydroxyl aluminum bond absorbing band in the standard sample set as independent variables, and the particle size surrogate quantity of the standard sample set as the adjustment independent variable. In the model, a monotonicity constraint of non-negative derivative is applied to the depth and band asymmetry of the hydroxyl aluminum bond absorbing band.

[0033] The residual squared error between the estimated clay proportion and the actual clay proportion of the standard sample is minimized as the target training model parameter, and the number of spline nodes and the penalty intensity are selected through cross-validation.

[0034] The model is tested for monotonicity by examining the distribution of the training residuals. If a local violation is found, the node density is automatically adjusted and the model is retrained until the monotonicity constraint is satisfied.

[0035] The depth, band asymmetry, and grain size proxy of the hydroxyl aluminum bond absorption band in the geological sample were normalized and input into the monotonic spline mapping model to obtain the clay proportion estimate.

[0036] Preferably, the endmember mixing weights are constructed based on the estimated clay proportion and the particle size surrogate, and the predicted values ​​and prediction uncertainties are obtained through the endmember atomic number table, including:

[0037] Obtain the endmember set and the atomic number table of the endmembers, wherein the endmember set is fixed to contain: clay, feldspar and quartz;

[0038] A basic weight vector is constructed based on the clay proportion estimate and the endmember set, where the basic weight of the clay endmember is equal to the clay proportion estimate, and the basic weight of the non-clay part is equal to one minus the clay proportion estimate.

[0039] Based on the prior feldspar ratio obtained from historical sample statistics, the non-clay part is divided into feldspar basic weight and quartz basic weight.

[0040] The particle size of the clay base weight and the non-clay base weight are adjusted according to the particle size proxy amount, while the feldspar base weight and the quartz base weight are adjusted with the same amplitude.

[0041] Normalize and combine the particle size-adjusted clay base weights, feldspar base weights, and quartz base weights to obtain the end-member mixed weights.

[0042] The endmember mixing weights are weighted and calculated with the endmember atomic number table to obtain the predicted value of the theoretical atomic number;

[0043] Error propagation is performed based on the variance and covariance of the endmember mixed weights to obtain the prediction uncertainty of the theoretical atomic order.

[0044] Preferably, the combined uncertainty is calculated based on the measured value and the measurement uncertainty, and the cross-modal inconsistency index is calculated based on the combined uncertainty, the predicted value, and the predicted uncertainty, including:

[0045] The combined uncertainty is obtained by adding the square of the effective atomic number measurement uncertainty to the square of the theoretical atomic number prediction uncertainty, and taking the arithmetic square root of the sum.

[0046] The cross-modal inconsistency index is obtained by normalizing the difference between the measured value of the effective atomic number and the predicted value of the theoretical atomic number based on the synthesis uncertainty.

[0047] The current sample is resampled using a bootstrap method to construct an empirical spatial distribution, and the significance score is obtained based on the proportion of cross-modal inconsistency index in the empirical spatial distribution.

[0048] The adaptive quantile threshold is determined based on the significance score of the current batch of samples.

[0049] Preferably, the acquisition and fusion scheme is designed based on particle size surrogate quantity, depth of the hydroxyl-containing aluminum bond absorption band, band shape asymmetry, and synthesis uncertainty, and the cross-modal inconsistency index is used as the optimization objective to obtain the optimized cross-modal inconsistency index, including:

[0050] Determine the range of values ​​for shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the sampling point;

[0051] A state vector is constructed based on the particle size surrogate quantity, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the synthesis uncertainty.

[0052] A design variable vector was constructed based on shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the site.

[0053] The surrogate model is trained based on a historical sample set that includes state vectors, design variable vectors, and a cross-modal inconsistency index.

[0054] Within the range of values ​​for the design variable vector, find the design variable vector that minimizes the output of the surrogate model to obtain the optimal retest configuration;

[0055] Based on the optimal retest configuration, perform shortwave infrared reflectance spectroscopy and X-ray fluorescence spectroscopy retests at the same physical point of the current sample, and recalculate the cross-modal inconsistency index and significance score;

[0056] If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the current state vector, the optimal retest configuration and the updated cross-modal inconsistency index are added to the historical sample set and the iteration continues.

[0057] Preferably, the short-wave infrared reflectance spectrum is demixed and anomaly determined based on the optimized cross-modal inconsistency index, including:

[0058] A prior component vector is constructed based on the estimated clay ratio and the prior feldspar ratio.

[0059] Align the short-wave infrared reflectance spectrum with the endmember spectrum corresponding to the endmember set at the same wavelength sampling points;

[0060] Endmember unmixing is performed with the sum of the reconstruction error between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum and the prior component vector deviation as the target, to obtain the final value of the component vector;

[0061] The unmixed residual sequence is obtained based on the difference between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum.

[0062] The in-band average residual and in-band variance are calculated based on the unmixed residual sequence, and anomaly detection is performed based on the optimized significance score and quantile threshold.

[0063] Compared with the prior art, the beneficial effects of the present invention are:

[0064] 1. This invention quantifies X-ray fluorescence spectroscopy and short-wave infrared reflectance spectroscopy into measured values ​​and measurement uncertainties of effective atomic numbers, as well as the depth, band asymmetry, and particle size surrogate of the absorption band containing hydroxyl aluminum bonds. It then uses monotonic spline regression and endmember hybrid calculation chains to obtain predicted values ​​and prediction uncertainties of theoretical atomic numbers. This effectively eliminates systematic interference from particle size, coating, and equipment drift in the determination process, thereby improving the accuracy and stability of component identification. Furthermore, by uniformly measuring synthesis uncertainty and normalizing the comparison of cross-modal inconsistency indices, the comparability and traceability of results from different batches and under different operating conditions are effectively guaranteed.

[0065] 2. This invention uses particle size surrogate quantity, depth of the hydroxyl aluminum bond absorption band, band shape asymmetry, and synthesis uncertainty as state variables, and uses the cross-modal inconsistency index as the optimization objective to jointly optimize the acquisition time, number of repeated acquisitions, and number of site micro-cleanings. This can effectively reduce the number of retesting rounds and resource consumption, suppress the impact of field fluctuations on the results, and thus improve the consistency and reproducibility across batches. By solving for the optimal retesting configuration within the feasible range and performing closed-loop verification, the rationality and statistical convergence of the time-space configuration are effectively guaranteed.

[0066] 3. This invention, by implementing endmember unmixing with consistency constraints after obtaining the optimized cross-modal inconsistency index and using the unmixing residual for anomaly determination, can effectively locate deviations caused by unmodeled endmembers, grain size anomalies, or surface coatings, eliminating the problem of difficulty in distinguishing based solely on overall error, thereby improving the accuracy and timeliness of anomaly detection. Attached Figure Description

[0067] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0068] Figure 1 A functional block diagram of a machine learning-based geological sample composition identification and anomaly detection system provided in an embodiment of the present invention;

[0069] Figure 2 This is a schematic diagram of the atomic number metrology module provided in an embodiment of the present invention;

[0070] Figure 3 This is a schematic diagram of the parameterized particle size extraction module provided in an embodiment of the present invention;

[0071] Figure 4 This is a schematic diagram of the clay ratio fitting module provided in an embodiment of the present invention;

[0072] Figure 5 This is a schematic diagram of the end-member weight calculation module provided in an embodiment of the present invention;

[0073] Figure 6 This is a schematic diagram of the cross-modal inconsistency module provided in an embodiment of the present invention; Detailed Implementation

[0074] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0075] Example: This example provides a machine learning-based geological sample composition identification and anomaly detection system. See [link to example]. Figure 1 Specifically, including:

[0076] The spectral acquisition module is used to acquire the X-ray fluorescence spectrum and short-wave infrared reflectance spectrum of geological samples;

[0077] In embodiments of the present invention, obtaining the X-ray fluorescence spectrum and short-wave infrared reflectance spectrum of geological samples includes:

[0078] X-ray fluorescence spectra and short-wave infrared reflectance spectra of collected geological samples;

[0079] Specifically, X-ray fluorescence spectroscopy (XRF) involves irradiating a sample with high-energy rays, exciting each element within the sample to emit fluorescent rays with characteristic energies, and recording the distribution of fluorescence intensity as a function of energy. Short-wave infrared reflectance spectroscopy (SIR) involves irradiating a sample in the short-wave infrared band (1-2.5 micrometers) and measuring the change in surface reflectance as a function of wavelength. XRF provides elemental-level evidence for obtaining measured values ​​of effective atomic numbers, while SIR provides mineralogical and structural-level evidence for determining the depth and band asymmetry of hydroxyl-containing aluminum bond absorption bands through continuum removal, and combines spectral slope as a particle size proxy. Comparing these two types of evidence within the same metrological framework allows for the construction of a cross-modal inconsistency index, revealing systematic deviations caused by particle size, water content, or surface coatings, enabling component identification and anomaly detection.

[0080] Specifically, an X-ray fluorescence spectrometer is used to irradiate the geological sample. This spectrometer emits primary X-rays of preset energy that act on the surface of the geological sample, exciting the elements in the sample to produce characteristic X-ray fluorescence. The detector of the X-ray fluorescence spectrometer collects the counting signals of X-ray fluorescence with different energies. The multichannel analyzer built into the instrument resolves and statistically analyzes the counting signals according to energy, generating an X-ray fluorescence energy spectrum including the Compton peak and Rayleigh peak. A short-wave infrared spectrometer is used to irradiate the detection point of the same geological sample. This short-wave infrared spectrometer emits continuous infrared light with a wavelength range of 1.0 micrometer to 2.5 micrometers. After the infrared light is reflected from the surface of the geological sample, the spectrometer's spectral system disperses the reflected light. The detector then converts the light signals of different wavelengths into electrical signals, which are then converted into short-wave infrared reflectance spectra for that detection point after analog-to-digital conversion.

[0081] Atomic number metrology module, see Figure 2 It is used to calculate the integrated intensity ratio of the Compton scattering peak and the Rayleigh scattering peak based on X-ray fluorescence spectrometry, and to obtain the measured value and measurement uncertainty based on the integrated intensity ratio;

[0082] In embodiments of the present invention, the integrated intensity ratio of the Compton scattering peak to the Rayleigh scattering peak is calculated based on X-ray fluorescence spectroscopy, and the measured value and measurement uncertainty are obtained based on the integrated intensity ratio, including:

[0083] Baseline correction and peak shape fitting were performed on the X-ray fluorescence spectrum to obtain the integrated intensity of the Compton scattering peak and the Rayleigh scattering peak.

[0084] Establish a monotonic calibration mapping relationship based on a standard sample with a known average atomic number;

[0085] The ratio of the integrated intensity of the Compton scattering peak to the integrated intensity of the Rayleigh scattering peak is converted by monotonic calibration mapping relationship to obtain the measured value of the effective atomic number.

[0086] The measurement uncertainty of the effective atomic number is obtained by propagating the error between the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak based on the covariance of the peak shape fitting.

[0087] Specifically, the integrated intensity of the Compton scattering peak is the area obtained by integrating the net peak shape function within its energy window after baseline correction and peak shape fitting, representing the inelastic scattering intensity; the integrated intensity of the Rayleigh scattering peak is the area obtained by integrating the net peak shape function within its energy window after baseline correction and peak shape fitting; the standard sample with average atomic number is a set of standard samples with known composition and proportions and which have undergone the same measurement and processing chain as the sample to be tested, and its average atomic number is calculated by weighting the atomic fraction of the elements.

[0088] Specifically, the original energy spectrum contains slowly varying background and peak overlap; directly taking the area will introduce systematic bias. Baseline correction is needed to remove the background, and peak shape fitting is required to separate and quantify the target spectral peaks to obtain measurable integrated intensities. When performing baseline correction and peak shape fitting on the X-ray fluorescence spectrum, a polynomial fitting method is used to fit the background region of the X-ray fluorescence spectrum to obtain a continuous baseline spectrum. This baseline spectrum is then subtracted from the original X-ray fluorescence spectrum to complete the baseline correction operation. For the Compton and Rayleigh scattering peaks, a Lorentz-Gaussian mixture function is used to fit the peak shape. The parameters of the mixture function are optimized using a nonlinear least squares method to achieve the best match between the mixture function and the peak shape region of the baseline-corrected energy spectrum. Integral calculations are then performed on the Compton and Rayleigh scattering peaks respectively to obtain the integrated intensities of the Compton and Rayleigh scattering peaks.

[0089] Specifically, different samples exhibit differences in geometry, density, and matrix effects. Therefore, standard samples consistent with the target conditions are needed to establish a stable correspondence between measurable quantities and compositional levels. Furthermore, monotonicity constraints ensure that signal changes correspond to compositional changes in only one direction. When establishing a monotonic calibration mapping relationship based on standard samples with known average atomic numbers, at least 10 standard samples with different average atomic numbers covering the possible atomic number range of the target geological sample are selected. Each standard sample undergoes the same X-ray fluorescence spectrometry acquisition and peak shape fitting processing as the geological sample to be tested, obtaining the ratio of the integrated intensity of the Compton scattering peak to the integrated intensity of the Rayleigh scattering peak for each standard sample. Using the known average atomic number of the standard sample as the dependent variable and the ratio of the corresponding integrated intensity of the Compton scattering peak to the Rayleigh scattering peak as the independent variable, a monotonic calibration mapping relationship is obtained using monotonic spline interpolation. This mapping relationship must satisfy the characteristic that the average atomic number monotonically increases as the ratio of the integrated intensity of the Compton scattering peak to the Rayleigh scattering peak increases. This establishes a reusable and traceable conversion channel between measurable intensity characteristics and effective atomic numbers, reducing drift caused by changes in environment and operating conditions.

[0090] Specifically, the ratio of the integrated intensity of the Compton scattering peak to that of the Rayleigh scattering peak can partially offset the influence of sample thickness, surface morphology, and overall gain. Inputting this ratio into a monotonic calibration mapping relationship yields the compositional hierarchy. When converting the ratio of the integrated intensity of the Compton scattering peak to that of the Rayleigh scattering peak using the monotonic calibration mapping relationship, the ratio of the integrated intensity of the Compton scattering peak to that of the Rayleigh scattering peak for the geological sample to be tested is calculated. This ratio is then input into a pre-established monotonic calibration mapping relationship. Using interpolation calculations based on the mapping relationship, the measured value of the effective atomic number of the geological sample to be tested is obtained. The output of the measured value of the effective atomic number, which is directly related to the sample composition and comparable across samples, provides a dimensionally unified measurement basis for subsequent consistency comparisons with predicted theoretical atomic numbers and for calculating cross-modal inconsistency indices.

[0091] Specifically, the measured value of the effective atomic number is obtained by transforming two integral intensities using a function. Its uncertainty stems from the uncertainty and correlation of the peak shape fitting parameters. Error propagation correctly transmits covariance information to the result quantity. The covariance matrix of the fitting parameters for the Compton and Rayleigh scattering peaks is obtained during the peak shape fitting process. This covariance matrix reflects the degree of error correlation between the fitting parameters. Based on the analytical expressions of the integral intensities of the Compton and Rayleigh scattering peaks with respect to the fitting parameters, and combined with the error propagation law, the covariance of the fitting parameters is transmitted to the ratio of the integral intensities of the Compton and Rayleigh scattering peaks, yielding the standard uncertainty of this ratio. Based on the derivative information of the monotonic calibration mapping relationship, the standard uncertainty of the ratio is further transmitted to the measured value of the effective atomic number, thus obtaining the measurement uncertainty of the effective atomic number. The measurement uncertainty of the effective atomic number can be subsequently combined with the predicted uncertainty of the theoretical atomic number to form a unified combined uncertainty, used as a scaling benchmark for the cross-modal inconsistency index, ensuring that comparisons and threshold controls are statistically significant.

[0092] See the parameterized particle size extraction module. Figure 3 It is used to remove the continuum from the short-wave infrared reflectance spectrum to obtain the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band. Based on the slope of the short-wave infrared reflectance spectrum, the particle size proxy is obtained.

[0093] In embodiments of the present invention, a continuum removal is performed on the short-wave infrared reflectance spectrum to obtain the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band. Based on the slope of the short-wave infrared reflectance spectrum, a particle size surrogate quantity is obtained, including:

[0094] Discrete approximations of the first and second derivatives of the reflectance sequence of short-wave infrared reflectance spectra as a function of wavelength are obtained by adjacent differences.

[0095] The local minimum point where the first derivative turns from negative to positive and the second derivative is positive is taken as the candidate point for the center of the hydroxyl-containing aluminum bond absorption band.

[0096] A processing window is generated based on the candidate points at the center of the absorption band containing hydroxyl aluminum bonds;

[0097] Specifically, the hydroxyl-containing aluminum bond absorption band refers to a characteristic absorption band in short-wave infrared reflectance spectra caused by the vibration of hydroxyl-containing aluminum bonds, the combination of hydroxyl stretching and aluminum-hydroxyl bending, or overtones, typically appearing around 2.2 micrometers. This absorption band mainly originates from the aluminum-hydroxyl structure in layered hydroxyl-containing aluminum silicate minerals, such as muscovite, illite, and kaolinite. Its central position is determined by the minimum reflectance after continuum removal, the band depth reflects the relative abundance of endmembers related to hydroxyl-containing aluminum bonds, and the asymmetry of the band shape is affected by endmember mixing and grain size. It can be used to distinguish different clay mineral assemblages or surface states.

[0098] Specifically, short-wave infrared reflectance spectra are discretely sampled data. Discrete approximations of the first and second derivatives are first obtained through adjacent differences. This allows the intensity curve to be converted into extrema and concavity / convexity information, providing a criterion for subsequent precise location of absorption band minima. For the discrete sequence of reflectance variation with wavelength in short-wave infrared reflectance spectra, the discrete approximation of the first derivative is calculated using the first-order forward difference formula, and the discrete approximation of the second derivative is calculated using the second-order forward difference formula. This yields the discrete approximation sequences of the first and second derivatives of the reflectance variation sequence with wavelength.

[0099] Specifically, the hydroxyl-containing aluminum bond absorption band exhibits extremely low local reflectivity near its center. Mathematically, its first derivative changes from negative to positive around this point, while the second derivative remains positive, concave upwards. Limiting the search to the instrument's usable range avoids edge noise and interference from invalid bands. Traversing all wavelengths within the instrument's usable shortwave infrared range, for each wavelength point, it is determined whether the discrete approximation of its corresponding first derivative satisfies the following conditions: the discrete approximation of the first derivative at the wavelength adjacent to its left is negative, and the discrete approximation of the first derivative at the wavelength adjacent to its right is positive. Simultaneously, it is determined whether the discrete approximation of the second derivative at this point is positive. If both conditions are met, the wavelength point is identified as a candidate center point for the hydroxyl-containing aluminum bond absorption band.

[0100] Specifically, the calculation of band depth and band shape asymmetry needs to be completed within a local window near the center, including continuum fitting and search of the shoulders on both sides. First, fix the processing window based on the central candidate point to ensure that subsequent parameters are calculated within the same scale and neighborhood, avoiding absorption by nearby water or other band shape disturbances. Searching outwards from both sides of the candidate center of the hydroxyl-containing aluminum bond absorption band, find the nearest position where the second derivative changes from positive to negative. Designate the left position as the left inflection point and the right position as the right inflection point. For each determined candidate center of the hydroxyl-containing aluminum bond absorption band, starting from that candidate center, check the discrete approximation value of the second derivative at each wavelength point in reverse order of decreasing wavelength. When the first discrete approximation value of the second derivative changes from positive to negative (i.e., less than or equal to zero) is encountered, that wavelength point is designated as the left inflection point. Then, starting from that candidate center, check the discrete approximation value of the second derivative at each wavelength point in the direction of increasing wavelength. When the first discrete approximation value of the second derivative changes from positive to negative (i.e., less than or equal to zero) is encountered, that wavelength point is designated as the right inflection point. Using the wavelength at the left inflection point as the left endpoint and the wavelength at the right inflection point as the right endpoint, directly use the wavelength value corresponding to the left inflection point as the left endpoint wavelength of the processing window, and the wavelength value corresponding to the right inflection point as the right endpoint wavelength of the processing window. The wavelength corresponding to the left endpoint is used as the starting wavelength of the window, and the wavelength corresponding to the right endpoint is used as the ending wavelength of the window. Simultaneously, the central candidate point is used as the central reference point for the absorption bands that need to be analyzed within this processing window, thus determining a wavelength range characteristic of hydroxyl-containing aluminum bond absorption bands as the processing window. This makes subsequent continuum removal and band parameter extraction more robust and improves the comparability of the depth and band shape asymmetry of hydroxyl-containing aluminum bond absorption bands.

[0101] Within the processing window, the extreme reflectance points at both ends of the processing window are connected by polynomial fitting to form a continuum baseline;

[0102] Divide the original reflectance within the processing window by the baseline reflectance of the corresponding band to obtain the continuum-removed spectrum;

[0103] The depth of the hydroxyl-containing aluminum bond absorption band is obtained by the difference in reflectance between the continuum baseline and the continuum removed spectrum at the center of the processing window.

[0104] Calculate the integral area from the left boundary to the center of the processing window and the integral area from the right boundary to the center of the window, and obtain the band shape asymmetry of the hydroxyl aluminum bond absorption band based on the ratio of the two areas.

[0105] Within the processing window, the slope of reflectivity with respect to wavelength is calculated using the least squares method, and the slope is used as a granular surrogate quantity.

[0106] Specifically, the continuum baseline is a smooth reference curve obtained by connecting the extreme reflectance points at both ends of the processing window using a polynomial. It is used to characterize the slowly varying background and brightness trends without the influence of absorption bands. The continuum-removed spectrum is a dimensionless spectrum obtained by dividing the original reflectance within the processing window point by point by the continuum baseline. The value is approximately one at the non-absorption points and less than one at the absorption points. The particle size surrogate is used to characterize the influence of particle size and surface roughness on scattering by the overall tilt of the short-wave infrared reflectance spectrum within the processing window.

[0107] Specifically, the absorption band is superimposed on a slowly undulating background and illumination trend. This slow change needs to be characterized by a continuum baseline to ensure that subsequent measurements of band shape and intensity reflect true absorption rather than baseline drift. Selecting the reflectance extrema at both ends of the window as anchor points avoids areas within the band affected by absorption. When fitting the continuum baseline within the processing window, the reflectance extrema corresponding to the left and right endpoints of the window are determined. The reflectance extrema at the left endpoint is the maximum reflectance near the left boundary of the processing window, and the reflectance extrema at the right endpoint is the maximum reflectance near the right boundary of the processing window. A cubic polynomial is used to fit the reflectance variation trend between these two reflectance extrema. The coefficients of this cubic polynomial are solved using the least squares method, ensuring that the fitted curve smoothly connects the two reflectance extrema, thus forming the continuum baseline. This reduces systematic bias caused by baseline rise or undulation.

[0108] Specifically, after obtaining the continuum baseline, the original reflectance value corresponding to each wavelength point within the processing window is divided by the baseline reflectance value corresponding to that wavelength point on the continuum baseline. This process is repeated point by point to obtain the continuum removal spectrum within the entire processing window. Scale the reflectance to the baseline to eliminate the multiplicative brightness and albedo differences within the window, ensuring that absorption is only manifested in a relatively weakened form.

[0109] Specifically, the relative attenuation at the center of the absorption band directly corresponds to the intensity information of the hydroxyl-containing aluminum bond-related endmembers. Measuring the center difference allows for accurate characterization of the absorption intensity. To calculate the depth of the hydroxyl-containing aluminum bond absorption band, first find the baseline reflectance value corresponding to the center of the processing window on the continuum baseline, and simultaneously find the corresponding reflectance value on the continuum removal spectrum. Then, subtract the reflectance value of the continuum removal spectrum in that band from the reflectance value of the continuum baseline in that band. The resulting difference represents the depth of the hydroxyl-containing aluminum bond absorption band, which is insensitive to illumination and the baseline, reflecting the relative abundance of hydroxyl-containing aluminum bond-related endmembers such as clay minerals.

[0110] Specifically, whether the absorption band shape is biased towards short or long wavelengths is often determined by both endmember mixing and grain size scattering. The area ratio can integrate the shape information from both sides and is more noise-resistant than single-point width or half-width at half-maximum. When calculating the band shape asymmetry of the hydroxyl-containing aluminum bond absorption band, the left boundary of the processing window is used as the integration start point and the center position as the integration end point. The trapezoidal integration method is used to integrate the reflectance curve of the continuum-removed spectrum within this interval, obtaining the integrated area from the left boundary to the center. Then, using the center position as the integration start point and the right boundary as the integration end point, the trapezoidal integration method is used again to integrate the reflectance curve of the continuum-removed spectrum within this interval, obtaining the integrated area from the right boundary to the center. The ratio of the integrated area from the left boundary to the center to the integrated area from the right boundary to the center is the band shape asymmetry of the hydroxyl-containing aluminum bond absorption band. This morphological quantity, band shape asymmetry, reflects the endmember combination and scattering effects, enhances the monotonic relationship with the clay ratio, and provides a basis for weight adjustment during endmember mixing.

[0111] Specifically, particle size and surface roughness cause an overall tilt in the spectral lines. A linear slope is used to compress this macroscopic trend into a single measurable quantity, making it easier to use as a regulating variable in subsequent models. When calculating the particle size surrogate, all wavelength points and their corresponding reflectance values ​​from the continuum-removed spectra are selected within the processing window. A linear fitting model of reflectance with respect to wavelength is constructed. The slope of this linear model is solved using the least squares method, minimizing the sum of squared residuals between the model-predicted reflectance and the actual continuum-removed spectrum reflectance. This slope is the reflectance slope with respect to wavelength and is used as the particle size surrogate.

[0112] Clay ratio fitting module, see Figure 4 It is used to perform monotonic spline regression fitting on the depth, band asymmetry and particle size surrogate of the hydroxyl aluminum bond absorption band to obtain the clay ratio estimate.

[0113] In embodiments of the present invention, monotonic spline regression fitting is performed on the depth, band asymmetry, and particle size surrogate of the hydroxyl aluminum bond absorbing band to obtain an estimated value of the clay proportion, including:

[0114] The actual clay ratio of each sample in the standard sample set was determined by experiments, and the short-wave infrared reflectance spectra of the samples were subjected to continuum removal and spectral slope calculation to obtain the depth, band asymmetry and particle size proxy of the hydroxyl aluminum bond absorption band.

[0115] Abnormal standard samples were removed and normalized for the depth, band asymmetry, and particle size proxy of the hydroxyl aluminum bond absorption band of each standard sample in the standard sample set;

[0116] Specifically, the actual clay proportion is the proportion of clay minerals in the total amount of the sample, such as clay end-members related to aluminum hydroxyl bonds, such as kaolinite, illite, and muscovite; the standard set is a set of standard samples with known composition and quantified actual clay proportion.

[0117] Specifically, the actual clay proportion is a necessary supervised objective for training the monotonic spline mapping model. The depth and band asymmetry of the hydroxyl aluminum bond absorption band can directly characterize the intensity and morphology information related to the clay endmembers. The spectral slope, as a particle size surrogate, is used to characterize the influence of scattering and roughness. When experimentally determining the actual clay proportion of each sample in the standard sample set, chemical analysis methods or existing standard mineral quantitative analysis methods are used, such as X-ray diffraction quantitative analysis combined with the Rietveld full-spectrum fitting method, to perform component analysis on each standard sample, accurately determining the mass proportion of clay minerals, such as kaolinite, illite, and muscovite, which contain hydroxyl aluminum bonds, thereby obtaining the actual clay proportion of each sample. A continuum removal operation is performed on the short-wave infrared reflectance spectrum of each sample, that is, according to the aforementioned method, the depth, band asymmetry, and particle size surrogate of the hydroxyl aluminum bond absorption band of each standard sample are calculated.

[0118] Specifically, standard sample collection may involve sporadic noise, labeling errors, or dimensional differences across batches; outliers can skew the model, and inconsistent unnormalized feature scales can lead to training instability. By removing outliers and normalizing to a unified scale, training data can be constrained to a physically reasonable and statistically stable range. When removing outliers from the standard sample set for the depth, band asymmetry, and granularity surrogate quantity of the hydroxyl-containing aluminum bond absorption band, Grubbs' test is used. A significance level is set, and the statistics corresponding to the depth, band asymmetry, and granularity surrogate quantity of the hydroxyl-containing aluminum bond absorption band for each standard sample are calculated. If the statistics exceed the critical value of the Grubbs' test, the standard sample is determined to be an outlier and removed from the standard sample set. After removing the abnormal standard samples, the depth of the hydroxyl-containing aluminum bond absorption band, the band shape asymmetry, and the particle size proxy of the remaining standard samples were normalized. The linear normalization method was used to map the value of each index to the interval between 0 and 1. That is, for each index, the current value of the index was subtracted from its minimum value in the remaining standard samples, and then divided by the difference between the maximum and minimum values ​​of the index in the remaining standard samples to obtain the normalized value.

[0119] A monotonic spline mapping model is constructed with the actual clay proportion of the standard sample set as the dependent variable, the depth and band asymmetry of the hydroxyl aluminum bond absorbing band in the standard sample set as independent variables, and the particle size surrogate quantity of the standard sample set as the adjustment independent variable. In the model, a monotonicity constraint of non-negative derivative is applied to the depth and band asymmetry of the hydroxyl aluminum bond absorbing band.

[0120] The residual squared error between the estimated clay proportion and the actual clay proportion of the standard sample is minimized as the target training model parameter, and the number of spline nodes and the penalty intensity are selected through cross-validation.

[0121] The model is tested for monotonicity by examining the distribution of the training residuals. If a local violation is found, the node density is automatically adjusted and the model is retrained until the monotonicity constraint is satisfied.

[0122] The depth, band asymmetry, and particle size surrogate of the hydroxyl aluminum bond absorption band in the geological sample were normalized and input into the monotonic spline mapping model to obtain the clay ratio estimate.

[0123] Specifically, the actual clay proportion and the depth and band asymmetry of the hydroxyl-containing aluminum bond absorption band exhibit a monotonic direction in mineralogical terms. The particle size surrogate alters the spectral morphology and intensity, requiring explicit inclusion as a moderating independent variable. By using monotonic splines and applying non-negative derivative constraints to the two main independent variables, the mineralogical priors can be solidified into computable directional constraints. When constructing the monotonic spline mapping model, the actual clay proportion of the standard sample set is used as the dependent variable, the depth and band asymmetry of the hydroxyl-containing aluminum bond absorption band in the standard sample set are used as independent variables, and the particle size surrogate in the standard sample set is used as a parameter for moderating the independent variable. A monotonic spline function is then used to establish the mapping relationship. In this model, a monotonicity constraint of non-negative derivatives is applied to the depth and band asymmetry of the hydroxyl-containing aluminum bond absorption band. This requires that the partial derivatives of the model with respect to the depth of the hydroxyl-containing aluminum bond absorption band and the partial derivatives with respect to the band asymmetry be non-negative. This ensures that the estimated clay proportion does not decrease as the depth or band asymmetry of the hydroxyl-containing aluminum bond absorption band increases, consistent with the physical correlation between clay minerals and spectral characteristics. This results in a model that both conforms to mineralogical directionality and adapts to grain size variations, avoiding non-physical inverse relationships and improving robustness across samples and batches.

[0124] Specifically, training with the residuals minimizing the squared error yields the optimal fit to the standard sample data. The number of spline nodes and the penalty strength jointly control the model's complexity and smoothness, and cross-validation can be used to achieve a balance between bias and variance. When training the model parameters, the objective function is to minimize the sum of squared residuals between the estimated clay proportion and the actual clay proportion in the standard sample. The model parameters are solved using an iterative optimization algorithm. Simultaneously, cross-validation is used to select the number of spline nodes and the penalty strength: the standard sample set is divided into multiple subsets, and different subsets are selected sequentially as the validation set, while the remaining subsets are used as the training set. The model is trained and its performance is validated. Based on the validation results, the optimal number of spline nodes and penalty strength for the model's generalization ability are determined, avoiding overfitting or underfitting.

[0125] Specifically, data sparsity or noise may lead to fitting results that violate physical orientations in local regions. By checking the residual distribution and adaptively increasing the local node density, local mismatches can be corrected while maintaining overall complexity. When checking the model's monotonicity constraints, the residual distribution generated during training is analyzed to determine whether the model violates the pre-applied monotonicity constraints. If the derivative of the model with respect to the depth or band shape asymmetry of the hydroxyl-containing aluminum bond absorption band is negative in a local region, indicating a local violation of the monotonicity constraint, the spline node density in that local region is automatically increased to refine the model's fitting accuracy in that region. The model is then retrained until it satisfies the non-negative derivative monotonicity constraints of the depth and band shape asymmetry of the hydroxyl-containing aluminum bond absorption band across the entire domain. This improves the model's consistency and reliability across the entire composition range and prevents training noise from propagating to subsequent calculation chains for endmember weights and theoretical atomic orders.

[0126] Specifically, when obtaining the estimated clay proportion of geological samples, the depth, band asymmetry, and grain size surrogate of the hydroxyl-containing aluminum bond absorption bands in the geological samples are first normalized using the same method as in the standard sample set processing, mapping these index values ​​to the same interval. Then, the normalized indexes are input into a trained monotonic spline mapping model, which calculates and outputs the estimated clay proportion of the geological sample. The estimated clay proportion is comparable across samples, providing a stable and directly usable input for subsequent steps using this estimate.

[0127] For the end-member weight calculation module, see [link / reference] Figure 5 It is used to construct end-member mixing weights based on the clay ratio estimate and particle size surrogate, and to obtain the predicted value and prediction uncertainty through the end-member atomic number table.

[0128] In embodiments of the present invention, endmember mixing weights are constructed based on the estimated clay proportion and particle size surrogate, and the predicted value and prediction uncertainty are obtained through the endmember atomic number table, including:

[0129] Obtain the endmember set and the atomic number table of the endmembers, wherein the endmember set is fixed to contain: clay, feldspar and quartz;

[0130] A basic weight vector is constructed based on the clay proportion estimate and the endmember set, where the basic weight of the clay endmember is equal to the clay proportion estimate, and the basic weight of the non-clay part is equal to one minus the clay proportion estimate.

[0131] Based on the prior feldspar ratio obtained from historical sample statistics, the non-clay part is divided into feldspar basic weight and quartz basic weight.

[0132] Specifically, the endmember set is a group of representative pure mineral components selected for component unmixing and theoretical atomic number calculation. In this embodiment, it is fixed to include three categories: clay, feldspar, and quartz. Each endmember corresponds to its standard short-wave infrared endmember spectrum and its matching representative atomic number. The atomic number table of the endmembers is a table that corresponds one-to-one between each endmember in the endmember set and its representative atomic number. The representative atomic number is derived from the known chemical composition of the endmember or a stable equivalent value measured by X-ray fluorescence.

[0133] Specifically, when obtaining the endmember set and the atomic number table of the endmembers, the endmember set used for geological sample composition analysis is predetermined. This endmember set is fixed to include three typical geological endmembers: clay, feldspar, and quartz. The average atomic number corresponding to each type of endmember is determined through experimental measurement, forming the atomic number table of the endmembers. The average atomic number of the clay endmember is obtained by statistical analysis of the atomic composition of common clay minerals such as kaolinite and montmorillonite. The average atomic number of the feldspar endmember is determined by statistical analysis of the atomic composition of feldspar minerals such as potassium feldspar and plagioclase. The average atomic number of the quartz endmember is calculated from the atomic composition of silicon dioxide.

[0134] Specifically, the estimated clay proportion directly reflects the share of clay endmembers. Using it as the basic weight for clay endmembers ensures a one-to-one correspondence between the proportions and physical quantities. Using its supplementary weight as the weight for the non-clay portion satisfies the non-negative and sum-to-one mass conservation laws, avoiding additional degrees of freedom. When constructing the basic weight vector based on the estimated clay proportion and the endmember set, the estimated clay proportion obtained through a monotonic spline mapping model is extracted. The basic weight of the clay endmembers is directly set to this estimated clay proportion, while the basic weight of the non-clay portion is obtained by subtracting the estimated clay proportion from 1. This constructs a basic weight vector containing the weights of both clay endmembers and the non-clay portion, where the sum of the clay endmember weights and the non-clay portion weights is 1, satisfying the normalization constraint of the component proportions. This provides a stable starting point for subsequent further subdivision and particle size adjustment within the non-clay portion and facilitates the propagation and auditing of uncertainties.

[0135] Specifically, the non-clay component is often dominated by feldspar and quartz. Directly introducing the prior feldspar ratio can maintain identifiability when information is insufficient or the banding is similar, avoiding overfitting; at the same time, it preserves geological rationality and facilitates subsequent synergistic adjustment under the influence of grain size. When dividing the non-clay component into feldspar and quartz base weights based on the prior feldspar ratio obtained from historical sample statistics, a large amount of analytical data from historical geological samples is collected. The composition of these historical samples is analyzed, and the proportion distribution of feldspar in the non-clay component is statistically analyzed to obtain the prior feldspar ratio from historical samples. Based on the weight of the non-clay component of the current geological sample, this non-clay component weight is multiplied by the prior feldspar ratio from historical samples to obtain the feldspar base weight; then, the non-clay component weight is multiplied by (1 minus the prior feldspar ratio from historical samples) to obtain the quartz base weight, thus completing the division of the non-clay component into feldspar and quartz base weights. Stably allocating the non-clay component into feldspar and quartz components improves the robustness and consistency of subsequent end-member mixing weight construction and theoretical atomic number calculation.

[0136] The particle size of the clay base weight and the non-clay base weight are adjusted according to the particle size proxy amount, while the feldspar base weight and the quartz base weight are adjusted with the same amplitude.

[0137] Normalize and combine the particle size-adjusted clay base weights, feldspar base weights, and quartz base weights to obtain the end-member mixed weights.

[0138] The endmember mixing weights are weighted and calculated with the endmember atomic number table to obtain the predicted value of the theoretical atomic number;

[0139] Error propagation is performed based on the variance and covariance of the endmember mixed weights to obtain the prediction uncertainty of the theoretical atomic order;

[0140] Specifically, the predicted theoretical atomic number is the average atomic number obtained by multiplying the endmember mixing weights by the atomic number tables of the endmembers one by one, and summing them up, under the premise that the endmember set is fixed as clay, feldspar and quartz. It represents the elemental hierarchy implied by the short-wave infrared endmember information under the current grain size conditions. The prediction uncertainty of the theoretical atomic number is the uncertainty corresponding to the predicted value of the theoretical atomic number, reflecting the sensitivity of the prediction to the fluctuation of the endmember mixing weights and the influence of the correlation between the weights.

[0141] Specifically, particle size and surface roughness alter the relative contributions of short-wave infrared scattering and absorption. Without correction, particle effects can be mistakenly interpreted as changes in end-member proportions. Adjusting the clay and non-clay base weights separately can compensate for the sensitivity of the two types of end-members to particle size. Using a common amplitude for the feldspar and quartz base weights maintains the relative proportions within the non-clay range without artificial distortion. When adjusting the clay and non-clay base weights based on particle size proxy, geological standards with different known particle sizes are selected. The particle size proxy and the corresponding adjustment coefficients for the clay, feldspar, and quartz base weights are measured for each standard. A model establishing the correspondence between the particle size proxy and the weight adjustment coefficients is then developed through statistical analysis or fitting methods. For clay-based weights, the grain size proxy of the current geological sample is input into the model to obtain an adjustment coefficient suitable for the clay-based weights. This adjustment coefficient is then used to scale the clay-based weights, achieving grain size adjustment. For non-clay-based weights, the grain size proxy is similarly input into the correspondence model to obtain an adjustment coefficient. Since feldspar and quartz-based weights require grain size adjustment with the same amplitude, this adjustment coefficient is used to scale both feldspar and quartz-based weights simultaneously, completing the grain size adjustment for non-clay-based weights. This approach reduces the system bias caused by grain size without disrupting mineralogical relationships, making subsequent end-member mixing weights more sensitive to changes in actual composition and more stable to physical disturbances.

[0142] Specifically, the endmember proportions must satisfy the mass conservation laws of non-negativity and a sum of 1. Particle size adjustment alters the magnitudes of each component, requiring normalization constraints to restore them to a physically interpretable weight space. When normalizing the particle size-adjusted clay, feldspar, and quartz base weights, the sum of their weights is first calculated. Then, each of the clay, feldspar, and quartz base weights is divided by this sum to ensure the sum of the three weights equals 1, satisfying the normalization requirements for component proportions. After normalization, the three weights are combined into a mixed endmember weight vector containing the respective weights of the clay, feldspar, and quartz endmembers. This provides standardized input for calculating the predicted theoretical atomic number and its uncertainty.

[0143] Specifically, end-member mixing weights characterize the stoichiometry, while the end-member atomic number table provides the representative atomic number for each end-member. A weighted sum of these two values ​​maps the mineral stoichiometry to elemental levels. When calculating the predicted theoretical atomic number by weighting the end-member mixing weights with the end-member atomic number table, the clay, feldspar, and quartz base weights (after grain size adjustment and normalization constraints) are extracted as the end-member mixing weights. The average atomic numbers of the clay, feldspar, and quartz end-members are obtained from the end-member atomic number table. The predicted theoretical atomic number is obtained by weighted summation of the clay base weight multiplied by the average atomic number of the clay end-members, the feldspar base weight multiplied by the average atomic number of the feldspar end-members, and the quartz base weight multiplied by the average atomic number of the quartz end-members. This lays the foundation for subsequent calculations of synthesis uncertainty and the measurement of the cross-modal inconsistency index.

[0144] Specifically, endmember mixed weights are derived from learning and estimation, and possess variance and covariance. Without downpropagation, the uncertainty of theoretical endmember quantities will be underestimated. Error propagation correctly transmits the uncertainty of the weights and their correlations to the theoretical atomic number level. When using error propagation based on the variance and covariance of the endmember mixed weights to obtain the predicted uncertainty of the theoretical atomic number, the variances of the clay, feldspar, and quartz endmember weights and their mutual covariances are determined, constructing variance and covariance matrices. According to the error propagation law, the variance and covariance of the endmember mixed weights are transmitted to the predicted theoretical atomic number, and the variance of the predicted theoretical atomic number is calculated. The square root of this variance is the predicted uncertainty of the theoretical atomic number. This can be synthesized into a combined uncertainty within a unified econometric framework, serving as a benchmark for cross-modal comparisons and significance assessments, improving the reliability of anomaly detection and optimization decisions.

[0145] Cross-modal inconsistency module, see Figure 6 It is used to calculate the combined uncertainty based on the measured value and the measurement uncertainty, and to calculate the cross-modal inconsistency index based on the combined uncertainty, the predicted value and the predicted uncertainty;

[0146] In embodiments of the present invention, a combined uncertainty is calculated based on measured values ​​and measurement uncertainties, and a cross-modal inconsistency index is calculated based on the combined uncertainty, predicted values, and predicted uncertainties, including:

[0147] The combined uncertainty is obtained by adding the square of the effective atomic number measurement uncertainty to the square of the theoretical atomic number prediction uncertainty, and taking the arithmetic square root of the sum.

[0148] The cross-modal inconsistency index is obtained by normalizing the difference between the measured value of the effective atomic number and the predicted value of the theoretical atomic number based on the synthesis uncertainty.

[0149] Specifically, the combined uncertainty is the total uncertainty obtained by combining the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number under the same metrological framework. It serves as a scale for subsequent comparison of the differences between measured and predicted quantities, ensuring that noise from different sources is fairly included and guaranteeing comparability and statistical validity within and across batches. The cross-modal inconsistency index is a dimensionless number that measures the relative deviation between the measured value of the effective atomic number on the X-ray fluorescence side and the predicted value of the theoretical atomic number on the shortwave infrared side. It reflects the significance of the deviation at the current uncertainty level and is a core objective quantity for design optimization and anomaly detection. The larger the value, the more inconsistent the information between the two modes, and the more necessary it is to retest, optimize, or conduct anomaly checks.

[0150] Specifically, the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number originate from the measurement chain and the model chain, respectively. Both contribute independently to the variance in terms of measurement. By combining these uncertainties according to the rule of summing the variances and then taking the square root, the uncertainties can be combined under a unified dimension, ensuring that the result is non-negative and accurately reflects the overall fluctuation level. When calculating the combined uncertainty, the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number are obtained; the measurement uncertainty of the effective atomic number is squared, and the prediction uncertainty of the theoretical atomic number is also squared; the results of these two squares are added together to obtain the sum of squares; then the arithmetic square root of this sum is taken to obtain the combined uncertainty.

[0151] Specifically, directly comparing the original differences between the two is affected by dimensions and heteroscedasticity, making cross-sample comparability difficult. Normalizing with the combined uncertainty is equivalent to calculating the standardized deviation, placing samples at different uncertainty levels on the same evaluation scale. When calculating the cross-modal inconsistency index, the measured value of the effective atomic number and the predicted value of the theoretical atomic number are obtained, and the difference between the two is calculated. This difference is then divided by the combined uncertainty. Through this normalization process, the cross-modal inconsistency index is obtained.

[0152] The current sample is resampled using a bootstrap method to construct an empirical spatial distribution, and the significance score is obtained based on the proportion of cross-modal inconsistency index in the empirical spatial distribution.

[0153] The adaptive quantile threshold is determined based on the significance score of the current batch of samples;

[0154] Specifically, the significance score is a probabilistic measure of whether the cross-modal inconsistency index is abnormally large for the current sample. The larger the value, the more significant it is. It reflects the probability that the observation deviation is caused by random fluctuations under the current uncertainty and operating conditions, and serves as a direct basis for subsequent optimization and judgment. The adaptive quantile threshold is a judgment threshold adaptively determined based on the significance score distribution of the current batch of samples. It is used to distinguish between samples with normal fluctuations and those that need to be retested or checked for abnormalities.

[0155] Specifically, under conditions of small field samples, heteroscedasticity, and non-Gaussian noise, the non-anomaly distribution of the cross-modal inconsistency index cannot be pre-defined. By maintaining the same site pairing structure and uncertainty, bootstrap resampling can yield the empirical spatial distribution of the current operating conditions without introducing parametric assumptions. The exceedance ratio measures the probabilistic degree of the observed values, forming a comparable significance score. When constructing the empirical spatial distribution for the current sample through bootstrap resampling, a subsample of the same size as the original dataset is repeatedly drawn with replacement from the original dataset. This drawing operation is repeated multiple times, and the cross-modal inconsistency index of the corresponding subsample is calculated after each drawing. All cross-modal inconsistency indices obtained from multiple drawings are summarized to construct the empirical spatial distribution of the cross-modal inconsistency index. The proportion by which the cross-modal inconsistency index actually calculated for the current sample exceeds this empirical spatial distribution is the significance score, reflecting the significance of the cross-modal inconsistency of the current sample. The significance score is entirely data-driven, adaptable to different instruments and batch noise structures, and can be directly compared across samples, providing a robust basis for subsequent optimization and judgment.

[0156] Specifically, different batches exhibit systematic differences in environment and collection conditions. Therefore, a threshold needs to be set based on the empirical distribution of significance scores for this batch, allowing the boundary between normal fluctuations and the need for retesting or anomaly verification to self-calibrate with the data. Collect the significance scores of all samples in the current batch to construct an empirical distribution of significance scores. Then, based on a preset quantile level, such as selecting the 95th quantile level which reflects the strictness of anomaly judgment, find the corresponding quantile level value in the empirical distribution of significance scores and determine this value as the adaptive quantile threshold for the current batch. This improves cross-batch comparability and judgment stability.

[0157] The optimization module is used to design the acquisition and fusion scheme based on the particle size proxy, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the synthesis uncertainty, and takes the cross-modal inconsistency index as the optimization objective to obtain the optimized cross-modal inconsistency index;

[0158] In embodiments of the present invention, a data acquisition and fusion scheme is designed based on particle size surrogate quantity, depth of the hydroxyl aluminum bond absorption band, band shape asymmetry, and synthesis uncertainty. The cross-modal inconsistency index is used as the optimization objective to obtain an optimized cross-modal inconsistency index, including:

[0159] Determine the range of values ​​for shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the sampling point;

[0160] A state vector is constructed based on the particle size surrogate quantity, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the synthesis uncertainty.

[0161] A design variable vector was constructed based on shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the site.

[0162] The surrogate model is trained based on a historical sample set that includes state vectors, design variable vectors, and a cross-modal inconsistency index.

[0163] Specifically, the number of micro-cleaning operations is the cumulative number of lightweight surface treatments performed at the same measurement point to remove loose dust, thin oxide films, or light surface coatings.

[0164] Specifically, the three controllable actions are limited to the feasible range allowed by the equipment and operation window to avoid unworkable solutions such as timeouts, overdoses, or excessive cleaning, while providing clear boundary conditions for subsequent optimization. When determining the value ranges for short-wave infrared acquisition time, the number of repeated X-ray fluorescence acquisitions, and the number of site micro-cleanings, for short-wave infrared acquisition time, considering the detector response characteristics of the short-wave infrared spectrometer and the on-site detection efficiency requirements, the signal-to-noise ratio of the spectrum at different acquisition times was tested through pre-experiments. The shortest acquisition time that satisfies the extraction of the hydroxyl-containing aluminum bond absorption band feature was set as the lower limit, and the longest allowable acquisition time by the equipment was set as the upper limit. For the number of repeated X-ray fluorescence acquisitions, based on the counting and statistical characteristics of X-ray fluorescence energy spectra, the coefficient of variation of the integral intensity of the Compton peak and Rayleigh peak under different repetitions of the same sample was analyzed through pre-experiments. The minimum number of acquisitions to ensure effective atomic number measurement accuracy was set as the lower limit, and the maximum number of acquisitions acceptable on-site was set as the upper limit. For the number of site micro-cleanings, based on the adhesion of interfering substances on the surface of geological samples, the stability of the spectrum and energy spectrum under different cleaning times was tested through pre-experiments. The minimum number of acquisitions that makes the feature tend to stabilize was set as the lower limit, and the maximum number of acquisitions before cleaning damages the sample was set as the upper limit. This determined the value ranges for each of the three values. This ensures that the optimization solution only returns the feasible optimal retest configuration, reducing invalid searches and on-site trial and error costs, and improving the reliability and security of online decision-making.

[0165] Specifically, when constructing the state vector based on the particle size surrogate quantity, the depth of the hydroxyl-containing aluminum bond absorption band, the band shape asymmetry, and the combined uncertainty, the particle size surrogate quantity obtained by the least squares slope of reflectance with respect to wavelength within the processing window, the depth of the hydroxyl-containing aluminum bond absorption band obtained by the continuum baseline and the continuum removal spectrum, and the band shape asymmetry are extracted from the short-wave infrared reflectance spectroscopy processing results. The combined uncertainty is the combined value of the effective atomic number measurement uncertainty and the theoretical atomic number prediction uncertainty. These four parameters are combined in sequence to form a state vector characterizing the multimodal detection state features of the sample.

[0166] Specifically, when constructing the design variable vector based on the shortwave infrared acquisition time, the number of repeated X-ray fluorescence acquisitions, and the number of site micro-cleanings, the adjustable operational parameters of the three acquisition and preprocessing stages—shortwave infrared acquisition time, the number of repeated X-ray fluorescence acquisitions, and the number of site micro-cleanings—are combined in sequence to form the design variable vector, which is used to model the impact of operational parameters on cross-modal inconsistencies.

[0167] Specifically, real-world evaluation is costly and time-consuming from data collection to processing and comparison. A data-driven surrogate model can approximate the expected value of the cross-modal inconsistency index from states and actions, and suppress overfitting through cross-validation and error assessment. When training the surrogate model based on a historical sample set containing state vectors, design variable vectors, and the cross-modal inconsistency index, multiple sets of historical sample data are collected. Each set contains the state vector of the corresponding sample, the design variable vector used, and the calculated cross-modal inconsistency index. After forming a training dataset from these data, a machine learning model such as Gaussian process regression or random forest is selected as the surrogate model. Using the state vector and design variable vector as input and the cross-modal inconsistency index as the prediction target, the model parameters are optimized by minimizing the mean squared error between the predicted and actual values. This allows the trained surrogate model to predict the cross-modal inconsistency index based on the input state and design parameters. This enables fast and low-cost online evaluation and optimization. Furthermore, the model's uncertainty can improve the stability and repeatability of decisions and promote cross-batch transfer.

[0168] Within the range of values ​​for the design variable vector, find the design variable vector that minimizes the output of the surrogate model to obtain the optimal retest configuration;

[0169] Based on the optimal retest configuration, perform shortwave infrared reflectance spectroscopy and X-ray fluorescence spectroscopy retests at the same physical point of the current sample, and recalculate the cross-modal inconsistency index and significance score;

[0170] If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the current state vector, the optimal retest configuration and the updated cross-modal inconsistency index are added to the historical sample set and the iteration continues.

[0171] Specifically, the surrogate model takes the state vector and design variables as input and outputs the expected value of the cross-modal inconsistency index. Minimizing this output within the feasible range allows for prioritizing the combination of operations that is expected to maximally reduce the cross-modal inconsistency index under resource constraints. Within the range of design variable vector values, to find the design variable vector that minimizes the surrogate model output to obtain the optimal retest configuration, numerical optimization algorithms such as sequential quadratic programming or genetic algorithms are used to search the design variable vector within the predetermined range of shortwave infrared acquisition time, X-ray fluorescence repeated acquisition number, and site micro-cleaning number. The surrogate model predicts the cross-modal inconsistency index based on the input design variable vector. The optimization algorithm continuously adjusts the values ​​of the design variable vector to find the design variable combination that minimizes the predicted cross-modal inconsistency index output by the surrogate model; this combination is the optimal retest configuration.

[0172] Specifically, retesting at the same physical point can eliminate interference from spatial heterogeneity, ensuring that changes mainly stem from configuration adjustments rather than sampling differences. After retesting, the cross-modal inconsistency index and significance score must be recalculated using the same metrology chain to obtain the true effect of the action. When performing short-wave infrared reflectance and X-ray fluorescence spectra retests at the same physical point of the current sample according to the optimal retest configuration, and recalculating the cross-modal inconsistency index and significance score, the spectrometer's acquisition duration is set according to the short-wave infrared acquisition time determined in the optimal retest configuration. Multiple X-ray fluorescence energy spectrum acquisitions are performed according to the number of X-ray fluorescence repeated acquisitions, and micro-cleaning operations are performed on the sample detection points according to the number of point micro-cleaning operations. For the short-wave infrared reflectance and X-ray fluorescence energy spectra obtained from the retests, the effective atomic number measured value, the theoretical atomic number predicted value, and the combined uncertainty of the two are recalculated according to the previous procedure to obtain the updated cross-modal inconsistency index. Based on the updated cross-modal inconsistency index, the significance score of the sample is recalculated according to the calculation method for the significance score.

[0173] Specifically, the adaptive quantile threshold represents the baseline of the current batch of data. If the significance score is not greater than this threshold, it indicates that the deviation is within the normal fluctuation range and the iteration can stop. If it does not meet the threshold, the triplet of state, action, and result for this round is added to the historical sample set, allowing subsequent solutions to continue optimization on data that is closer to the current working conditions. When determining the termination of the iteration, the updated significance score is compared with the adaptive quantile threshold. If the updated significance score is not greater than the adaptive quantile threshold, it means that the consistency of cross-modal data has met the requirements. At this time, the iteration process is terminated, and the current optimal retest configuration and the updated cross-modal inconsistency index are output. If the updated significance score is greater than the adaptive quantile threshold, it means that cross-modal consistency still needs to be optimized. At this time, the state vector of this round, including granularity surrogate quantity, depth of hydroxyl aluminum bond absorption band, band shape asymmetry and synthesis uncertainty, optimal retest configuration, and updated cross-modal inconsistency index, is added to the historical sample set. Then, the process returns to the step of solving for the optimal retest configuration within the range of the design variable vector values, and iterative optimization continues until the termination condition is met.

[0174] The component unmixing module is used to perform component unmixing and anomaly detection on short-wave infrared reflectance spectra based on the optimized cross-modal inconsistency index.

[0175] In an embodiment of the present invention, component unmixing and anomaly determination of short-wave infrared reflectance spectra based on an optimized cross-modal inconsistency index include:

[0176] A prior component vector is constructed based on the estimated clay ratio and the prior feldspar ratio.

[0177] Align the short-wave infrared reflectance spectrum with the endmember spectrum corresponding to the endmember set at the same wavelength sampling points;

[0178] Endmember unmixing is performed with the sum of the reconstruction error between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum and the prior component vector deviation as the target, to obtain the final value of the component vector;

[0179] Specifically, the clay proportion estimate directly provides the physical share of clay endmembers. The non-clay portion is easily indistinguishable when information is insufficient. Feldspar and quartz have similar band shapes and significant noise; therefore, the non-clay portion can be further divided into feldspar and quartz using the prior feldspar proportion. The clay proportion estimate is used as the proportion of clay endmembers in the component vector, and the prior feldspar proportion is used as the proportion of feldspar endmembers. Then, by subtracting the sum of the clay proportion estimate and the prior feldspar proportion from 1, the proportion of quartz endmembers in the component vector is obtained. Following the order of clay, feldspar, and quartz, these three proportions are combined into a prior component vector with an element-wise sum of 1, serving as the prior constraint for subsequent endmember unmixing. This improves the identifiability and stability of the unmixing process, while reducing the degrees of freedom and the risk of overfitting.

[0180] Specifically, endmember unmixing requires the observation vector and the endmember matrix to be at the same sampling grid. If the sampling grids are inconsistent, interpolation errors and displacement errors will be incorrectly attributed to component differences. When aligning the shortwave infrared reflectance spectrum with the endmember spectra corresponding to the endmember set at the same wavelength sampling points, first determine the wavelength sampling range and resolution of the current geological sample's shortwave infrared reflectance spectrum, and then extract the standard spectra of clay, feldspar, and quartz endmembers in the endmember set. Perform interpolation or resampling operations on the endmember spectra to ensure that the wavelength sampling points of the endmember spectra completely match the wavelength sampling points of the current geological sample's shortwave infrared reflectance spectrum. This ensures that each sampling point in the wavelength dimension corresponds one-to-one, eliminating artifacts introduced by sampling inconsistencies, ensuring that the reconstruction error only reflects component deviation rather than numerical alignment issues, and improving unmixing accuracy.

[0181] Specifically, the reconstruction error term ensures a good fit to the observed data, while the prior bias term introduces the prior component vector formed in the previous step as a soft constraint, providing directional guidance when there is noise, granular scattering, or similar band shapes. When performing endmember demixing with the sum of the reconstruction error between the weighted sum of short-wave infrared reflectance spectra and endmember spectra, and the prior component vector deviation, the objective function is first constructed as follows: The first part is the reconstruction error of the weighted sum of short-wave infrared reflectance spectra and endmember spectra, calculated using Euclidean distance. That is, for each wavelength sampling point, the difference between the actual reflectance spectral value and the predicted value of the weighted sum of each endmember spectrum according to the component proportion is calculated, and then the sum of the squares of the differences of all sampling points is obtained. The second part is the prior component vector deviation, calculated using Euclidean distance with the sum of the squares of the differences between the current component vector to be solved and the prior component vector. A regularization coefficient is introduced, which can be adaptively adjusted according to the magnitude of the cross-modal inconsistency index. The larger the cross-modal inconsistency index, the higher the weight of the prior component vector. The two parts are combined into the overall objective function in the form of reconstruction error plus regularization coefficient multiplied by prior component vector deviation. An optimization algorithm constrained by the non-negativity of each element of the component vector and a sum of 1, such as the constrained least squares method, is used to find the component vector that minimizes the overall objective function. This component vector is the final value of the component vector.

[0182] The unmixed residual sequence is obtained based on the difference between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum.

[0183] The in-band average residual and in-band variance are calculated based on the unmixed residual sequence, and anomaly labels are obtained by using the optimized significance score and quantile threshold as thresholds to determine anomalies.

[0184] Specifically, the unmixed residual sequence is a point-by-point difference sequence formed by subtracting the reconstructed reflectance obtained by weighting the final values ​​of the endmember spectra according to the component vectors from the observed reflectance of the shortwave infrared reflectance spectrum at the same wavelength sampling point. This sequence corresponds one-to-one with the wavelength sequence, and the values ​​represent the positive and negative deviations on the reflectance scale, reflecting the degree of wavelength-by-wavelength matching of the model to the observed spectrum.

[0185] Specifically, the reconstructed spectrum obtained by weighting and summing the endmember spectra according to the final values ​​of the component vectors represents the expected response of the model under the endmember ratio. Subtracting the observed shortwave infrared reflectance spectrum from this reconstructed spectrum wavelength by wavelength transforms underfitting and unmodeled effects into a measurable error sequence. When obtaining the unmixing residual sequence based on the difference between the shortwave infrared reflectance spectrum and the weighted sum of the endmember spectra, the final values ​​of the component vectors obtained after endmember unmixing are acquired. Then, the spectrum of each endmember is weighted and summed according to the corresponding proportion in the component vector to obtain the weighted sum of the endmember spectra. The reflectance value of the shortwave infrared reflectance spectrum at each wavelength sampling point is subtracted from the reflectance value of the weighted sum of the endmember spectra at the corresponding wavelength sampling point. Arranging the differences of all wavelength sampling points in sequence forms the unmixing residual sequence. Decomposing the unmixing quality from the overall error into wavelength-by-wavelength deviations facilitates locating the source of the problem near the target absorption band, providing a directly traceable data foundation for subsequent in-band statistics and anomaly detection.

[0186] Specifically, single-point residuals are susceptible to noise. Intra-band average residuals are used to characterize system deviations, and intra-band variance is used to characterize structural fluctuations. This can stably reflect the mechanistic anomalies of the target absorption band. By combining the statistics with the optimized significance score and adaptive quantile threshold, the judgment can be automatically calibrated according to the uncertainty level of the current batch. After generating the unmixed residual sequence, the wavelength set of the target absorption band is determined according to the aforementioned processing window. This set is defined as the in-band region. Within the in-band region, the arithmetic mean of the unmixed residual sequence is calculated point by point to obtain the in-band average residual. Simultaneously, the squared mean of the in-band residuals is calculated with the in-band average residual as the center to obtain the in-band variance. To avoid interference from differences in brightness and sampling density among different samples, the residuals are first normalized using the continuum baseline before calculating the above statistics, so that the expectation of the non-absorption region is close to zero, thus ensuring that the statistics can be compared across samples. The in-band average residual and the in-band variance are used as particulate residual measures. The optimized significance score and adaptive quantile threshold are simultaneously retrieved as the judgment threshold. When the optimized significance score is greater than the adaptive quantile threshold and the absolute value of the in-band average residual or the in-band variance is within the abnormal range of the empirical distribution of this batch, the current sample is judged to have a cross-modal inconsistency anomaly, and an anomaly label is output. When the above conditions are not met, a normal label is output. Under a unified metrological framework, a judgment with both statistical basis and physical orientation is achieved, reducing over-detection or under-detection caused by fixed thresholds.

[0187] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A machine learning-based system for geological sample composition identification and anomaly detection, characterized in that, include: The spectral acquisition module is used to acquire the X-ray fluorescence spectrum and short-wave infrared reflectance spectrum of geological samples; The atomic number metrology module is used to calculate the integrated intensity ratio of the Compton scattering peak and the Rayleigh scattering peak based on X-ray fluorescence spectrometry, and to obtain the measured value and measurement uncertainty based on the integrated intensity ratio; The parameter-based particle size extraction module is used to perform continuum removal on the short-wave infrared reflectance spectrum to obtain the depth and band asymmetry of the hydroxyl aluminum bond absorption band. Based on the slope of the short-wave infrared reflectance spectrum, the particle size distribution is obtained. The clay proportion fitting module is used to perform monotonic spline regression fitting on the depth, band asymmetry and particle size surrogate of the hydroxyl aluminum bond absorption band to obtain the clay proportion estimate. The end-member weight calculation module is used to construct end-member mixing weights based on the estimated clay ratio and particle size surrogate, and to obtain the predicted value and prediction uncertainty through the end-member atomic number table. The cross-modal inconsistency module is used to calculate the combined uncertainty based on the measured values ​​and measurement uncertainties, and to calculate the cross-modal inconsistency index based on the combined uncertainty, the predicted value, and the predicted uncertainty. The optimization module is used to design the acquisition and fusion scheme based on the particle size proxy, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the synthesis uncertainty, and takes the cross-modal inconsistency index as the optimization objective to obtain the optimized cross-modal inconsistency index. The component unmixing module is used to perform component unmixing and anomaly detection on short-wave infrared reflectance spectra based on the optimized cross-modal inconsistency index. Continuum removal was performed on the short-wave infrared reflectance spectrum to obtain the depth and band shape asymmetry of the hydroxyl-containing aluminum bond absorption band. Based on the slope of the short-wave infrared reflectance spectrum, the particle size distribution was obtained, including: Discrete approximations of the first and second derivatives of the reflectance sequence of short-wave infrared reflectance spectra as a function of wavelength are obtained by adjacent differences. The local minimum point where the first derivative turns from negative to positive and the second derivative is positive is taken as the candidate point for the center of the hydroxyl-containing aluminum bond absorption band. A processing window is generated based on the candidate points at the center of the absorption band containing hydroxyl aluminum bonds; Within the processing window, the extreme reflectance points at both ends of the processing window are connected by polynomial fitting to form a continuum baseline; Divide the original reflectance within the processing window by the baseline reflectance of the corresponding band to obtain the continuum-removed spectrum; The depth of the hydroxyl-containing aluminum bond absorption band is obtained by the difference in reflectance between the continuum baseline and the continuum removed spectrum at the center of the processing window. Calculate the integral area from the left boundary to the center of the processing window and the integral area from the right boundary to the center of the window, and obtain the band shape asymmetry of the hydroxyl aluminum bond absorption band based on the ratio of the two areas. Within the processing window, the slope of reflectivity with respect to wavelength is calculated using the least squares method, and the slope is used as a granular surrogate quantity. Monotonic spline regression was performed on the depth, band asymmetry, and particle size surrogate of the hydroxyl-containing aluminum bond absorption band to obtain estimated values ​​of the clay proportion, including: The actual clay ratio of each sample in the standard sample set was determined by experiments, and the short-wave infrared reflectance spectra of the samples were subjected to continuum removal and spectral slope calculation to obtain the depth, band asymmetry and particle size proxy of the hydroxyl aluminum bond absorption band. Abnormal standard samples were removed and normalized for the depth, band asymmetry, and particle size proxy of the hydroxyl aluminum bond absorption band of each standard sample in the standard sample set; A monotonic spline mapping model is constructed with the actual clay proportion of the standard sample set as the dependent variable, the depth and band asymmetry of the hydroxyl aluminum bond absorbing band in the standard sample set as independent variables, and the particle size surrogate quantity of the standard sample set as the adjustment independent variable. In the model, a monotonicity constraint of non-negative derivative is applied to the depth and band asymmetry of the hydroxyl aluminum bond absorbing band. The residual squared error between the estimated clay proportion and the actual clay proportion of the standard sample is minimized as the target training model parameter, and the number of spline nodes and the penalty intensity are selected through cross-validation. The model is tested for monotonicity by examining the distribution of the training residuals. If a local violation is found, the node density is automatically adjusted and the model is retrained until the monotonicity constraint is satisfied. The depth, band asymmetry, and particle size surrogate of the hydroxyl aluminum bond absorption band in the geological sample were normalized and input into the monotonic spline mapping model to obtain the clay ratio estimate. The combined uncertainty is calculated based on the measured values ​​and measurement uncertainties, and the cross-modal inconsistency index is calculated based on the combined uncertainty, predicted values, and predicted uncertainties, including: The combined uncertainty is obtained by adding the square of the effective atomic number measurement uncertainty to the square of the theoretical atomic number prediction uncertainty, and taking the arithmetic square root of the sum. The cross-modal inconsistency index is obtained by normalizing the difference between the measured value of the effective atomic number and the predicted value of the theoretical atomic number based on the synthesis uncertainty. The current sample is resampled using a bootstrap method to construct an empirical spatial distribution, and the significance score is obtained based on the proportion of cross-modal inconsistency index in the empirical spatial distribution. The adaptive quantile threshold is determined based on the significance score of the current batch of samples.

2. The geological sample composition identification and anomaly detection system based on machine learning according to claim 1, characterized in that, The integrated intensity ratio of the Compton scattering peak to the Rayleigh scattering peak was calculated based on X-ray fluorescence spectroscopy, and the measured value and measurement uncertainty were obtained from the integrated intensity ratio, including: Baseline correction and peak shape fitting were performed on the X-ray fluorescence spectrum to obtain the integrated intensity of the Compton scattering peak and the Rayleigh scattering peak. Establish a monotonic calibration mapping relationship based on a standard sample with a known average atomic number; The ratio of the integrated intensity of the Compton scattering peak to the integrated intensity of the Rayleigh scattering peak is converted by monotonic calibration mapping relationship to obtain the measured value of the effective atomic number. The measurement uncertainty of the effective atomic number is obtained by propagating the error between the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak based on the covariance of the peak shape fitting.

3. The geological sample composition identification and anomaly detection system based on machine learning according to claim 1, characterized in that, Endmember mixing weights are constructed based on the estimated clay proportion and particle size surrogate, and the predicted values ​​and prediction uncertainties are obtained through the endmember atomic number table, including: Obtain the endmember set and the atomic number table of the endmembers, wherein the endmember set is fixed to contain: clay, feldspar and quartz; A basic weight vector is constructed based on the clay proportion estimate and the endmember set, where the basic weight of the clay endmember is equal to the clay proportion estimate, and the basic weight of the non-clay part is equal to one minus the clay proportion estimate. Based on the prior feldspar ratio obtained from historical sample statistics, the non-clay part is divided into feldspar basic weight and quartz basic weight. The particle size of the clay base weight and the non-clay base weight are adjusted according to the particle size proxy amount, while the feldspar base weight and the quartz base weight are adjusted with the same amplitude. Normalize and combine the particle size-adjusted clay base weights, feldspar base weights, and quartz base weights to obtain the end-member mixed weights. The endmember mixing weights are weighted and calculated with the endmember atomic number table to obtain the predicted value of the theoretical atomic number; Error propagation is performed based on the variance and covariance of the endmember mixed weights to obtain the prediction uncertainty of the theoretical atomic order.

4. The geological sample composition identification and anomaly detection system based on machine learning according to claim 1, characterized in that, A data acquisition and fusion scheme was designed based on particle size surrogate quantity, depth of the hydroxyl-containing aluminum bond absorption band, band shape asymmetry, and synthesis uncertainty. The cross-modal inconsistency index was used as the optimization objective to obtain the optimized cross-modal inconsistency index, which includes: Determine the range of values ​​for shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the sampling point; A state vector is constructed based on the particle size surrogate quantity, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the synthesis uncertainty. A design variable vector was constructed based on shortwave infrared acquisition time, number of repeated X-ray fluorescence acquisitions, and number of micro-cleaning operations at the site. The surrogate model is trained based on a historical sample set that includes state vectors, design variable vectors, and a cross-modal inconsistency index. Within the range of values ​​for the design variable vector, find the design variable vector that minimizes the output of the surrogate model to obtain the optimal retest configuration; Based on the optimal retest configuration, perform shortwave infrared reflectance spectroscopy and X-ray fluorescence spectroscopy retests at the same physical point of the current sample, and recalculate the cross-modal inconsistency index and significance score; If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the current state vector, the optimal retest configuration and the updated cross-modal inconsistency index are added to the historical sample set and the iteration continues.

5. The geological sample composition identification and anomaly detection system based on machine learning according to claim 4, characterized in that, Compositional demixing and anomaly detection of short-wave infrared reflectance spectra based on the optimized cross-modal inconsistency index include: A prior component vector is constructed based on the estimated clay ratio and the prior feldspar ratio. Align the short-wave infrared reflectance spectrum with the endmember spectrum corresponding to the endmember set at the same wavelength sampling points; Endmember unmixing is performed with the sum of the reconstruction error between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum and the prior component vector deviation as the target, to obtain the final value of the component vector; The unmixed residual sequence is obtained based on the difference between the weighted sum of the short-wave infrared reflectance spectrum and the endmember spectrum. The in-band average residual and in-band variance are calculated based on the unmixed residual sequence, and anomaly detection is performed based on the optimized significance score and quantile threshold.

Citation Information

Patent Citations

  • Tunnel rock mineral identification method and system cooperating with multi-element spectrum

    CN117828397A

  • Cultivated land soil environment detection method and system

    CN119375177A