Geological sample component identification and anomaly detection system based on machine learning
By using a machine learning system to process X-ray fluorescence spectra and short-wave infrared reflectance spectra, the problem of inconsistent measurement results in geological sample composition identification and anomaly detection was solved, achieving more accurate and stable composition identification and anomaly detection.
Patent Information
- Application Number
- CN202511605748.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-05
AI Technical Summary
In existing technologies, the measurement results of X-ray fluorescence spectra and short-wave infrared reflectance spectra of geological samples lack a unified metrological scale, are easily affected by particle size and surface condition, leading to deviations in composition determination and making it difficult to meet the needs of rapid and reliable composition identification and anomaly detection.
A machine learning-based geological sample composition identification and anomaly detection system is adopted. The system acquires X-ray fluorescence energy spectrum and short-wave infrared reflectance spectrum through the spectral acquisition module. Combined with atomic number metrology module, parameterized particle size extraction module, clay ratio fitting module, endmember weight calculation module and cross-mode inconsistency module, the data is processed and optimized. Finally, the composition is determined by the composition demixing module.
It effectively eliminates systematic interference from particle size, coating, and equipment drift in the determination, improves the accuracy and stability of component identification, ensures the consistency and comparability of results, reduces the number of retesting rounds and resource consumption, and improves the accuracy and timeliness of anomaly detection.
Smart Images

Figure CN121068531A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of geological sample composition identification and detection, and particularly relates to a geological sample composition identification and anomaly detection system based on machine learning. BACKGROUND
[0002] In the scenes of mine development, exploration sampling, geological scientific research and laboratory rapid screening, short-wave infrared reflectance spectrum and ray fluorescence spectrum need to be obtained at the same measuring point to determine the rock and mineral composition and mark the anomaly point. The depth and shape of the aluminum hydroxyl absorption band in the short-wave infrared reflectance spectrum will change due to the influence of factors such as particle size difference, surface dust or film, water content change, lighting and geometric conditions. The Compton and Rayleigh scattering intensities in the ray fluorescence spectrum will also fluctuate with the average atomic number and matrix effect changes. This disturbance will be transmitted along the detection chain, so that the two sensing channels give inconsistent results for the same sample, thereby causing deviation in composition identification and anomaly judgment.
[0003] The existing technology mainly adopts single-channel analysis or simple parallel comparison. For example, only the depth threshold of short-wave infrared in the fixed window is used for judgment, or only the intensity ratio of ray fluorescence is used for empirical conversion and manual comparison with the former. Or directly linear unmixing without correcting the overall slope change caused by particle size and coating, lacking self-correction of sampling geometry and equipment drift. More importantly, the existing method generally lacks the step of normalizing the measurement and model errors into the same metrological framework, and cannot consider the measurement uncertainty of the effective atomic number and the uncertainty of the theoretical atomic number derived from the end-member ratio. Therefore, the difference between the two results lacks a comparable scale, and can only rely on fixed thresholds or empirical rules, making it difficult to adaptively give reliable thresholds with batches and working conditions, and difficult to meet the needs of rapid and reliable judgment on site. SUMMARY
[0004] The purpose of the present application is to solve the problem of lack of unified metrological scale of measurement results of ray fluorescence spectrum channel and short-wave infrared reflectance spectrum channel in the prior art, which is prone to composition determination deviation caused by particle size and surface state interference. A geological sample composition identification and anomaly detection system based on machine learning is proposed.
[0005] To solve the problems existing in the prior art, the present application adopts the following technical scheme: A geological sample composition identification and anomaly detection system based on machine learning, comprising: A spectrum acquisition module for acquiring ray fluorescence spectrum and short-wave infrared reflectance spectrum of a geological sample; An atomic number metrology module for calculating the integral intensity ratio of Compton scattering peak and Rayleigh scattering peak based on the ray fluorescence spectrum, and obtaining the measured value and measurement uncertainty according to the integral intensity ratio. a band participation extraction module configured to remove continuum from the short-wave infrared reflectance spectrum to obtain a depth of the hydroxyl aluminum bond absorption band and a band shape asymmetry, and to obtain a proxy of particle size according to a slope of the short-wave infrared reflectance spectrum; a clay proportion fitting module configured to perform a monotonic spline regression fitting on the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the proxy of particle size to obtain an estimated value of clay proportion; an end-member weight calculation module configured to construct an end-member mixing weight according to the estimated value of clay proportion and the proxy of particle size, and to obtain a predicted value and a predicted uncertainty by an end-member periodic table; a cross-modal inconsistency module configured to calculate a combined uncertainty based on the measured value and the measurement uncertainty, and to calculate a cross-modal inconsistency index according to the combined uncertainty, the predicted value, and the predicted uncertainty; an optimization module configured to design a collection and fusion scheme based on the proxy of particle size, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the combined uncertainty, and to obtain an optimized cross-modal inconsistency index by taking the cross-modal inconsistency index as an optimization target; a component unmixing module configured to perform component unmixing and anomaly judgment on the short-wave infrared reflectance spectrum based on the optimized cross-modal inconsistency index.
[0006] Preferably, an integral intensity ratio of the Compton scattering peak and the Rayleigh scattering peak is calculated based on the ray fluorescence spectrum, and the measured value and the measurement uncertainty are obtained according to the integral intensity ratio, including: baseline correction and peak shape fitting are performed on the ray fluorescence spectrum to obtain the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak; a monotonic calibration mapping relationship is established according to a standard sample with a known average atomic number; the ratio of the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak is converted by the monotonic calibration mapping relationship to obtain a measured value of the effective atomic number; error propagation is performed on the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak according to the covariance of the peak shape fitting to obtain a measurement uncertainty of the effective atomic number.
[0007] Preferably, continuum is removed from the short-wave infrared reflectance spectrum to obtain a depth of the hydroxyl aluminum bond absorption band and a band shape asymmetry, and a proxy of particle size is obtained according to a slope of the short-wave infrared reflectance spectrum, including: a first-order and a second-order derivative of a reflectivity of the short-wave infrared reflectance spectrum with respect to wavelength are obtained by adjacent difference; a local minimum point at which the first-order derivative is converted from negative to positive and the second-order derivative is positive is taken as a candidate point of the center of the hydroxyl aluminum bond absorption band; a processing window is generated according to the candidate point of the center of the hydroxyl aluminum bond absorption band; In the processing window, the continuum baseline is formed by connecting the reflectance extremum points at both ends of the processing window through polynomial fitting; The original reflectance in the processing window is divided by the baseline reflectance of the corresponding waveband to obtain the continuum-removed spectrum; The depth of the hydroxyl aluminum bond absorption band is obtained according to the difference between the reflectance of the continuum baseline and the continuum-removed spectrum at the center of the processing window; The integrated area from the left boundary to the center of the processing window and the integrated area from the right boundary to the center of the window are calculated, and the band shape asymmetry of the hydroxyl aluminum bond absorption band is obtained according to the ratio of the two areas; In the processing window, the slope of the reflectance with respect to the wavelength is calculated by the least square method, and the slope is taken as the granularity proxy.
[0008] Preferably, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the granularity proxy are monotonically spline regression fitted to obtain the clay proportion estimate, including: The actual clay proportion of each sample in the standard sample set is determined by experiment, and the continuum removal and spectral slope calculation are performed on the short-wave infrared reflectance spectrum of the sample to obtain the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the granularity proxy; The depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the granularity proxy of each standard sample in the standard sample set are subjected to outlier sample rejection and normalization processing; A monotonous spline mapping model is constructed with the actual clay proportion of the standard sample set as the dependent variable, the depth of the hydroxyl aluminum bond absorption band and the band shape asymmetry of the standard sample set as the independent variables, and the granularity proxy of the standard sample set as the adjustment independent variable, and the monotonous constraint of non-negative derivative is imposed on the depth of the hydroxyl aluminum bond absorption band and the band shape asymmetry in the model; The residual error of the clay proportion estimate and the actual clay proportion value of the standard sample is minimized as the target to train the model parameters, and the number of spline knots and the penalty strength are selected by cross-validation; The model is tested for violation of the monotonicity constraint by the distribution of the training residual error, and if there is local violation, the knot density is automatically adjusted and retrained until the monotonicity constraint is satisfied; The depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the granularity proxy of the geological sample are normalized and input into the monotonous spline mapping model to obtain the clay proportion estimate.
[0009] Preferably, the endmember mixing weight is constructed according to the clay proportion estimate and the granularity proxy, and the predicted value and the prediction uncertainty are obtained from the atomic sequence table of the endmember, including: The endmember set and the atomic sequence table of the endmember are obtained, wherein the endmember set fixedly contains: clay, feldspar and quartz; constructing a basic weight vector based on the clay proportion estimation value and the endmember set, wherein the basic weight of the clay endmember is equal to the clay proportion estimation value, and the basic weight of the non-clay part is equal to one minus the clay proportion estimation value; dividing the non-clay part into a feldspar basic weight and a quartz basic weight based on the prior feldspar proportion obtained by statistical analysis of historical samples; adjusting the granularity of the clay basic weight and the non-clay basic weight according to the granularity proxy, wherein the feldspar basic weight and the quartz basic weight are adjusted by a common amplitude of granularity adjustment; normalizing and combining the clay basic weight, the feldspar basic weight and the quartz basic weight after granularity adjustment to obtain endmember mixing weights; performing weighted calculation on the endmember mixing weights and the endmember atomic number table to obtain predicted values of theoretical atomic numbers; performing error propagation based on the variance and covariance of the endmember mixing weights to obtain the predicted uncertainty of the theoretical atomic numbers.
[0010] Preferably, the synthetic uncertainty is calculated based on the measured value and the measurement uncertainty, and the cross-modal inconsistency index is calculated according to the synthetic uncertainty, the predicted value and the predicted uncertainty, including: adding the square of the effective atomic number measurement uncertainty to the square of the theoretical atomic number predicted uncertainty, taking the arithmetic square root of the sum to obtain the synthetic uncertainty; performing normalized processing on the difference between the measured value of the effective atomic number and the predicted value of the theoretical atomic number based on the synthetic uncertainty to obtain the cross-modal inconsistency index; performing bootstrap resampling on the current sample to construct an empirical empty distribution, and obtaining a significance score according to the exceeding proportion of the cross-modal inconsistency index in the empirical empty distribution; determining an adaptive quantile threshold value according to the significance score of the current batch of samples.
[0011] Preferably, the design of the acquisition and fusion scheme is based on the granularity proxy, the depth of the aluminum bond absorption band containing hydroxyl, the band shape asymmetry and the synthetic uncertainty, and the cross-modal inconsistency index is taken as the optimization target to obtain the optimized cross-modal inconsistency index, including: determining the value range of the short-wave infrared acquisition time, the number of repeated acquisitions of the X-ray fluorescence and the number of point micro-cleaning; constructing a state vector based on the granularity proxy, the depth of the aluminum bond absorption band containing hydroxyl, the band shape asymmetry and the synthetic uncertainty; constructing a design variable vector based on the short-wave infrared acquisition time, the number of repeated acquisitions of the X-ray fluorescence and the number of point micro-cleaning; training a proxy model based on a historical sample set containing the state vector, the design variable vector and the cross-modal inconsistency index; In the value range of the design variable vector, the design variable vector that minimizes the output of the surrogate model is solved to obtain an optimal retest configuration; According to the optimal retest configuration, short-wave infrared reflectance spectroscopy and X-ray fluorescence energy spectrum retests are performed at the same physical point of the current sample, and the cross-modal inconsistency index and the significance score are recalculated; If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated, and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the state vector of this round, the optimal retest configuration and the updated cross-modal inconsistency index are added to the historical sample set and the iteration is continued.
[0012] Preferably, based on the optimized cross-modal inconsistency index, component unmixing and anomaly judgment of the short-wave infrared reflectance spectroscopy are performed, including: An a priori component vector is constructed according to the clay proportion estimate value and the a priori feldspar proportion; The short-wave infrared reflectance spectroscopy and the endmember spectrum corresponding to the endmember set are aligned at the same wavelength sampling point; The endmember unmixing is performed to obtain the final value of the component vector, with the reconstruction error between the short-wave infrared reflectance spectroscopy and the weighted sum of the endmember spectrum and the deviation of the a priori component vector as the target; The unmixing residual sequence is obtained according to the difference between the short-wave infrared reflectance spectroscopy and the weighted sum of the endmember spectrum; The in-band average residual and the in-band variance are calculated based on the unmixing residual sequence, and the anomaly judgment is performed according to the optimized significance score and the quantile threshold as the threshold.
[0013] Compared with the prior art, the present application has the following advantages: 1、The present application can effectively eliminate the systematic interference of particle size, coating and equipment drift on the judgment by quantifying the X-ray fluorescence energy spectrum and the short-wave infrared reflectance spectroscopy into the measured value and the measurement uncertainty of the effective atomic number, and the depth, band shape asymmetry and particle size proxy of the aluminum hydroxyl bond absorption band, and obtaining the predicted value and the predicted uncertainty of the theoretical atomic number through the monotone spline regression and the endmember mixing calculation chain, thereby improving the accuracy and stability of component identification; through the unified measurement of combined uncertainty and the normalized comparison of cross-modal inconsistency index, the comparability and traceability of results under different batches and different working conditions are effectively guaranteed.
[0014] 2、The application can effectively reduce the retest rounds and resource consumption, inhibit the influence of on-site fluctuations on the results, thereby improving the consistency and reproducibility across batches, by taking the particle size proxy, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the synthetic uncertainty as state variables, and taking the cross-modal inconsistency index as an optimization target to carry out joint optimization of the acquisition time, the number of repeated acquisitions and the number of point micro-cleaning times.
[0015] 3、The application can effectively locate the deviation caused by unmodeled end members, particle size abnormalities or surface coating, eliminate the problem that it is difficult to distinguish only by the overall error, thereby improving the accuracy and timeliness of anomaly detection, by implementing end member unmixing with consistency constraint after obtaining the optimized cross-modal inconsistency index, and performing anomaly judgment with unmixing residual. BRIEF DESCRIPTION OF DRAWINGS
[0016] The drawings described herein are used to provide further understanding of the application, form part of the application, the illustrative embodiments of the application and the description thereof are used to explain the application, and do not constitute improper limitation on the application. In the drawings: Figure 1 A functional module diagram of a geological sample composition identification and anomaly detection system based on machine learning provided by an embodiment of the application; Figure 2 An atomic order metrology module flowchart provided by an embodiment of the application; Figure 3 A band parameter particle size extraction module flowchart provided by an embodiment of the application; Figure 4 A clay proportion fitting module flowchart provided by an embodiment of the application; Figure 5 An end member weight calculation module flowchart provided by an embodiment of the application; Figure 6 A cross-modal inconsistency module flowchart provided by an embodiment of the application; DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments.
[0018] Embodiment: The embodiment provides a geological sample composition identification and anomaly detection system based on machine learning, referring to Figure 1 , specifically, comprising: A spectrum acquisition module for acquiring the ray fluorescence spectrum and short wave infrared reflectance spectrum of the geological sample; In the embodiments of the present application, the ray fluorescence spectrum and the short-wave infrared reflectance spectrum of the geological sample are acquired, comprising: Acquiring the ray fluorescence spectrum and the short-wave infrared reflectance spectrum of the geological sample; Specifically, the ray fluorescence spectrum is obtained by recording the distribution of the fluorescence intensity with the energy change after the sample is irradiated by high-energy rays, and each element in the sample is excited and emits fluorescence rays with characteristic energy; the short-wave infrared reflectance spectrum is obtained by measuring the surface reflectivity with the wavelength change after the sample is irradiated in the short-wave infrared band of one micrometer to two point five micrometers. Among them, the ray fluorescence spectrum provides element level evidence, which is used to obtain the measured value of the effective atomic number; the short-wave infrared reflectance spectrum provides mineralogy and structure level evidence, which is used to remove the continuous unit to obtain the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band, and the spectrum slope is combined as a particle size proxy. Comparing the two types of evidence in the same measurement framework can construct a cross-modal inconsistency index to reveal the systematic deviation caused by particle size, water content or surface coating, and realize composition identification and anomaly detection.
[0019] Specifically, the X-ray fluorescence spectrometer is used to irradiate the geological sample, the X-ray fluorescence spectrometer emits primary X-rays with a preset energy to act on the surface of the geological sample, so that each element in the geological sample is excited to generate characteristic X-ray fluorescence, the count signals of different energy characteristic X-ray fluorescence are collected by the detector of the X-ray fluorescence spectrometer, the count signals are distinguished and counted by the multi-channel analyzer built in the instrument according to the energy, and the ray fluorescence spectrum containing the Compton peak and the Rayleigh peak is generated; the same geological sample is irradiated by the short-wave infrared spectrometer, the short-wave infrared spectrometer emits continuous infrared light with a wavelength range of 1.0 micrometer to 2.5 micrometers, the infrared light is reflected by the surface of the geological sample, the reflected light is spectrally resolved by the spectrometer system of the short-wave infrared spectrometer, and then the light signals of different wavelengths are converted into electrical signals by the detector, and the short-wave infrared reflectance spectrum of the detection point is generated after analog-to-digital conversion.
[0020] The atomic number measurement module, see Figure 2 is used to calculate the integral intensity ratio of the Compton scattering peak and the Rayleigh scattering peak based on the ray fluorescence spectrum, and to obtain the measured value and the measurement uncertainty according to the integral intensity ratio; In the embodiments of the present application, the integral intensity ratio of the Compton scattering peak and the Rayleigh scattering peak is calculated based on the ray fluorescence spectrum, and the measured value and the measurement uncertainty are obtained according to the integral intensity ratio, comprising: Baseline correction and peak fitting are performed on the ray fluorescence spectrum to obtain the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak; A monotonic calibration mapping relationship is established according to the standard sample with a known average atomic number; The ratio of the integral intensity of the Compton scattering peak to the integral intensity of the Rayleigh scattering peak is converted by a monotonic calibration mapping relationship to obtain a measured value of the effective atomic number; The measurement uncertainty of the effective atomic number is obtained by error propagation of the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak according to the covariance of the peak shape fitting; Specifically, the integral intensity of the Compton scattering peak is the area obtained by integrating the net peak shape function in the energy window of the Compton scattering peak after baseline correction and peak shape fitting, which represents the inelastic scattering intensity; the integral intensity of the Rayleigh scattering peak is the area obtained by integrating the net peak shape function in the energy window of the Rayleigh scattering peak after baseline correction and peak shape fitting; and the average atomic number standard sample is a standard sample set with known composition and ratio, which is measured and processed in the same way as the sample to be measured, and the average atomic number is calculated according to the element atomic fraction.
[0021] Specifically, the original energy spectrum contains a slow varying background and peak overlap, and direct area taking will introduce systematic bias; background removal by baseline correction and separation and quantification of target spectral peaks by peak shape fitting are required to obtain quantifiable integral intensity. When performing baseline correction and peak shape fitting on the X-ray fluorescence spectrum, a polynomial fitting method is used to fit the background region of the X-ray fluorescence spectrum to obtain a continuous baseline spectrum; then the original X-ray fluorescence spectrum is subtracted from the baseline spectrum to complete the baseline correction operation. For the Compton scattering peak and the Rayleigh scattering peak, a Lorentz-Gaussian mixed function is used to expand and fit the peak shape, the parameters of the mixed function are optimized by a nonlinear least squares method to achieve the best matching of the mixed function and the peak shape region of the baseline corrected energy spectrum, and the integral operation is performed on the Compton scattering peak and the Rayleigh scattering peak respectively to obtain the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak.
[0022] Specifically, different samples have differences in geometry, density and matrix effect, and a standard sample consistent with the target working condition is needed to establish a stable corresponding relationship between the measurable quantity and the composition level quantity; and the monotonicity constraint can ensure that the signal change only corresponds to the composition change in one direction. According to the known average atomic number of the standard sample, a monotonic calibration mapping relationship is established by selecting no less than 10 standard samples with different average atomic numbers and covering the possible atomic number range of the target geological sample. Each standard sample is subjected to the same X-ray fluorescence spectrum acquisition and peak fitting process as the geological sample to be measured, and the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity of each standard sample is obtained; the known average atomic number of the standard sample is used as the dependent variable, and the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity corresponding to the standard sample is used as the independent variable, and a monotonic spline interpolation method is used to fit the monotonic calibration mapping relationship, which needs to meet the characteristics of the average atomic number monotonically increasing with the increase of the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity. A reusable and traceable conversion channel is formed from the measurable intensity feature to the effective atomic number, reducing the drift caused by environmental and working condition changes.
[0023] Specifically, the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity can partially offset the influence of sample thickness, surface topography and overall gain, and inputting the ratio into the monotonic calibration mapping relationship can obtain the composition level quantity. When converting the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity through the monotonic calibration mapping relationship, the ratio of the Compton scattering peak integral intensity to the Rayleigh scattering peak integral intensity of the geological sample to be measured is calculated, and then the ratio is input into the pre-established monotonic calibration mapping relationship. By means of interpolation calculation of the mapping relationship, the measured value of the effective atomic number of the geological sample to be measured is obtained. The measured value of the effective atomic number directly related to the composition of the sample and comparable across samples is output, which lays a foundation for subsequent consistency comparison and calculation of the cross-modal inconsistency index between the predicted value of the theoretical atomic number and the measured value of the effective atomic number.
[0024] Specifically, the measured value of effective atomic number is obtained by function transformation of two integral intensities, and the uncertainty thereof is derived from the uncertainty and correlation of the peak shape fitting parameters; error propagation can correctly transfer the covariance information to the result quantity; the covariance matrix of the fitting parameters of the Compton scattering peak and the Rayleigh scattering peak is obtained from the peak shape fitting process, which reflects the error correlation degree between the fitting parameters; based on the analytical expressions of the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak with respect to the fitting parameters, and combined with the error propagation law, the covariance of the fitting parameters is transferred to the ratio of the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak, and the standard uncertainty of the ratio is obtained; according to the derivative information of the monotone calibration mapping relationship, the standard uncertainty of the ratio is further transferred to the measured value of the effective atomic number, so that the measurement uncertainty of the effective atomic number is obtained. The measurement uncertainty of the effective atomic number can be combined with the prediction uncertainty of the theoretical atomic number into a unified combined uncertainty, which is used as a scale reference of the cross-modal inconsistency index to ensure that the comparison and threshold control have statistical significance.
[0025] The parametric particle size extraction module, as shown in Figure 3 , is used for continuum removal of the short-wave infrared reflectance spectrum to obtain the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band, and the particle size proxy is obtained according to the slope of the short-wave infrared reflectance spectrum. In the embodiments of the present application, the continuum removal is performed on the short-wave infrared reflectance spectrum to obtain the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band, and the particle size proxy is obtained according to the slope of the short-wave infrared reflectance spectrum, comprising: The first and second derivatives of the reflectivity of the short-wave infrared reflectance spectrum with respect to wavelength are obtained by adjacent difference; The local minimum point at which the first derivative changes from negative to positive and the second derivative is positive is taken as the candidate point of the center of the hydroxyl aluminum bond absorption band; A processing window is generated according to the candidate point of the center of the hydroxyl aluminum bond absorption band; Specifically, the hydroxyl aluminum bond absorption band refers to a characteristic absorption band caused by the vibration of the hydroxyl aluminum bond, the combination or frequency multiplication of the hydroxyl stretching and the aluminum-hydroxyl bending in the short-wave infrared reflectance spectrum, which usually appears near 2.2 microns; the absorption band is mainly derived from the aluminum-hydroxyl structure in layered hydroxyl aluminum silicate minerals such as muscovite, illite, and kaolinite. The center position is determined by the minimum value of the reflectivity after continuum removal, the band depth reflects the relative abundance of the hydroxyl aluminum bond related endmember, and the band shape asymmetry is affected by endmember mixing and particle size, which can be used to distinguish different clay mineral combinations or surface states.
[0026] Specifically, the wave infrared reflectance spectrum is discrete sampling data, and first-order derivative and second-order derivative discrete approximations are obtained by using adjacent difference, so as to convert the intensity curve into extreme value and concave-convex information, and provide criteria for subsequent accurate positioning of the absorption band minimum point. For the discrete sequence of the reflectivity of the short wave infrared reflectance spectrum changing with the wavelength, the first-order derivative discrete approximation value is calculated by using the first-order forward difference formula, and the second-order derivative discrete approximation value is calculated by using the second-order forward difference formula, so as to obtain the first-order derivative discrete approximation sequence and the second-order derivative discrete approximation sequence of the sequence of the reflectivity changing with the wavelength.
[0027] Specifically, the hydroxyl-containing aluminum bond absorption band exhibits local reflectivity minimum near the center position, and the mathematical characteristics are that the first-order derivative changes from negative to positive at the left and right of the point, and the second-order derivative is positive and concave upward; the search is limited to the instrument usable range, so as to avoid edge noise and invalid band interference. All wavelength points in the instrument usable short wave infrared range are traversed, for each wavelength point, it is judged whether the corresponding first-order derivative discrete approximation value satisfies the conditions that the first-order derivative discrete approximation value of the left adjacent wavelength point is negative, and the first-order derivative discrete approximation value of the right adjacent wavelength point is positive, and whether the corresponding second-order derivative discrete approximation value is positive; if the two conditions are satisfied at the same time, the wavelength point is determined as the center candidate point of the hydroxyl-containing aluminum bond absorption band.
[0028] Specifically, the band depth and band shape asymmetry calculation needs to complete the continuum fitting and the two sides shoulder search in the local window near the center; First, fix the processing window based on the center candidate point, ensure that the subsequent parameters are calculated in the same scale and the same neighborhood, avoid being absorbed by the adjacent water or other band disturbances. From the two sides of the hydroxyl aluminum bond absorption band center candidate point, search outward to find the position where the second derivative changes from positive to non-positive, take the left side position as the left inflection point and the right side position as the right inflection point. For each determined hydroxyl aluminum bond absorption band center candidate point, from the center candidate point, check the discrete approximation value of the second derivative of each wavelength point in reverse order as the wavelength decreases. When the first second derivative discrete approximation value changes from positive to non-positive, that is, the wavelength point is less than or equal to zero, the wavelength point is determined as the left inflection point. Then, from the center candidate point, check the discrete approximation value of the second derivative of each wavelength point in the direction of increasing wavelength. When the first second derivative discrete approximation value changes from positive to non-positive, that is, the wavelength point is less than or equal to zero, the wavelength point is determined as the right inflection point. When the wavelength at the left inflection point is taken as the left endpoint and the wavelength at the right inflection point is taken as the right endpoint, the wavelength value corresponding to the searched left inflection point is directly taken as the left endpoint wavelength of the processing window, and the wavelength value corresponding to the searched right inflection point is taken as the right endpoint wavelength of the processing window. Take the wavelength corresponding to the left endpoint as the starting wavelength of the window, and take the wavelength corresponding to the right endpoint as the ending wavelength of the window. At the same time, take the center candidate point as the absorption band center reference point that needs to be analyzed in the processing window, so as to determine a wavelength interval of a hydroxyl aluminum bond absorption band feature as a processing window. Make the subsequent continuum removal and band parameter extraction more robust, and improve the comparability of the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band.
[0029] In the processing window, the continuum baseline is formed by connecting the reflectance extreme points at both ends of the processing window through polynomial fitting; Divide the original reflectance in the processing window by the baseline reflectance corresponding to the waveband to obtain the continuum removal spectrum; According to the difference between the reflectance of the continuum baseline and the continuum removal spectrum at the center of the processing window, the depth of the hydroxyl aluminum bond absorption band is obtained; Calculate the integral area from the left boundary to the center of the processing window and the integral area from the right boundary to the center of the window, and obtain the band shape asymmetry of the hydroxyl aluminum bond absorption band according to the ratio of the two areas; In the processing window, the slope of the reflectance with respect to the wavelength is calculated by the least square method, and the slope is taken as the granularity proxy; Specifically, the continuum baseline is a smooth reference curve obtained by connecting the reflectance extreme points at both ends of the processing window with a polynomial, which is used to depict the slowly varying background and brightness trend without the influence of absorption bands; the continuum-removed spectrum is a dimensionless spectrum obtained by dividing the original reflectance in the processing window by the continuum baseline, and the value is approximately one at the non-absorption position and less than one at the absorption position; the granularity proxy is the overall inclination of the short-wave infrared reflectance spectrum in the processing window, which is used to represent the influence of particle size and surface roughness on scattering.
[0030] Specifically, the absorption band is superimposed on the slowly varying background and illumination trend, and the continuum baseline is used to depict the slow variation, ensuring that the subsequent measurement of band shape and intensity only reflects the true absorption and not the baseline drift. The reflectance extreme points at both ends of the window are selected as anchor points, which can avoid the area affected by absorption within the band. When fitting the continuum baseline in the processing window, the reflectance extreme point corresponding to the left end point of the processing window is the maximum value point of the reflectance near the left boundary of the processing window, and the reflectance extreme point corresponding to the right end point of the processing window is the maximum value point of the reflectance near the right boundary of the processing window; and a cubic polynomial is used to fit the reflectance trend between the two reflectance extreme points, and the coefficients of the cubic polynomial are solved by the least squares method, so that the fitting curve can smoothly connect the two reflectance extreme points, thereby forming the continuum baseline. This reduces the systematic deviation caused by baseline uplift or fluctuation.
[0031] Specifically, after obtaining the continuum baseline, the original reflectance value at each wavelength point in the processing window is divided by the baseline reflectance value corresponding to the wavelength point on the continuum baseline, and the continuum-removed spectrum in the entire processing window is obtained by point-by-point calculation. Scaling the reflectance by the baseline can eliminate the multiplicative brightness and albedo differences within the window, so that the absorption is only in the form of relative weakening.
[0032] Specifically, the relative weakening amplitude at the center of the absorption band directly corresponds to the intensity information of the end member containing the aluminum bond related to hydroxyl, and the central difference measurement can accurately represent the absorption intensity. When calculating the depth of the absorption band containing the aluminum bond related to hydroxyl, the baseline reflectance value corresponding to the center of the processing window is found on the continuum baseline, and the reflectance value corresponding to the center of the processing window is found on the continuum-removed spectrum, and then the reflectance value of the continuum baseline at the wavelength band is subtracted from the reflectance value of the continuum-removed spectrum at the wavelength band. The difference obtained is the depth of the absorption band containing the aluminum bond related to hydroxyl, which is insensitive to illumination and baseline, and reflects the relative abundance of clay minerals and other end members related to the aluminum bond related to hydroxyl.
[0033] Specifically, whether the absorption band shape is skewed to the short-wave or long-wave side is commonly determined by end-member mixing and particle size scattering, and the area ratio can comprehensively integrate the shape information on both sides and is more resistant to noise than single-point width or half-height width. When calculating the band shape asymmetry of the aluminum bond containing hydroxyl absorption band, the left boundary of the processing window is taken as the integral starting point, the center position is taken as the integral termination point, the trapezoidal integral method is used to integrate the reflectance curve of the continuum-removed spectrum in the interval to obtain the integral area from the left boundary to the center; then the center position is taken as the integral starting point, the right boundary is taken as the integral termination point, and the trapezoidal integral method is also used to integrate the reflectance curve of the continuum-removed spectrum in the interval to obtain the integral area from the right boundary to the center; the integral area from the left boundary to the center is divided by the integral area from the right boundary to the center, and the obtained ratio is the band shape asymmetry of the aluminum bond containing hydroxyl absorption band. The morphological quantity band shape asymmetry can reflect the shape of end-member combination and scattering effect, enhance the monotonic relationship with the clay ratio, and provide a basis for weight adjustment in end-member mixing.
[0034] Specifically, particle size and surface roughness can cause overall spectral line tilting, and a linear slope is used to compress this macro trend into a single measurable quantity, which is convenient to use as an adjustment variable in subsequent models. When calculating the particle size proxy, all wavelength points and their corresponding reflectance values of the continuum-removed spectrum in the processing window are selected to construct a linear fitting model of reflectance with respect to wavelength, and the slope of the linear model is solved by least squares method, so that the residual sum of squares between the predicted reflectance of the model and the actual reflectance of the continuum-removed spectrum is minimized. The slope is the slope of the reflectance with respect to the wavelength, which is used as the particle size proxy.
[0035] The clay ratio fitting module, see Figure 4 is used for monotonous spline regression fitting of the depth, band shape asymmetry and particle size proxy of the aluminum bond containing hydroxyl absorption band to obtain a clay ratio estimation value; In the embodiments of the present application, the depth, band shape asymmetry and particle size proxy of the aluminum bond containing hydroxyl absorption band are monotonously spline regression fitted to obtain a clay ratio estimation value, comprising: The actual clay ratio of each sample in the standard sample set is determined by experiment, and the continuum-removed spectrum and spectral slope calculation are performed on the short-wave infrared reflectance spectrum of the sample to obtain the depth, band shape asymmetry and particle size proxy of the aluminum bond containing hydroxyl absorption band; The depth, band shape asymmetry and particle size proxy of the aluminum bond containing hydroxyl absorption band of each standard sample in the standard sample set are subjected to outlier sample rejection and normalization processing; Specifically, the actual clay ratio is the proportion of clay minerals related to the aluminum bond containing hydroxyl in the total amount of the sample, such as kaolinite, illite, muscovite and the like; the standard sample set is a set of standard samples with known composition and quantified actual clay ratio.
[0036] Specifically, the actual clay proportion is a supervised target necessary for training the monotonic spline mapping model, the depth of the hydroxyl aluminum bond absorption band and the band shape asymmetry can directly represent the intensity and morphological information related to the clay end member, and the spectral slope is used as a proxy for particle size to characterize the influence of scattering and roughness. When determining the actual clay proportion of each sample in the standard sample set, chemical analysis or existing standard mineral quantitative analysis methods, such as X-ray diffraction quantitative analysis combined with Rietveld full spectrum fitting method, are used to analyze the composition of each standard sample, accurately determine the mass proportion of clay minerals such as kaolinite, illite, and muscovite, which contain hydroxyl aluminum, and thus obtain the actual clay proportion of each sample. The continuum removal operation is performed on the short-wave infrared reflectance spectrum of each sample, i.e., the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the particle size proxy of each standard sample are calculated according to the method described above.
[0037] Specifically, the standard sample collection may have occasional noise, label errors, or cross-batch dimensional differences; abnormal samples will deviate the model, and inconsistent feature scales will make the training unstable; by removing abnormal samples and normalizing the scale, the training data can be constrained within a physically reasonable and statistically stable range. When removing abnormal samples from the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the particle size proxy of each standard sample in the standard sample set, Grubbs test is used to set the significance level, calculate the statistics corresponding to the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the particle size proxy of each standard sample, and if the statistics exceed the Grubbs test critical value, the standard sample is determined to be an abnormal sample and is removed from the standard sample set. After completing the removal of abnormal samples, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry, and the particle size proxy of the remaining standard samples are normalized, and the linear normalization method is used to map the values of each index to the interval of 0 to 1, i.e., for each index, the current value of the index is subtracted from the minimum value of the index in the remaining standard samples, and then divided by the difference between the maximum and minimum values of the index in the remaining standard samples to obtain the normalized value.
[0038] A monotonic spline mapping model is constructed with the actual clay proportion of the standard sample set as the dependent variable, the depth of the hydroxyl aluminum bond absorption band and the band shape asymmetry of the standard sample set as the independent variables, and the particle size proxy of the standard sample set as the adjusting independent variable, and the depth of the hydroxyl aluminum bond absorption band and the band shape asymmetry are subjected to non-negative derivative monotonicity constraint in the model; The residual error between the estimated clay proportion and the actual clay proportion of the standard sample is minimized as the square error of the target training model parameters, and the number of spline knots and the penalty strength are selected through cross-validation; The distribution of the training residual is used to test whether the model violates the monotonicity constraint, and if there is a local violation, the node density is automatically adjusted and retrained until the monotonicity constraint is satisfied; The depth, band shape asymmetry and granularity proxy of the hydroxyl aluminum bond absorption band of the geological sample are normalized and input into a monotone spline mapping model to obtain the clay proportion estimation value; Specifically, the actual clay proportion and the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band have a monotone direction in mineralogy, and the granularity proxy changes the spectral form and intensity, and needs to be explicitly included as an adjusting independent variable. By applying a non-negative derivative constraint to the two main independent variables, the mineralogical prior can be solidified into a computable directional constraint. When constructing the monotone spline mapping model, the actual clay proportion of the standard sample set is used as the dependent variable, the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band of the standard sample set are used as the independent variables, and the granularity proxy of the standard sample set is used as the parameter of the adjusting independent variable. A monotone spline function is used to establish a mapping relationship. In this model, the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band are subjected to a non-negative derivative monotonicity constraint, that is, the partial derivative of the model with respect to the depth of the hydroxyl aluminum bond absorption band is non-negative, and the partial derivative with respect to the band shape asymmetry is also non-negative, ensuring that as the depth or band shape asymmetry of the hydroxyl aluminum bond absorption band increases, the clay proportion estimation value does not decrease, which conforms to the physical correlation law of clay minerals and spectral characteristics. A model that conforms to the mineralogical direction and can adapt to changes in granularity is obtained, avoiding non-physical inverse relationships and improving the robustness across samples and batches.
[0039] Specifically, the residual minimum square error training can obtain the optimal fitting of the standard sample data; the number of spline knots and the penalty strength jointly control the complexity and smoothness of the model, and cross-validation can be used to balance the bias and variance. When training the model parameters, the residual sum of squares of the clay proportion estimation value and the actual clay proportion value of the standard sample is minimized as the objective function, and the model parameters are solved by an iterative optimization algorithm. At the same time, the cross-validation method is used to select the number of spline knots and the penalty strength: the standard sample set is divided into multiple subsets, different subsets are selected as the validation set in turn, the remaining subsets are used as the training set, the model is trained and the performance is verified, and the number of spline knots and the penalty strength that make the generalization ability of the model optimal are determined according to the verification result, to avoid overfitting or underfitting of the model.
[0040] Specifically, data sparsity or noise may cause the fitting result in a local region to violate the physical direction; by residual distribution test and adaptive increase of local node density, local mismatch can be corrected while keeping the overall complexity controllable. When testing the monotonicity constraint of the model, the residual distribution generated during the training process is analyzed to determine whether the model violates the pre-imposed monotonicity constraint. If the derivative of the model with respect to the depth or band shape asymmetry of the aluminum bond containing hydroxyl group absorption band in a local region is negative, that is, the local monotonicity constraint is violated, the spline node density of the local region is automatically increased, the fitting accuracy of the model in the region is refined, and then the model is retrained until the model satisfies the non-negative derivative monotonicity constraint of the depth and band shape asymmetry of the aluminum bond containing hydroxyl group absorption band in the global range. The consistency and reliability of the model in the full composition range are improved, and the propagation of training noise to the subsequent calculation chain of end member weight and theoretical atomic number is avoided.
[0041] Specifically, when obtaining the clay proportion estimate value of the geological sample, the depth, band shape asymmetry and particle size proxy of the aluminum bond containing hydroxyl group absorption band of the geological sample are normalized, and the normalization method is consistent with that of the standard sample set. The values of these indicators are mapped to the same interval; then the normalized indicators are input into the trained monotonic spline mapping model, and the clay proportion estimate value of the geological sample is calculated and output by the model. The clay proportion estimate value is comparable across samples, providing stable and directly callable input for subsequent stages using the estimate.
[0042] The end member weight calculation module, as shown in Figure 5 , is configured to construct an end member mixing weight according to the clay proportion estimate value and the particle size proxy, and obtain a predicted value and a predicted uncertainty through an end member atomic number table; In the embodiment of the present application, the end member mixing weight is constructed according to the clay proportion estimate value and the particle size proxy, and the predicted value and the predicted uncertainty are obtained through the end member atomic number table, comprising: The end member set and the atomic number table of the end member are obtained, wherein the end member set fixedly contains clay, feldspar and quartz; A basic weight vector is constructed based on the clay proportion estimate value and the end member set, wherein the basic weight of the clay end member is equal to the clay proportion estimate value, and the basic weight of the non-clay part is equal to one minus the clay proportion estimate value; The non-clay part is divided into a feldspar basic weight and a quartz basic weight based on the prior feldspar proportion obtained by statistical analysis of historical samples; Specifically, the endmember set is a set of representative pure mineral components selected for atomic order calculation in component unmixing theory, and in this embodiment, it is fixed to contain three types of clay, feldspar and quartz. Each endmember corresponds to its standard short-wave infrared endmember spectrum and the representative atomic order matched therewith; the atomic order table of the endmember is a correspondence table in which each endmember in the endmember set corresponds to its representative atomic order, which is derived from the known chemical composition of the endmember or the stable equivalent value measured by X-ray fluorescence.
[0043] Specifically, when the endmember set and the atomic order table of the endmember are obtained, the endmember set for component analysis of the geological sample is determined in advance, and the endmember set is fixed to contain three types of typical geological endmembers, i.e., clay, feldspar and quartz. The average atomic order corresponding to each type of endmember is determined by experimental measurement to form the atomic order table of the endmember. The average atomic order of the clay endmember is obtained by statistical calculation of the atomic composition of common clay minerals such as kaolinite and montmorillonite. The average atomic order of the feldspar endmember is determined by statistical calculation of the atomic composition of feldspar minerals such as potassium feldspar and plagioclase. The average atomic order of the quartz endmember is calculated from the atomic composition of silicon dioxide.
[0044] Specifically, the clay proportion estimate directly reflects the share of the clay endmember, and using it as the basic weight of the clay endmember can make the proportion one-to-one correspond to the physical quantity. Using its supplement as the weight of the non-clay part can satisfy the mass conservation of non-negativity and sum of one, avoiding additional degrees of freedom. When the clay proportion estimate and the endmember set are used to construct the basic weight vector, the clay proportion estimate obtained by the monotone spline mapping model is extracted. The basic weight of the clay endmember is directly set to the clay proportion estimate, and the basic weight of the non-clay part is obtained by subtracting the clay proportion estimate from 1. In this way, the basic weight vector containing the weight of the clay endmember and the weight of the non-clay part is constructed, wherein the sum of the weight of the clay endmember and the weight of the non-clay part is 1, satisfying the normalization constraint of the component proportion. This provides a solid starting point for subsequent subdivision and particle size adjustment of the non-clay part, and facilitates the transmission and audit of the uncertainty.
[0045] Specifically, the non-clay part is usually dominated by feldspar and quartz, and the introduction of the prior feldspar proportion can maintain distinguishability in the case of insufficient information or similar band shapes, avoid overfitting, and maintain geological rationality for subsequent coordinated adjustment under the influence of particle size. When the prior feldspar proportion obtained based on historical sample statistics is used to divide the non-clay part into a feldspar basic weight and a quartz basic weight, a large number of analysis data of historical geological samples are collected, the historical samples are subjected to composition detection, the distribution of the proportion of feldspar in the non-clay part in the historical samples is counted, and the prior feldspar proportion obtained based on historical sample statistics is obtained. Multiply the non-clay part weight of the current geological sample by the prior feldspar proportion obtained based on historical sample statistics to obtain the feldspar basic weight. Multiply the non-clay part weight by (1 minus the prior feldspar proportion obtained based on historical sample statistics) to obtain the quartz basic weight, thereby completing the division of the non-clay part into the feldspar basic weight and the quartz basic weight. The non-clay part is stably allocated into two items of feldspar and quartz, which improves the robustness and consistency of subsequent endmember mixing weight construction and theoretical atomic number calculation.
[0046] The particle size of the clay basic weight and the non-clay basic weight is adjusted according to the particle size proxy, and the particle size of the feldspar basic weight and the quartz basic weight is adjusted by a common amplitude; The clay basic weight, the feldspar basic weight and the quartz basic weight after particle size adjustment are normalized and combined to obtain the endmember mixing weight; The endmember mixing weight and the endmember atomic number table are weighted to obtain the predicted value of the theoretical atomic number; The error propagation is performed based on the variance and covariance of the endmember mixing weight to obtain the predicted uncertainty of the theoretical atomic number; Specifically, the predicted value of the theoretical atomic number is the average atomic number obtained by multiplying the endmember mixing weight and the atomic number table of the endmember one by one and summing up under the premise that the endmember set is fixed as clay, feldspar and quartz, which represents the element level quantity implied by the short-wave infrared endmember information under the current particle size condition. The predicted uncertainty of the theoretical atomic number is the uncertainty corresponding to the predicted value of the theoretical atomic number, which reflects the sensitivity of the prediction to the fluctuation of the endmember mixing weight and the influence of the correlation between the weights.
[0047] Specifically, the particle size and surface roughness will change the relative contribution of short-wave infrared scattering and absorption, and if not corrected, the particle effect will be mistaken for endmember proportion change. By adjusting the clay basis weight and non-clay basis weight respectively, the sensitivity of the two types of endmembers to particle size can be compensated. By using a common amplitude for feldspar basis weight and quartz basis weight, the relative proportion within non-clay can be maintained without being artificially distorted. When adjusting the clay basis weight and non-clay basis weight according to the particle size proxy, select geological samples with different known particle sizes, measure the particle size proxy and the adjustment coefficient of the corresponding clay, feldspar and quartz basis weight of each sample, and establish a corresponding relationship model between the particle size proxy and the weight adjustment coefficient through statistical analysis or fitting method. For the clay basis weight, input the particle size proxy of the current geological sample into the model to obtain the adjustment coefficient of the adapted clay basis weight, and use the adjustment coefficient to scale the clay basis weight to realize the particle size adjustment of the clay basis weight. For the non-clay basis weight, input the particle size proxy into the corresponding relationship model to obtain the adjustment coefficient. Since the feldspar basis weight and the quartz basis weight need to be scaled by a common amplitude of particle size, the adjustment coefficient is used to scale the feldspar basis weight and the quartz basis weight at the same time to complete the particle size adjustment of the non-clay basis weight. Under the premise of not destroying the mineralogical relationship, the system bias caused by particle size is reduced, making the subsequent endmember mixing weight more sensitive to real composition changes and more stable to physical interference.
[0048] Specifically, the endmember proportion must satisfy the mass conservation of non-negative and sum to one, and the particle size adjustment will change the magnitude of each component, which needs to be normalized to restore the weight space that can be physically interpreted. When normalizing the clay basis weight, feldspar basis weight and quartz basis weight after particle size adjustment, first calculate the weight sum of the three, and then divide the clay basis weight, feldspar basis weight and quartz basis weight by the sum, so that the weight sum of the three is 1, satisfying the normalization requirement of the proportion of components. After normalization, the three weights are combined into an endmember mixing weight vector containing the weights of clay, feldspar and quartz endmembers. Provide a standard input for the predicted value and prediction uncertainty of the theoretical atomic number.
[0049] Specifically, the end member mixing weight depicts the proportion, and the end member atomic sequence table provides the representative atomic sequence corresponding to each end member; the weighted sum of both can map the mineral proportion to the element level quantity. When the end member mixing weight and the end member atomic sequence table are weighted to obtain the predicted value of the theoretical atomic sequence, the clay basis weight, the feldspar basis weight and the quartz basis weight after the granularity adjustment and the normalization constraint are extracted as the end member mixing weight; the average atomic sequence of the clay end member, the feldspar end member and the quartz end member is obtained from the end member atomic sequence table; the predicted value of the theoretical atomic sequence is obtained through the weighted sum operation of the clay basis weight multiplied by the average atomic sequence of the clay end member, the feldspar basis weight multiplied by the average atomic sequence of the feldspar end member and the quartz basis weight multiplied by the average atomic sequence of the quartz end member. This lays a foundation for subsequent synthetic uncertainty calculation and cross-modal inconsistency index measurement.
[0050] Specifically, the end member mixing weight comes from learning and estimation, and there are variance and covariance. If it is not propagated downward, the uncertainty of the theoretical end quantity will be underestimated. Error propagation can correctly transfer the uncertainty of the weight and its correlation to the theoretical atomic sequence level. When the variance and covariance of the end member mixing weight are used for error propagation to obtain the predicted uncertainty of the theoretical atomic sequence, the variance of the clay, feldspar and quartz end member weight in the end member mixing weight and the covariance between them are determined to construct the variance and covariance matrix; according to the error propagation law, the variance and covariance of the end member mixing weight are transferred to the predicted value of the theoretical atomic sequence, and the variance of the predicted value of the theoretical atomic sequence is calculated. The square root of the variance is the predicted uncertainty of the theoretical atomic sequence. The synthetic uncertainty can be synthesized under the unified measurement framework, used as a scale benchmark for cross-modal comparison and significance evaluation, and the reliability of anomaly detection and optimization decision is improved.
[0051] The cross-modal inconsistency module, see Figure 6 , is used for calculating the synthetic uncertainty based on the measured value and the measurement uncertainty, and calculating the cross-modal inconsistency index according to the synthetic uncertainty, the predicted value and the predicted uncertainty; In the embodiments of the present application, the synthetic uncertainty is calculated based on the measured value and the measurement uncertainty, and the cross-modal inconsistency index is calculated according to the synthetic uncertainty, the predicted value and the predicted uncertainty, which comprises: The square of the effective atomic sequence measurement uncertainty is added to the square of the theoretical atomic sequence predicted uncertainty, and the arithmetic square root of the sum is obtained to obtain the synthetic uncertainty; The difference between the measured value of the effective atomic sequence and the predicted value of the theoretical atomic sequence is normalized based on the synthetic uncertainty to obtain the cross-modal inconsistency index; Specifically, the combined uncertainty of the effective atomic number is the total uncertainty obtained by combining the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number in the same measurement framework, which is used as a scale for comparing the difference between the measured value and the predicted value, so that different sources of noise are fairly accounted for, ensuring comparability within and across batches and statistical effectiveness; the cross-modal inconsistency index is a dimensionless number that measures the relative deviation between the measured value of the effective atomic number on the side of the X-ray fluorescence and the predicted value of the theoretical atomic number on the side of the short-wave infrared, and reflects the significance of the deviation at the current uncertainty level. It is the core target quantity for design optimization and anomaly judgment. The larger the value, the more inconsistent the information from the two modalities, and the more likely it is that the sample needs to be retested, optimized, or subjected to anomaly checking.
[0052] Specifically, the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number come from the measurement chain and the model chain, respectively, and both contribute variance independently in metrology. According to the combination rule of sum of variance and square root, the uncertainties can be combined in the same dimension to ensure that the result is non-negative and can truly reflect the overall fluctuation level. When calculating the combined uncertainty, the measurement uncertainty of the effective atomic number and the prediction uncertainty of the theoretical atomic number are obtained; the measurement uncertainty of the effective atomic number is squared, and the prediction uncertainty of the theoretical atomic number is also squared; the square of the two results is added to obtain the sum of squares; and the arithmetic square root of the sum is taken to obtain the combined uncertainty.
[0053] Specifically, directly comparing the original difference between the two values will be affected by the dimension and heteroscedasticity, making it difficult to compare across samples. Normalizing with the combined uncertainty is equivalent to calculating the standardized deviation, which puts samples with different uncertainty levels on the same evaluation scale. When calculating the cross-modal inconsistency index, the measured value of the effective atomic number and the predicted value of the theoretical atomic number are obtained, and the difference between the two values is calculated; the difference is divided by the combined uncertainty to obtain the cross-modal inconsistency index through such normalization.
[0054] Self-sampling is performed on the current sample to construct an empirical null distribution, and the significance score is obtained based on the exceedance proportion of the cross-modal inconsistency index in the empirical null distribution; An adaptive quantile threshold is determined based on the significance score of the current batch of samples; Specifically, the significance score is a probabilistic quantity that measures whether the cross-modal inconsistency index is significantly large on the current sample. The larger the value, the more significant it is, reflecting the likelihood of observing a deviation caused by random fluctuations at the current uncertainty and operating conditions. It is a direct basis for subsequent optimization and judgment; the adaptive quantile threshold is a judgment threshold determined adaptively based on the significance score distribution of the current batch of samples, which is used to distinguish between normal fluctuations and samples that need to be retested or subjected to anomaly checking.
[0055] Specifically, the cross-modal inconsistency index cannot be preset under the conditions of small sample, heteroscedastic and non-Gaussian noise on the spot, the self-sampling is kept with the same site pairing structure and uncertainty, the empirical null distribution of the current working condition is obtained without introducing parameter assumptions, the probability degree of the observation value is exceeded by the proportion measurement, and a comparable significance score is formed. When self-sampling is performed on the current sample to construct the empirical null distribution, a sub-sample with the same size as the original data set is repeatedly extracted from the original data set of the current sample in a way with replacement, the extraction operation is repeated multiple times, and the cross-modal inconsistency index of the corresponding sub-sample is calculated after each extraction. All cross-modal inconsistency indexes obtained by multiple extractions are summarized to construct the empirical null distribution of the cross-modal inconsistency index. The proportion of the cross-modal inconsistency index actually calculated in the current sample that exceeds the empirical null distribution is the significance score, which reflects the significance of the cross-modal inconsistency of the current sample. The significance score is completely data-driven and can adapt to the noise structure of different instruments and batches, and can be directly compared across samples to provide a robust basis for subsequent optimization and judgment.
[0056] Specifically, there are systematic differences in the environment and collection conditions of different batches, and the threshold needs to be set according to the empirical distribution of the significance score of the current batch, so that the boundary between normal fluctuations and retesting or abnormal checking is self-calibrated with data. Collect the significance scores of all samples in the current batch to construct the empirical distribution of the significance score; then according to the preset quantile level, such as selecting the 95% quantile level that can reflect the strictness of abnormal judgment, find the value corresponding to the quantile level in the empirical distribution of the significance score, and determine the value as the adaptive quantile threshold of the current batch. The cross-batch comparability and judgment stability are improved.
[0057] The optimization module is used for designing the collection and fusion scheme based on the particle size proxy, the depth of the absorption band containing the aluminum bond with a hydroxyl group, the band shape asymmetry and the synthesis uncertainty, and taking the cross-modal inconsistency index as an optimization target to obtain an optimized cross-modal inconsistency index. In the embodiments of the present application, the collection and fusion scheme is designed based on the particle size proxy, the depth of the absorption band containing the aluminum bond with a hydroxyl group, the band shape asymmetry and the synthesis uncertainty, and the cross-modal inconsistency index is taken as an optimization target to obtain an optimized cross-modal inconsistency index, including: Determine the value range of the short-wave infrared acquisition time, the ray fluorescence repeated acquisition times and the point site micro-cleaning times; Construct a state vector based on the particle size proxy, the depth of the absorption band containing the aluminum bond with a hydroxyl group, the band shape asymmetry and the synthesis uncertainty; Construct a design variable vector based on the short-wave infrared acquisition time, the ray fluorescence repeated acquisition times and the point site micro-cleaning times; training a proxy model based on a historical sample set comprising a state vector, a design variable vector, and a cross-modality inconsistency index; Specifically, the point micro-cleaning frequency is the cumulative number of light surface treatments performed at the same measurement point to remove loose dust, thin oxide film or slight surface coating. Specifically, the three controllable actions are limited within the feasible region allowed by the equipment and the operation window, avoiding the generation of unexecutable schemes such as timeout, overdosage or excessive cleaning, while providing clear boundary conditions for subsequent optimization. When determining the value range of short-wave infrared acquisition time, X-ray fluorescence repeated acquisition times and point micro-cleaning frequency, for short-wave infrared acquisition time, combining the detector response characteristics of the short-wave infrared spectrometer and the on-site detection efficiency requirements, through pre-experiment test of spectral signal-to-noise ratio under different time lengths, the shortest time length that can meet the extraction of hydroxyl aluminum bond absorption band characteristics is set as the lower limit, and the longest time length allowed by the equipment is set as the upper limit; for X-ray fluorescence repeated acquisition times, according to the counting statistical characteristics of X-ray fluorescence spectrum, through pre-experiment analysis of the variation coefficient of Compton peak and Rayleigh peak integral intensity under different repetition times of the same sample, the minimum number of times that can ensure the effective atomic number measurement accuracy is set as the lower limit, and the maximum number of times that can be accepted on site is set as the upper limit; for point micro-cleaning frequency, according to the attachment of surface interference of geological samples, through pre-experiment test of the stability of spectrum and energy spectrum under different cleaning times, the minimum number of times that makes the characteristics stable is set as the lower limit, and the maximum number of times before cleaning destroys the sample is set as the upper limit, so as to determine the value range of the three respectively. Ensure that the optimization solution only returns the optimal retest configuration that can be implemented, reduce invalid search and on-site trial and error cost, and improve the reliability and safety of online decision-making.
[0058] Specifically, when constructing the state vector based on the granularity proxy, the depth of the hydroxyl aluminum bond absorption band, the band shape asymmetry and the combined uncertainty, from the short-wave infrared reflectance spectrum processing result, the granularity proxy obtained from the least squares slope of reflectivity about wavelength in the processing window, the depth of the hydroxyl aluminum bond absorption band obtained from the continuous baseline and the continuous spectrum removal spectrum, and the band shape asymmetry are extracted; from the cross-modality inconsistency index calculation process, the combined value of effective atomic number measurement uncertainty and theoretical atomic number prediction uncertainty, i.e. combined uncertainty, is extracted; these four parameters are combined in order to form a state vector representing the state characteristics of sample multi-modal detection.
[0059] Specifically, when constructing the design variable vector based on short-wave infrared acquisition time, X-ray fluorescence repeated acquisition times and point micro-cleaning frequency, the adjustable operation parameters of short-wave infrared acquisition time, X-ray fluorescence repeated acquisition times and point micro-cleaning frequency in the acquisition and preprocessing links are combined in order to form a design variable vector, which is used for subsequent modeling of the influence of operation parameters on cross-modality inconsistency.
[0060] Specifically, the real evaluation has high cost and long time delay from collection to processing to comparison. The data-driven proxy model can approximately map the state and action to the expected value of the cross-modal inconsistency index, and suppress overfitting through cross-validation and error evaluation. When training the proxy model based on the historical sample set containing the state vector, the design variable vector, and the cross-modal inconsistency index, multiple sets of historical sample data are collected, each set containing the state vector, the design variable vector, and the calculated cross-modal inconsistency index of the corresponding sample. After these data are combined into a training data set, a machine learning model such as Gaussian process regression or random forest is selected as the proxy model, with the state vector and the design variable vector as input and the cross-modal inconsistency index as the prediction target. The model parameters are optimized by minimizing the mean square error of the predicted value and the actual value, so that the trained proxy model can predict the cross-modal inconsistency index based on the input state and design parameters. Fast and low-cost online evaluation and optimization can be achieved, and the stability and repeatability of the decision can be improved by combining model uncertainty, which promotes cross-batch migration.
[0061] Within the value range of the design variable vector, the design variable vector that minimizes the output of the proxy model is solved to obtain the optimal retest configuration; According to the optimal retest configuration, short-wave infrared reflectance spectroscopy and X-ray fluorescence spectrum retest are performed at the same physical point of the current sample, and the cross-modal inconsistency index and the significance score are recalculated; If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the state vector of this round, the optimal retest configuration, and the updated cross-modal inconsistency index are added to the historical sample set and the iteration is continued; Specifically, the proxy model takes the state vector and the design variable as input and the expected value of the cross-modal inconsistency index as output. Within the feasible value range, the operation combination that is expected to reduce the cross-modal inconsistency index by the largest margin can be prioritized under resource constraints. When solving the design variable vector that minimizes the output of the proxy model to obtain the optimal retest configuration, numerical optimization algorithms such as sequential quadratic programming or genetic algorithms are used to search for the design variable vector within the predetermined value range of the short-wave infrared acquisition time, the X-ray fluorescence repeated acquisition times, and the point position micro-cleaning times. The proxy model will predict the cross-modal inconsistency index based on the input design variable vector. The optimization algorithm adjusts the value of the design variable vector to find the design variable combination that minimizes the predicted value of the cross-modal inconsistency index output by the proxy model. This combination is the optimal retest configuration.
[0062] Specifically, the same physical point retest can eliminate the interference of spatial heterogeneity, ensure that the change mainly comes from configuration adjustment rather than sampling difference, and the cross-modal inconsistency index and significance score must be recalculated with the same measurement chain after retest to obtain the real effect of action. According to the optimal retest configuration, the short-wave infrared reflectance spectrum and the X-ray fluorescence spectrum are retested at the same physical point of the current sample, and the cross-modal inconsistency index and the significance score are recalculated. When the cross-modal inconsistency index and the significance score are recalculated, the spectrometer is set to the acquisition time length according to the short-wave infrared acquisition time set in the optimal retest configuration, the X-ray fluorescence spectrum is acquired multiple times according to the X-ray fluorescence repeated acquisition times, and the sample detection point is micro-cleaned according to the point micro-cleaning times; the short-wave infrared reflectance spectrum and the X-ray fluorescence spectrum obtained by retest are recalculated according to the previous process to obtain the effective atomic sequence measured value, the theoretical atomic sequence predicted value and the combined uncertainty of the two, and then the updated cross-modal inconsistency index is obtained; based on the updated cross-modal inconsistency index, the significance score of the sample is recalculated according to the calculation method of the significance score.
[0063] Specifically, the adaptive quantile threshold represents the data baseline of the current batch, and the significance score not greater than the threshold indicates that the deviation has been within the normal fluctuation range, and the iteration can be stopped; if the threshold is not met, the state, action and result triplets of the current round are added to the historical sample set, so that the subsequent solution can continue to optimize on the data closer to the current working condition. When the iteration termination condition is judged, the updated significance score is compared with the adaptive quantile threshold. If the updated significance score is not greater than the adaptive quantile threshold, it indicates that the consistency of the cross-modal data has met the requirements, and the iteration process is terminated at this time, and the current optimal retest configuration and the updated cross-modal inconsistency index are output; if the updated significance score is greater than the adaptive quantile threshold, it indicates that the cross-modal consistency still needs to be optimized, and at this time the state vector of the current round including the granularity proxy quantity, the depth of the absorption band containing the aluminum hydroxyl bond, the band shape asymmetry, the combined uncertainty, the optimal retest configuration and the updated cross-modal inconsistency index are added to the historical sample set, and then the step of solving the optimal retest configuration in the value range of the design variable vector is returned to continue iteration optimization until the termination condition is met.
[0064] The component unmixing module is used for component unmixing and abnormality judgment of the short-wave infrared reflectance spectrum based on the optimized cross-modal inconsistency index; In the embodiments of the present application, the component unmixing and abnormality judgment of the short-wave infrared reflectance spectrum based on the optimized cross-modal inconsistency index comprises: The prior component vector is constructed according to the clay proportion estimation value and the prior feldspar proportion; The short-wave infrared reflectance spectrum and the end-member spectrum corresponding to the end-member set are aligned at the same wavelength sampling point; The end member unmixing is performed by taking the sum of the reconstruction error between the short-wave infrared reflectance spectrum and the weighted sum of the end member spectra and the deviation of the prior component vector as the target, to obtain the final value of the component vector. Specifically, the clay proportion estimation value directly gives the physical share of the clay end member, and the non-clay part is easily indistinguishable when information is insufficient. The feldspar and quartz have similar band shapes and large noise, and the non-clay part can be further divided into feldspar and quartz by the prior feldspar proportion. The clay proportion estimation value is taken as the proportion of the clay end member in the component vector, the prior feldspar proportion is taken as the proportion of the feldspar end member in the component vector, and the proportion of the quartz end member in the component vector is obtained by calculating 1 minus the sum of the clay proportion estimation value and the prior feldspar proportion. The three proportions are combined in the order of clay, feldspar and quartz to form a prior component vector with an element sum of 1, which is used as a prior constraint basis for subsequent end member unmixing. The distinguishability and stability of unmixing are improved, and the degree of freedom and overfitting risk are reduced.
[0065] Specifically, the end member unmixing requires that the observation vector and the end member matrix are at the same sampling grid point. If the sampling grid is inconsistent, the interpolation error and displacement error will be wrongly attributed to the component difference. When aligning the end member spectra corresponding to the short-wave infrared reflectance spectrum and the end member set at the same wavelength sampling point, first determine the wavelength sampling range and resolution of the short-wave infrared reflectance spectrum of the current geological sample, and then extract the standard spectra of the clay, feldspar and quartz end members in the end member set. Interpolate or resample the end member spectra to completely match the wavelength sampling points of the end member spectra and the wavelength sampling points of the short-wave infrared reflectance spectrum of the current geological sample, ensure one-to-one correspondence between the two at each sampling point in the wavelength dimension, eliminate the pseudo-difference introduced by inconsistent sampling, ensure that the reconstruction error only reflects the component deviation rather than the numerical alignment problem, and improve the unmixing accuracy.
[0066] Specifically, the reconstruction error term ensures the fitting of the observation data, and the prior deviation term introduces the prior component vector formed in the last step as a soft constraint to provide directional guidance when the noise, granularity scattering or band shape is similar. When the endmember unmixing is performed by taking the sum of the reconstruction error between the short-wave infrared reflectance spectrum and the endmember spectrum weighted sum and the prior component vector deviation as the target, a target function is first constructed: the first part is the reconstruction error of the short-wave infrared reflectance spectrum and the endmember spectrum weighted sum, and the Euclidean distance is used to calculate, that is, for each wavelength sampling point, the actual reflectance spectrum value is subtracted from the predicted value obtained by weighting and summing each endmember spectrum according to the component proportion, and then the square sum of the difference values of all sampling points is calculated; the second part is the prior component vector deviation, and the Euclidean distance is used to calculate the square sum of the difference between the current to-be-solved component vector and the prior component vector; a regularization coefficient is introduced, which can be adaptively adjusted according to the size of the cross-modal inconsistency index, and the larger the cross-modal inconsistency index, the higher the weight of the prior component vector; the two parts are combined into the total target function in the form of the reconstruction error plus the regularization coefficient multiplied by the prior component vector deviation; through an optimization algorithm with the constraints of non-negativity and sum of 1 of the elements of the component vector, such as a constrained least squares method, the component vector that minimizes the total target function is solved, which is the final value of the component vector.
[0067] A residual error sequence is obtained according to the difference between the short-wave infrared reflectance spectrum and the endmember spectrum weighted sum; The in-band average residual and the in-band variance are calculated based on the residual error sequence, and the abnormality is determined according to the optimized significance score and the quantile threshold as a threshold to obtain an abnormal label; Specifically, the residual error sequence is a point-by-point difference sequence formed by subtracting the reconstructed reflectance obtained by weighting the endmember spectrum according to the final value of the component vector from the observed reflectance of the short-wave infrared reflectance spectrum at the same wavelength sampling point. The sequence corresponds to the wavelength sequence one by one, and the numerical value is the positive or negative deviation on the reflectance scale, reflecting the wavelength-by-wavelength matching degree of the model to the observation spectrum.
[0068] Specifically, the reconstructed spectrum obtained by weighting and summing the endmember spectra according to the final values of the component vectors represents the expected response of the model under the endmember proportion. Subtracting the observed short-wave infrared reflectance spectrum from the reconstructed spectrum at each wavelength can convert the insufficient fitting and unmodeled effects into a measurable error sequence. When obtaining the residual error sequence by subtracting the weighted sum of the endmember spectra from the observed short-wave infrared reflectance spectrum at each wavelength, the final values of the component vectors obtained after endmember unmixing are obtained, and the spectrum of each endmember is weighted and summed according to the corresponding proportion in the component vector to obtain the weighted sum of the endmember spectra. Subtracting the reflectance value of the short-wave infrared reflectance spectrum at each wavelength sampling point from the reflectance value of the weighted sum of the endmember spectra at the corresponding wavelength sampling point, and arranging the difference values of all wavelength sampling points in order to form the residual error sequence. Decomposing the unmixing quality into a wavelength-by-wavelength deviation from the overall error facilitates locating the problem source near the target absorption band, and provides a directly traceable data basis for subsequent in-band statistics and anomaly judgment.
[0069] Specifically, single-point residuals are susceptible to noise, in-band average residuals are used to depict systematic deviations, and in-band variances are used to depict structural fluctuations, which can stably reflect mechanism-based anomalies of the target absorption band. Combining the statistics with the optimized significance score and the adaptive quantile threshold can automatically calibrate the judgment according to the uncertainty level of the current batch. After generating the residual error sequence, the wavelength set of the target absorption band is determined according to the aforementioned processing window, which is defined as the in-band region. The in-band average residual is obtained by calculating the arithmetic mean of the residual error sequence at each point in the in-band region, and the in-band variance is obtained by calculating the square average of the in-band residual centered on the in-band average residual. To avoid the interference of brightness and sampling density differences of different samples, the residuals are first normalized by using a continuous baseline before calculating the above statistics, so that the expected value of the non-absorption region is close to zero, thereby ensuring that the statistics can be compared across samples. The in-band average residual and the in-band variance are used as the localized residual measure, and the optimized significance score and the adaptive quantile threshold are simultaneously called as the judgment threshold. When the optimized significance score is greater than the adaptive quantile threshold and the absolute value of the in-band average residual or the in-band variance is within the abnormal range of the empirical distribution of the current batch, it is determined that the current sample has a cross-modal inconsistency anomaly, and an abnormal label is output. When the above conditions do not hold, a normal label is output. The judgment has both statistical basis and physical directionality under a unified metrology framework, and reduces the over-detection or missed detection caused by fixed thresholds.
[0070] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A machine learning based geologic sample composition identification and anomaly detection system, comprising: The method comprises the following steps: a spectrum acquisition module for acquiring the X-ray fluorescence spectrum and the short-wave infrared reflectance spectrum of the geological sample; an atomic number metrology module for obtaining the measured value and the measurement uncertainty based on the X-ray fluorescence spectrum; a band parameter extraction module for preprocessing the short-wave infrared reflectance spectrum to obtain the spectral characteristic parameters of the absorption band and the particle size proxy; a clay proportion fitting module for regression fitting the spectral characteristic parameters of the absorption band and the particle size proxy to obtain the clay proportion estimate value; an end member weight calculation module for constructing the end member mixing weight according to the clay proportion estimate value and the particle size proxy, and obtaining the prediction value and the prediction uncertainty through the end member atomic number table; a cross-modal inconsistency module for calculating the combined uncertainty based on the measured value and the measurement uncertainty, and calculating the cross-modal inconsistency index according to the combined uncertainty, the prediction value and the prediction uncertainty; an optimization module for optimizing the cross-modal inconsistency index based on the particle size proxy, the spectral characteristic parameters of the absorption band and the combined uncertainty; a component unmixing module for component unmixing and anomaly judgment of the short-wave infrared reflectance spectrum based on the optimized cross-modal inconsistency index. 2.The machine learning based geological sample composition identification and anomaly detection system of claim 1, wherein, The method comprises the following steps: baseline correction and peak shape fitting are performed on the X-ray fluorescence spectrum to obtain the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak; a monotonic calibration mapping relationship is established according to the standard sample with known average atomic number; the ratio of the integral intensity of the Compton scattering peak to the integral intensity of the Rayleigh scattering peak is converted through the monotonic calibration mapping relationship to obtain the measured value of the effective atomic number; error propagation is performed on the integral intensity of the Compton scattering peak and the integral intensity of the Rayleigh scattering peak according to the covariance of the peak shape fitting to obtain the measurement uncertainty of the effective atomic number. 3.The machine learning based geological sample composition identification and anomaly detection system of claim 1, wherein, The method comprises the following steps: the first and second derivatives of the reflectivity of the short-wave infrared reflectance spectrum with respect to wavelength are obtained through adjacent difference; the local minimum point where the first derivative changes from negative to positive and the second derivative is positive is taken as the candidate point of the absorption band center of the aluminum bond containing hydroxyl group; a processing window is generated according to the candidate point of the absorption band center of the aluminum bond containing hydroxyl group; in the processing window, the extreme points of the reflectivity at both ends of the processing window are connected to form a continuous baseline through polynomial fitting; the original reflectivity in the processing window is divided by the baseline reflectivity of the corresponding waveband to obtain the continuous uniformity removal spectrum; the depth of the absorption band of the aluminum bond containing hydroxyl group is obtained according to the difference in reflectivity at the center of the processing window between the continuous baseline and the continuous uniformity removal spectrum; the integral area from the left boundary to the center of the processing window and the integral area from the right boundary to the center of the window are calculated, and the band shape asymmetry of the absorption band of the aluminum bond containing hydroxyl group is obtained according to the ratio of the two areas; in the processing window, the slope of the reflectivity with respect to wavelength is calculated through the least square method, and the slope is taken as the particle size proxy. 4.The machine learning based geological sample composition identification and anomaly detection system of claim 3, wherein, The method comprises the following steps: The actual clay proportion of each sample in the calibration sample set is determined by experiment, and the depth, band shape asymmetry and granularity proxy of the hydroxyl aluminum bond absorption band of the sample are calculated by continuum removal and spectral slope calculation on the short wave infrared reflectance spectrum of the sample; The depth, band shape asymmetry and granularity proxy of the hydroxyl aluminum bond absorption band of each sample in the calibration sample set are subjected to outlier sample elimination and normalization processing; A monotonic spline mapping model is constructed, taking the actual clay proportion of the calibration sample set as the dependent variable, the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band of the calibration sample set as the independent variable, and the granularity proxy of the calibration sample set as the adjusting independent variable, and the depth and band shape asymmetry of the hydroxyl aluminum bond absorption band are subjected to monotonicity constraint of non-negative derivative in the model; The residual minimum square error of the clay proportion estimation value and the actual clay proportion value of the sample is minimized as the target to train the model parameters, and the number of spline knots and the penalty strength are selected by cross-validation; The distribution of the training residual is tested to determine whether the model violates the monotonicity constraint, and if there is a local violation, the node density is automatically adjusted and retrained until the monotonicity constraint is satisfied; The depth, band shape asymmetry and granularity proxy of the hydroxyl aluminum bond absorption band of the geological sample are normalized and input into the monotonic spline mapping model to obtain the clay proportion estimation value. 5.The machine learning based geological sample composition identification and anomaly detection system of claim 4, wherein, The endmember mixing weight is constructed according to the clay proportion estimation value and the granularity proxy, and the prediction value and prediction uncertainty are obtained from the endmember atomic number table, including: Obtain the endmember set and the atomic number table of the endmember, wherein the endmember set fixedly contains: clay, feldspar and quartz; Based on the clay proportion estimation value and the endmember set, a basic weight vector is constructed, wherein the basic weight of the clay endmember is equal to the clay proportion estimation value, and the basic weight of the non-clay part is equal to one minus the clay proportion estimation value; The non-clay part is divided into feldspar basic weight and quartz basic weight based on the prior feldspar proportion obtained by statistical analysis of historical samples; The clay basic weight and the non-clay basic weight are adjusted according to the granularity proxy, wherein the feldspar basic weight and the quartz basic weight are adjusted by a common amplitude of granularity; The clay basic weight, the feldspar basic weight and the quartz basic weight after granularity adjustment are normalized and combined to obtain the endmember mixing weight; The endmember mixing weight and the endmember atomic number table are weighted to obtain the prediction value of the theoretical atomic number; Based on the variance and covariance of the endmember mixing weight, error propagation is performed to obtain the prediction uncertainty of the theoretical atomic number.
6. The machine learning based geological sample composition identification and anomaly detection system of claim 1, wherein, The synthetic uncertainty is calculated based on the measured value and the measurement uncertainty, and the cross-modal inconsistency index is calculated based on the synthetic uncertainty, the prediction value and the prediction uncertainty, including: The square of the effective atomic number measurement uncertainty and the square of the theoretical atomic number prediction uncertainty are added, and the arithmetic square root of the sum is taken to obtain the synthetic uncertainty; The difference between the measured value of the effective atomic number and the prediction value of the theoretical atomic number is normalized based on the synthetic uncertainty to obtain the cross-modal inconsistency index; The current sample is resampled to construct an empirical empty distribution, and the significance score is obtained according to the exceedance ratio of the cross-modal inconsistency index in the empirical empty distribution; An adaptive quantile threshold is determined according to the significance score of the current batch of samples.
7. The machine learning based geological sample composition identification and anomaly detection system of claim 6, wherein, The cross-modal inconsistency index is optimized based on the granularity proxy, spectral characteristic parameters of the absorption band, and synthetic uncertainty, including: The value range of the short-wave infrared acquisition time, the X-ray fluorescence repeated acquisition times, and the point micro-cleaning times is determined; The state vector is constructed based on the granularity proxy, the depth of the absorption band containing the aluminum-hydroxyl bond, the band shape asymmetry, and the synthetic uncertainty; The design variable vector is constructed based on the short-wave infrared acquisition time, the X-ray fluorescence repeated acquisition times, and the point micro-cleaning times; The agent model is trained based on the historical sample set containing the state vector, the design variable vector, and the cross-modal inconsistency index; The design variable vector that minimizes the output of the agent model is solved within the value range of the design variable vector, and the optimal retest configuration is obtained; The short-wave infrared reflectance spectrum and the X-ray fluorescence spectrum are retested at the same physical point of the current sample according to the optimal retest configuration, and the cross-modal inconsistency index and the significance score are recalculated; If the updated significance score is not greater than the adaptive quantile threshold, the iteration is terminated, and the optimal retest configuration and the updated cross-modal inconsistency index are output; otherwise, the state vector, the optimal retest configuration, and the updated cross-modal inconsistency index of this round are added to the historical sample set, and the iteration is continued.
8. The machine learning based geological sample composition identification and anomaly detection system of claim 7, wherein, Based on the optimized cross-modal inconsistency index, the short-wave infrared reflectance spectrum is unmixing and anomaly judgment, including: The prior component vector is constructed according to the clay proportion estimate and the prior feldspar proportion; The short-wave infrared reflectance spectrum and the end-member spectrum corresponding to the end-member set are aligned at the same wavelength sampling points; The end-member unmixing is performed to obtain the final value of the component vector, with the reconstruction error between the weighted sum of the short-wave infrared reflectance spectrum and the end-member spectrum and the sum of the deviation of the prior component vector as the target; The unmixing residual sequence is obtained according to the difference between the short-wave infrared reflectance spectrum and the weighted sum of the end-member spectrum; The in-band average residual and the in-band variance are calculated based on the unmixing residual sequence, and the anomaly judgment is performed according to the optimized significance score and the quantile threshold as the threshold.
Citation Information
Patent Citations
Tunnel rock mineral identification method and system cooperating with multi-element spectrum
CN117828397A
Cultivated land soil environment detection method and system
CN119375177A
Coal gangue AI identification method and system based on infrared spectrum measurement
CN119688637A
Method for realizing carbon powder purity intelligent analysis by using deep learning
CN119961572A