Mining compound structure intelligent inversion system and method fused with spectroscopic big data
The intelligent inversion system for the structure of mineral compounds, which integrates spectroscopic big data, has achieved standardized acquisition and processing of full-band spectroscopic data. By combining multi-dimensional databases and structure-property relationship training through machine learning and deep learning, it has solved the problems of insufficient data quality and verification basis in the inversion of mineral compound structures, and improved the accuracy and intelligence of the inversion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MASCH NUMBER INSTR (ZHEJIANG) CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies for inverting the structure of mineral compounds suffer from problems such as poor data quality, one-sided feature learning, and lack of verification basis, resulting in inaccurate inversion results and insufficient intelligence.
The intelligent inversion system for the structure of mineral compounds, which integrates spectroscopic big data, includes a chemical spectrum data acquisition module, a database construction module, a model training module, and an inversion module. It optimizes the model by acquiring full-band spectroscopic data, standardizing preprocessing, constructing a multi-dimensional database, training the structure-property relationship by integrating machine learning and deep learning, and combining laboratory test data.
It improves the efficiency, accuracy, and adaptability of mineral compound structure inversion, meets the practical application needs of online detection in green mines, and ensures the reliability and intelligence of inversion results.
Smart Images

Figure CN121963935A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of compound structure inversion technology, specifically to an intelligent inversion system and method for the structure of mining compounds that integrates spectroscopic big data. Background Technology
[0002] In the scenario of mineral compound structure inversion, mineral samples cover three states: solid, liquid, and gas. Operating parameters such as temperature, humidity, and dust concentration in mining areas fluctuate significantly. Spectroscopic data are easily affected by system noise, baseline drift, and laser energy fluctuations. Furthermore, there is a complex nonlinear relationship between elemental characteristic spectra and compound structures. Existing technologies reveal the following shortcomings: First, existing technologies lack adaptive correction mechanisms for operating condition fluctuations, and interference such as noise and baseline drift in the raw data are not effectively removed, resulting in low-quality basic data used for inversion and difficulty in accurately reflecting the true spectral characteristics of compounds. Second, existing technologies mostly use a single algorithm model for structure-activity relationship training, which cannot comprehensively explore the linear correlation and local nonlinearity between spectral features and structural parameters, nor can it capture long-range dependencies across parameters. This one-sided feature learning leads to incomplete inversion logic. Finally, existing technologies lack unified comparison benchmarks and quantitative verification criteria in the inversion process, and the validity of the inversion results depends on subjective experience, which is prone to misjudgment of structural parameters or deviations in component content estimation.
[0003] Therefore, there is an urgent need for an intelligent inversion system and method that integrates multi-step standardized data processing, working condition adaptation correction, multi-algorithm collaborative training, and multi-dimensional database support to solve the problems of poor data quality, one-sided feature learning, and lack of verification basis in existing technologies, and to improve the intelligence and accuracy of mineral compound structure inversion. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an intelligent inversion system and method for the structure of mineral compounds that integrates spectroscopic big data, solving the problems of insufficient intelligence and accuracy in the inversion of the structure of mineral compounds.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data, comprising: The spectral data acquisition module is used to acquire full-band spectral data of mining samples and perform standardized preprocessing to obtain standardized spectral data.
[0006] The database construction module is used to integrate standardized spectroscopic data with spectroscopic theoretical simulation data to construct a spectroscopic standard database. The spectroscopic standard database includes a molecular spectral standard data sub-database, a compound structural spectral data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation association data sub-database.
[0007] The model training module is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components based on a standard spectroscopic database and by combining machine learning and deep learning algorithms, and outputs an initial structure-activity relationship model between spectra and compound structures.
[0008] The inversion module is used to input standardized spectroscopic data acquired in real time into the structure-activity relationship model, call the spectroscopic standard database to compare molecular spectral standard data, perform spectral feature analysis and compound structure inversion, and output the inversion results of the target compound.
[0009] The model optimization module is used to calculate the deviation between the inversion results and the laboratory test data, and to feed the deviation results back to the structure-activity relationship model to optimize the structure-activity relationship model.
[0010] A smart inversion method for the structure of mineral compounds integrating spectroscopic big data includes the following steps: Full-band spectroscopic data were acquired from the mineral samples, and standardized preprocessing was performed to obtain standardized spectroscopic data.
[0011] Standardized spectroscopic data is integrated with spectroscopic theoretical simulation data to construct a spectroscopic standard database, which includes a molecular spectral standard data sub-database, a compound structural spectroscopic data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation correlation data sub-database.
[0012] Based on a standard spectroscopic database, a fusion of machine learning and deep learning algorithms is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components, and output an initial structure-activity relationship model between spectra and compound structures.
[0013] The standardized spectroscopic data collected in real time are input into the structure-activity relationship model, and the standard spectroscopic database is called to compare the molecular spectral standard data, analyze the spectral features and invert the compound structure, and output the inversion results of the target compound.
[0014] The deviation between the inversion results and the laboratory test data is calculated, and the deviation results are fed back to the structure-activity relationship model to optimize the structure-activity relationship model.
[0015] The present invention has the following beneficial effects: This invention achieves standardized acquisition of full-band spectroscopic data from mineral samples of different morphologies through a chemical spectrum data acquisition module. Combined with a database construction module, it integrates standardized spectroscopic data with spectroscopic theoretical simulation data to form a multi-dimensional standard spectroscopic database. Utilizing a structure-property relationship model and inversion module that integrates machine learning and deep learning, it completes spectral feature analysis and intelligent inversion of compound structures. Furthermore, a model optimization module optimizes the model based on feedback from laboratory test data, improving the efficiency, accuracy, and adaptability of mineral compound structure inversion. This meets the practical application needs of online detection in green mines and solves the problems of insufficient intelligence and accuracy in mineral compound structure inversion.
[0016] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0017] Figure 1 This is a flowchart of the intelligent inversion system for the structure of mineral compounds that integrates spectroscopic big data, as described in this invention.
[0018] Figure 2 This is a flowchart of the intelligent inversion method for the structure of mineral compounds that integrates spectroscopic big data, as described in this invention. Detailed Implementation
[0019] Please see Figure 1 The present invention provides a technical solution: an intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data, comprising: The spectral data acquisition module is used to acquire full-band spectral data of mining samples and perform standardized preprocessing to obtain standardized spectral data.
[0020] The specific process is as follows: A specialized sample-bearing device adapted to solid, liquid, and gaseous mineral forms is used to position and fix mineral samples of different types. Solid mineral samples are fixed using a wear-resistant sample stage with anti-slip positioning grooves; liquid mineral samples are guided through a transparent, corrosion-resistant flow cell with built-in temperature control components and a flow rate regulating pump; and gaseous mineral samples are purified and guided through a sealed flow chamber with a pressure balancing valve and a dust removal filter. Simultaneously, a condition sensing unit collects real-time environmental parameters such as temperature, humidity, and dust concentration in the mining area, outputting the adapted mineral sample to be tested and its corresponding real-time parameters.
[0021] Based on real-time operating parameters, the parameters of the laser emission module corresponding to the laser-induced breakdown spectroscopy technology are adjusted. The laser emission module adopts a pulse width of 10 fs-10 ns and an energy density of 10. 6 -10 9The W / cm² ultrashort pulse laser generator automatically switches between 1064nm, 532nm, or 266nm wavelengths according to the matrix composition of the mineral sample to be tested. The ultrashort pulse laser is focused onto the surface of the mineral sample after being adapted through the optical path collimation component, which excites the material on the sample surface to produce an ablation effect and form a plasma plume.
[0022] Wavelength and intensity spectra of the plasma plume were acquired to obtain raw spectroscopic data. A high-resolution spectrometer with a spectral resolution ≤0.01nm and a detection range covering the entire ultraviolet-visible-near-infrared band of 190nm-1100nm was used for acquisition. The characteristic spectral line signals of all elements emitted by atomic and ion transitions during plasma cooling were captured at an acquisition frequency of 100 points / second. The spectral line signals were transmitted to the data buffer module through a high-speed data transmission interface. Sample information, operating parameters and acquisition timestamps were stored in binary format, and a two-dimensional matrix containing wavelength and intensity was output.
[0023] The original spectroscopic data were decomposed into wavelets at scales of 3-5, and after noise removal, baseline correction was performed to obtain the baseline-corrected spectroscopic data.
[0024] Wavelet decomposition selects the db4 wavelet as the decomposition basis function. The high-frequency coefficients after decomposition are processed by soft thresholding (the threshold is 1.5-2.0 times the standard deviation of the original spectral data) to remove system noise and random noise. The processed high-frequency coefficients and low-frequency coefficients are then reconstructed to output the denoised spectral data.
[0025] The background wavelength range without characteristic peaks is selected as the baseline reference range. The baseline drift curve is generated by fitting a 3rd to 5th order polynomial. The intensity value of each wavelength point in the denoised spectral data is subtracted from the corresponding baseline drift curve value to eliminate baseline drift interference and output the baseline-corrected spectral data.
[0026] The baseline-corrected spectroscopic data were normalized and outlier removed to obtain standardized spectroscopic data.
[0027] The specific process is as follows: using the characteristic spectral peaks of the matrix elements in the baseline-corrected spectroscopic data as the standard reference peaks, the peak intensity of the reference peaks is calculated as the normalization reference value. The intensity values of all wavelength points in the baseline-corrected spectroscopic data are divided by the normalization reference value to eliminate the influence of excitation efficiency differences and laser energy fluctuations, and the normalized spectroscopic data is output.
[0028] The 3σ criterion is used to identify outliers in normalized spectroscopic data. The mean μ and standard deviation σ of the intensity values at each wavelength point in the normalized spectroscopic data are calculated. Intensity values outside the range of [μ-3σ, μ+3σ] are identified as outliers. Outliers are filled in by linear interpolation (the interpolation is based on the mean intensity of the three consecutive wavelength points before and after the outlier). Finally, normalized spectroscopic data in the form of a two-dimensional effective data matrix of wavelength-normalized intensity is output.
[0029] By defining a standardized process for acquiring and preprocessing spectroscopic data, the laser emission module parameters are first adjusted to excite the sample and form a plasma plume. Then, the raw data is acquired through high-resolution acquisition. Subsequently, noise is removed by wavelet decomposition, baseline is corrected by polynomial fitting, matrix peaks are normalized, and outliers are removed by the 3σ criterion. This systematically eliminates interference factors such as system noise, baseline drift, and laser energy fluctuations, ensuring the accuracy, consistency, and effectiveness of the spectroscopic data. This provides high-quality, standardized basic data support for the subsequent construction of a standard spectroscopic database, training of structure-activity relationship models, and compound structure inversion, avoiding the negative impact of low-quality data on subsequent processes.
[0030] The database construction module is used to integrate standardized spectroscopic data with spectroscopic theoretical simulation data to construct a spectroscopic standard database. The spectroscopic standard database includes a molecular spectral standard data sub-database, a compound structural spectral data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation association data sub-database.
[0031] The specific process is as follows: Spectroscopic simulation data based on quantum chemical calculations were obtained and then fused with standardized spectroscopic data after format unification, association and annotation of sample and simulation information, and removal of redundant and invalid entries to obtain a fused data set.
[0032] Historical operating condition parameters are retrieved, and the correlation between operating condition parameters and historical feature parameters is extracted. A mapping model between operating conditions and operating condition adaptation correction parameters is established, and the data is classified and stored to form a three-dimensional correlation dataset. The operating condition adaptation correlation data sub-library is output. The historical feature parameters include historical molecular spectral feature parameters and historical compound structural parameters. Specifically: Historical monitoring data from the mining area is retrieved, and historical operating parameters (including temperature, humidity, dust concentration, and pressure) and corresponding historical characteristic parameters (historical molecular spectral characteristic parameters: characteristic peak wavelength, peak intensity, peak area, and full width at half maximum; historical compound structural parameters: bond length, bond angle, functional group type, molecular configuration, and crystal structure parameters) are extracted. These are then associated with the corresponding mineral types, sample morphology, and data acquisition timestamps to form the original historical dataset. Abnormal data in the original historical dataset is identified and removed using the 3σ criterion, and missing data is filled using linear interpolation. The historical operating parameters and historical characteristic parameters are normalized to eliminate dimensional differences, resulting in a standardized historical dataset.
[0033] Based on a standardized historical dataset, sensitive characteristic parameters that are sensitive to changes in operating conditions are selected (by analyzing the coefficient of variation of characteristic parameters in different operating condition ranges, molecular spectral characteristic parameters and compound structural parameters with coefficients of variation greater than a set threshold are selected as sensitive characteristic parameters). Pearson correlation analysis and Spearman rank correlation analysis are used to calculate the correlation strength between each historical operating condition parameter and the sensitive characteristic parameter, constructing a correlation strength matrix. Based on the correlation strength matrix, operating condition-sensitive characteristic parameter combinations with correlation strength greater than a set threshold are selected, clarifying the influence trend (positive / negative correlation) and influence weight of different operating condition parameters on the sensitive characteristic parameter, and outputting a set of correlation patterns between operating conditions and sensitive characteristic parameters.
[0034] Multiple regression, gradient boosting tree, or neural network algorithms are selected as the core algorithms for the mapping model. Standardized historical operating condition parameters are used as the input vector, and the deviation compensation amount of sensitive feature parameters is used as the target output vector to construct a mapping model between operating conditions and operating condition adaptation correction parameters. The operating condition adaptation correction parameters include spectral correction coefficients (including peak intensity correction coefficients and peak area correction coefficients) to compensate for the impact of operating condition changes on molecular spectral feature parameters, and structural correction parameters (including bond length deviation compensation values and bond angle deviation compensation values) to compensate for the impact of operating condition changes on compound structural parameters. The standardized historical dataset is divided into a training set and a validation set according to a set ratio. These sets are input into the mapping model for iterative training. By adjusting the model hyperparameters (such as the number of decision trees in the gradient boosting tree and the number of hidden layer nodes in the neural network), the mean square error between the model's predicted values and the actual deviation compensation amounts is minimized until the prediction accuracy of the validation set meets the set requirements. The trained mapping model is then output.
[0035] Based on the trained mapping model, standard operating parameters for different operating condition intervals (continuous operating condition intervals are divided according to the value range of the operating parameters) are input. The model calculates the corresponding spectral correction coefficients and structural correction parameters for each operating condition interval, forming a table of correspondence between operating conditions and correction parameters. The generation logic of the correction parameters is as follows: based on the correlation between operating conditions and sensitive feature parameters, the quantitative correspondence between changes in operating conditions and deviations in feature parameters is obtained through model fitting, and then the value of the correction parameter used to offset the deviation is derived in reverse.
[0036] Using the working condition range, correction parameter, and sensitive feature parameter type as a three-dimensional dimension, the working condition-correction parameter correspondence table and the working condition-sensitive feature parameter association rule set are merged to form a three-dimensional association dataset. The three-dimensional association dataset is classified and stored according to mineral type and sample morphology, and multi-dimensional indexes (working condition range index, mineral type index, and feature parameter type index) are established. Following preset metadata specifications, data sources, association rule basis, and model training accuracy information are labeled. A unified data retrieval and calling interface is constructed, and a working condition-adapted association data sub-library is output.
[0037] Molecular spectral feature parameters are extracted from the fused dataset, and deviation calibration and correction are performed based on the working condition adaptation correction parameters. The data are then classified and stored according to mineral type and molecular type. After data verification and removal of abnormal entries, a molecular spectral standard data sub-library is obtained.
[0038] Compound structural parameters are extracted from the fused dataset, corrected by operating condition adaptation, and then stored hierarchically after establishing parameter mapping relationships, resulting in a sub-library of compound structural spectroscopic data.
[0039] By integrating the molecular spectroscopy standard data sub-library, the compound structure spectroscopy data sub-library, the elemental characteristic spectroscopy data sub-library, and the operating condition adaptation association data sub-library, a multi-dimensional unified indexing system is constructed. After setting metadata specifications, establishing a dynamic update mechanism, and ensuring data security in accordance with the FAIR principle, a spectroscopy standard database is output.
[0040] The construction logic of the spectroscopic standard database was refined. By integrating standardized spectroscopic data with theoretical simulation data from quantum chemical calculations, the comprehensiveness of data sources was enriched. A mapping model between operating conditions and condition-adaptive correction parameters was established, achieving precise compensation for the impact of operating condition changes on characteristic parameters. The molecular spectroscopy standard data sub-database and the compound structure spectroscopic data sub-database were corrected and categorized for storage, ensuring data standardization and relevance. A multi-dimensional unified indexing system constructed by integrating multiple sub-databases improved the efficiency of data retrieval and retrieval. The spectroscopic standard database constructed by this scheme combines data integrity, scientific correction, and ease of use, providing a reliable data foundation for training structure-activity relationship models and a standardized basis for data comparison and verification during the inversion process.
[0041] The model training module is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components based on a standard spectroscopic database and by combining machine learning and deep learning algorithms, and outputs an initial structure-activity relationship model between spectra and compound structures.
[0042] The specific process is as follows: Extract elemental characteristic spectral parameters (including characteristic peak wavelength, peak intensity, peak area, and half-width at half-maximum) from the molecular spectral standard data library; extract corresponding compound structural parameters (including bond length, bond angle, functional group type, molecular configuration, and crystal structure parameters) from the compound structural spectral data library; and extract operating condition correction coefficients from the operating condition adaptation association data library.
[0043] A one-to-one correspondence between elemental characteristic spectral parameters and compound structural parameters is established based on the unique sample identifier. Deviation calibration is performed by combining the working condition correction coefficient. The data is divided into training set, validation set and test set in a stratified sampling method at a ratio of 7:2:1. Outliers and missing values in the dataset are removed, and training dataset, validation dataset and test dataset are output.
[0044] The machine learning module employs a parallel combination of random forest and support vector machine algorithms to uncover the linear correlation and local feature mapping patterns between elemental spectral parameters and compound structural parameters. The deep learning module utilizes a Long Short-Term Memory (LSTM) network structure to capture the sequence association features of compound structural parameters. A feature-level fusion interface is set up to concatenate the local feature vectors output by the machine learning module with the deep feature vectors output by the deep learning module, constructing a unified feature fusion matrix and outputting the machine learning and deep learning fusion algorithm framework.
[0045] Configure hyperparameters for the machine learning and deep learning fusion algorithm framework: For the Random Forest algorithm, set the number of decision trees to 100-200 and the maximum depth to 10-20 layers; for the Support Vector Machine algorithm, set the kernel function to radial basis function and the penalty coefficient C to a range of 0.1-10; for the LSTM network, set 2-4 hidden layers and 128-512 hidden units. Set the training batch size to 32-128, the learning rate to 0.001-0.01, and the number of iterations to 50-100 epochs; initialize the network weights to random values following a normal distribution.
[0046] Based on the random forest algorithm and the support vector machine algorithm, feature processing is performed on the corrected elemental spectral parameters and compound structural parameters in the training dataset to obtain local feature mapping vectors.
[0047] Based on the LSTM network, the corrected elemental spectral parameters and compound structural parameters in the training dataset are processed to obtain the structural time-series feature vector.
[0048] The local feature mapping vector and the structural temporal feature vector are concatenated to output a global fusion feature matrix.
[0049] The global fusion feature matrix is input into the fully connected layer, and the Sigmoid activation function is used to output the predicted values of the compound structure components. The error between the predicted values and the actual values of the compound structure components is calculated using the mean squared error as the loss function. The Adam optimizer is used for iterative training. When the error value of the validation dataset converges to below the set value, the number of neurons in the fully connected layer, the fusion weights, and the regularization parameter (L2 regularization coefficient 0.001) are adjusted to suppress overfitting, and the optimized fusion model is output.
[0050] The optimized fusion model is applied to the test dataset for performance evaluation. Once the performance meets the target, the model structure (random forest + SVM + LSTM fusion architecture), parameters (number of convolutional kernels, number of LSTM hidden units, fusion weights, etc.) and input / output format are solidified to obtain the initial structure-activity relationship model between spectroscopic and compound structures.
[0051] The training process of the structure-activity relationship (SPR) model was clearly defined. A stratified sampling method was used to rationally divide the dataset, ensuring the representativeness of the training, validation, and testing data. An algorithmic framework integrating machine learning and deep learning was adopted, utilizing random forests and SVMs to mine linear correlations and local features, while also leveraging LSTMs to capture sequence correlation features, achieving comprehensive learning of the structure-activity relationship between elemental spectral features and compound structural components. Overfitting was suppressed through fully connected layer optimization and regularization parameter adjustment, improving the model's generalization ability. The trained SPR model exhibits high prediction accuracy and strong stability, accurately characterizing the correspondence between spectral features and compound structures, providing core algorithmic support for subsequent target compound structure inversion and ensuring the reliability of the inversion results.
[0052] The process of performing feature processing on the corrected elemental spectral parameters and compound structural parameters in the training dataset based on the random forest algorithm and support vector machine algorithm to obtain local feature mapping vectors is as follows: The corrected elemental characteristic spectral parameters are input into the random forest algorithm, and the characteristic importance score of each parameter is calculated based on the out-of-bag data error.
[0053] Select the elemental feature spectral parameters whose feature importance scores are greater than a set score threshold, and output a subset of elemental feature spectral parameters.
[0054] The subset of elemental characteristic spectral parameters and the corrected compound structural parameters are concatenated column-wise according to the unique sample identifier to form a joint feature matrix.
[0055] The joint feature matrix and the corresponding compound structural component labels are input into the support vector machine algorithm. The joint feature matrix and labels are fitted and trained by minimizing structural risk. The linear correlation weights and local nonlinear feature mapping rules between core element features and corrected structural parameters are explored. The output is a 256-dimensional local feature mapping vector, which contains the linear correlation coefficients of each core parameter, the binary encoding of local features, and the nonlinear mapping feature values.
[0056] The feature processing of the machine learning module was refined. A random forest algorithm was used to calculate feature importance scores, identifying core spectral parameters that significantly influence the compound's structural composition, effectively eliminating redundant information and reducing computational complexity. The core parameters were concatenated with the compound's structural parameters to form a joint feature matrix. After fitting and training using an SVM algorithm, the linear correlation weights and local nonlinear mapping patterns between the core features and structural parameters were accurately discovered. The output 256-dimensional local feature mapping vector ensured both feature effectiveness and condensed feature representation, laying a high-quality foundation for subsequent fusion with features output from the deep learning module, and improving the overall model's training efficiency and the targeted nature of feature learning.
[0057] The process of obtaining the structural time-series feature vector by performing feature processing on the corrected elemental spectral parameters and compound structural parameters in the training dataset based on the LSTM network is as follows: Based on the unique identifier of the sample, a one-to-one correspondence is established between the corrected elemental characteristic spectral parameters (including characteristic peak wavelength, corrected peak intensity, corrected peak area, and corrected full width at half maximum) and the compound structural parameters (including corrected bond length, corrected bond angle, functional group type code, and corrected molecular configuration parameters), forming a joint parameter set for each sample.
[0058] The set of joint parameters is serialized and arranged in a set order to obtain fused sequence data.
[0059] Following the logical order of characteristic peak wavelength → corrected peak intensity → corrected peak area → corrected full width at half maximum (FWHM) → corrected bond length → corrected bond angle → functional group type encoding → corrected molecular configuration parameters, the spectral-structural joint parameter set of each sample is sequentially arranged. Each sequence has a fixed length of 64 (zero padding is used when there are fewer than 64 parameters, and the first 64 are truncated when there are more than 64 parameters). A fusion sequence data with a dimension of [number of samples × 64 × 8] (8 is the number of parameter types) is constructed. This data integrates the correlation information between spectral features and structural parameters, laying the foundation for temporal correlation mining.
[0060] The fused sequence data is input into an LSTM network (2 hidden layers, 256 hidden units per layer). The first LSTM layer uses a forget gate, a sigmoid activation function, and a threshold of 0.3 to filter and retain effective historical information (such as the correlation traces between characteristic peak wavelengths and subsequent bond lengths). The input gate updates the cell state to incorporate the current spectral-structural joint parameters. The cell state is updated by fusing historical and current information through a tanh activation function. The output gate controls the current output, and the output dimension is a primary fused temporal feature map of [number of samples × 64 × 256]. This feature map captures the local correlation patterns between spectral parameters and structural parameters in the sequence dimension.
[0061] The output primary fusion temporal feature map is input into the second-layer LSTM. Using the same gating mechanism, long-range dependencies across time steps are deeply mined (e.g., the chain reaction of characteristic peak intensity changes on molecular configuration via multiple parameters). The output dimension is [sample number × 64 × 256]. The output vector of the last time step of the second-layer LSTM (which aggregates the temporal correlation information of the entire spectrum-structure fusion sequence) is taken as the fusion temporal feature vector for that sample. The output dimension is [sample number × 256], which contains the local correlation weights between elemental characteristic spectra and compound structural parameters, cross-parameter long-range dependency encoding, and global cooperative variation characteristics.
[0062] The specific logic of LSTM network feature processing was clarified. By constructing spectral-structure fusion sequence data in a predetermined order, the organic integration of two types of key parameters was achieved, preserving the natural correlation between parameters. A two-layer LSTM network was used to process the fusion sequence data, capturing both local temporal correlations between parameters and uncovering long-range dependencies across time steps, comprehensively extracting the synergistic variation features of spectral and structural parameters. The output structural temporal feature vector enriches the feature dimensions, compensating for the limitations of single features in characterizing structure-activity relationships, and helping the fusion algorithm framework to more comprehensively learn the deep correlation between elemental spectra and compound structures, improving the feature representation capability and prediction accuracy of the structure-activity relationship model.
[0063] The inversion module is used to input standardized spectroscopic data acquired in real time into the structure-activity relationship model, call the spectroscopic standard database to compare molecular spectral standard data, perform spectral feature analysis and compound structure inversion, and output the inversion results of the target compound.
[0064] The specific process is as follows: Laser-induced breakdown spectroscopy (LAS) is used to emit ultrashort laser pulses with pulse widths of 10 fs to 10 ns to excite plasma on the surface of target mineral samples in the mining area, capturing the full-band ultraviolet-visible-near-infrared spectral signal of 190 nm to 1100 nm to obtain raw spectroscopic data. Noise is removed from the raw spectroscopic data using a db4 wavelet 3-5 scale decomposition wavelet transform denoising algorithm. Baseline correction is performed through 3-5 order polynomial fitting. Peak normalization is performed based on the characteristic peaks of the matrix elements. Outliers are identified using the 3σ criterion and completed using linear interpolation. Finally, standardized spectroscopic data in the form of a wavelength-normalized intensity two-dimensional matrix is output.
[0065] The system retrieves the spectroscopic standard database (including the molecular spectral standard data sub-database, the compound structural spectral data sub-database, and the working condition adaptation correlation data sub-database), and collects the current working condition parameters such as temperature, humidity, dust concentration, and pressure in the mining area in real time through the working condition sensing unit. Based on the working condition parameters, it matches the corresponding working condition interval in the working condition adaptation correlation data sub-database to obtain working condition adaptation correction parameters such as spectral intensity correction coefficient and structural parameter deviation compensation value.
[0066] The standardized spectroscopic data output from the previous step is calibrated using the operating condition adaptation correction parameters to obtain the standardized spectroscopic data after operating condition adaptation correction. The similarity between the standardized spectroscopic data after adaptation correction and the molecular spectral characteristic parameters such as characteristic peak wavelength, corrected peak intensity, and corrected peak area in the molecular spectral standard data sub-library is calculated. A similarity threshold of ≥90% is set to screen out the candidate spectral feature set that meets the matching degree, and the standardized spectroscopic data after operating condition adaptation correction and the candidate spectral feature set are output.
[0067] The candidate spectral feature set is input into the structure-activity relationship model. A random forest algorithm is used to perform secondary screening based on feature importance, retaining core spectral feature parameters such as peak wavelength, corrected peak intensity, corrected peak area, and corrected full width at half maximum (FWHM) with a feature importance score ≥ 0.7. These core spectral feature parameters are then input into an LSTM network, where a gating mechanism captures the cooperative variation patterns between parameters, outputting the initial inversion results for the target compound.
[0068] Extract corrected standard structural parameters (standard bond lengths, standard bond angles, standard crystal structure parameters) and standard component content ranges that match the target compound's structural type. Cross-validate the preliminary inversion results with a sub-library of compound structural spectroscopic data, and calculate the structural parameter deviations and relative errors of component content.
[0069] If the structural parameter deviation is no greater than 5% and the relative error of component content is no greater than 3%, the inversion result is deemed valid. It is then standardized and formatted to generate the inversion result, which is output in real-time through the green mine online detection and analysis system output interface. If the criteria are not met, the process returns to the previous steps to readjust the model feature screening threshold and analytical parameters, and the spectral feature analysis and compound structure inversion process is executed again.
[0070] The complete process of compound structure inversion has been refined. By real-time acquisition of operating parameters and matching corresponding operating condition adaptation correction parameters, operating condition deviation calibration of real-time spectral data has been achieved, improving the compatibility of data with standard databases. Candidate spectral feature sets with satisfactory matching are screened through similarity calculations, reducing the interference of invalid features on the inversion results. Core features are analyzed using a structure-activity relationship model to output initial inversion results, which are then cross-validated using a compound structure spectroscopic data sub-library to ensure the accuracy of the inversion results. The design of re-executing the inversion process when standards are not met further improves the reliability of the results. This process achieves standardized processing across the entire chain from data calibration to result output, balancing inversion efficiency and accuracy, and meeting the practical application needs of online detection of compound structures in the mining industry.
[0071] The process of cross-validating the preliminary inversion results with a sub-library of compound structure spectrometry data and calculating the structural parameter deviations and relative errors of component contents is as follows: Corrected standard structural parameters and standard component content ranges that are consistent with the target compound structure type (inorganic compound / organic compound / complex) in the preliminary inversion results are extracted from the compound structure spectroscopic database. The standard structural parameters include standard bond length, standard bond angle, and standard crystal structure parameters. The standard component content range is the industry standard or laboratory calibration effective range of the content of each component of this type of compound.
[0072] The inverted bond length, inverted bond angle, and inverted crystal structure parameters in the preliminary inversion results are selected as the structural parameters to be verified. They are matched one by one with the corresponding standard bond length, standard bond angle, and standard crystal structure parameters, and the single parameter deviation value of each structural parameter is calculated.
[0073] Single parameter deviation value = | Value of the structural parameter to be verified - Value of the corresponding standard structural parameter| / Value of the corresponding standard structural parameter × 100%.
[0074] The maximum value among all single-parameter deviation values is taken as the structural parameter deviation value of the target compound.
[0075] Extract the inversion content percentage of each component from the preliminary inversion results, use the median of the standard component content range as the standard content benchmark value, and calculate the single-component relative error of each component.
[0076] Single-component relative error = |Percentage of inverted content of a certain component - Benchmark value of standard content of that component| / Benchmark value of standard content of that component × 100%.
[0077] The maximum value among all individual component relative errors is taken as the relative error of the component content of the target compound.
[0078] The specific calculation methods for cross-validation, bias, and relative error were clarified. By extracting standard parameters and content ranges consistent with the target compound's structural type, the uniformity of the validation benchmark was ensured. A comprehensive and rigorous evaluation of the inversion results was achieved by calculating the single-parameter bias and single-component relative error through one-to-one matching and taking the maximum value as the judgment index. This scheme provides a quantitative basis and standardized process for determining the validity of the inversion results, avoids the uncertainty of subjective judgment, accurately identifies biases in the inversion results, ensures the accuracy and reliability of the output inversion results, and provides a clear direction for bias feedback for subsequent model optimization.
[0079] The model optimization module is used to calculate the deviation between the inversion results and the laboratory test data, and to feed the deviation results back to the structure-activity relationship model to optimize the structure-activity relationship model.
[0080] The specific process is as follows: Mineral samples from the same batch and under the same working conditions as the target compound were collected and inverted. Laboratory testing methods conforming to industry or national standards (such as X-ray diffraction, high-performance liquid chromatography, and elemental analysis) were used to determine the standard structural parameters (including standard bond length, standard bond angle, standard functional group type, standard molecular configuration, and standard crystal structure parameters) and standard component content data (including mass fraction and mole fraction of each component) of the target compound. Environmental working condition parameters during the testing process were recorded to form laboratory testing data, which served as the benchmark for deviation calculation.
[0081] Based on the unique sample identifier, a one-to-one correspondence is established between the inversion results and the laboratory standard test dataset, and the deviation values of structural parameters and the relative errors of component content are calculated respectively.
[0082] Set deviation judgment thresholds, where the average relative deviation threshold for structural parameters is ≤3%, the maximum relative deviation threshold for structural parameters is ≤5%, the average relative error threshold for component content is ≤2%, and the maximum relative error threshold for component content is ≤3%. If the deviation values of structural parameters and the relative errors of component content in the inversion results both meet the deviation judgment threshold requirements, no optimization is needed. Otherwise, mark the deviation quantification results (including deviations of each single parameter, average deviation, and maximum deviation) and the corresponding inversion sample data (including real-time spectral data, operating parameters, and inversion parameters) as the deviation correction sample set.
[0083] The structure-property relationship model was retrained based on the bias-corrected sample set.
[0084] The specific process is as follows: the bias correction sample set is merged into the original training dataset, the training set and validation set are re-divided (the ratio is maintained at 7:2), and the core parameters of the structure-effect relationship model are adjusted based on the bias quantification results: (1) Machine learning module parameter adjustment: according to the bias weights of structural parameters and component content, the number of decision trees of the random forest algorithm (within the range of 100-200 trees) and the maximum depth (within the range of 10-20 layers) are adjusted, and the feature importance screening threshold is optimized. The penalty coefficient C of the support vector machine algorithm is adjusted (within the range of 0.1-10), and the gamma value of the radial basis function is dynamically corrected based on the bias type. (2) Deep learning module parameter adjustment: the number of hidden layers of the LSTM network (within the range of 2-4 layers) and the number of hidden units (within the range of 128-512) are adjusted, and the forget gate threshold is optimized (within the range of 0.2-0.4). (3) Adjustment of fusion weights and regularization parameters: Based on the bias feedback results, the fusion weights of the local feature mapping vector and the structural temporal feature vector are redistributed (the total weights are 1), and the L2 regularization coefficient of the fully connected layer is adjusted (within the range of 0.0005-0.0015) to suppress overfitting.
[0085] Model Retraining and Performance Validation: The fusion algorithm framework is retrained iteratively using the adjusted parameters. The batch size is set to 32-128, the learning rate to 0.001-0.01, and the number of iterations to 50-100. Mean squared error is used as the loss function, and the Adam optimizer is used for parameter updates. After each training round, the loss value is monitored using the validation set. Training is stopped when the validation set loss value converges to below 1e-5 and shows no decrease for five consecutive rounds. The retrained model is then applied to the original test dataset and the bias correction sample set for bias metric review. If all bias metrics meet the threshold after review, the adjusted model structure and parameters are fixed, and the optimized spectroscopic-compound structure-activity relationship model is output. If the threshold is not met, the steps are repeated until the bias metrics meet the requirements.
[0086] The optimization process of the structure-property relationship model was refined. By establishing the correspondence between the inversion results and laboratory standard test data, the deviation values of structural parameters and the relative errors of component content were accurately quantified. Substandard samples were marked as deviation correction samples and incorporated into the original training data, enabling targeted iteration of the model. The model's learning ability for deviation features was optimized by adjusting the algorithm hyperparameters, fusion weights, and regularization coefficients. This scheme constructs a closed-loop mechanism of inversion-verification-feedback-optimization, which can continuously improve the model's inversion accuracy and generalization ability, enabling the model to adapt to the inversion needs of different working conditions and different types of mineral compounds, extending the effective lifespan of the model, and ensuring the long-term reliability of the system.
[0087] Intelligent inversion method for the structure of mineral compounds integrating spectroscopic big data, such as Figure 2 As shown, it includes the following steps: Full-band spectroscopic data were acquired from the mineral samples, and standardized preprocessing was performed to obtain standardized spectroscopic data.
[0088] Standardized spectroscopic data is integrated with spectroscopic theoretical simulation data to construct a spectroscopic standard database, which includes a molecular spectral standard data sub-database, a compound structural spectroscopic data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation correlation data sub-database.
[0089] Based on a standard spectroscopic database, a fusion of machine learning and deep learning algorithms is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components, and output an initial structure-activity relationship model between spectra and compound structures.
[0090] The standardized spectroscopic data collected in real time are input into the structure-activity relationship model, and the standard spectroscopic database is called to compare the molecular spectral standard data, analyze the spectral features and invert the compound structure, and output the inversion results of the target compound.
[0091] The deviation between the inversion results and the laboratory test data is calculated, and the deviation results are fed back to the structure-activity relationship model to optimize the structure-activity relationship model.
[0092] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0093] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0094] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0095] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0096] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0097] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A smart inversion system for the structure of mineral compounds integrating spectroscopic big data, characterized in that, include: The spectral data acquisition module is used to acquire full-band spectral data of mineral samples and perform standardized preprocessing to obtain standardized spectral data. The database construction module is used to integrate standardized spectroscopic data with spectroscopic theoretical simulation data to construct a spectroscopic standard database. The spectroscopic standard database includes a molecular spectral standard data sub-database, a compound structural spectroscopic data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation association data sub-database. The model training module is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components based on a standard spectroscopic database and by combining machine learning and deep learning algorithms, and outputs an initial structure-activity relationship model between spectra and compound structures. The inversion module is used to input the standardized spectroscopic data acquired in real time into the structure-activity relationship model, call the spectroscopic standard database to compare molecular spectral standard data, perform spectral feature analysis and compound structure inversion, and output the inversion results of the target compound. The model optimization module is used to calculate the deviation between the inversion results and the laboratory test data, and to feed the deviation results back to the structure-activity relationship model to optimize the structure-activity relationship model.
2. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 1, characterized in that, The process of acquiring full-band spectroscopic data from mineral samples and performing standardized preprocessing to obtain standardized spectroscopic data is as follows: By adjusting the parameters of the laser emission module, the ultrashort pulse laser is focused onto the surface of the mineral sample to be tested through the optical path collimation component, which excites the material on the sample surface to produce an ablation effect and form a plasma plume. Wavelength and intensity spectra of the plasma plume were acquired to obtain raw spectroscopic data; The original spectroscopic data is decomposed by wavelet, noise is removed, and baseline correction is performed to obtain the baseline-corrected spectroscopic data. The baseline-corrected spectroscopic data were normalized and outlier removed to obtain standardized spectroscopic data.
3. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 1, characterized in that, The process of fusing standardized spectroscopic data with spectroscopic theoretical simulation data to construct a standard spectroscopic database is as follows: Spectroscopic simulation data based on quantum chemical calculations were obtained and then fused with standardized spectroscopic data after format unification, association and annotation of sample and simulation information, and removal of redundant and invalid entries to obtain a fused data set. Historical operating condition parameters are retrieved, the correlation between operating condition parameters and historical feature parameters is extracted, a mapping model between operating conditions and operating condition adaptation correction parameters is established, and the output operating condition adaptation correlation data sub-library is classified and stored. The historical feature parameters include historical molecular spectral feature parameters and historical compound structure parameters. Molecular spectral feature parameters are extracted from the fused dataset, and deviation calibration and correction are performed based on the working condition adaptation correction parameters. The data are classified and stored according to mineral type and molecular type. After data verification and removal of abnormal entries, a molecular spectral standard data sub-library is obtained. Compound structural parameters are extracted from the fused dataset, corrected by working condition adaptation, and stored hierarchically after establishing parameter mapping relationships, outputting a sub-library of compound structural spectroscopic data. By integrating the molecular spectroscopy standard data sub-library, the compound structure spectroscopy data sub-library, the elemental characteristic spectroscopy data sub-library, and the operating condition adaptation association data sub-library, a multi-dimensional unified indexing system is constructed, and a spectroscopy standard database is output.
4. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 1, characterized in that, Based on a standard spectroscopic database, and employing a fusion of machine learning and deep learning algorithms, the process of training the structure-activity relationship between elemental characteristic spectra and compound structural components, and outputting an initial spectroscopic and compound structure structure-activity relationship model, is as follows: Extract elemental characteristic spectral parameters from the molecular spectral standard data sub-library, extract corresponding compound structural parameters from the compound structural spectral data sub-library, and extract operating condition correction coefficients from the operating condition adaptation association data sub-library. A one-to-one correspondence between elemental characteristic spectral parameters and compound structural parameters is established based on the unique sample identifier. Deviation calibration is performed by combining the working condition correction coefficient. Data is divided by stratified sampling and training dataset, validation dataset and test dataset are output. Based on the random forest algorithm and the support vector machine algorithm, feature processing is performed on the corrected elemental spectral parameters and compound structural parameters in the training dataset to obtain local feature mapping vectors. Based on the LSTM network, feature processing is performed on the corrected elemental spectral parameters and compound structural parameters in the training dataset to obtain the structural time-series feature vector; The local feature mapping vector and the structural temporal feature vector are concatenated to output a global fusion feature matrix; The global fusion feature matrix is input into the fully connected layer, and the Sigmoid activation function is used to output the predicted values of the compound structure components. The error between the predicted values and the actual values of the compound structure components is calculated using the mean squared error as the loss function. The Adam optimizer is used for iterative training. When the error value of the validation dataset converges to below the set value, the number of neurons in the fully connected layer, the fusion weights, and the regularization parameters are adjusted to suppress overfitting, and the optimized fusion model is output. The optimized fusion model is applied to the test dataset for performance evaluation. Once the performance meets the standards, the model structure, parameters, and input / output formats are solidified to obtain the initial structure-activity relationship model of spectroscopy and compound structure.
5. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 4, characterized in that, The process of performing feature processing on the corrected elemental spectral parameters and compound structural parameters in the training dataset based on the random forest algorithm and support vector machine algorithm to obtain local feature mapping vectors is as follows: The corrected elemental characteristic spectral parameters are input into the random forest algorithm, and the feature importance score of each parameter is calculated based on the out-of-bag data error. Filter out elemental feature spectral parameters whose feature importance scores are greater than a set score threshold, and output a subset of elemental feature spectral parameters; The subset of elemental characteristic spectral parameters and the corrected compound structural parameters are concatenated column-wise according to the unique sample identifier to form a joint feature matrix; The joint feature matrix and the corresponding compound structural component labels are input into a support vector machine algorithm. The algorithm is trained to fit the joint feature matrix and labels by minimizing structural risk, and outputs a local feature mapping vector.
6. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 4, characterized in that, The process of performing feature processing on the corrected elemental spectral parameters and compound structural parameters in the training dataset based on the LSTM network to obtain the structural time-series feature vector is as follows: Based on the unique sample identifier, a one-to-one correspondence is established between the corrected elemental characteristic spectral parameters and the compound structural parameters, forming a joint parameter set for each sample; The set of joint parameters is serialized and arranged in a set order to obtain fused sequence data. The fused sequence data is input into the LSTM network.
7. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 1, characterized in that, The process of inputting standardized spectroscopic data acquired in real time into the structure-activity relationship model, comparing molecular spectral standard data with spectroscopic standard databases, analyzing spectral features and retrieving compound structures, and outputting the retrieving results of the target compound is as follows: Real-time acquisition of operating condition parameters; matching the corresponding operating condition range in the operating condition adaptation associated data sub-database based on the operating condition parameters; obtaining operating condition adaptation correction parameters. The standardized spectroscopic data were calibrated for deviation using the operating condition adaptation correction parameters to obtain the standardized spectroscopic data after operating condition adaptation correction. The similarity between the corrected standardized spectroscopic data and the molecular spectral standard data sub-library was calculated to screen out the candidate spectral feature set with the matching degree meeting the standard. Input the candidate spectral feature set into the structure-activity relationship model to output the initial inversion results of the target compound; The preliminary inversion results were cross-validated with the compound structure spectroscopic data sub-library, and the structural parameter deviation values and relative errors of component content were calculated. If the deviation of structural parameters is no greater than 5% and the relative error of component content is no greater than 3%, the inversion result is deemed valid, and it is standardized and formatted to generate the inversion result. If the target is not met, the spectral feature analysis and compound structure inversion process will be executed again.
8. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 7, characterized in that, The process of cross-validating the preliminary inversion results with the compound structure spectroscopic data sub-library and calculating the structural parameter deviation value and the relative error of component content is as follows: Extract corrected standard structural parameters and standard component content ranges that are consistent with the target compound structure type in the preliminary inversion results from the compound structure spectroscopic database; Select the structural parameters to be verified from the preliminary inversion results, match them one by one with the corresponding structural parameters, and calculate the single parameter deviation value of each structural parameter. The maximum value among all single-parameter deviation values is taken as the structural parameter deviation value of the target compound. Extract the inversion content percentage of each component from the preliminary inversion results, use the median of the standard component content range as the standard content benchmark value, and calculate the single-component relative error of each component. The maximum value among all individual component relative errors is taken as the relative error of the component content of the target compound.
9. The intelligent inversion system for the structure of mineral compounds integrating spectroscopic big data according to claim 1, characterized in that, The process of calculating the deviation between the inversion results and the laboratory test data, and then feeding the deviation results back into the structure-property relationship model to optimize the structure-property relationship model is as follows: Based on the unique sample identifier, a one-to-one correspondence between the inversion results and the laboratory standard test dataset is established, and the structural parameter deviation value and the relative error of component content are calculated respectively. Set a deviation judgment threshold. If the deviation values of structural parameters and the relative errors of component content in the inversion results meet the deviation judgment threshold requirements, no optimization is needed; otherwise, mark the deviation quantification results and the corresponding inversion sample data as the deviation correction sample set. The structure-property relationship model was retrained based on the bias-corrected sample set.
10. A method for intelligent inversion of the structure of mineral compounds by integrating spectroscopic big data, characterized in that, Includes the following steps: Full-band spectroscopic data were acquired from the mineral samples, and standardized preprocessing was performed to obtain standardized spectroscopic data. Standardized spectroscopic data is integrated with spectroscopic theoretical simulation data to construct a spectroscopic standard database, which includes a molecular spectral standard data sub-database, a compound structural spectroscopic data sub-database, an elemental characteristic spectral data sub-database, and a working condition adaptation correlation data sub-database. Based on a standard spectroscopic database, a fusion of machine learning and deep learning algorithms is used to train the structure-activity relationship between elemental characteristic spectra and compound structural components, and output an initial structure-activity relationship model between spectra and compound structures. The standardized spectroscopic data collected in real time are input into the structure-activity relationship model, and the standard spectroscopic database is called to compare the molecular spectral standard data, analyze the spectral features and invert the compound structure, and output the inversion results of the target compound. The deviation between the inversion results and the laboratory test data is calculated, and the deviation results are fed back to the structure-activity relationship model to optimize the structure-activity relationship model.