Chemical evolution prediction method based on spectroscopy inversion neural network algorithm

By constructing a chemical and materials database and a spectroscopic inversion neural network algorithm, molecular descriptors are generated, multi-source data are integrated, and a spectroscopic structure-activity relationship prediction model is constructed. This solves the problems of data dispersion and cumbersome experimental verification in existing technologies, and realizes synchronous and accurate prediction of multi-dimensional parameters of chemical evolution.

CN122024902APending Publication Date: 2026-05-12北京机数小来智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京机数小来智能科技有限公司
Filing Date
2026-02-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In the existing technology for predicting chemical evolution, data is scattered, inconsistent in format, and weakly correlated, resulting in fragmented operation procedures and cumbersome experimental verification. It is impossible to directly deduce the chemical structure changes and evolution laws from spectroscopic measurement data. Multiple hypothesis-experiment-correction processes are required, and a single model can only predict parameters in a single dimension.

Method used

A chemical and materials database is constructed to generate molecular descriptors. A spectroscopic structure-activity relationship prediction model is built through a spectroscopic inversion neural network algorithm. Multi-source data are integrated to achieve synchronous and accurate learning of spectrum-structure, structure-activity, and spectrum-activity correlations. The output data includes chemical structure change parameters, molecular interaction laws, and property evolution trends.

Benefits of technology

It enables simultaneous and accurate prediction of multi-dimensional parameters, reduces the cumbersome steps of experimental verification, improves the efficiency and accuracy of chemical evolution prediction, provides high-quality data support, and ensures data traceability and chemical rationality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024902A_ABST
    Figure CN122024902A_ABST
Patent Text Reader

Abstract

The invention discloses a chemical evolution prediction method based on a spectroscopy inversion neural network algorithm, and relates to the technical field of inversion prediction. According to the method, a chemical and material database is firstly constructed, so that data support is provided for subsequent prediction; the molecular descriptor fusing the spectral characteristics, the microstructure and the physical property associated information is generated, and the defects that a traditional molecular descriptor is single in information and insufficient in representativeness are overcome; the molecular descriptor is used as input, a spectrum structure-effect relationship prediction model with common feature extraction and multi-branch special prediction capabilities is constructed, and spectrum-structure, structure-effect and spectrum-effect associated synchronous precise learning is realized; the compatibility and reliability of the input data and the model are ensured by carrying out noise reduction and standardization preprocessing on the target spectroscopic data subsequently, and finally, chemical structure change parameters, molecular interaction rules and physical property evolution trend data are directly output through model inversion, so that the defects that a traditional method needs multi-step splitting and experimental verification is tedious are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of inversion prediction technology, specifically to a chemical evolution prediction method based on a spectroscopic inversion neural network algorithm. Background Technology

[0002] In the scenario of predicting chemical evolution, the spectroscopic data, structural parameters, and physical properties involved in chemical processes are highly dimensional and strongly correlated, and the experimental data are sparse and have poor repeatability. Existing technologies have revealed the following shortcomings: First, in existing technologies, chemical data is scattered across multiple sources, including literature, theoretical calculations, and experimental measurements. This necessitates manually breaking down data collection, format conversion, and correlation matching into independent steps, as well as conducting separate experiments for structural analysis and property testing. This fragmented workflow significantly increases time costs. Second, existing technologies cannot directly deduce the laws governing changes and evolution of chemical structures from spectroscopic measurement data. They require multiple cycles of hypothesis-experiment-correction to verify structural hypotheses. Furthermore, a single model can only predict a single-dimensional parameter (e.g., predicting only structure or only property), necessitating the construction of multiple independent models and cross-validation of results through experiments. This results in repetitive and cumbersome experimental verification processes.

[0003] Therefore, there is an urgent need for an integrated method that combines multi-source data association and storage, spectroscopic inversion intelligent modeling, and multi-dimensional synchronous prediction to solve the problems of fragmented processes and reliance on repeated experimental verification in existing technologies, and to improve the efficiency and accuracy of chemical evolution prediction. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a chemical evolution prediction method based on a spectroscopic inversion neural network algorithm, which solves the problems of fragmented processes and reliance on repeated experimental verification in existing technologies.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a chemical evolution prediction method based on a spectroscopic inversion neural network algorithm, comprising the following steps: Constructing a chemistry and materials database; Based on the quantum chemical Raman spectroscopy theoretical data and experimental spectroscopic data in the aforementioned chemical and materials database, molecular descriptors are generated for the microstructure and physical properties of chemicals. Using the molecular descriptors as input features and combining them with the chemical and materials database, a spectroscopic structure-activity relationship prediction model is constructed through neural network training. Acquire the target spectroscopic measurement data of the chemical to be predicted, and after noise reduction and standardization preprocessing, input the target spectroscopic measurement data into the spectroscopic structure-activity relationship prediction model; The preprocessed target spectroscopic measurement data are inverted and analyzed using the aforementioned spectroscopic structure-activity relationship prediction model, and the chemical structure change parameters, molecular interaction rules, and property evolution trend data of the chemical to be predicted are output.

[0006] The present invention has the following beneficial effects: This invention first constructs a chemical and materials database to provide data support for subsequent predictions. Then, it generates a molecular descriptor that integrates spectroscopic features, microstructure, and property correlation information, overcoming the shortcomings of traditional molecular descriptors that are limited in information and lack representativeness. Subsequently, using this molecular descriptor as input, a spectroscopic structure-activity relationship prediction model is constructed, which combines common feature extraction and multi-branch specific prediction capabilities, achieving simultaneous and accurate learning of spectroscopic-structure, structure-activity, and spectroscopic-activity correlations. Further, by performing noise reduction and standardization preprocessing on the target spectroscopic data, the compatibility and reliability of the input data and the model are ensured. Finally, the model inversion directly outputs chemical structure change parameters, molecular interaction laws, and property evolution trend data, avoiding the drawbacks of traditional methods that require multiple steps and cumbersome experimental verification.

[0007] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0008] Figure 1 This is a flowchart of the chemical evolution prediction method based on the spectroscopic inversion neural network algorithm of the present invention. Detailed Implementation

[0009] Please see Figure 1 The present invention provides a technical solution: a chemical evolution prediction method based on a spectroscopic inversion neural network algorithm, comprising the following steps: Construct a database of chemistry and materials.

[0010] It should be noted that the specific process is as follows: The scope of data collection was determined and multi-source data collection was carried out. The collected data included publicly available literature data composed of domestic and foreign academic literature and technical patent literature related to chemistry and materials, quantum chemical theoretical data composed of quantum chemical Raman spectroscopy theoretical calculation data, experimental measured data composed of chemical synthesis process parameters, sample characterization spectroscopic data, material structure data, and physical property test data, as well as application-specific targeted data composed of special experimental data for energy catalysts, battery electrolytes, battery positive and negative electrodes, and optical thin films. Using machine intelligence reading algorithms and natural language processing technology, keyword recognition is performed on publicly available literature data from multiple sources to establish unstructured intermediate data containing entity association information; Specifically, the process first identifies key entities in the literature, such as chemical substance names, experimental condition parameters, functional properties of substances, physicochemical properties of substances, and experimental operation actions. Then, through entity matching and semantic association analysis, the logical relationships between these key entities are mined to obtain unstructured intermediate data containing entity association information.

[0011] Based on the ontology library in the field of chemistry, a structured data field system is defined. The field system includes a unique substance identifier field, an experimental condition description field, a performance index quantification field, a substance structural parameter field, and an associated reaction type field. Unstructured intermediate data is mapped and filled according to the structured data field system to generate standardized structured literature data. The quantum chemical Raman spectroscopy theoretical data in multi-source data is converted using a unified data format protocol, and the parameters such as spectrum wavelength and intensity are converted into standard numerical formats to form a standard theoretical spectrum dataset. Outliers in the experimental measured data from multi-source data were removed using the 3σ criterion, including synthesis process parameters, characterization spectroscopic data, structural data, and physical property data. Missing data were filled in using linear interpolation. Dimensions were unified according to internationally accepted dimensional standards. A unique material identifier code was assigned to each standardized experimental data point. A mapping between experimental data and corresponding materials was established to obtain a standardized experimental measured dataset. Construct a chemical knowledge graph that includes rules on element conservation, rules on the rationality of reaction conditions, rules on the correspondence between spectroscopic features and material structure, and rules on the correlation and constraint of physical properties; It should be noted that the specific construction process of the chemical knowledge graph is as follows: Based on the principles of chemometrics, and using the unique identification code of substances as the correlation benchmark, this paper defines the quantitative judgment criteria for the conservation of the number of atoms of each element in reactants and products, clarifies the element type identification threshold and the allowable range of atomic number error, establishes a one-to-one correspondence verification logic between the element composition set of reactants and the element composition set of products, and outputs the structured rules for element conservation.

[0012] By integrating parameters such as reaction temperature, pressure, catalyst type, and solvent system from structured literature data and standardized experimental data, and combining chemical thermodynamics and kinetics theories, the effective range of condition parameters for different reaction types (such as synthesis reactions, catalytic reactions, and decomposition reactions) is divided, logical constraint relationships between parameters are defined (such as the matching relationship between high-temperature reactions and specific catalysts), and structured rules for the rationality of reaction conditions are output.

[0013] Based on the characterization spectroscopic data (including Raman spectroscopy, ultraviolet spectroscopy, and gas chromatography data) in the standard theoretical spectrum dataset and standardized experimental measured data, the structural parameters (bond length, bond angle, and functional group type) of the corresponding substances are associated to establish a mapping relationship model between spectroscopic characteristic parameters (peak position, peak intensity, and peak shape) and material structural parameters. The characteristic spectral peak intervals and intensity thresholds corresponding to different structural units are clarified, and the structural rules corresponding to the spectroscopic features and material structures are output.

[0014] Based on physical property test data (such as melting point, boiling point, conductivity, and catalytic activity) from standardized experimental data, and combined with the property trend of similar substances, we define the correlation constraints between physical property parameters (such as the positive correlation threshold range between molecular polarity and solubility) and the corresponding constraints between physical properties and structural parameters (such as the quantitative correlation between functional group type and melting point), and output the structured rules of physical property correlation constraints.

[0015] By integrating the above four types of structured rules, a rule set with a unified format is formed. Each rule in the rule set contains a unique rule identifier, applicable scenario, judgment logic, quantification threshold, and associated data type field.

[0016] Import the rule set into the data matching engine and perform traversal matching on structured literature data, standard theoretical spectrum datasets, and standardized experimental measurement datasets respectively; For structured literature data, reaction formulas, reaction conditions, material structure descriptions, and physical property descriptions are extracted and matched with element conservation structure rules and reaction condition rationality structure rules. Examples of material-reaction condition-product associations and material structure-physical property associations that conform to the rules are extracted. For the spectroscopic data and corresponding material structure parameters in the standard theoretical spectrum dataset and the standardized experimental measurement dataset, the spectroscopic feature parameter-material structure correspondence structure rule is matched, and the associated instances of spectroscopic feature parameter-material structure parameter are extracted. For the physical property data in the standardized experimental measured dataset, match it with the structured rules of physical property association constraints, and extract physical property parameter-physical property parameter and physical property parameter-structural parameter association instances; All extracted associated instances are deduplicated and integrated into a rule-based associated instance set using the unique material identifier code as the link.

[0017] Define the entity types in the knowledge graph, including material entities (related to the unique identifier code, chemical name, and molecular formula of the substance), spectroscopic feature entities (related to spectroscopic type, peak position, peak intensity, and peak shape parameters), structural parameter entities (related to bond length, bond angle, functional groups, and crystal structure parameters), physical property parameter entities (related to parameters and values ​​such as melting point, boiling point, conductivity, and catalytic activity), and reaction condition entities (related to parameters and values ​​such as temperature, pressure, catalyst, and solvent).

[0018] Based on the rule set and the rule association instance set, the relationship types between entities are defined as follows: element composition association (corresponding to element conservation rule), reaction condition constraint association (corresponding to reaction condition rationality rule), spectrum-structure correspondence association (corresponding to spectrum feature and material structure correspondence rule), and physical property association (corresponding to physical property association constraint rule).

[0019] Using the unique identifier code of a substance as the association key, a mapping model between each entity and its corresponding relationship is established, clarifying the entity attributes, relationship attributes and association thresholds, and forming a knowledge graph topology model.

[0020] The rule set is transformed into executable code that the rule engine can recognize, embedded in the knowledge graph topology model, and the rule triggering conditions are defined (triggered when data is entered into the database, triggered when related queries are performed, and triggered when data is updated). For the data to be verified, the rule engine calls the corresponding structured rules to perform element conservation verification on the material entity-reaction condition entity-product material entity, perform rationality verification on the reaction type-reaction condition entity, perform spectrum-structure correspondence verification on the spectroscopic feature entity-structural parameter entity, and perform association constraint verification on the physical property parameter entity-physical property parameter entity / structural parameter entity, and output the verification results (pass / redundancy / error).

[0021] Select a portion of multi-source data (accounting for 10%-20% of the total data volume) as a test set, import it into a knowledge graph topology model with a rule engine for trial verification; collect erroneous data and rule mismatch instances in the verification results, and analyze the reasons (such as unreasonable rule thresholds, missing entity relationship definitions, and data anomalies); based on the analysis results, adjust the quantization thresholds in the rule set, supplement entity relationship types, optimize the verification logic, and update the knowledge graph topology model; repeat the steps until the verification accuracy is ≥95%, and obtain the optimized chemical knowledge graph model.

[0022] A graph database (such as Neo4j) is used to store the chemical knowledge graph model, storing entity data, relation data, rule sets, and verification logic in corresponding data nodes and relation edges, respectively. The unique identifier of a substance is used as the primary key index, and the unique identifier of a rule, entity type, and relation type are used as auxiliary indexes to ensure efficient execution of data retrieval and rule verification. After the index is built, a chemical knowledge graph is formed that includes rules on element conservation, rules on the rationality of reaction conditions, rules on the correspondence between spectroscopic features and substance structure, and rules on the correlation of physical properties.

[0023] Based on unique substance identifiers and substance association codes, we establish spectral-structure-effect relationships among structured literature data, standard theoretical spectrum datasets, and standardized experimental measurement data. Furthermore, the process of establishing the spectrum-structure-effect relationship is as follows: Using the unique identifier code of a substance as the association key, structured literature data, standard theoretical spectrum datasets, and standardized experimental measurement datasets are traversed and matched to aggregate various types of data corresponding to the same substance, forming a single-substance multi-source data set. The single-substance multi-source data set includes the substance's literature structured information (substance structure description, physical property description, reaction condition information), theoretical spectroscopic data (Raman spectrum wavelength, intensity parameters), experimental spectroscopic data (Raman spectrum, ultraviolet spectrum, gas chromatography, and other characterization data), substance structure parameters (bond length, bond angle, functional group type, crystal structure parameters), physical property test data (melting point, boiling point, conductivity, catalytic activity, and other parameters), and reaction condition data (temperature, pressure, catalyst type, solvent system parameters).

[0024] Spectrum-structure-effect relationship types and attribute definitions: Three types of correlation sub-relationships are defined: spectrum-structure correlation (correlation between spectroscopic characteristic parameters and material structure parameters), structure-effect correlation (correlation between material structure parameters and physical property test data), and spectrum-effect correlation (correlation between spectroscopic characteristic parameters and physical property test data). Define the attributes of each type of association: including the type of association parameter (e.g., spectroscopic parameters such as peak position / peak intensity / peak shape, structural parameters such as bond length / functional group, physical property parameters such as melting point / catalytic activity), parameter matching threshold (allowable parameter error range), data source identifier (distinguishing between theoretical data / experimental data / literature data), and association confidence calculation method (association confidence = number of parameter matches that meet the rules / total number of parameters in this type of association).

[0025] Construction of a multi-dimensional association mapping model: Spectrum-structure correlation mapping construction: Theoretical spectroscopic feature parameters, experimental spectroscopic feature parameters, and material structure parameters are extracted from a multi-source dataset of a single substance. The correspondence rules between spectroscopic features and material structure in the chemical knowledge graph are called to establish a one-to-one correspondence mapping between spectroscopic feature parameters (peak position range, peak intensity value, peak shape characteristics) and material structure parameters (specific functional groups, bond length and bond angle range, crystal structure type). The correlation confidence of each mapping is calculated, and effective spectrum-structure correlation mappings with a confidence of ≥90% are retained. Structure-activity relationship mapping construction: Extract the structural parameters and physical property test data of substances from multi-source datasets of single substances, establish a quantitative relationship mapping between structural parameters and physical property parameters based on the physical property relationship constraint rules in the chemical knowledge graph, clarify the correspondence model between changes in structural parameters (such as the magnitude of bond length increase or decrease, and the type of functional group substitution) and changes in physical property parameters (such as the range of melting point increase or decrease, and the change in catalytic activity), and label the applicable constraints of the mapping (such as the reaction temperature range and the type of solvent system). Spectrum-Effect Correlation Mapping Construction: Based on the established spectrum-structure correlation mapping and structure-effect correlation mapping, an indirect correlation mapping between spectroscopic feature parameters and physical property test data is constructed through logical transitive reasoning; at the same time, directly corresponding spectroscopic feature parameters and physical property test data are extracted from single-substance multi-source datasets to supplement the direct correlation mapping, and the rationality of the mapping is verified by combining chemical knowledge graph rules to form a complete spectrum-effect correlation mapping.

[0026] The mapping data of the three types of related sub-relationships are imported into the rule engine of the chemical knowledge graph to verify whether the spectrum-structure correlation mapping conforms to the rule of correspondence between spectroscopic features and material structure, whether the structure-effect correlation mapping conforms to the property correlation constraint rule, and whether the spectrum-effect correlation mapping satisfies the logical transitivity and chemical thermodynamic / kinetic related rules. The verification results (pass / redundant / abnormal) are output. Based on the unique identification code of the substance, the type of related sub-relationship, and the correlation parameter pair as the joint judgment criteria, duplicate correlation mapping entries are identified and deleted, and the effective mapping with the highest correlation confidence is retained. For correlation mappings that are judged as abnormal, the original data in the single substance multi-source data set is traced back, and the correlation reasoning ability of the chemical knowledge graph and the effective correlation data of similar substances are combined for correction. Abnormal correlation mappings that cannot meet the rule requirements through correction are directly removed.

[0027] The optimized spectral-structure association mapping, structure-effect association mapping, and spectral-effect association mapping are structurally integrated. Using the unique substance identifier code as an index, a unique association identifier is assigned to each association, generating a spectral-structure-effect association dataset containing fields such as association unique identifier, substance unique identifier code, association sub-relationship type, association parameter pair, association confidence level, data source identifier, and applicable constraints. Through the substance unique identifier code, a traceability link is established between the spectral-structure-effect association dataset and structured literature data, standard theoretical spectrum dataset, and standardized experimental measurement dataset, clarifying the original data entries corresponding to each association, ensuring that the associations are traceable and verifiable, and finally completing the establishment of spectral-structure-effect associations among the three types of datasets.

[0028] A chemical knowledge graph is used to store structured literature data, standard theoretical spectrum datasets, standardized experimental measurement datasets, and the spectrum-structure-effect relationships among the three, forming a chemical and materials database.

[0029] By defining the scope of multi-source data collection, structuring publicly available literature data, standardizing theoretical and experimental data, constructing a chemical knowledge graph, establishing spectral-structure-activity relationship relationships, and storing the data using the knowledge graph, a complete chemical and materials database with integrity, relevance, and reliability was built. Its advantages lie in effectively solving the problems of scattered, inconsistent, weakly correlated, and redundant chemical data in existing systems. It provides high-quality, highly correlated data support for subsequent molecular descriptor generation and training of spectral-structure-activity relationship prediction models, ensuring data traceability and chemical rationality, and laying a solid data foundation for the accuracy of the entire prediction method.

[0030] Based on the quantum chemical Raman spectroscopy theoretical data and experimental spectroscopy data in the aforementioned chemical and materials database, molecular descriptors are generated for the microstructure and physical properties of chemicals.

[0031] Using the unique identifier code of a substance as the search key, a set of target substances containing both theoretical data and experimental spectroscopic data of quantum chemical Raman spectroscopy is screened from a chemical and materials database. The experimental spectroscopic data includes characterization data such as Raman spectroscopy, ultraviolet spectroscopy, and gas chromatography, which are derived from standardized experimental measurement datasets.

[0032] Based on the spectral-structure-effect correlation in the chemistry and materials database, the microstructural parameters (bond length, bond angle, functional group type, crystal structure parameters) and physical property data (melting point, boiling point, conductivity, catalytic activity, reactivity) of the target substance are associated through the unique identification code of the substance, forming a four-dimensional associated data set of quantum chemical Raman spectroscopy theoretical data, experimental spectral data, microstructural parameters and physical property data.

[0033] The theoretical and experimental data of quantum chemical Raman spectroscopy were standardized. The theoretical data underwent format parsing, extracting fundamental parameters such as wavelength, peak intensity, and peak width. These parameters were then converted into a standard numerical matrix according to a unified data format protocol, and redundant noise signals from theoretical calculations were removed to obtain standardized quantum chemical Raman spectroscopy theoretical data. For the experimental data, an adaptive wavelet denoising algorithm was used to remove instrument noise and environmental interference. Baseline correction eliminated background drift, and the 3σ criterion was used to remove abnormal peak positions. Dimensions were then unified according to internationally accepted spectroscopic data standards to obtain a standardized spectroscopic dataset. Based on the correspondence rules between spectroscopic features and material structures in a chemical knowledge graph, the consistency between the standardized quantum chemical Raman spectroscopy theoretical and experimental data was verified. Spectral similarity was calculated, and four-dimensional related data groups with similarity greater than a set threshold were retained to form a valid dataset.

[0034] For the standardized spectroscopic dataset in the effective dataset, local and global spectroscopic features are extracted and fused to form an initial spectroscopic feature vector with unified dimensions. Each initial spectroscopic feature vector is associated with the corresponding unique material identifier code and microstructure parameters and physical property data labels. Convolutional Neural Network (CNN) is used to extract local spectral features, and Long Short-Term Memory Network (LSTM) is used to extract global spectral features. Local spectral features include feature peak position coordinates, peak intensity relative ratios, peak shape fitting parameters (Gaussian fitting coefficients, Lorentz fitting width), and feature peak combination patterns, while global spectral features include the distribution range of spectral peak positions, feature peak intensity distribution entropy, and overall spectral symmetry parameters.

[0035] The microstructure parameters are structured and encoded to obtain the microstructure quantization vector, and the physical property data are classified and quantized to obtain the physical property feature quantization vector. Based on the physical property association constraint rules in the chemical knowledge graph, the correlation between the microstructure quantization vector and the physical property feature quantization vector is verified, and conflicting data is eliminated to ensure the chemical rationality of the encoded data.

[0036] Specifically: functional group types are mapped to one-hot encoded vectors, continuous parameters such as bond length and bond angle are normalized (normalization interval [0,1]), crystal structure parameters are converted into numerical representations of lattice constants and space group identifiers, and integrated to form a fixed-dimensional microstructure quantization vector.

[0037] Continuous physical property parameters such as melting point and boiling point are standardized into standard scores using Z-score, while hierarchical physical property parameters such as catalytic activity and reactivity are mapped into ordered numerical labels to form a quantitative vector of physical property characteristics.

[0038] A spectroscopic pre-trained model is constructed and trained based on the initial spectroscopic feature vector, microstructure quantization vector, and physical property feature quantization vector, and outputs a molecular descriptor.

[0039] Molecular descriptors are generated through a complete process: screening target substance sets, constructing four-dimensional correlation datasets, standardizing spectroscopic data, extracting and fusing spectroscopic features, quantifying and encoding structural and physical property parameters, and training a pre-trained model. The advantage lies in the fact that the generated molecular descriptors deeply integrate the potential correlation information between spectroscopic features, microstructure, and physical properties. This solves the problems of traditional molecular descriptors having limited information and insufficient correlation, providing highly discriminative and representative input features for spectroscopic structure-activity relationship prediction models. This ensures that the model can capture the core laws of chemical evolution processes, improving the accuracy and generalization ability of subsequent predictions.

[0040] Furthermore, the process of outputting the molecular descriptor is as follows: The model is constructed with a feature encoding branch and a supervised constraint branch. The feature encoding branch takes the initial spectroscopic feature vector as input, performs feature mapping and dimensionality compression through a multilayer perceptron, and outputs an intermediate feature vector. The supervised constraint branch predicts the microstructure quantization vector based on the intermediate feature vector, using mean squared error (MSE) as the loss function. It also predicts the physical property feature quantization vector based on the intermediate feature vector, using a weighted sum of cross-entropy loss function (for categorical physical properties) and mean squared error loss function (for continuous physical properties) as the joint loss function. Using the initial spectroscopic feature vector as input, and the corresponding microstructure quantization vector and physical property feature quantization vector as supervision labels, the spectroscopic pre-trained model is constructed and iteratively trained. After each round of training, the spectroscopic-structure-effect association rules of the chemical knowledge graph are called to verify the model output. When the joint loss function value converges to the preset threshold (≤0.001) and the verification accuracy is greater than the set accuracy threshold, the training stops and a converged spectroscopic pre-trained model is obtained. The intermediate feature vector output layer data of the feature encoding branch in the converged spectroscopic pre-trained model is extracted as the initial molecular descriptor, which contains fused information of spectroscopic features, microstructure correlation features, and physical property correlation features. Based on the spectrum-structure-activity relationship in the chemistry and materials database, the association confidence of the initial molecular descriptor with the microstructure parameters and physical property data of the corresponding substance is calculated, and feature dimensions with an association confidence greater than the set confidence threshold are retained. Principal component analysis (PCA) is used to optimize the dimensionality of the retained features, removing feature redundancy and forming a standardized vector with fixed dimensions (such as 256 or 512 dimensions), which is the molecular descriptor.

[0041] By constructing a bi-branch spectroscopic pre-trained model, employing multi-loss function constraints for training, extracting initial molecular descriptors, and then filtering and optimizing them using confidence levels, the generation process of molecular descriptors was further refined. The advantages are that supervised constraint branching ensures a strong correlation between molecular descriptors and their structure and properties, confidence level filtering eliminates redundant and invalid features, and dimensionality reduction optimization improves computational efficiency. Ultimately, molecular descriptors with effectiveness, simplicity, and generalization ability are obtained, ensuring the efficiency of subsequent model training and strengthening the model's ability to capture spectral-structure-effect relationships, thus laying a high-quality feature foundation for accurately predicting chemical evolution.

[0042] Using the molecular descriptors as input features and combining them with the chemical and materials database, a spectroscopic structure-activity relationship prediction model is constructed through neural network training.

[0043] Specifically: Obtain the initial fused feature vector and preprocess it to obtain the standardized input feature matrix; The specific process is as follows: Using the unique identifier code of a substance as the search key, a target training substance set is selected from the chemistry and materials database that simultaneously contains spectroscopic pre-training molecular descriptors, spectroscopic data (quantum chemical Raman spectroscopy theoretical data, standardized experimental spectroscopic data), microstructure parameters, and physical property data. The microstructure parameters include bond length, bond angle, functional group type, and crystal structure parameters, and the physical property data includes melting point, boiling point, electrical conductivity, catalytic activity, and reactivity. Based on the spectrum-structure-effect correlation in the chemistry and materials database, an alignment mapping is established through the unique identifier encoding of substances to obtain training data pairs corresponding to molecular descriptors and labels: the pre-trained spectroscopic molecular descriptors are used as input features and associated with spectrum-structure labels (correspondence between spectroscopic feature parameters and microstructure parameters), structure-effect labels (correspondence between microstructure parameters and physical property data), and spectrum-effect labels (correspondence between spectroscopic feature parameters and physical property data) to form training data pairs of input features and triple labels; The rule set of the chemical knowledge graph (the rule of correspondence between spectroscopic features and material structure, and the rule of constraint on the association of physical properties) is called to verify the rationality of the training data pairs. Abnormal data pairs that violate the rules are removed (such as mismatch between spectroscopic features and structural parameters, and conflict between the association of structure and physical properties). Valid training data pairs are retained and the training set, validation set and test set are divided in a ratio of 8:1:1. Molecular descriptors are extracted from the training set and supplemented with corresponding features from the chemistry and materials database, including quantified reaction condition parameters (temperature, pressure, catalyst type, numerical characterization of solvent system) and molecular topological features (molecular connectivity index, number of rings, number of substituents), to form an initial fusion feature vector. Z-score standardization (mean = 0, standard deviation = 1) is used to eliminate dimensional differences. Redundancy between features is calculated using mutual information entropy. Feature dimensions with redundancy ≥ 80% are removed, and effective features are retained to form a standardized input feature matrix with fixed dimensions (such as 512-dimensional or 1024-dimensional).

[0044] A training model for predicting the spectrum-structure-effect relationship is constructed, comprising a shared feature layer and a branch prediction layer. The shared feature layer is used to extract common information across tasks from the input features, and the branch prediction layer learns a single spectrum-structure-effect relationship task based on the common information. The shared feature layer takes the standardized input feature matrix as input, uses the ReLU activation function, and combines an attention mechanism to output a common feature vector; Specifically: The first fully connected layer takes a standardized input feature matrix as input, performs feature mapping through 2048 neurons, and outputs a 2048-dimensional primary feature vector using the ReLU activation function. The second fully connected layer takes the 2048-dimensional primary feature vector output from the first fully connected layer as input, performs feature compression and enhancement through 1024 neurons, and outputs a 1024-dimensional intermediate feature vector using the ReLU activation function. The third fully connected layer takes the 1024-dimensional intermediate feature vector output by the second fully connected layer as input, further refines the information through 512 neurons, and outputs a 512-dimensional high-level feature vector using the ReLU activation function. The attention mechanism layer (multi-head self-attention module, number of heads = 8) takes as input a 512-dimensional high-level feature vector output from the third fully connected layer. By strengthening the key dimension weights of spectroscopic correlation features, structural correlation features, and physical property correlation features, it outputs a 512-dimensional common feature vector, which serves as the unified input to the branch prediction layer.

[0045] The branch prediction layer includes a spectrum-structure correlation prediction branch, a structure-effect correlation prediction branch, and a spectrum-effect correlation prediction branch, all of which take a common feature vector as input. The spectrum-structure correlation prediction branch outputs predicted values ​​of microstructure parameters, the structure-effect correlation prediction branch outputs predicted values ​​of physical property data, and the spectrum-effect correlation prediction branch outputs numerical values ​​of continuous physical property parameters and ordered labels of hierarchical physical property parameters. Specifically: The spectral-structural correlation prediction branch inputs a 512-dimensional common feature vector from the shared feature layer output; The first fully connected layer uses 256 neurons to perform targeted feature transformation and uses the ReLU activation function to output a 256-dimensional spectral-specific feature vector. The second fully connected layer takes a 256-dimensional spectral-structure-specific feature vector as input, refines the features through 128 neurons, and outputs predicted values ​​of microstructure parameters (numerical prediction results of bond length and bond angle, and numerical encoding results of functional group type and crystal structure parameters) using the Sigmoid activation function.

[0046] The 512-dimensional common feature vector is obtained from the shared feature layer output of the structure-effect correlation prediction branch input; The first fully connected layer: performs targeted feature transformation through 256 neurons, and outputs a 256-dimensional construct-specific feature vector using the ReLU activation function; The Dropout layer takes a 256-dimensional feature vector as input and randomly discards some neurons with a dropout probability of 0.3 to avoid overfitting. The output is a 256-dimensional feature vector with redundancy removed. The second fully connected layer takes a 256-dimensional de-redundant feature vector as input and refines the features through 128 neurons. Continuous physical property parameters are activated by a Linear activation function to output specific values, while hierarchical physical property parameters are activated by a Softmax activation function to output ordered labels, ultimately yielding the predicted values ​​of the physical property feature data.

[0047] The spectral-effect correlation prediction branch inputs a 512-dimensional common feature vector from the shared feature layer output; A single fully connected layer directly maps the association between features and physical properties through 128 neurons. Continuous physical property parameters are output with specific values ​​using the Linear activation function, while hierarchical physical property parameters are output with ordered labels using the Softmax activation function, forming a spectrum-effect direct association prediction channel.

[0048] The loss values ​​of the three types of branch tasks are combined with the chemical rule constraint penalty term as a joint loss function, and the formula is: ; Where: Loss_struct is the mean squared error (MSE) loss for spectrum-structure association prediction, Loss_prop is the mixed loss for structure-effect association prediction (MSE for continuous properties and cross-entropy for hierarchical properties), Loss_spec-prop is the mean squared error loss for spectrum-effect association prediction, Loss_rule is the penalty term for rule violation in the chemical knowledge graph (the penalty value is accumulated according to the degree of violation when a rule is violated); α, β, and γ are the branch loss weights (all initially set to 1.0 and dynamically adjusted during training), and λ is the rule penalty weight (set to 0.5). The adaptive momentum estimation algorithm (Adam) is used as the optimizer, the initial learning rate is set to 1e-4, the learning rate decay strategy is adopted (decaying to 90% of the original learning rate every 100 rounds), the training batch size is set to 32, and the maximum number of training rounds is set to 1000 rounds. The standardized input feature matrix and the corresponding triple label are input into the neural network in batches to start iterative training. After each round of training, the Loss_total value of the training set and the validation set is calculated.

[0049] Intermediate validation is performed after every 50 rounds of training: the rule engine of the chemical knowledge graph is called to validate the model output results of the validation set, including whether the spectrum-structure association results conform to the rules of correspondence between spectroscopic features and material structure, whether the structure-effect association results conform to the rules of physical property association, whether the spectrum-effect association results satisfy logical transitivity, and the statistical rule compliance rate.

[0050] If the rule compliance rate is less than 85%, the value of λ is increased (by 0.1 each time) and the training parameters of the previous round are backtracked for retraining; if the value of Loss_total in the validation set does not decrease for 50 consecutive rounds (fluctuation range ≤1e-5), the early stopping mechanism is triggered and the current training is stopped.

[0051] The grid search method was used to optimize the hyperparameters (number of neurons in each fully connected layer, number of attention heads, dropout probability, and initial learning rate). The optimal combination of hyperparameters was selected with the comprehensive performance index of the validation set (comprehensive performance index = 0.4 × spectral-construction prediction accuracy + 0.4 × construct-effect prediction accuracy + 0.2 × spectral-effect prediction accuracy) as the optimization objective. Based on the weight parameters of the optimal model, the importance score of each input feature dimension is calculated using the permutation importance analysis method. Redundant features with an importance score <0.01 are removed, the input feature matrix is ​​reconstructed, and the model is retrained to improve inference efficiency.

[0052] The test set is input into the optimized spectral structure-activity relationship prediction model, and the performance indicators of the three types of association prediction tasks are calculated respectively. If the target is met, the model is trained; otherwise, the spectral structure-activity relationship prediction model is optimized again.

[0053] Specifically: When the accuracy of the spectral-structure correlation prediction is ≥90%, the average error of the structure-activity correlation prediction is ≤5% (for continuous properties) / the accuracy is ≥88% (for hierarchical properties), and the confidence level of the spectral-activity correlation prediction is ≥85%, the prediction results of the test set are cross-validated by calling the chemical knowledge graph to ensure that all prediction results comply with the rules of element conservation, the rules of reasonable reaction conditions, etc., and the rule compliance rate is ≥92%. For the prediction results that do not meet the standards, error analysis is performed, the model branch loss weights (α, β, γ) are adjusted, and training data in the corresponding fields (such as special data in the fields of energy catalysts and battery electrolytes) are added. The training is iterated again until all performance indicators meet the standards, and finally a stable spectral-structure-activity relationship prediction model is formed.

[0054] By employing feature preprocessing, constructing a "shared feature layer - branch prediction layer" model architecture, designing a joint loss function, optimizing training strategies, and optimizing hyperparameters, a complete spectral structure-activity relationship prediction model was built. Its advantages lie in the model architecture's ability to efficiently extract common and specific features across tasks, the integration of chemical rule constraints into the joint loss function to ensure the chemical rationality of the prediction results, and the further improvement of model performance through hyperparameter optimization. This model addresses the problems of limited prediction and poor generalization ability in traditional single-task models, achieving simultaneous and accurate prediction of spectral-structure, structure-activity, and spectral-activity relationships, providing efficient and reliable model support for multi-dimensional parameter output of chemical evolution.

[0055] The target spectroscopic measurement data of the chemical to be predicted is obtained, and after noise reduction and standardization preprocessing, the target spectroscopic measurement data is input into the spectroscopic structure-activity relationship prediction model.

[0056] It should be noted that the specific process is as follows: Characterization methods (including Raman spectroscopy, ultraviolet spectroscopy, and gas chromatography) that are consistent with experimental spectroscopic data in the chemistry and materials database are used to perform spectroscopic measurements on the chemical to be predicted, collect raw target spectroscopic measurement data, and simultaneously record instrument parameters (such as excitation wavelength, detection accuracy, and scanning range) and experimental environmental conditions (such as temperature, humidity, and pressure) during the measurement process.

[0057] A unique temporary substance identification code is assigned to the chemical to be predicted, and associated with known basic information (such as known elemental composition, approximate reaction system, and target application scenario) to form a data package that links the temporary substance identification code, the original target spectroscopic measurement data, the measurement conditions, and the known basic information.

[0058] An adaptive wavelet denoising algorithm is used to remove instrument noise and environmental interference signals from the original target spectroscopic measurement data, while retaining characteristic spectral peak information. The 3σ criterion is used to identify and remove abnormal peak positions and abnormal intensity data in the original target spectroscopic measurement data, while retaining effective features that conform to the distribution law of spectroscopic data. Following the unified data format protocol for experimental spectroscopic data in the Chemistry and Materials Database, the preprocessed effective features are converted into a standard numerical matrix to obtain standardized target spectroscopic data. By calling the rules for the correspondence between spectroscopic features and material structures in the chemical knowledge graph, the standardized target spectroscopic data is subjected to basic validity verification (such as whether the characteristic peak positions are within a reasonable range). After removing invalid data (if the verification fails, it is returned to be measured again), the standardized target spectroscopic data is input into the spectroscopic structure-activity relationship prediction model.

[0059] The preprocessed target spectroscopic measurement data are inverted and analyzed using the aforementioned spectroscopic structure-activity relationship prediction model, and the chemical structure change parameters, molecular interaction rules, and property evolution trend data of the chemical to be predicted are output.

[0060] Furthermore, the specific process is as follows: Based on the preprocessed target spectroscopic measurement data, a standardized input feature vector is obtained. The structure-activity relationship prediction model is used to process the standardized input feature vector to obtain the predicted values ​​of the microstructure parameters, physical property characteristics, and auxiliary physical property characteristics of the chemical to be predicted. The standardized input feature vector is sequentially passed through the first fully connected layer (2048 neurons, ReLU activation) of the model's shared feature layer to output a 2048-dimensional primary feature vector, the second fully connected layer (1024 neurons, ReLU activation) to output a 1024-dimensional intermediate feature vector, and the third fully connected layer (512 neurons, ReLU activation) to output a 512-dimensional high-level feature vector. Finally, the key feature weights are strengthened through a multi-head self-attention module (number of heads = 8), outputting a 512-dimensional common feature vector containing cross-task correlation information of spectroscopy, structure, and physical properties.

[0061] Spectrum-structure correlation prediction branch: The 512-dimensional common feature vector is input into this branch, and is converted into a 256-dimensional spectrum-structure specific feature vector through the first fully connected layer (256 neurons, ReLU activation). Then, the initial microstructure parameter prediction values ​​of the chemical to be predicted (including bond length, bond angle values, functional group type, and numerical encoding of crystal structure parameters) are output through the second fully connected layer (128 neurons, Sigmoid activation).

[0062] Structure-activity relationship prediction branch: The 512-dimensional common feature vector is input into this branch. After passing through the first fully connected layer (256 neurons, ReLU activation), a 256-dimensional structure-activity specific feature vector is generated. After redundancy is removed by the dropout layer (probability=0.3), the second fully connected layer uses the Linear activation function (continuous physical property) and the Softmax activation function (hierarchical physical property) respectively to output the initial physical property feature prediction values ​​(including melting point, boiling point, conductivity, catalytic activity, etc.).

[0063] Spectrum-Effect Correlation Prediction Branch: The 512-dimensional common feature vector is input into this branch, and the auxiliary prediction value of the physical property feature is directly mapped through a single fully connected layer (128 neurons) to serve as the cross-branch verification benchmark.

[0064] Multi-level cross-validation is performed on the predicted values ​​of microstructure parameters, physical properties, and auxiliary physical properties of the chemical to be predicted, to obtain the corrected predicted values ​​of microstructure parameters and unified physical properties. Based on the known basic information of the chemical to be predicted (initial elemental composition, preset reaction system), basic indicators such as bond length variation, bond angle offset, functional group addition and subtraction type and quantity, and crystal structure distortion parameters are extracted from the predicted values ​​of the microstructure parameters after verification and correction. By matching the structural evolution patterns of similar chemicals with unique substance identification codes, and combining the characteristic changes of standardized target spectroscopic data, a corresponding model of spectroscopic characteristic changes and microstructural changes is established. Based on the extracted basic indicators, the chemical structure change parameters of the chemical to be predicted in the target reaction process are derived, including the structural change rate and the spectroscopic characteristic thresholds corresponding to key structural transformation nodes. Based on the verified and corrected predicted values ​​of microstructure parameters (bond length, bond angle, functional group type), and combined with the molecular interaction energy calculation logic in quantum chemical Raman spectroscopy theoretical data, the chemical bond strength and intermolecular force types (hydrogen bond, van der Waals force, electrostatic force, etc.) within the molecules of the chemical to be predicted are analyzed. The trend of the changes in the unified physical property prediction values ​​(solubility, catalytic activity, etc.) after verification and correction of the associated output is combined with the output chemical structure change parameters (bond length change range, functional group increase or decrease, etc.) to deduce the quantitative correlation law between molecular interaction strength and physical property parameters, clarify the influence mechanism of molecular interaction enhancement and weakening on the chemical evolution process, and output molecular interaction law data (including interaction type, interaction strength range, dynamic change law, as the core input for subsequent physical property evolution trend prediction). Based on the unified predicted physical property characteristics after verification and correction, and combined with known reaction condition parameters (temperature, pressure, reaction time), an initial evolution model of physical property parameters is constructed. Chemical structure change parameters (structural change rate, key transition nodes) and molecular interaction law data (dynamic change of interaction intensity) are incorporated to correct the initial evolution model and clarify the driving weight of structural change and molecular interaction on physical property parameters (calibrated based on the property trend law of similar substances in the chemical knowledge graph). Using the modified evolution model, the direction (increase / decrease / stabilization), rate of evolution, and final stable value of the physical properties of the chemical to be predicted at different reaction stages are predicted. At the same time, based on the historical verification error between the auxiliary prediction value of physical property characteristics and the unified prediction value of physical property characteristics, the confidence interval (confidence level ≥ 85%) of the prediction results of physical properties at each stage is given, and complete physical property evolution trend data is output.

[0065] By employing multi-branch parallel prediction, multi-level cross-validation, derivation of chemical structure change parameters, mining of molecular interaction patterns, and modification of property evolution models, the inversion analysis and result output process has been comprehensively refined. Its advantages lie in achieving accurate inversion across the entire chain, from spectroscopic data to core parameters of chemical evolution. Multi-branch prediction and cross-validation ensure consistency of results, while structured derivation and model modification uncover the essential laws of evolution. This solves the problem of traditional inversion methods that only output single parameters and lack in-depth correlation analysis, providing comprehensive guidance for chemical research and development on structural changes, molecular interactions, and property trends, helping researchers accurately grasp the chemical evolution process.

[0066] The process of obtaining the verified and corrected predicted values ​​of microstructure parameters and unified physical property characteristics is as follows: The atomic composition rationality of the predicted initial microstructure parameters is verified by the element conservation rule; the matching between the initial microstructure parameters and the standardized target spectroscopic data is verified by the rule of correspondence between spectroscopic features and material structure; and the adaptability of the two types of physical property predicted values ​​to the known reaction system is verified by the rule of rationality of reaction conditions. Calculate the relative error between the initial predicted physical property characteristics and the auxiliary predicted physical property characteristics, and at the same time verify the correlation between the initial predicted microstructure parameters and the two types of predicted physical property values ​​(based on the physical property correlation constraint rules). If the rule compliance rate of a certain branch prediction result is less than 90% (such as the initial microstructure parameters violating element conservation), the prediction value of that branch is corrected based on the chemical knowledge graph association reasoning ability and the spectrum-structure-activity relationship data of similar substances in the database. If the relative error between the initial physical property prediction value and the auxiliary physical property prediction value is ≥10%, the initial physical property prediction value is used as the basis. The trend information of the auxiliary physical property prediction value and the spectrum-effect correlation features in the molecular descriptor pre-trained by spectrometry are fused together, and a unified physical property prediction value is obtained by correcting it through a weighted fusion algorithm (the weights are allocated based on the rule conformity rate of the two types of prediction values). If the corrected microstructure parameters and the predicted values ​​of the unified physical property characteristics still violate the physical property correlation constraint rules, backtrack and adjust the initial predicted values ​​of the microstructure parameters until all data meet the rule requirements; The output includes the verified and corrected predicted values ​​of microstructure parameters and the verified and corrected predicted values ​​of unified physical property characteristics (both of which are the core input data for subsequent steps and have been incorporated into the verification logic of the auxiliary predicted values ​​of physical property characteristics).

[0067] The verification and correction process is refined by using multiple rules to validate initial prediction results, calculate relative errors in property predictions, hierarchically correct conflicting data, and retrospectively adjust logical conflicts. Its advantage lies in establishing a rigorous result verification and correction mechanism, effectively eliminating conflicts between prediction results from different branches, correcting abnormal data that violate chemical rules, and ensuring the accuracy, consistency, and chemical rationality of predicted values ​​for microstructure parameters and physical properties. This provides reliable foundational data for subsequent derivation of chemical structure change parameters, mining of molecular interaction laws, and prediction of property evolution trends, further improving the output quality and reliability of the entire prediction method.

[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0072] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0073] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A chemical evolution prediction method based on a spectroscopic inversion neural network algorithm, characterized in that, Includes the following steps: Constructing a chemistry and materials database; Based on the quantum chemical Raman spectroscopy theoretical data and experimental spectroscopic data in the aforementioned chemical and materials database, molecular descriptors are generated for the microstructure and physical properties of chemicals. Using the molecular descriptors as input features and combining them with the chemical and materials database, a spectroscopic structure-activity relationship prediction model is constructed through neural network training. Acquire the target spectroscopic measurement data of the chemical to be predicted, and after noise reduction and standardization preprocessing, input the target spectroscopic measurement data into the spectroscopic structure-activity relationship prediction model; The preprocessed target spectroscopic measurement data are inverted and analyzed using the aforementioned spectroscopic structure-activity relationship prediction model, and the chemical structure change parameters, molecular interaction rules, and property evolution trend data of the chemical to be predicted are output.

2. The chemical evolution prediction method based on spectroscopic inversion neural network algorithm according to claim 1, characterized in that, The process of constructing the chemistry and materials database is as follows: Determine the data collection scope and perform multi-source data collection; Using machine intelligence reading algorithms and natural language processing technology, keyword recognition is performed on publicly available literature data from multiple sources to establish unstructured intermediate data containing entity association information; Based on the ontology library in the field of chemistry, a structured data field system is defined, and unstructured intermediate data is mapped and filled according to the structured data field system to generate standardized structured literature data. Quantum chemical Raman spectroscopy theoretical data from multi-source data are converted using a unified data format protocol to form a standard theoretical spectrum dataset. Outliers were removed from the experimental data from the multi-source dataset using the 3σ criterion, missing data were filled in using linear interpolation, and the units were standardized according to internationally accepted unit standards. A unique material identifier code was assigned to each standardized experimental data point, and a mapping between experimental data and corresponding materials was established to obtain a standardized experimental data set. Construct a chemical knowledge graph that includes rules on element conservation, rules on the rationality of reaction conditions, rules on the correspondence between spectroscopic features and material structure, and rules on the correlation and constraint of physical properties; Based on unique substance identifiers and substance association codes, we establish spectral-structure-effect relationships among structured literature data, standard theoretical spectrum datasets, and standardized experimental measurement data. A chemical knowledge graph is used to store structured literature data, standard theoretical spectrum datasets, standardized experimental measurement datasets, and the spectrum-structure-effect relationships among the three, forming a chemical and materials database.

3. The chemical evolution prediction method based on spectroscopic inversion neural network algorithm according to claim 2, characterized in that, The process of establishing the rules of element conservation, the rules of rationality of reaction conditions, the rules of correspondence between spectroscopic features and material structure, and the rules of constraint of correlation of physical properties is as follows: Based on the principles of chemometrics, and using the unique identification code of substances as the correlation benchmark, we define the quantitative judgment criteria for the conservation of the number of atoms of each element in reactants and products, clarify the element type identification threshold and the allowable range of atomic number error, establish a one-to-one correspondence verification logic between the element composition set of reactants and the element composition set of products, and output the structured rules for element conservation. By integrating conditional parameters from structured literature data and standardized experimental data, and combining chemical thermodynamics and kinetics theories, we can divide the effective range of conditional parameters for different reaction types, define the logical constraint relationships between parameters, and output structured rules for the rationality of reaction conditions. Based on the characterization spectroscopic data in the standard theoretical spectrum dataset and standardized experimental measurement data, the structural parameters of the corresponding substances are associated to establish a mapping relationship model between spectroscopic feature parameters and material structural parameters, clarify the characteristic spectral peak intervals and intensity thresholds corresponding to different structural units, and output the structural rules corresponding to the spectroscopic features and material structures. Based on the physical property test data in the standardized experimental data, and combined with the property change law of similar materials, we define the correlation constraints between physical property parameters and the corresponding constraints between physical property and structural parameters, and output the structured rules of physical property correlation constraints.

4. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 1, characterized in that, The process of generating molecular descriptors based on the quantum chemical Raman spectroscopy theoretical data and experimental spectroscopic data from the aforementioned chemical and materials database, targeting the microstructure and physical properties of chemicals, is as follows: Using the unique identifier code of a substance as the search key, a set of target substances is selected from a chemical and materials database; Based on the spectral-structure-effect correlation in the chemistry and materials database, the microstructure parameters and physical property data corresponding to the target substance are linked by the unique identification code of the substance, forming a four-dimensional linked data group of quantum chemical Raman spectroscopy theoretical data, experimental spectroscopy data, microstructure parameters and physical property data; The theoretical and experimental spectroscopic data of quantum chemical Raman spectroscopy are standardized, and the consistency between the standardized theoretical and experimental spectroscopic data of quantum chemical Raman spectroscopy is verified based on the correspondence rules between spectroscopic features and material structures in the chemical knowledge graph. The spectrum similarity is calculated, and four-dimensional related data groups with similarity greater than a set threshold are retained to form an effective data set. For the standardized spectroscopic dataset in the effective dataset, local and global spectroscopic features are extracted and fused to form an initial spectroscopic feature vector with unified dimensions. Microstructure parameters are structured and encoded to obtain microstructure quantization vectors, and physical property characteristic data are classified and quantized to obtain physical property characteristic quantization vectors. A spectroscopic pre-trained model is constructed and trained based on the initial spectroscopic feature vector, microstructure quantization vector, and physical property feature quantization vector, and outputs a molecular descriptor.

5. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 4, characterized in that, The process of constructing a spectroscopic pre-trained model and training it based on the initial spectroscopic feature vector, microstructure quantization vector, and physical property feature quantization vector to output a molecular descriptor is as follows: Construct the feature encoding branch and the supervision constraint branch of the spectroscopic pre-trained model. The feature encoding branch takes the initial spectroscopic feature vector as input, performs feature mapping and dimensionality compression through a multilayer perceptron, and outputs an intermediate feature vector. The supervision and constraint branch predicts the microstructure quantization vector based on the intermediate feature vector, using mean square error as the loss function. It also predicts the physical property quantization vector based on the intermediate feature vector, using a weighted sum of cross-entropy loss function and mean square error loss function as the joint loss function. Using the initial spectroscopic feature vector as input, and the corresponding microstructure quantization vector and physical property feature quantization vector as supervision labels, the spectroscopic pre-training model is constructed and iteratively trained. When the joint loss function value converges to the preset threshold and the verification accuracy is greater than the set accuracy threshold, the training is stopped, and the converged spectroscopic pre-training model is obtained. Extract the intermediate feature vector output layer data of the feature encoding branch in the converged spectroscopic pre-trained model as the initial molecular descriptor; Based on the spectrum-structure-activity relationship in the chemistry and materials database, the association confidence of the initial molecular descriptor with the microstructure parameters and physical property data of the corresponding substance is calculated, and feature dimensions with an association confidence greater than the set confidence threshold are retained. Principal component analysis is used to optimize the dimensionality of the retained features, removing feature redundancy and forming a standardized vector, which is the molecular descriptor.

6. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 1, characterized in that, The process of using the molecular descriptor as input features, combining it with the chemical and materials database, and training a neural network to construct a spectroscopic structure-activity relationship prediction model is as follows: Obtain the initial fused feature vector and preprocess it to obtain the standardized input feature matrix; Construct a trainingable spectral structure-function relationship prediction model that includes a shared feature layer and a branch prediction layer, wherein: The shared feature layer takes the standardized input feature matrix as input, uses the ReLU activation function, and combines an attention mechanism to output a common feature vector; The branch prediction layer includes a spectrum-structure correlation prediction branch, a structure-effect correlation prediction branch, and a spectrum-effect correlation prediction branch, all of which take a common feature vector as input. The spectrum-structure correlation prediction branch outputs predicted values ​​of microstructure parameters, the structure-effect correlation prediction branch outputs predicted values ​​of physical property data, and the spectrum-effect correlation prediction branch outputs numerical values ​​of continuous physical property parameters and ordered labels of hierarchical physical property parameters. The loss values ​​of the three types of branch tasks are combined with the chemical rule constraint penalty term as a joint loss function; An adaptive momentum estimation algorithm was used as the optimizer, with an initial learning rate of 1e-4, a learning rate decay strategy, a training batch size of 32, and a maximum training epoch of 1000. The grid search method is used to optimize the hyperparameters, and the optimal combination of hyperparameters is selected with the comprehensive performance index of the validation set as the optimization objective. The test set is input into the optimized spectral structure-activity relationship prediction model, and the performance indicators of the three types of association prediction tasks are calculated respectively. If the target is met, the model is trained; otherwise, the spectral structure-activity relationship prediction model is optimized again.

7. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 6, characterized in that, The process of obtaining the initial fused feature vector and preprocessing it to obtain the standardized input feature matrix is ​​as follows: Using the unique identifier code of a substance as the search key, a set of target training substances is selected from a chemistry and materials database; Based on the spectrum-structure-effect relationship in the chemistry and materials database, an alignment mapping is established through the unique identification code of substances to obtain training data pairs corresponding to set molecular descriptors and labels; The rule set of the chemical knowledge graph is invoked to verify the rationality of the training data pairs, and abnormal data pairs that violate the rules are removed. The training set, validation set, and test set are divided proportionally. Molecular descriptors are extracted from the training set and supplemented with corresponding features from the chemistry and materials database to form an initial fused feature vector. Z-score standardization is used to eliminate dimensional differences; effective features are preserved through mutual information entropy to form a standardized input feature matrix.

8. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 1, characterized in that, The process of acquiring the target spectroscopic measurement data of the chemical to be predicted, performing noise reduction and standardization preprocessing on the target spectroscopic measurement data, and then inputting it into the spectroscopic structure-activity relationship prediction model is as follows: An adaptive wavelet denoising algorithm is used to remove instrument noise and environmental interference signals from the original target spectroscopic measurement data, while retaining characteristic spectral peak information. The 3σ criterion is used to identify and remove abnormal peak positions and abnormal intensity data in the original target spectroscopic measurement data, while retaining effective features that conform to the distribution law of spectroscopic data. Following the unified data format protocol for experimental spectroscopic data in the Chemistry and Materials Database, the preprocessed effective features are converted into a standard numerical matrix to obtain standardized target spectroscopic data. By calling the rules for the correspondence between spectroscopic features and material structures in the chemical knowledge graph, the basic validity of the standardized target spectroscopic data is verified. After removing invalid data, the standardized target spectroscopic data is input into the spectroscopic structure-activity relationship prediction model.

9. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 1, characterized in that, The process of performing inversion analysis on the preprocessed target spectroscopic measurement data using the spectroscopic structure-activity relationship prediction model, and outputting the chemical structure change parameters, molecular interaction rules, and property evolution trend data of the chemical to be predicted is as follows: Based on the preprocessed target spectroscopic measurement data, a standardized input feature vector is obtained. The structure-activity relationship prediction model is used to process the standardized input feature vector to obtain the predicted values ​​of the microstructure parameters, physical property characteristics, and auxiliary physical property characteristics of the chemical to be predicted. Multi-level cross-validation is performed on the predicted values ​​of microstructure parameters, physical properties, and auxiliary physical properties of the chemical to be predicted, to obtain the corrected predicted values ​​of microstructure parameters and unified physical properties. Based on the known basic information of the chemical to be predicted, basic indicators are extracted from the predicted values ​​of microstructure parameters after verification and correction. By matching the structural evolution patterns of similar chemicals with unique substance identification codes, and combining the characteristic changes of standardized target spectroscopic data, a corresponding model of spectroscopic characteristic changes and microstructural changes is established. Based on the extracted basic indicators, the chemical structural change parameters of the chemical to be predicted in the target reaction process are derived. Based on the verified and corrected predicted values ​​of microstructure parameters, and combined with the molecular interaction energy calculation logic in quantum chemical Raman spectroscopy theoretical data, the strength of chemical bonds and the type of intermolecular forces within the molecules of the chemical to be predicted are analyzed. The changing trend of the unified physical property prediction values ​​after verification and correction of the associated output, combined with the output chemical structure change parameters, is used to deduce the quantitative correlation law between molecular interaction strength and physical property parameters, and output molecular interaction law data. Based on the unified physical property prediction value after verification and correction, and combined with known reaction condition parameters, an initial evolution model of physical property parameters is constructed. Chemical structure change parameters and molecular interaction law data are then incorporated to correct the initial evolution model. The modified evolution model predicts the direction, rate, and final stable value of the physical properties of the chemical under test at different reaction stages. It also outputs complete physical property evolution trend data based on the historical verification error between the auxiliary prediction value and the unified physical property prediction value.

10. The chemical evolution prediction method based on the spectroscopic inversion neural network algorithm according to claim 9, characterized in that, The process of obtaining the verified and corrected predicted values ​​of microstructure parameters and unified physical property characteristics is as follows: The atomic composition rationality of the predicted initial microstructure parameters is verified by the element conservation rule; the matching between the initial microstructure parameters and the standardized target spectroscopic data is verified by the rule of correspondence between spectroscopic features and material structure; and the adaptability of the two types of physical property predicted values ​​to the known reaction system is verified by the rule of rationality of reaction conditions. Calculate the relative error between the initial predicted physical property characteristics and the auxiliary predicted physical property characteristics, and at the same time verify the correlation between the initial predicted microstructure parameters and the two types of predicted physical property characteristics; If the rule compliance rate of a certain branch prediction result is <90%, the prediction value of that branch is corrected based on the chemical knowledge graph association reasoning ability and the spectrum-structure-activity relationship data of similar substances in the database. If the relative error between the initial predicted physical property value and the auxiliary predicted physical property value is ≥10%, the initial predicted physical property value is used as the basis. The trend information of the auxiliary predicted physical property value and the spectrum-effect correlation features in the pre-trained molecular descriptor are fused together, and a weighted fusion algorithm is used to correct and obtain a unified predicted physical property value. If the corrected microstructure parameters and the predicted values ​​of the unified physical property characteristics still violate the physical property correlation constraint rules, backtrack and adjust the initial predicted values ​​of the microstructure parameters until all data meet the rule requirements; Output the verified and corrected predicted values ​​of microstructure parameters and the verified and corrected predicted values ​​of uniform physical properties.