Molecular spectroscopy prediction method and system based on solvent effect

CN122715840APending Publication Date: 2026-09-08DIVAMICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610798036.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0003]发明人在实现本发明的过程中发现,现有有机分子吸收光谱预测技术存在难以克服的核心缺陷:深度学习类预测模型仅以孤立单分子结构作为输入,完全忽略实际应用中溶剂环境对分子电子结构和能级的影响,导致预测结果与实验测量值偏差较大,难以满足真实场景下光谱预测的高精度、高可靠性需求

Benefits of technology

[0015]According to the technical solution of the present invention, by obtaining the molecular structure of organic molecules, the target light wavelength, and the solvent descriptor vector, wavelength feature vector, solvent feature vector, and molecular feature vector are generated through multiple encoding pathways. The wavelength feature vector and molecular feature vector are then fused and predicted using dual-path conditional modulation of the solvent feature vector. This helps to achieve accurate prediction of the absorption spectrum of organic molecules under solvent effects, improves the accuracy, reliability, and generalization ability of spectral prediction in complex solvent environments, and provides stable and reliable technical support for the research and development of organic photovoltaic materials, performance evaluation of optoelectronic devices, and analysis of molecular photophysical properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122715840A_ABST
    Figure CN122715840A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of molecular spectroscopy prediction method and system based on solvent effect.The molecular structure of organic molecule, target light wavelength and solvent descriptor vector are obtained, wavelength feature vector, solvent feature vector and molecular feature vector are generated through multiple coding paths, and wavelength feature vector and molecular feature vector are modulated and fused after two-way conditioning using solvent feature vector, which helps to realize the accurate prediction of organic molecule absorption spectrum under solvent effect, improve the accuracy, reliability and generalization ability of spectrum prediction in complex solvent environment, and provide stable and reliable technical support for the fields of organic photovoltaic material research and development, optoelectronic device performance evaluation and molecular optical physical property analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a molecular spectral prediction method and system based on solvent effects. Background Technology

[0002] The absorption spectra of organic molecules are core data for measuring their photophysical properties and assessing their application potential in optoelectronic devices such as organic photovoltaics. They also have significant reference value in materials research and development, performance screening, and other fields. Existing methods for predicting the absorption spectra of organic molecules are mainly divided into two categories: one relies on complex calculations based on quantum chemical theory, and the other uses deep learning models to directly predict data. Both methods have formed relatively mature technical paths and can meet the spectral prediction needs in conventional scenarios.

[0003] In the process of realizing this invention, the inventors discovered that existing organic molecule absorption spectroscopy prediction technology has a core defect that is difficult to overcome: deep learning prediction models only take isolated single-molecule structures as input, completely ignoring the influence of the solvent environment on the molecular electronic structure and energy levels in actual applications, resulting in a large deviation between the prediction results and experimental measurements, making it difficult to meet the high accuracy and high reliability requirements of spectral prediction in real scenarios. Summary of the Invention

[0004] Based on this, and to address the aforementioned problems, this invention provides a molecular spectral prediction method and system based on solvent effects. By acquiring the molecular structure of organic molecules, the target light wavelength, and the solvent descriptor vector, wavelength feature vectors, solvent feature vectors, and molecular feature vectors are generated through multiple encoding pathways. The solvent feature vectors are then used to perform dual-path conditional modulation on the wavelength feature vectors and molecular feature vectors before fusion prediction. This helps to achieve accurate prediction of the absorption spectra of organic molecules under solvent effects, improves the accuracy, reliability, and generalization ability of spectral prediction in complex solvent environments, and provides stable and reliable technical support for fields such as organic photovoltaic material research and development, optoelectronic device performance evaluation, and molecular photophysical property analysis.

[0005] In a first aspect, the present invention provides a molecular spectral prediction method based on solvent effects. The method includes: acquiring the molecular structure of the organic molecule to be predicted, the target light wavelength, and the target solvent descriptor vector of the solvent in which the organic molecule is located; inputting the target light wavelength into a first encoding pathway of an absorbance prediction model to obtain a wavelength feature vector; inputting the target solvent descriptor vector into a solvent encoding pathway of the absorbance prediction model to obtain a solvent feature vector; inputting the molecular structure into a second encoding pathway of the absorbance prediction model to obtain a molecular feature vector; the absorbance prediction model conditionally modulating the wavelength feature vector and the molecular feature vector according to the solvent feature vector to generate a solvent-sensing wavelength feature vector and a solvent-sensing molecular feature vector; fusing the solvent-sensing wavelength feature vector and the solvent-sensing molecular feature vector to obtain a fused feature vector; and predicting the absorbance of the organic molecule at the target light wavelength based on the fused feature vector.

[0006] Optional physicochemical parameters include static dielectric constant, dynamic dielectric constant, hydrogen bond donor capacity parameter, hydrogen bond acceptor capacity parameter, and surface tension. These parameters can comprehensively characterize the essential properties of solvents from multiple dimensions such as polarity, polarization, hydrogen bonding, and interfacial characteristics, avoiding the one-sidedness of single-parameter characterization. They can accurately capture the differentiated regulatory effects of different solvents on the electronic structure and energy levels of molecules, significantly enhancing the model's adaptability and discriminative power in complex solvent environments, and effectively improving the accuracy, reliability, and generalization ability of spectral predictions under solvent effects.

[0007] Optionally, the wavelength feature vector is conditionally modulated based on the solvent feature vector, including: concatenating the wavelength feature vector and the solvent feature vector, generating a gate vector using an activation function; and multiplying the gate vector element-wise with the wavelength feature vector to obtain the solvent-sensing wavelength feature vector. This solvent-sensing wavelength feature vector generation method can adjust the feature response intensity corresponding to different wavelengths according to the differences in the solvent feature vector, accurately simulating the selective influence of the solvent on each band of the spectrum. This conditional modulation method has clear physical meaning, strong interpretability, effectively improves the fit between the wavelength feature vector and the solvent environment, enhances the model's ability to express solvent effects, and thus improves the prediction accuracy of organic molecule absorption spectra.

[0008] Optionally, the molecular feature vector is conditionally modulated based on the solvent feature vector, including: generating a scaling factor and bias based on the solvent feature vector using a fully connected network; and performing an affine transformation on the molecular feature vector based on the scaling factor and bias to obtain a solvent-sensing molecular feature vector. This solvent-sensing molecular feature vector generation method can systematically adjust the molecular feature vector according to the solvent environment, accurately simulating the regulatory effect of the solvent field on the molecular electronic structure. This conditional modulation method has high parameter efficiency and clear physical meaning, effectively enhancing the ability of the molecular feature vector to characterize the solvent effect, thereby improving the prediction accuracy and reliability of organic molecule absorption spectra.

[0009] Optionally, the molecular structure is input into the second encoding pathway of the absorbance prediction model to obtain molecular feature vectors. This includes: converting the molecular structure into atomic encoding tensors and pairwise encoding tensors; inputting the atomic encoding tensors and pairwise encoding tensors into a pre-trained molecular representation model to extract molecular feature vectors composed of the combination of individual atomic feature vectors and pairwise feature vectors. This molecular feature vector extraction method can simultaneously and completely represent both the intrinsic properties of atoms and the spatial interaction information between atoms. The pre-trained molecular representation model can incorporate rich prior knowledge of molecular structure, enhancing the expressive power and information completeness of the molecular feature vectors, and providing a reliable foundation for subsequent conditional modulation and feature fusion of solvent feature vectors.

[0010] Optionally, the solvent-sensing wavelength feature vector and the solvent-sensing molecular feature vector are fused to obtain a fused feature vector. This includes: concatenating the solvent-sensing wavelength feature vector and the solvent-sensing molecular feature vector to obtain a concatenated feature tensor; and integrating the features of the concatenated feature tensor to generate the fused feature vector. This fused feature vector generation method can fully integrate the wavelength response information represented by the solvent-sensing wavelength feature vector and the molecular structure information represented by the solvent-sensing molecular feature vector, strengthening the correlation between wavelength information, molecular structure information, and solvent environment information. This fusion method has high information utilization and sufficient feature interaction, effectively improving the comprehensive characterization ability of the fused feature vector and providing a reliable basis for accurately predicting the absorption spectra of organic molecules.

[0011] Optionally, the method further includes training the absorbance prediction model, including: constructing a training dataset containing absorption spectral curves of different organic molecules under different solvent descriptor vectors; wherein the absorption spectral curves consist of multiple sets of wavelength-absorbance sampling points; and training the absorbance prediction model based on the training dataset. This absorbance prediction model training method enables the absorbance prediction model to fully learn the intrinsic correlation between different organic molecules, different solvent environments, and absorption spectral responses; effectively establishes the mapping relationship between solvent descriptor vectors, molecular structures, and absorbance, ensuring that the absorbance prediction model has the ability to accurately predict the absorption spectra of organic molecules under various solvent conditions, and improving the reliability and generalization performance of the model prediction results.

[0012] Secondly, the present invention also provides a molecular spectral prediction system based on solvent effects. This system includes: an input module for acquiring the molecular structure of the organic molecule to be predicted, the target light wavelength, and the target solvent descriptor vector of the solvent in which the organic molecule is located; a first encoding pathway module for inputting the target light wavelength into the first encoding pathway of an absorbance prediction model to obtain a wavelength feature vector; a solvent encoding pathway module for inputting the target solvent descriptor vector into the solvent encoding pathway of the absorbance prediction model to obtain a solvent feature vector; a second encoding pathway module for inputting the molecular structure into the second encoding pathway of the absorbance prediction model to obtain a molecular feature vector; a conditional modulation module for conditionally modulating the wavelength feature vector and the molecular feature vector according to the solvent feature vector, respectively, to generate a solvent-sensing wavelength feature vector and a solvent-sensing molecular feature vector; a feature fusion module for fusing the solvent-sensing wavelength feature vector and the solvent-sensing molecular feature vector to obtain a fused feature vector; and an absorbance output module for predicting and outputting the absorbance of the organic molecule at the target light wavelength based on the fused feature vector.

[0013] Thirdly, embodiments of the present invention also provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the steps in any of the above implementations.

[0014] Fourthly, embodiments of the present invention provide an electronic device, the electronic device including a memory and a processor, the memory storing program instructions, and the processor reading and running the program instructions, executing the steps in any of the above implementations.

[0015] According to the technical solution of the present invention, by obtaining the molecular structure of organic molecules, the target light wavelength, and the solvent descriptor vector, wavelength feature vector, solvent feature vector, and molecular feature vector are generated through multiple encoding pathways. The wavelength feature vector and molecular feature vector are then fused and predicted using dual-path conditional modulation of the solvent feature vector. This helps to achieve accurate prediction of the absorption spectrum of organic molecules under solvent effects, improves the accuracy, reliability, and generalization ability of spectral prediction in complex solvent environments, and provides stable and reliable technical support for the research and development of organic photovoltaic materials, performance evaluation of optoelectronic devices, and analysis of molecular photophysical properties. Attached Figure Description

[0016] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the molecular spectral prediction method based on solvent effect provided in the embodiments of the present invention; Figure 2 This is a flowchart of the solvent sensing wavelength feature vector generation method provided in the embodiments of the present invention; Figure 3 This is a flowchart of the solvent-sensing molecule feature vector generation method provided in the embodiments of the present invention; Figure 4 This is a flowchart of the molecular feature vector extraction method provided in the embodiments of the present invention; Figure 5 This is a flowchart of the fusion feature vector generation method provided in the embodiments of the present invention; Figure 6 This is a flowchart of the absorbance prediction model training method provided in the embodiments of the present invention; Figure 7 This is a schematic diagram of the molecular spectral prediction system based on solvent effect provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will now be described with reference to the accompanying drawings. For example, the flowcharts and block diagrams in the drawings illustrate the architecture, functions, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or part of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions. In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0019] Absorption spectra of organic molecules are core data for evaluating their photophysical properties and assessing their application potential in optoelectronic devices such as organic photovoltaics. Currently, the prediction of organic molecule absorption spectra mainly relies on two technical approaches: one is quantum chemical theoretical calculation, such as time-dependent density functional theory, which, although theoretically accurate, suffers from high computational complexity, long processing time, and low efficiency, making it unsuitable for large-scale rapid molecule screening; the other is the deep learning prediction method developed in recent years, which directly predicts absorbance by using light wavelength and molecular structure as input, significantly improving prediction efficiency and becoming the mainstream research direction.

[0020] While existing deep learning prediction methods are highly efficient, they suffer from significant technical limitations in real-world applications. In real-world scenarios, organic molecules are often in solvent environments such as solutions or solid films. The polarity, hydrogen bonding ability, and polarizability of the solvent significantly alter the molecular electronic structure and vibrational energy levels, triggering solvochromic effects and causing marked changes in the position, shape, and intensity of absorption spectrum peaks. For example, the measured absorption peaks of the same organic photovoltaic material molecule in polar and non-polar solvents can differ by tens of nanometers. However, existing deep learning models only use isolated single-molecule structures as input, completely ignoring the influence of the solvent environment. This results in large deviations between predicted and experimental values, failing to meet the accuracy requirements of real-world scenarios. Furthermore, the feature interaction mechanisms of existing deep learning models are relatively simple, directly splicing and fusing wavelength and molecular structure codes. They lack a fine-grained characterization of the relationship between wavelength, molecular structure, and solvent environment, and have not established a conditional modulation mechanism for solvent features on wavelength and molecular features. This makes it difficult for the models to effectively characterize photophysical processes under complex solvent effects, limiting prediction accuracy and reliability.

[0021] Based on the technical solution of this invention, by obtaining the molecular structure of organic molecules, the target light wavelength, and the solvent descriptor vector, wavelength feature vector, solvent feature vector, and molecular feature vector are generated through multiple encoding pathways. Then, the wavelength feature vector and molecular feature vector are fused and predicted after dual-path conditional modulation using the solvent feature vector. This helps to achieve accurate prediction of the absorption spectrum of organic molecules under solvent effects, improves the accuracy, reliability, and generalization ability of spectral prediction in complex solvent environments, and provides stable and reliable technical support for the research and development of organic photovoltaic materials, performance evaluation of optoelectronic devices, and analysis of molecular photophysical properties.

[0022] Please refer to Figure 1 , Figure 1 This is a flowchart of a solvent-effect-based molecular spectral prediction method provided in an embodiment of the present invention. The molecular spectral prediction method includes: Step S100: Obtain the molecular structure of the organic molecule to be predicted, the target light wavelength, and the target solvent descriptor vector of the solvent in which the organic molecule is located.

[0023] In step S100 above, the molecular structure of the organic molecule to be predicted can be obtained, including the atomic sequence of the molecule, the atomic type of each atom, and the three-dimensional spatial coordinate information, so as to fully characterize the spatial configuration of the molecule; the target light wavelength to be predicted can be obtained, which can be a single wavelength or multiple discrete wavelengths in the ultraviolet to near-infrared band; the target solvent descriptor vector corresponding to the actual solvent in which the organic molecule is located can be obtained. This vector is composed of at least one physicochemical parameter of the solvent, preferably composed of at least two parameters from static dielectric constant, dynamic dielectric constant, hydrogen bond donor ability parameter, hydrogen bond acceptor ability parameter, and surface tension, which are used to comprehensively characterize the polarity, polarization, hydrogen bonding and interfacial properties of the solvent. In this way, the synchronous acquisition of three core input information, namely molecular structure, target light wavelength and target solvent descriptor vector, is completed, providing a complete data foundation for subsequent feature generation, conditional modulation and spectral prediction.

[0024] Step S200: Input the target light wavelength into the first encoding path of the absorbance prediction model to obtain the wavelength feature vector.

[0025] In step S200 above, the obtained target light wavelength is input into the first encoding path of the absorbance prediction model. The first encoding path can perform nonlinear encoding processing on the target light wavelength using a preset Gaussian kernel function to generate a Gaussian encoded vector. Then, the Gaussian encoded vector is subjected to dimensionality reduction processing by a fully connected layer to obtain a dimensionality-reduced vector. Subsequently, the dimensionality-reduced vector is normalized by a normalization layer to generate a wavelength feature vector with uniform dimension and stable representation, thus completing the conversion of the target light wavelength into structured features and providing basic feature data for subsequent conditional modulation.

[0026] It should be noted that the absorbance prediction model is an end-to-end prediction model built on deep learning. Its structure is adapted to the requirements of solvent effect modeling and has high inference efficiency. The absorbance prediction model has three parallel encoding paths: the first encoding path, the solvent encoding path, and the second encoding path. It is also equipped with a conditional modulation module, a feature fusion module, and an output module. It can convert the three types of input information—target light wavelength, solvent descriptor vector, and molecular structure—into feature vectors and complete absorbance prediction under solvent effect.

[0027] Step S300: Input the target solvent descriptor vector into the solvent encoding path of the absorbance prediction model to obtain the solvent feature vector.

[0028] In step S300 above, the obtained target solvent descriptor vector can be completely input into the solvent encoding pathway of the absorbance prediction model. The solvent encoding pathway performs feature mapping processing on the target solvent descriptor vector through a fully connected network to complete the vector dimension unification and feature information enhancement, generating a solvent feature vector with standardized dimensions that can comprehensively characterize the physicochemical properties of the solvent. This feature vector will serve as the core basis for the subsequent conditional modulation step, providing support for the control of the wavelength feature vector and molecular feature vector by the solvent effect.

[0029] Step S400: Input the molecular structure into the second encoding pathway of the absorbance prediction model to obtain the molecular feature vector.

[0030] In step S400 above, the obtained organic molecular structure can be input into the second encoding pathway of the absorbance prediction model; the molecular structure is transformed to generate atomic encoding tensors and paired encoding tensors, and then the two types of encoding tensors are input into the pre-trained molecular representation model. The molecular representation model completes feature extraction and integration to obtain a molecular feature vector composed of the feature vectors of each atom and the paired feature vectors. This vector fully carries the structural information of the molecule itself, providing basic data for subsequent conditional modulation and feature fusion.

[0031] It should be noted that the second encoding pathway is the functional pathway used by the absorbance prediction model to analyze molecular structure information. The pre-trained molecular representation model built into the second encoding pathway is a general deep learning model for molecular structure learning. This model has been pre-trained with a large number of organic molecular samples and has the ability to accurately extract inherent molecular features such as atomic properties and inter-atomic relationships. It can stably output highly expressive molecular features. Furthermore, this pre-trained molecular representation model can be the UniMol model. The UniMol model is a pre-trained molecular model based on 3D coordinates. It can directly use the three-dimensional spatial structure information of molecules for representation learning, effectively capturing the influence of spatial conformation on electronic transitions. It is different from the traditional representation methods based on one-dimensional sequences of SMILES or two-dimensional molecular diagrams, avoiding the loss of spatial information due to the lack of dimension, and is more suitable for the prediction scenarios of photophysical properties of organic molecules.

[0032] Step S500: The absorbance prediction model conditionally modulates the wavelength feature vector and molecular feature vector based on the solvent feature vector to generate solvent-sensing wavelength feature vector and solvent-sensing molecular feature vector.

[0033] In step S500 above, the solvent feature vector and wavelength feature vector obtained above are retrieved, and the two sets of vectors are concatenated and then used to generate a gated vector by an activation function. The gated vector and the wavelength feature vector are multiplied element-wise to obtain the solvent-sensing wavelength feature vector. At the same time, the solvent feature vector and molecular feature vector are retrieved, and a scaling factor and bias are generated using the solvent feature vector through a fully connected network. The molecular feature vector is subjected to an affine transformation based on the scaling factor and bias to obtain the solvent-sensing molecular feature vector. The feature vectors obtained after the two modulations are completed are both integrated with solvent environment information, laying the foundation for subsequent feature fusion and absorbance prediction.

[0034] Step S600: Fuse the solvent-sensing wavelength feature vector with the solvent-sensing molecule feature vector to obtain a fused feature vector.

[0035] In step S600 above, the generated solvent-sensing wavelength feature vector and solvent-sensing molecule feature vector are first spliced ​​together to obtain a spliced ​​feature tensor. Then, the spliced ​​feature tensor is processed to fully explore the correlation information between different features to generate a fused feature vector. This vector contains all the effective information of wavelength, molecular structure and solvent environment, and can comprehensively reflect the light absorption characteristics of organic molecules under solvent effect.

[0036] Step S700: Based on the fused feature vector, predict the absorbance of the output organic molecule at the target light wavelength.

[0037] In step S700 above, the fused feature vector is input into the output module of the absorbance prediction model. The fused feature vector is analyzed and processed. Combining the intrinsic relationship between molecular structure, light wavelength, solvent environment and light absorption characteristics learned in the previous stage of the absorbance prediction model, numerical deduction and calculation are completed. The predicted absorbance value of organic molecules at the corresponding target light wavelength is output, thus completing the entire spectral prediction process. The results fully take into account the influence of solvent effect and have higher accuracy and practicality.

[0038] Please refer to Figure 2 , Figure 2 The flowchart below shows a solvent-sensing wavelength feature vector generation method provided in an embodiment of the present invention; the solvent-sensing wavelength feature vector generation method includes: Step S20: Concatenate the wavelength feature vector with the solvent feature vector and generate a gated vector using an activation function.

[0039] In step S20 above, the wavelength feature vector output through the first encoding path and the solvent feature vector output through the solvent encoding path are retrieved. A vector concatenation operation is performed on the two sets of feature vectors to integrate the feature information of two different dimensions and different semantics into a whole vector. Then, a nonlinear operation is performed on the concatenated vector through a preset activation function to complete feature selection and weight allocation, and generate a gated vector that can characterize the effect of solvent on wavelength feature regulation.

[0040] Step S21: Multiply the gate vector and the wavelength feature vector element by element to obtain the solvent-sensing wavelength feature vector.

[0041] In step S21 above, the generated gated vector is matched element by element with the initial wavelength feature vector and multiplied. The feature values ​​of each dimension in the wavelength feature vector are adjusted by using the weight coefficients of the gated vector, so that the wavelength features are integrated with solvent property information to generate a solvent-sensing wavelength feature vector. This solvent-sensing wavelength feature vector can truly reflect the response characteristics of light waves under the action of the solvent environment.

[0042] Please refer to Figure 3 , Figure 3 This is a flowchart of a solvent-sensing molecule feature vector generation method provided in an embodiment of the present invention; the solvent-sensing molecule feature vector generation method includes: Step S30: Generate scaling factor and bias based on solvent feature vector using a fully connected network.

[0043] In step S30 above, the solvent feature vector obtained through the solvent encoding path is input into a preset fully connected network. The fully connected network combines the parameters learned by the absorbance prediction model to perform depth calculation and dimension transformation on the solvent feature vector, and extracts the scaling factor and bias that can reflect the solvent regulation effect. The two types of parameters are used to realize feature amplitude adjustment and feature offset correction, respectively.

[0044] Step S31: Perform an affine transformation on the molecular feature vector based on the scaling factor and bias to obtain the solvent-sensing molecular feature vector.

[0045] In step S31 above, the generated scaling factor, bias, and molecular feature vector output by the second encoding pathway are retrieved. The molecular feature vector is processed dimension by dimension according to the affine transformation operation rules. The feature amplitude is adjusted by the scaling factor, and the feature offset compensation is completed by combining the bias. The solvent-sensing molecular feature vector with solvent property information is obtained, which accurately reflects the control effect of the solvent on the molecular structure features.

[0046] Please refer to Figure 4 , Figure 4This is a flowchart of a molecular feature vector extraction method provided in an embodiment of the present invention; the molecular feature vector extraction method includes: Step S40: Convert the molecular structure into atomic encoding tensors and paired encoding tensors.

[0047] In step S40 above, the complete structural data of the organic molecule to be predicted is read, and the type, properties and spatial distribution information of all atoms in the molecule are encoded according to the preset encoding rules to generate atomic encoding tensors. At the same time, the connection relationship and spatial interaction relationship between atoms in the molecule are encoded to generate paired encoding tensors. The two types of tensors completely retain the atomic monomer information and the inter-atomic correlation information, respectively.

[0048] Step S41: Input the atomic encoding tensor and the pairwise encoding tensor into the pre-trained molecular representation model to extract the molecular feature vector composed of the feature vectors of each atomic sub-feature vector and the pairwise feature vectors.

[0049] In step S41 above, the obtained atomic encoding tensor and paired encoding tensor are input into the pre-trained molecular representation model. The molecular representation model relies on the ability formed by learning from massive samples in the early stage to extract atomic sub-feature vectors and paired feature vectors from the two types of tensors respectively, and integrates and splices all feature vectors to form a complete molecular feature vector. This molecular feature vector fully carries the structural characteristics of the molecule itself, providing reliable feature support for the subsequent conditional modulation stage.

[0050] Please refer to Figure 5 , Figure 5 This is a flowchart of a fusion feature vector generation method provided in an embodiment of the present invention; the fusion feature vector generation method includes: Step S50: Concatenate the solvent-sensing wavelength feature vector with the solvent-sensing molecule feature vector to obtain the concatenated feature tensor.

[0051] In step S50 above, the solvent sensing wavelength feature vector and the solvent sensing molecular feature vector are retrieved, and the two sets of vectors are spliced ​​according to the preset arrangement rules. The feature information representing the wavelength response and the feature information representing the molecular structure are integrated into a unified whole, thereby forming a spliced ​​feature tensor, which provides a complete data carrier for subsequent deep feature interaction and integration.

[0052] Step S51: Perform feature integration on the concatenated feature tensor to generate a fused feature vector.

[0053] In step S51 above, feature integration processing can be performed on the spliced ​​feature tensor to sort out the intrinsic relationship between wavelength information, molecular structure information and solvent environment information, eliminate redundant features and enhance the expression of effective features to generate a fused feature vector. This fused feature vector can comprehensively reflect the absorption characteristics of organic molecules to target light waves under solvent effects, providing a complete feature basis for subsequent absorbance prediction.

[0054] Please refer to Figure 6 , Figure 6 This is a flowchart of an absorbance prediction model training method provided in an embodiment of the present invention; the absorbance prediction model training method includes: Step S60: Construct a training dataset containing absorption spectrum curves of different organic molecules under different solvent descriptor vectors.

[0055] In step S60 above, various organic molecules with different structures and solvents with different physicochemical properties are collected. The actual absorption spectrum curves composed of multiple wavelength-absorbance sampling points are collected for each combination. The molecular structure, target light wavelength, and target solvent descriptor vector composed of the corresponding solvent physicochemical parameters are extracted for each sample. All sample information is organized, classified and labeled in a unified format to build a training dataset. This training dataset can comprehensively cover different molecules, different solvent environments and corresponding spectral data, providing sufficient sample support for the model to learn the intrinsic mapping relationship.

[0056] Step S61: Train the absorbance prediction model based on the training dataset.

[0057] In step S61 above, the completed training dataset is input into the absorbance prediction model in batches. The absorbance prediction model sequentially performs feature extraction, conditional modulation, feature fusion, and absorbance prediction. By comparing the predicted value output by the absorbance prediction model with the actual absorbance value in the dataset, the loss error is calculated. Based on the backpropagation mechanism of the loss error, the internal parameters of the model are continuously updated iteratively. The model is repeatedly trained until the loss value of the absorbance prediction model tends to stabilize and the prediction accuracy reaches the preset requirements, thus completing the model training. This enables the absorbance prediction model to accurately predict absorption spectra under different molecular and solvent conditions.

[0058] For example, an absorbance prediction model is first constructed, with three preset inputs: the first input receives the light wavelength x, the second input receives the molecular structure M, and the third input receives the solvent descriptor vector S; the molecular structure M may include the atomic sequence of the molecule and the atomic type and three-dimensional coordinates of each atom, and the solvent descriptor vector S may be composed of multiple physicochemical parameters of the solvent, including the static dielectric constant ε and the dynamic dielectric constant n. 2The parameters are: hydrogen bond donor capacity α, hydrogen bond acceptor capacity β, and surface tension γ. .

[0059] The absorbance prediction model includes a first input encoding pathway, a third input solvent encoding pathway, a second input molecular encoding pathway, and a feature fusion and output module.

[0060] The first input encoding path includes a first embedding encoding module, a fully connected layer, and a normalization layer. The first embedding encoding module uses D0 preset Gaussian kernel functions to encode the wavelength x of the light wave based on the following formula: in, The scalar output of the k-th Gaussian kernel forms the Gaussian encoded vector A; x is the wavelength of the input target light wave; u k σ represents the center position of the k-th Gaussian kernel (the center on the wavelength axis); k Let be the standard deviation of the k-th Gaussian kernel, controlling the kernel width; D0 is the number of Gaussian kernels, equal to the dimension of the Gaussian encoding vector A; k is the Gaussian kernel index, ranging from 1 to D0; this yields a D0-dimensional Gaussian encoding vector A; subsequently, a fully connected layer reduces A to a D0 / 2-dimensional reduced encoding vector B, and a normalization layer performs LayerNorm processing on B to obtain a normalized encoding vector C; the third input solvent encoding path includes a solvent encoding module and a conditional gating fusion module. The solvent encoding module consists of a fully connected network with at least one hidden layer, taking a solvent descriptor vector S as input and outputting a solvent feature vector C with the same dimension as C. s The conditional gating fusion module receives C and C. s First, the two are concatenated, then passed through a fully connected layer and a sigmoid function to obtain the gated vector G, as shown in the formula: Where G is the gate vector used to modulate the wavelength characteristics; σ( ) is the Sigmoid activation function, with an output range of (0, 1); W is the weight matrix of the fully connected layer (trainable parameters); This is a vector concatenation operation, combining C and C... s Connect the beginning and end; C is the normalized wavelength eigenvector; C s b is the solvent feature vector; b is the bias vector of the fully connected layer (a trainable parameter); then C is modulated by element-wise multiplication, and the solvent-sensing wavelength feature vector is obtained based on the following formula. : in, G is the solvent-sensing wavelength feature vector; G is the gate vector. For element-wise multiplication (Hadamard product); C is the original normalized wavelength eigenvector.

[0061] The second input molecular coding pathway includes a second embedding coding module, a pre-trained UniMol model, and a solvent conditioning module. The second embedding coding module performs atom-type one-thermal encoding on each atom of the molecular structure M according to the UniMol model's coding rules, generating a shape... The atomic encoding tensor E is used to generate a shape based on the three-dimensional coordinates of the atoms. Each unit is a pairwise encoded tensor P of D2-dimensional vectors, where Given the total number of atoms, the pre-trained UniMol model extracts and fuses features from E and P, outputting a shape of... The atomic characteristic tensor Q (composed of sub-characteristic vectors) Composition) and shape The paired feature tensor V, which represents the sub-feature vector of each atom. It is sequentially concatenated with its corresponding row feature vector (composed of all paired feature vectors in that row) to obtain the result from... Sub-feature vectors The molecular characteristic tensor H of each Dimensions The solvent conditioning module is based on C s For each Affine transformation is performed through two small fully connected networks. and Based on the following formula from C s Generate scaling factor and bias: in, The solvent modulated molecular feature vector for the i-th atom; For solvent characteristic C s The learned scaling factor vector; For solvent characteristic C s The learned bias vector; Let i be the original molecular feature vector of the i-th atom; i is the atom index; thus, the solvent-modulated molecular feature vector is obtained. The feature fusion and output module includes a feature fusion module and an MLP model. The feature fusion module converts the solvent-sensing wavelength feature vector... Modulate the molecular feature vector of each solvent Perform vector concatenation to obtain a vector with dimension . fusion vector ,all Composition shape is The fusion tensor R is used as the input of the MLP model. After processing through several fully connected layers, the final output is the predicted absorbance y in scalar form.

[0062] Furthermore, the constructed absorbance prediction model is trained; a training dataset is collected, with each data record including the molecular structure of the organic molecule, the solvent descriptor vector of the solvent in which the molecule is located, and the experimentally measured and discretized absorption spectrum curve under the given conditions as a label. The dataset is divided into a training set and an evaluation set according to a preset ratio; then, the training loop is entered, and for each record in the training set, the wavelength of each sampling point of the label curve is... The molecular structure M and the solvent descriptor vector S are input into the model to obtain the predicted absorbance. and the label absorbance The mean squared error loss is calculated based on the following formula: in, is the mean squared error loss value; N is the total number of wavelength sampling points for a single spectral sample; j is the sampling point index; Predict the absorbance of the model at the j-th wavelength; Let be the experimental label absorbance at the j-th wavelength; the model parameters can be updated by backpropagation using the Adam or SGD optimizer; after one round of training, the root mean square error (RMSE) is calculated using the evaluation set. Training stops when the RMSE meets the preset range, otherwise it continues to iterate until the absorbance prediction model converges and reaches the preset accuracy requirement.

[0063] Based on the trained absorbance prediction model, the system receives the target wavelength sequence (e.g., 300nm to 800nm, with 1nm intervals), target molecular structure, and target solvent descriptor vector provided by the user. For each target wavelength in the sequence, it is input into the trained absorbance prediction model along with the molecular structure and solvent descriptor vector to obtain the corresponding predicted absorbance. All sampling points are sorted by wavelength and curve-fitted in a two-dimensional coordinate plane. Finally, the system outputs and feeds back the complete absorption spectrum curve of the organic molecule under the specified solvent environment, thus achieving accurate prediction of the absorption spectrum of organic molecules under solvent effects.

[0064] Please refer to Figure 7 , Figure 7This is a schematic diagram of the molecular spectral prediction system based on solvent effect provided in an embodiment of the present invention. The system includes an input module 10, a first encoding pathway module 20, a solvent encoding pathway module 30, a second encoding pathway module 40, a conditional modulation module 50, a feature fusion module 60, and an absorbance output module 70. The input module 10 is used to acquire the molecular structure of the organic molecule to be predicted, the target light wavelength, and the target solvent descriptor vector of the solvent in which the organic molecule is located. The first encoding pathway module 20 is used to input the target light wavelength into the first encoding pathway of the absorbance prediction model to obtain the wavelength feature vector. The solvent encoding pathway module 30 is used to input the target solvent descriptor... The solvent encoding pathway of the absorbance prediction model is used to obtain a solvent feature vector; the second encoding pathway module 40 is used to input the molecular structure into the second encoding pathway of the absorbance prediction model to obtain a molecular feature vector; the conditional modulation module 50 is used to conditionally modulate the wavelength feature vector and the molecular feature vector according to the solvent feature vector to generate a solvent-sensing wavelength feature vector and a solvent-sensing molecular feature vector; the feature fusion module 60 is used to fuse the solvent-sensing wavelength feature vector and the solvent-sensing molecular feature vector to obtain a fused feature vector; the absorbance output module 70 is used to predict and output the absorbance of the organic molecule at the target light wavelength based on the fused feature vector.

[0065] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the steps in any of the above implementations.

[0066] Based on the same inventive concept, the present invention also provides an electronic device, which includes a memory and a processor. The memory stores program instructions, and when the processor reads and runs the program instructions, it executes the steps in any of the above implementation methods.

[0067] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A molecular spectral prediction method based on solvent effects, characterized in that, include: The molecular structure, target light wavelength, and target solvent descriptor vector of the organic molecule to be predicted are obtained; wherein the target solvent descriptor vector is composed of at least one physicochemical parameter of the solvent. The target light wavelength is input into the first encoding path of the absorbance prediction model to obtain the wavelength feature vector; The target solvent descriptor vector is input into the solvent encoding path of the absorbance prediction model to obtain the solvent feature vector; The molecular structure is input into the second encoding pathway of the absorbance prediction model to obtain the molecular feature vector; The absorbance prediction model conditionally modulates the wavelength feature vector and the molecular feature vector based on the solvent feature vector to generate a solvent-sensing wavelength feature vector and a solvent-sensing molecular feature vector. The solvent-sensing wavelength feature vector is fused with the solvent-sensing molecule feature vector to obtain a fused feature vector; Based on the fused feature vector, the absorbance of the organic molecule at the target light wavelength is predicted and output.

2. The method according to claim 1, characterized in that, The physicochemical parameters include static dielectric constant, dynamic dielectric constant, hydrogen bond donor capability parameter, hydrogen bond acceptor capability parameter, and surface tension.

3. The method according to claim 1, characterized in that, The conditional modulation of the wavelength feature vector based on the solvent feature vector includes: The wavelength feature vector is concatenated with the solvent feature vector, and an activation function is used to generate a gated vector. The solvent-sensing wavelength feature vector is obtained by multiplying the gate vector element-wise with the wavelength feature vector.

4. The method according to claim 1, characterized in that, The conditional modulation of molecular feature vectors based on solvent feature vectors includes: A scaling factor and bias are generated based on the solvent feature vector using a fully connected network. Based on the scaling factor and bias, the molecular feature vector is subjected to an affine transformation to obtain the solvent-sensing molecular feature vector.

5. The method according to claim 1, characterized in that, The step of inputting the molecular structure into the second encoding pathway of the absorbance prediction model to obtain the molecular feature vector includes: The molecular structure is converted into an atomic encoding tensor and a pairwise encoding tensor; The atomic encoding tensor and the paired encoding tensor are input into a pre-trained molecular representation model to extract a molecular feature vector composed of the feature vectors of each atomic sub-feature vector and the paired feature vectors.

6. The method according to claim 1, characterized in that, The step of fusing the solvent-sensing wavelength feature vector with the solvent-sensing molecule feature vector to obtain a fused feature vector includes: The solvent-sensing wavelength feature vector and the solvent-sensing molecule feature vector are concatenated to obtain the concatenated feature tensor. The spliced ​​feature tensor is integrated to generate a fused feature vector.

7. The method according to claim 1, characterized in that, The method further includes training the absorbance prediction model, including: A training dataset is constructed containing absorption spectrum curves of different organic molecules under different solvent descriptor vectors; wherein the absorption spectrum curves are composed of multiple sets of wavelength-absorbance sampling points; The absorbance prediction model is trained based on the training dataset.

8. A molecular spectral prediction system based on solvent effects, characterized in that, include: The input module is used to acquire the molecular structure of the organic molecule to be predicted, the target light wavelength, and the target solvent descriptor vector of the solvent in which the organic molecule is located; wherein the target solvent descriptor vector is composed of at least one physicochemical parameter of the solvent; The first encoding path module is used to input the target light wavelength into the first encoding path of the absorbance prediction model to obtain the wavelength feature vector. The solvent encoding pathway module is used to input the target solvent descriptor vector into the solvent encoding pathway of the absorbance prediction model to obtain the solvent feature vector; The second encoding pathway module is used to input the molecular structure into the second encoding pathway of the absorbance prediction model to obtain the molecular feature vector. The conditional modulation module is used to conditionally modulate the wavelength feature vector and the molecular feature vector according to the solvent feature vector, respectively, to generate a solvent-sensing wavelength feature vector and a solvent-sensing molecular feature vector; The feature fusion module is used to fuse the solvent-sensing wavelength feature vector with the solvent-sensing molecule feature vector to obtain a fused feature vector. An absorbance output module is used to predict and output the absorbance of the organic molecule at the target light wavelength based on the fused feature vector.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, perform the steps of the method according to any one of claims 1-7.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores program instructions, and when the processor executes the program instructions, it performs the steps of the method according to any one of claims 1-7.