Mixed solution Raman spectrum generation method and device, medium and equipment
By constructing a mixed solution Raman spectrum generation model, using dual encoders and QKV attention mechanism for feature fusion, and combining multi-level lightweight upsampling and hybrid attention mechanism, the problems of low efficiency and insufficient accuracy in the acquisition of mixed solution Raman spectrum data are solved, and high-precision spectrum reconstruction and proportional control are achieved.
Patent Information
- Application Number
- CN202510596523.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology has problems in acquiring Raman spectroscopy data of mixed solutions, such as long time consumption, high resource consumption, sensitivity to environmental interference and limited data collection. In addition, the simulation synthesis method has insufficient accuracy and the experimental collection method has low efficiency.
A mixed solution Raman spectrum generation method is adopted. By obtaining the Raman spectrum data and mixing ratio parameters of the first solution and the second solution, a mixed solution Raman spectrum generation model is constructed. The dual encoder structure is combined with the QKV attention mechanism for feature fusion, and the spectrum is reconstructed through multi-level lightweight upsampling and hybrid attention mechanism. The mixing ratio parameters and multiple loss function optimization are introduced.
High-precision modeling of the Raman spectrum of mixed solutions was achieved, the signal-to-noise ratio was improved, the mixing ratio was accurately controlled, the mean square error was reduced, and the problems of low efficiency and insufficient accuracy were solved.
Smart Images

Figure CN120673886A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of Raman spectroscopy data measurement and analysis, and specifically relates to a method, device, medium and equipment for generating Raman spectra of a mixed solution. Background Art
[0002] Currently, the acquisition of Raman spectral data for mixed solutions primarily relies on two methods: simulated synthesis and experimental acquisition. The simulated synthesis method primarily generates theoretical spectral data by linearly superimposing the collected pure substances and setting the superposition ratio. However, due to its linear superposition, this method cannot accurately simulate the peak shift characteristics of real spectra and therefore has significant limitations. The experimental acquisition method primarily utilizes specialized experimental equipment to directly measure the Raman spectra of mixed solutions to obtain actual data. This method has the following drawbacks: 1. It is time-consuming: Each data acquisition requires meticulous preparation of experimental conditions and operation of complex equipment, and the measurement process is time-consuming, making large-scale data acquisition inefficient. 2. It consumes a significant amount of resources: It involves expensive experimental equipment, large amounts of chemical reagents, and energy consumption, and long-term use of the equipment increases wear and maintenance costs. 3. It is sensitive to environmental interference: Environmental factors such as temperature, humidity, and light can easily interfere with the data, affecting its accuracy and reliability. 4. It has a limited data collection capacity: It struggles to meet the large-scale dataset requirements of modern data analysis and processing.
[0003] In view of the shortcomings of the above two methods, it is necessary to propose a new method for generating Raman spectra of mixed solutions. Summary of the Invention
[0004] In view of the deficiencies in the prior art, the main purpose of this application is to provide a method, device, medium and equipment for generating Raman spectra of mixed solutions. This application aims to obtain more accurate Raman spectra data of mixed solutions.
[0005] To achieve the above objectives, this application provides the following technical solutions:
[0006] A method for generating a Raman spectrum of a mixed solution comprises: obtaining first Raman spectrum data of a first solution, second Raman spectrum data of a second solution, and any mixing ratio parameter of the first solution and the second solution; preprocessing the first Raman spectrum data, the second Raman spectrum data, and the any mixing ratio parameter; constructing a mixed solution Raman spectrum generation model and training the model; and inputting the preprocessed first Raman spectrum data, the second Raman spectrum data, and the any mixing ratio parameter into the mixed solution Raman spectrum generation model to obtain Raman spectrum data of a mixed solution consisting of the first solution and the second solution.
[0007] Optionally, preprocessing the first Raman spectral data, the second Raman spectral data, and the mixing ratio parameter includes: performing data expansion on the first Raman spectral data, the second Raman spectral data, and the mixing ratio parameter, respectively; and performing data sorting on the first Raman spectral data, the second Raman spectral data, and the mixing ratio parameter after data expansion.
[0008] Optionally, the mixed solution Raman spectrum generation model includes: an input layer, a first encoder, a second encoder, a feature fusion layer, a splicing layer and a decoder, wherein the input layer includes a first branch and a second branch, the first branch is used to input the spliced first Raman spectrum data and the second Raman spectrum data, and the second branch is used to input the third Raman spectrum data of the mixed solution composed of the first solution and the second solution; the first encoder is used to encode the first Raman spectrum data and the second Raman spectrum data to obtain the pure substance characteristics of the first solution and the second solution; the second encoder is used to encode the third Raman spectrum data to obtain the mixture characteristics of the mixed solution; the feature fusion layer is used to fuse the pure substance characteristics of the first solution and the second solution and the mixture characteristics of the mixed solution to obtain a fused feature; the splicing layer is used to splice the fused feature with the mixing ratio parameter to obtain an output vector; and the decoder is used to decode the output vector to obtain the Raman spectrum data of the mixed solution.
[0009] Optionally, the mixed solution Raman spectrum generation model is trained by the following steps: collecting Raman spectrum data of the first solution and the second solution in different mixing ratios and Raman spectrum data of a mixed solution composed of the first solution and the second solution in different mixing ratios, and dividing them into a training set and a validation set; setting training parameters, and training the model through the training set. During the training process, when the loss function converges, the model training is completed; verifying the trained model through the validation set. If the mean absolute error, mean square error, and root mean square error, which are model performance evaluation indicators, are all less than a threshold, the model verification is passed; otherwise, adjusting the training parameters or expanding the training set samples to retrain the model until the model verification is passed.
[0010] The present application also provides a mixed solution Raman spectrum generating device, which includes: an acquisition module for acquiring first Raman spectrum data of a first solution, second Raman spectrum data of a second solution, and a mixing ratio parameter of the first solution and the second solution; a preprocessing module for preprocessing the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter; a model construction and training module for constructing a mixed solution Raman spectrum generation model and training the model; and a generation module for inputting the preprocessed first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter into the mixed solution Raman spectrum generation model to generate Raman spectrum data of a mixed solution consisting of the first solution and the second solution.
[0011] Optionally, the preprocessing module includes: an expansion submodule for performing data expansion on the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter respectively; and a sorting submodule for performing data sorting on the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter after the data expansion.
[0012] The present application also provides a storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the method as described in any of the preceding items.
[0013] The present application also provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the program.
[0014] This application can bring the following beneficial effects:
[0015] This application achieves high-precision modeling of the Raman spectrum of mixed solutions through a dual encoder structure combined with a feature fusion layer of the QKV attention mechanism; the decoder effectively restores the peak morphology and suppresses noise interference through multi-level lightweight upsampling, a hybrid attention mechanism, and a spectral peak positioning module, thereby significantly improving the signal-to-noise ratio (SNR) of the generated spectrum; at the same time, the introduction of mixing ratio parameters and joint optimization of multiple loss functions in model training can accurately control the mixing ratio while ensuring the restoration of spectral details. Compared with the traditional linear superposition method, it can reduce the mean square error and solve the problems of low efficiency of the experimental acquisition method and insufficient accuracy of the simulation synthesis method. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a method for generating a Raman spectrum of a mixed solution provided in one embodiment of the present application;
[0017] Figure 2 is a schematic diagram of the ratio structure of the first solution and the second solution;
[0018] Figure 3 This is a schematic diagram of constructing positive and negative sample pairs provided by another embodiment of the present application;
[0019] Figure 4 This is a schematic structural diagram of a mixed solution Raman spectrum generation model provided by another embodiment of the present application;
[0020] Figure 5 This is a schematic diagram of the structure of block_base and block_deform provided in another embodiment of the present application;
[0021] Figure 6 1 is a schematic diagram of a comparison between a spectrum obtained based on the method described in the present application and a true spectrum, provided in another embodiment of the present application;
[0022] Figure 7 This is a schematic diagram comparing the spectrum obtained by the linear superposition method and the true spectrum;
[0023] Figure 8 is a schematic structural diagram of a storage medium provided by another embodiment of the present application;
[0024] Figure 9 This is a structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0027] In the present invention, unless otherwise specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0028] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme in which A and B are satisfied at the same time. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in this field to implement. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0029] Figure 1 An exemplary embodiment of the present application provides a method for generating Raman spectra of a mixed solution with controllable ratio, such as Figure 1 Said method comprises the following steps:
[0030] S100: Acquire first Raman spectrum data of a first solution, second Raman spectrum data of a second solution, and any mixing ratio parameter of the first solution and the second solution;
[0031] In this step, the first solution can be recorded as A, the second solution can be recorded as B, and the mixed solution of the first solution and the second solution is recorded as AB, and the mixing ratio of the two is recorded as P_AB. In this embodiment, the mixing ratio of the two is increased by 10% (e.g. Figure 2 As shown, the ratio of the first solution to the second solution can be set to any one of 100:0, 10:90, 20:80, 30:70, 40:60, 50:50, 60:40, 70:30, 80:20, 90:10, 100:0). Alternatively, an interval of 5% or less can be used to generate a more refined mixing result as required. By configuring multiple groups of samples with different mixing ratios, M groups of data are constructed, each group containing [(A, B, AB, P_AB), (P_AB, GT_AB), where AB is used as both an input feature and the corresponding true label GT_AB, and finally a training set is formed as shown in the figure. Figure 3 The positive and negative sample pairs shown.
[0032] S200: Preprocessing the first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter;
[0033] S300: Constructing a Raman spectrum generation model for mixed solutions and training the model;
[0034] S400: Inputting the pre-processed first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter into the mixed solution Raman spectrum generation model to obtain Raman spectrum data of the mixed solution consisting of the first solution and the second solution.
[0035] In another exemplary embodiment, in step S200, preprocessing the first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter includes the following steps:
[0036] S201: performing data expansion on the first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter respectively;
[0037] S202: performing data sorting on the first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter after data expansion.
[0038] In another exemplary embodiment, in step S300, the mixed solution Raman spectrum generation model includes: an input layer, a first encoder, a second encoder, a feature fusion layer, a splicing layer, and a decoder.
[0039] The input layer includes two branches, the first branch is used to input the spliced first Raman spectrum data and the second Raman spectrum data, and the second branch is used to input the third Raman spectrum data of the mixed solution composed of the first solution and the second solution. Figure 4 As shown, the first Raman spectrum data is expressed as [a1, a2, a3, ..., a1024] (each data represents a spectrum intensity), and the second Raman spectrum data is expressed as [b1, b2, b3, ..., b1024] (each data represents a spectrum intensity).
[0040] The first encoder corresponds to the first branch of the input layer, and the first encoder includes 6 layers of block_base (such as Figure 5As shown, each block_base has the same structure, including a one-dimensional standard convolution layer + a batch normalization layer + a ReLU activation function. The six one-dimensional standard convolution layers have a convolution kernel of 5, a stride of 2, and the number of channels is 128, 256, 512, 512, 1024, and 1024, respectively. The first encoder takes the first Raman spectrum data and the second Raman spectrum data as input, and gradually compresses the input length and expands the number of channels through six layers of block_base. The first three layers use a convolution kernel with a stride of 2 (kernel = 5) to quickly reduce the dimensionality of the input first Raman spectrum data and the second Raman spectrum data to 160×512. The last three layers alternately use convolution kernels with stride of 1 and stride of 2 to further compress the reduced dimensionality of the first Raman spectrum data and the second Raman spectrum data to 1024×5. Among them, each one-dimensional standard convolutional layer captures the local features (such as peak shape and noise distribution) and global patterns (such as the correlation between peaks) of the first and second Raman spectral data in turn, and the final output dimension is 1024×5 (1024 represents the high-order spectral features of the encoded pure substance (such as molecular bond vibration mode), and 5 represents the discretized representation of the corresponding spectral key frequency bands (such as the 5 main spectral peak areas)). The latent vector {c1,c2,c3,...,cN,N=1024} (pure substance characteristics of the first solution and the second solution).
[0041] The second encoder corresponds to the second branch of the input layer. The structure of the second encoder is similar to that of the first encoder, except that the fourth and fifth layers of block_base in the first encoder are replaced by block_deform (deformable convolution layer). The second encoder takes the third Raman spectrum data as input, and the first three layers use block_base to extract the mixture features in the third Raman spectrum data. The fourth and fifth layers are replaced by block_deform, as shown in the following figure. Figure 5As shown, the block_deform first generates a horizontal offset (activation function: tanh) through an independent convolutional layer Conv1D (3×1) to dynamically adjust the position of the feature map; second, it uses average interpolation to perform offset correction on the original feature map to capture the nonlinear offset characteristics of the spectral peak. The second encoder finally outputs a potential feature vector {d1, d2, d3,…, dN, N=1024} (mixture characteristics of the mixed solution) with a dimension of 1024×5 (1024 represents the complex spectral characteristics of the encoded mixture (such as spectral peak offset and overlapping peak morphology), 5 represents the frequency band distribution characteristics under the mixing action, and complements the pure substance characteristics of the first encoder). It should be noted that by replacing both the fourth and fifth layers of block_base in the first encoder with block_deform, the second encoder achieves the following technical benefits: 1. By introducing a learnable offset, block_deform allows the convolution kernel's sampling position to be dynamically adjusted based on the local features of the input data, thereby overcoming the limitations of traditional convolution's fixed geometric structure. In Raman spectroscopy, the spectral peaks of different samples may undergo slight shifts or shape changes due to experimental conditions (such as temperature and concentration). block_deform automatically aligns these deformations, thereby improving the robustness of feature extraction. 2. Fixed-size convolution kernels (such as 3×1) can only capture local patterns within a fixed range and have limited ability to characterize complex spectral peaks (such as overlapping peaks and broad peaks). Based on the local characteristics of the input spectrum (such as peak intensity gradients), the offset of block_deform guides the convolution kernel to focus on key areas (such as peak tops and baseline inflection points). 3. Raman spectra are often accompanied by noise (such as fluorescence background and random noise), and traditional convolution amplifies this noise due to its fixed sampling position. Block_deform can avoid high-noise areas (such as areas with drastic baseline fluctuations) through offsets and prioritize sampling areas with high signal-to-noise ratios (such as sharp spectral peaks).
[0042] The feature fusion layer includes three parallel 1D convolution branches, each using a convolution kernel of different sizes (3×1, 5×1, 7×1), with a step size of 1, and the output channel remains consistent (for example, 256). The inputs of the three parallel 1D convolution branches are the splicing results of the output of the first encoder (pure substance features) and the output of the second encoder (mixture features) in the channel dimension. By capturing narrow peak details (3×1), main peak morphology (5×1) and wide peak contour (7×1) through different convolution kernels, the coverage capability of the frequency band can be improved, and local and global patterns can be constructed. In addition, the mixing ratio parameter is first expanded to an embedding vector P1 (dimension 1024) through the first embedding layer (Embedding_1), that is, P1=Embedding_1(P AB), and then input into the feature fusion layer. The feature fusion layer also includes a multi-scale feature channel fusion module, which is used to splice the outputs of the three parallel 1D convolution branches in the channel dimension to obtain multi-scale fusion features (number of channels = 256×3 = 768), and then compress the number of channels to the original dimension (256) through 1×1 convolution to reduce the computational complexity of the model. The compressed multi-scale features are spliced with the embedding vector P1 in the channel dimension to form the final fusion features. The feature fusion layer adopts the QKV attention mechanism, which takes the final fusion features as input. On the one hand, the QKV attention mechanism first splices the pure substance features and the vector P1 in the channel dimension to obtain the spliced features, that is, C⊕P1. Then, the spliced features are compressed to the target dimension (such as 256) through the fully connected layer through the following formula to obtain the query matrix Q, which is used to retrieve key information in the mixture features. The query matrix Q is expressed as:
[0043]
[0044] On the other hand, the QKV attention mechanism extracts "keys" and "values" from the mixture features of the second encoder for matching and weighted fusion with the query matrix Q. Specifically, the output of the second encoder generates keys (K) and values (V) through two fully connected layers, namely:
[0045] K=Dense(D), V=Dense(D)
[0046] Among them, the key (K) is used to characterize the importance of each position in the mixture feature; the value (V) includes the specific information of the mixture feature and is used for weighted summation.
[0047] Furthermore, the QKV attention mechanism fuses the query matrix Q, key (K) and value (V) to obtain the fused feature r through the following formula:
[0048]
[0049] in, Represents a normalization factor, which is used to prevent the dot product result from being too large, causing the gradient to be unstable.
[0050] Furthermore, the fused feature r is mapped to a probability distribution to enhance the diversity of generation. Specifically, the mean mu = {m1,m2,m3,…,m5*N} and variance sigma {s1,s2,s3,…,s5*N} of each position of the final feature r (i.e., r1,r2,r3,…,r5*N) are first calculated through the fully connected layer (Dense):
[0051] mu=Dense(r)
[0052] sigma=(r)
[0053] Second, we sample the noise ε from a normal distribution to generate a latent vector Z:
[0054] Z=mu+σ·sigma Z={z1,z2,z3,…,zN}:
[0055] Wherein, mu represents mean; sigma represents variance; ε represents noise; and N = 1024.
[0056] The concatenation layer is used to convert the one-dimensional data obtained by sampling into a potential vector Z (dimension: 1024×5) and a mixing ratio parameter (dimension: 1024×1, generated by the second embedding layer (Embedding_2)) and concatenate them in the channel dimension to obtain an output vector (dimension: 2048×5) as the input of the decoder.
[0057] The decoder includes a first lightweight upsampling layer (PixelShuffle, upsampling factor = 2) + adaptive example normalization (AdaIN) module, a 1×1 convolution layer, a first transposed convolution layer (convolution kernel = 5, step size = 2) + hybrid attention mechanism, a second lightweight upsampling layer (PixelShuffle, upsampling factor = 2) + spectral peak location module, a second transposed convolution layer (convolution kernel = 5, step size = 2) + hybrid attention mechanism, a third lightweight upsampling layer (PixelShuffle, upsampling factor = 2) + adaptive example normalization (AdaIN) module, a third transposed convolution layer (convolution kernel = 5, step size = 2) + hybrid attention mechanism, a fourth transposed convolution layer (convolution kernel = 5, step size = 2) and an output layer. Below, this application describes in detail the decoding process of the decoder for the input vector:
[0058] First, the output vector obtained after splicing is upsampled by the first lightweight upsampling layer to obtain a feature map with a dimension of 512×10, and the adaptive instance normalization (AdaIN) is used to dynamically inject the mixing ratio parameter into the feature map to achieve normalization and layer-by-layer modulation of the feature map feature distribution, and output a feature map with a dimension of 512×10; then, the number of channels of the feature map with a dimension of 512×10 is compressed from 512 to 256 through 1×1 convolution, and then upsampled by the first transposed convolution layer to obtain a feature map with a dimension of 128×20, and the feature map is weighted optimized using the hybrid attention mechanism (hybrid attention The power mechanism includes channel attention and spatial attention, among which channel attention is used to calculate the importance weight of each channel and suppress irrelevant frequency bands; spatial attention is used to generate a spatial weight map and focus on the spectral peak area), and outputs a feature map of dimension 128×20 after weighted optimization; the second lightweight upsampling layer upsamples the feature map of dimension 128×20 to obtain a feature map of dimension 64×40. At the same time, the spectrum peak positioning module uses 1D convolution to extract the spectrum peak position in the feature map of dimension 64×40, generate a spectrum peak mask, and output a feature map of dimension 64×40 to preliminarily locate the key spectrum peak; the second transposed convolution layer upsamples the feature map of dimension 128×20 to obtain a feature map of dimension 64×40. The feature map with a dimension of 64×40 is transposed convolutionally performed, and optimized by the hybrid attention mechanism, and the output feature map with a dimension of 32×80 is output; the third lightweight upsampling layer upsamples the feature map with a dimension of 32×80, and normalizes and modulates layer by layer through adaptive example normalization, and outputs a feature map with a dimension of 16×160; the third transposed convolutional layer transposes the feature map with a dimension of 16×160, and optimizes it by the hybrid attention mechanism, and outputs a feature map with a dimension of 8×320; the fourth transposed convolutional layer transposes the feature map with a dimension of 8×320, and outputs a feature map with a dimension of 1×1280 Feature map; the output layer outputs high-resolution spectral data through the linear activation function Sigmoid (Output_1 = {new_ab1, new_ab2, …, new_ab1024}, representing the mixed spectral data of the first solution and the second solution generated by the model), and at the same time extracts global information from the features, and predicts the mixing ratio parameters through global average pooling and fully connected layers (Output_2 = {new_ab1, new_ab2, …, new_ab1024} representing the mixing ratio of the first solution and the second solution), achieving the dual-objective output of spectrum generation and ratio control.It should be noted that the prediction of Output_2 is not redundant, but part of the self-supervision mechanism. For example, during the training phase, the model can detect potential data deviations or noise interference by comparing the predicted ratio Output_2 with the actual ratio (input of step S100), thereby optimizing the feature extraction capability of the encoder-decoder. That is, Output_2 serves as a supervisory signal during the training phase, rather than a necessary output in actual applications. During the data preprocessing phase, if there is noise or error in the mixing ratio parameter, the prediction task of Output_2 can enhance the robustness of the model to proportion disturbances. In summary, the mixing ratio parameter (known input) is a necessary condition for the model to generate a spectrum, and the predicted output (Output_2) is an internal verification and constraint of the input ratio. For example, assuming the input mixing ratio is 50:50, the spectrum generated by the model should reflect this ratio, and the predicted value of Output_2 should be close to 50:50. If the two are inconsistent, the model parameters can be adjusted by a loss function (such as a component ratio loss function) to ensure alignment of the input and output.
[0059] The decoder gradually reconstructs the low-dimensional latent vector (encoding the high-order spectral features of pure substances) into a high-resolution spectral output by alternating stacking of multi-level lightweight upsampling layers (PixelShuffle) and transposed convolution layers, combined with the dynamic feature distribution alignment of the adaptive example normalization (AdaIN) module, the key area focusing of the hybrid attention mechanism, and the physical prior constraints of the spectral peak positioning module. Its core function is to accurately restore the spectral peak shape, frequency band position and baseline trend, while suppressing noise interference and background fluctuations. Among them, the lightweight upsampling layer achieves efficient resolution improvement through channel rearrangement, which can avoid the blurring artifacts of traditional interpolation; the AdaIN module is based on the dynamic feature of the input feature. The normalization parameters are adjusted in a state to ensure the consistency of spectral styles of different molecular structures; the hybrid attention mechanism can enhance the detailed expression of key spectral peak areas and weaken irrelevant noise through the joint weighting of channel and spatial dimensions; the spectral peak positioning module explicitly embeds the physical properties of the spectrum (such as peak shape symmetry and frequency band continuity), which can constrain the reconstruction process to conform to the distribution law of the real spectrum; the multi-level transposed convolution layer gradually expands the receptive field through large kernel convolution, which can capture the global correlation between broad peak profile and baseline drift, and ultimately achieve high-fidelity and highly robust spectral reconstruction, thereby balancing computational efficiency and detail restoration capabilities, and is suitable for tasks such as spectral recovery, mixed unmixing and cross-device migration in complex noise environments.
[0060] In another exemplary embodiment, in step S300, the mixed solution Raman spectrum generation model is trained by the following steps:
[0061] S301: Collecting Raman spectral data of a first solution and a second solution at different mixing ratios and Raman spectral data of a mixed solution composed of the first solution and the second solution at different mixing ratios, and dividing the data into a training set and a validation set, for example, with a division ratio of 7:3;
[0062] S302: Set training parameters. For example, booster selects the tree-based model gbtree by default, sets learning_rate to 0.01, and max_depth to 3. The model is trained using the training set. During the training process, when the loss function converges, the model training is completed.
[0063] S303: The trained model is verified through the validation set. If the mean absolute error (MAE), mean square error (MSE) and root mean square error (RMSE), which are the model performance evaluation indicators, are all less than the threshold (the threshold of MAE is set to 0.05, and the threshold of MSE is set to 0.001), the model verification is passed; otherwise, the training parameters are adjusted (for example, learning_rate can be adjusted to 0.005, and max_depth can be adjusted to 5) or the training set samples are expanded (for example, the division ratio is adjusted to 8:2) to retrain the model until the model verification is passed.
[0064] In this step, the loss function of the model training consists of three parts. The first part is the reconstruction loss function L recon , expressed as:
[0065]
[0066] Where N represents the length of the spectral data, i represents the wave number sequence, represents the spectral intensity value of wave number i, Spectral data representing the true labels.
[0067] The second part is the KL divergence, which is used to control the data distribution in the latent space and is expressed as follows:
[0068] L KL =KL(Q(z|x)||P(z))
[0069] Among them, Q(z|x) is the posterior distribution, and P(z) uses the normal distribution.
[0070] The third part is the component ratio loss function, which is expressed as:
[0071]
[0072] Where N represents the length of the spectral data, Represents the predicted proportion result, Represents true proportional data.
[0073] The final loss function is expressed as follows:
[0074] L total =α·L pf +β·L KL +γ·L pf
[0075] Among them, α, β, and γ represent hyperparameters used to balance the impact of each loss function. For example, they can be set to 1, 0.005, and 0.001 respectively.
[0076] Below, the present application uses methanol and ethanol as the first solution and the second solution respectively to exemplify the scheme described in the present application.
[0077] 1. Sample preparation and data collection
[0078] Under 785nm laser excitation, the Raman spectrum of methanol was collected, as shown in Table 1:
[0079] Table 1
[0080] <![CDATA[Wave number (cm- 1 )]]> Strength (analog value) Characteristic peak description 500 0.12±0.08 Baseline noise ... ... ... 1035 2.85±0.15 CO stretching vibration (main peak) 1450 0.45±0.10 CH bending vibration 2845 1.60±0.12 CH symmetrical stretching vibration ... ... ... 3100 0.10±0.05 Baseline noise
[0081] Under the same conditions, the Raman spectrum of ethanol was collected, as shown in Table 2:
[0082] Table 2
[0083] <![CDATA[Wave number (cm- 1 )]]> Strength (analog value) Characteristic peak description 500 0.15±0.07 Baseline noise ... ... ... 880 1.90±0.13 CCO stretching vibration 1050 3.10±0.18 CO stretching vibration (main peak) 1450 0.60±0.11 <![CDATA[CH2 bending vibration]]> 2920 2.20±0.15 CH stretching vibration ... ... ... 3100 0.12±0.06 Baseline noise
[0084] Preparation of mixed solution: Prepare a methanol-ethanol mixed solution (volume ratio 1:1) and collect its mixed spectrum, as shown in Table 3:
[0085] Table 3
[0086] <![CDATA[Wavenumber (cm- 1 )]]> Strength (analog value) Characteristic peak description 500 0.14±0.12 Mixed baseline noise ... ... ... 880 0.95±0.15 Ethanol CCO peak (intensity halved) 1035 1.40±0.20 Methanol CO peak (overlapping with ethanol peak) 1050 1.55±0.22 Ethanol CO peak (overlapping with methanol peak) 2845 0.80±0.18 Methanol CH peak (partially submerged) 2920 1.10±0.17 Ethanol CH peak (partially submerged) ... ... ... 3100 0.11±0.08 Mixed baseline noise
[0087] 2. Compare with standard data
[0088] The standard data are shown in Table 4:
[0089] Table 4
[0090] substance <![CDATA[Key peak position (cm- 1 )]]> Theoretical strength Peak shape Methanol 1035 3.00 Gauss Peak 2845 1.80 Symmetrical peak ethanol 880 2.00 Gauss Peak 1050 3.20 Slight right deviation 2920 2.40 Symmetrical peak
[0091] 3. The model reconstruction results (after unmixing the mixed spectrum) are shown in Table 5:
[0092] Table 5
[0093]
[0094] 4. Conclusion
[0095] 1. Methanol (1035cm -1 ) and ethanol (1050cm -1 ) in the overlapping region of CO peaks, the model successfully separated the two, and the peak spacing restoration error was only 0.5cm-1 (the actual 15cm -1 →Reconstruction 14.5cm -1 ).
[0096] 2. The original mixed spectrum signal-to-noise ratio (SNR=12dB) is improved to SNR=25dB or more after reconstruction by the decoder (e.g. ethanol 2920cm -1 Peak District).
[0097] 3. Reconstruction of the CO peak of methanol at half maximum width (FWHM = 12 cm -1 ) and standard data (FWHM=11.5cm -1 ) deviation is <5%, which is in line with the resolution limit of the spectrometer.
[0098] Figure 6 1 is a schematic diagram of a comparison between a spectrum obtained based on the method described in the present application and a true spectrum, provided in another embodiment of the present application; Figure 7 This is a schematic diagram comparing the spectrum obtained by the linear superposition method with the true spectrum; based on Figure 6 It can be seen that the peak position error of the mixed spectrum generated by the method described in this application (such as the methanol CO peak at 1035 cm-1 and the ethanol CO peak at 1050 cm-1) compared with the true spectrum is only 0.03% (about 0.3 cm-1), and the peak height error is ≤3.3%. This shows that the model described in this application, through the dual encoder structure and QKV attention mechanism, can accurately capture the slight shift and intensity changes of the spectral peaks, overcoming the systematic bias caused by the traditional linear superposition method that ignores nonlinear effects (such as intermolecular interactions and solvent effects). Figure 7 It is shown that the linear superposition method simply weights the spectra of pure substances, resulting in the shift of key peak positions (such as the methanol CO peak shifts to 1040 cm-1, with an error of 5 cm-1) and severe distortion of peak intensity (error ≥ 25%). The fundamental reason is that the nonlinear superposition characteristics of spectral peaks in the actual mixing process (such as peak broadening and baseline drift) are not taken into account.
[0099] Below, this application performs MSE mean statistics on the validation set of all 9 ratios. The comparison of the results of several methods is shown in Table 6:
[0100] Table 6
[0101]
[0102] Table 6 compares the performance of five different structures on error evaluation metrics. The results show that the "Deformable Convolution + QKV Fusion" method has the lowest error value (0.00030), significantly outperforming the other methods, indicating higher accuracy in feature extraction. In contrast, the traditional linear method has the highest error (0.0016), demonstrating its limitations in representing complex features. Meanwhile, the "dual-branch" structure (0.00053) also demonstrates superior performance, outperforming the "single pure substance branch" (0.00090) and "single mixture branch" (0.00087), demonstrating the positive effects of fusing multiple types of information in the task.
[0103] In another exemplary embodiment, the present application also provides a device for generating a Raman spectrum of a mixed solution with controllable ratio, the device comprising: an acquisition module 100 for acquiring first Raman spectrum data of a first solution, second Raman spectrum data of a second solution, and a mixing ratio parameter of the first solution and the second solution; a preprocessing module 200 for preprocessing the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter; a model construction and training module 300 for constructing a mixed solution Raman spectrum generation model and training the model; a generation module 400 for inputting the preprocessed first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter into the mixed solution Raman spectrum generation model to generate Raman spectrum data of a mixed solution consisting of the first solution and the second solution.
[0104] Based on the above embodiments, Figure 8 , for an explanation of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 8 The computer-readable storage medium shown is a CD-ROM 40, on which a computer program (i.e., a program product) is stored. When executed by a processor, the computer program implements each step described in the above method embodiment, for example, obtaining first Raman spectral data of a first solution, second Raman spectral data of a second solution, and any mixing ratio parameter of the first solution and the second solution; preprocessing the first Raman spectral data, the second Raman spectral data, and any mixing ratio parameter; constructing a mixed solution Raman spectral generation model and training the model; and inputting the preprocessed first Raman spectral data, the second Raman spectral data, and any mixing ratio parameter into the mixed solution Raman spectral generation model to obtain Raman spectral data of a mixed solution composed of the first solution and the second solution. The specific implementation methods of each step are not repeated here.
[0105] It should be noted that the computer-readable storage medium includes but is not limited to phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which are not listed here one by one.
[0106] Based on the above embodiments, the present application also provides an electronic device, which is described below with reference to Figure 9 An electronic device for file downloading according to an exemplary embodiment of the present application is described.
[0107] Figure 9 A block diagram of an exemplary electronic device 50 suitable for implementing the embodiments of the present application is shown. The electronic device 50 may be a computer system or a cloud server. Figure 5 The electronic device 50 shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0108] like Figure 9 As shown, the electronic device 50 includes but is not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).
[0109] The electronic device 50 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 50, including volatile and non-volatile media, removable and non-removable media.
[0110] The system memory 502 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022. The electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example only, the ROM 5023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 9 is not shown in the , usually referred to as "hard drive"). Although not in Figure 9As shown in FIG, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 503 via one or more data medium interfaces. System memory 502 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of various embodiments of the present application.
[0111] A program / utility 5025 having a set (at least one) of program modules 5024 may be stored, for example, in system memory 502. Such program modules 5024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 5024 generally implement the functions and / or methods of the embodiments described herein.
[0112] The electronic device 50 can also communicate with one or more external devices 504 (such as a keyboard, a pointing device, a display, etc.). Such communication can be performed through an input / output (I / O) interface 505. In addition, the electronic device 50 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network such as the Internet) through a network adapter 506. Figure 9 As shown, the network adapter 506 communicates with other modules (such as the processing unit 501, etc.) of the electronic device 50 via the bus 503. It should be understood that although Figure 9 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 50 .
[0113] The processing unit 501 executes various functional applications and data processing by running programs stored in the system memory 502. For example, it obtains first Raman spectral data of a first solution, second Raman spectral data of a second solution, and any mixing ratio parameter of the first and second solutions; preprocesses the first and second Raman spectral data, and any mixing ratio parameter; constructs and trains a mixed solution Raman spectral generation model; and inputs the preprocessed first and second Raman spectral data, and any mixing ratio parameter, into the mixed solution Raman spectral generation model to obtain Raman spectral data of a mixed solution composed of the first and second solutions. The specific implementation of each step will not be repeated here. It should be noted that although the detailed description above mentions several units / modules or sub-units / sub-modules of the concurrent file download device, this division is merely exemplary and not mandatory. In practice, according to embodiments of the present application, the features and functions of two or more units / modules described above may be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above may be further divided and embodied by multiple units / modules.
[0114] In the description of this application, it should be noted that the terms "first", "second" and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance.
[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0118] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0119] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a cloud server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0120] The above embodiments are intended only to illustrate the technical concepts and features of this application. Their purpose is to enable those familiar with the art to understand the content of this application and implement it accordingly. They are not intended to limit the scope of protection of this application. Any equivalent changes or modifications made in accordance with the spirit of this application shall be included in the scope of protection of this application.
Claims
1. A method for generating Raman spectra of a mixed solution, characterized in that: The method comprises: Acquire first Raman spectrum data of the first solution, second Raman spectrum data of the second solution, and any mixing ratio parameter of the first solution and the second solution; preprocessing the first Raman spectrum data, the second Raman spectrum data, and any mixing ratio parameter; Construct a Raman spectrum generation model for mixed solutions and train the model; The preprocessed first Raman spectrum data, second Raman spectrum data and any mixing ratio parameter are input into the mixed solution Raman spectrum generation model to obtain Raman spectrum data of the mixed solution consisting of the first solution and the second solution.
2. The method according to claim 1, characterized in that Preprocessing the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter includes: performing data expansion on the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter respectively; The first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter after data expansion are sorted.
3. The method according to claim 1, characterized in that The mixed solution Raman spectrum generation model includes: Input layer, first encoder, second encoder, feature fusion layer, concatenation layer and decoder, where The input layer includes a first branch and a second branch, the first branch is used to input the spliced first Raman spectrum data and the second Raman spectrum data, and the second branch is used to input the third Raman spectrum data of the mixed solution composed of the first solution and the second solution; The first encoder is used to encode the first Raman spectrum data and the second Raman spectrum data to obtain pure substance characteristics of the first solution and the second solution; The second encoder is used to encode the third Raman spectrum data to obtain a mixture characteristic of the mixed solution; The feature fusion layer is used to fuse the pure substance features of the first solution and the second solution and the mixture features of the mixed solution to obtain a fusion feature; The splicing layer is used to splice the fusion features and the mixing ratio parameters to obtain an output vector; The decoder is used to decode the output vector to obtain Raman spectrum data of the mixed solution.
4. The method according to claim 1, wherein The mixed solution Raman spectrum generation model is trained by the following steps: Collecting Raman spectral data of the first solution and the second solution at different mixing ratios and Raman spectral data of a mixed solution composed of the first solution and the second solution at different mixing ratios, and dividing the data into a training set and a validation set; Set the training parameters and train the model using the training set. During the training process, when the loss function converges, the model training is completed; The trained model is verified using the validation set. If the mean absolute error, mean square error, and root mean square error, which are the model performance evaluation indicators, are all less than the threshold, the model verification is passed; otherwise, the training parameters are adjusted or the training set samples are expanded to retrain the model until the model verification is passed.
5. A mixed solution Raman spectrum generating device, characterized in that: The device comprises: an acquisition module, configured to acquire first Raman spectrum data of the first solution, second Raman spectrum data of the second solution, and a mixing ratio parameter of the first solution and the second solution; a preprocessing module, configured to preprocess the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter; Model building and training module, used to build a mixed solution Raman spectrum generation model and train the model; A generation module is used to input the preprocessed first Raman spectrum data, second Raman spectrum data and any mixing ratio parameter into the mixed solution Raman spectrum generation model to generate Raman spectrum data of a mixed solution consisting of the first solution and the second solution.
6. The device according to claim 5, characterized in that The pre-processing module comprises: an expansion submodule, configured to perform data expansion on the first Raman spectrum data, the second Raman spectrum data, and the mixing ratio parameter respectively; The arranging submodule is used to arrange the first Raman spectrum data, the second Raman spectrum data and the mixing ratio parameter after data expansion.
7. A storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, enable the computer to execute the method according to any one of claims 1 to 4.
8. An electronic device, characterized in that: The electronic device comprises: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.