Detection and analysis system for oil-control soothing extract
Through multi-spectral synchronous acquisition and advanced data processing technology, the problem of poor selectivity of UV-vis spectrophotometry in the detection of complex plant extracts is solved, and high-precision and high-selectivity detection of active ingredients is achieved.
Patent Information
- Application Number
- CN202510216109.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-27
AI Technical Summary
The existing ultraviolet-visible spectrophotometry is poorly selective in the detection of complex plant extracts, resulting in signal crossing and interference, affecting the accuracy of quantitative analysis and the reliability of results.
Multi-spectral synchronous acquisition technology is used to capture the UV, infrared and Raman spectral data of the samples, fuse it into a high-dimensional data framework, and the signals of the target active ingredient are separated through precise calibration, sparse decomposition, cross-modal verification, dynamic convolution and attention mechanisms, and converted into concentration data through segmented modeling and geometric constraints.
It improves the detection accuracy and selectivity of active ingredients in complex plant extracts, overcomes the problem of signal cross-interference, and ensures the accuracy of quantitative analysis and the credibility of results.
Smart Images

Figure CN120043983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of extract detection, and more particularly, to a detection and analysis system for oil-control and soothing extracts. Background Art
[0002] In the detection process of oil-control and soothing extracts, ultraviolet-visible spectrophotometry (UV-Vis), as a commonly used analytical method, determines its chemical components or active substances by measuring the absorption of ultraviolet or visible light at specific wavelengths by the sample. However, the selectivity of the UV-Vis spectroscopy is poor, and this defect is particularly obvious in the analysis of complex plant extracts. Plant extracts contain a variety of polar and non-polar compounds, and their absorption spectra may overlap or produce similar absorption characteristics, making it difficult to accurately distinguish the contributions of different components.
[0003] Specifically, the absorption wavelengths of substances such as flavonoids, phenolic acids, and triterpenoids are relatively close. In the same spectral range, multiple compounds may absorb simultaneously and cause signal crossover, resulting in interference between components. This phenomenon not only affects the accuracy of quantitative analysis of active ingredients but also reduces the reliability of experimental results. In addition, due to changes in solvents, pH values, and other environmental factors, the UV-Vis spectroscopy also has a solvent effect, further affecting the accuracy of the measurement results.
[0004] These deficiencies make the detection technology based on UV-Vis face great challenges in dealing with complex plant extracts, and there is an urgent need for a new method to improve its selectivity and resolution ability to ensure accuracy and credibility.
[0005] To solve the above problems, a technical solution is provided now. Summary of the Invention
[0006] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a detection and analysis system for oil-control and soothing extracts. By synchronously collecting multi-spectral data to capture the full-range characteristics of the sample - the ultraviolet spectrum reveals absorbance changes, the infrared spectrum locks in functional group vibrations, and the Raman spectrum tracks molecular backbone displacements - and fusing this information into a high-dimensional data framework to avoid the limitations of a single spectrum. Then, through precise calibration to remove noise, combined with sparse decomposition and cross-modal verification, it is ensured that the extracted component signals truly reflect chemical properties and are not masked by redundant interference. On this basis, dynamic convolution and attention mechanisms are used to further isolate the clear signals of the target active ingredients and eliminate the influence of irrelevant backgrounds. Finally, through piecewise modeling and geometric constraints, the absorbance changes are accurately converted into concentration data, which remains stable even under background fluctuations or non-linear conditions. This progressive approach not only solves the problems of signal crossover and poor selectivity but also significantly improves the quantitative accuracy, making the detection results reliable in complex systems and providing a practical and efficient solution for the component analysis of oil-control and soothing extracts to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A detection and analysis system for oil-control and soothing extracts, comprising:
[0009] A full-spectrum co-acquisition module, a matrix deconstruction module, a spectral correlation analysis module, and a curvature inversion module;
[0010] The full-spectrum co-acquisition module: Synchronously collects various spectral data of the sample using a combined ultraviolet, infrared, and Raman spectroscopy, and generates a multi-dimensional spatio-temporal feature tensor through comprehensive weighting;
[0011] The matrix deconstruction module: After performing multivariate baseline calibration on the multi-dimensional spatio-temporal feature tensor, executes sparse-constrained non-negative matrix factorization to obtain a preliminary screened component spectrum, and optimizes the decomposition result through cross-modal absorption peak symmetry correlation verification and absorption continuity topological closure detection;
[0012] The matrix deconstruction module: Based on the optimized component spectrum, constructs a correlation matrix of ultraviolet, infrared, and Raman spectra, extracts multi-spectral feature cross-response values using deformable convolution kernels, and generates an active ingredient separation spectrum through attention probability density;
[0013] The curvature inversion module: Performs spectral band probability binning on the active ingredient separation spectrum, establishes a differential mapping relationship between the absorbance gradient field and the concentration, and uses an adaptive fern classifier for variable-weight local linear fitting, and finally fuses the hyperbolic curvature constraint term to generate the concentration time course of the target component.
[0014] In a preferred embodiment, the process of obtaining the multi-dimensional spatio-temporal feature tensor is as follows:
[0015] Collect the absorbance of the full ultraviolet-visible spectrum band, the functional group vibration characteristics of the Fourier transform infrared spectrum, and the molecular skeleton displacement signals of the Raman spectrum; after data collection, record the time-varying matrix of ultraviolet absorbance through a quartz cuvette, use the dynamic compensation algorithm and multiple scans averaging to improve the quality of the infrared signal, and apply Lorentz fitting to extract the Raman peak position parameters; subsequently, rely on the FPGA clock signal to align the time axis and establish a non-linear mapping table to associate the three-spectrum features, and calibrate the parametric features through the first derivative, second derivative, and full width at half maximum respectively, generating a structured parameter library containing absorbance, vibration derivative values, and displacement gradients; on this basis, construct a three-level molecular fingerprint index, and use the KL divergence algorithm to weight and fuse the three-spectrum features to form a multi-dimensional spatio-temporal feature tensor composed of time, spectral parameters, and feature weights.
[0016] In a preferred embodiment, after performing multivariate baseline calibration on the multi-dimensional spatio-temporal feature tensor, perform sparse constraint-based non-negative matrix factorization to obtain the preliminary screened component spectra, and the specific process is as follows:
[0017] First, based on the multi-dimensional spatio-temporal feature tensor database, perform baseline correction on each time slice in a per-modal manner; after completing the baseline correction, expand the three-dimensional spectral data matrix into a three-dimensional data matrix by time slice, reconstruct it into a generalized two-dimensional matrix, and use sparse constraint-based non-negative matrix factorization to extract the preliminary candidate component spectra and concentration time courses.
[0018] In a preferred embodiment, the specific steps of cross-modal absorption peak symmetry degree correlation verification are as follows:
[0019] For each candidate component spectrum, expand a predetermined length on both sides of the maximum value position of its absorption peak, and calculate its left and right gradient fields; at the main peak position encoded by the vibration band, extract the first derivative of the normalized intensity within the range of ±10 cm -1 as the gradient value of the vibration feature; by projecting the gradient fields of the ultraviolet, infrared, and Raman modalities into a shared embedding space, construct a cross-modal symmetric manifold matrix and calculate its co-monotonicity score:
[0020] Intra-modal monotonicity weight: m = UV, IR, Raman, where w m represents the monotonicity weight of modality m; n is the number of characteristic peaks in each modality; i represents the index of each gradient point participating in the calculation; g m,left and g m,right represent the left gradient and right gradient of modality m respectively; corr is the Pearson correlation coefficient; UV, IR, and Raman represent the ultraviolet spectrum, infrared spectrum, and Raman spectrum respectively;
[0021] Cross-modal cross-spectrum monotonicity index: where is the cross-modal cross-spectrum monotonicity index; w m and w m′ represent the symmetry weights of modes m and m′, reflecting the strength of each mode in terms of gradient symmetry; JS is the Jensen-Shannon divergence; overlap(g m ,g m′ ) calculates the overlapping region of the signals of modes m and m′ on the wavelength or frequency axis.
[0022] In a preferred embodiment, the specific steps for absorption continuity topological closure detection are as follows:
[0023] For each time slice, calculate its second derivative along the wavelength and vibration band Use the Vietoris-Rips complex to generate a persistence diagram and extract the birth-death radius pairs (b l ,d l ) of the residual structure; Gradient field dispersion index: Gradient field dispersion index: where l is the index of the hole; C is the gradient field dispersion index; is the gradient field dispersion index; δ is the smoothing factor; entropy(p(d l )) is the entropy value of the hole lifetime distribution; p(d l ) represents the probability distribution of the hole duration, describing the lifetime of each hole.
[0024] In a preferred embodiment, finally, by integrating the cross-spectrum monotonicity index and the gradient field dispersion index construct a judgment coefficient Λ; if the judgment coefficient is greater than the critical threshold, it is considered that the preliminary screening component spectrum satisfies cross-modal consistency and low residual interference, and the decomposition result is retained; if the judgment coefficient is less than or equal to the critical threshold, non-convex optimization recalibration is triggered.
[0025] In a preferred embodiment, the construction process of the correlation matrix is as follows:
[0026] Concatenate the ultraviolet, infrared, and Raman eigenvectors of each time slice in sequence into a long vector. To solve the problem of inconsistent dimensions, adjust the feature length through a one-dimensional convolutional layer, fill the missing parts with zeros and align the band ranges; then, apply a set of learnable weight matrices to the concatenated vector for linear transformation. Each set of weight matrices generates a latent eigenvalue with a fixed dimension, and after transformation, it is processed using a smoothing non-linear function based on the Gaussian error distribution to enhance the feature expression ability; finally, a correlation matrix is formed, including the number of time slices, the total number of concatenated features, and the latent space dimension, for subsequent cross-modal analysis.
[0027] In a preferred embodiment, the multi-spectral feature cross-response values are extracted by a deformable convolution kernel, and the specific process is as follows:
[0028] To achieve dynamic sampling, first, the front and rear local windows of each time slice are extracted from the correlation matrix, and the offset is predicted through a two-layer linear network. The first layer compresses the hidden feature dimension, and the second layer outputs the adjustment values of each convolution kernel sampling point in three directions, forming a five-dimensional data structure describing the offset field. Then, for each time slice and ultraviolet wavelength point, the sampling position is adjusted according to the offset field, and data is extracted from the correlation matrix. When the sampling point falls on a non-integer position, the exact value is calculated by bilinear interpolation. Then, the sampling values are weighted and summed by a set of learnable 3×3×3 weights to generate the cross-response value. The cross-response value reflects the coupling strength between the ultraviolet wavelength point and the infrared and Raman features within a local time, and finally forms a two-dimensional data table with the dimension of the number of time slices multiplied by the number of ultraviolet wavelengths.
[0029] In a preferred embodiment, the separation spectrum of the active ingredient is generated by the attention probability density, and the specific process is as follows:
[0030] The values of each time slice and ultraviolet wavelength point are extracted from the cross-response value, and feature representations are generated through two linear transformations: one transformation maps a single response value into a high-dimensional vector, and the other transformation maps all response values of the corresponding time slice into a high-dimensional matrix. Then, the dot product of the two is calculated and scaled, and a normalization function is applied to obtain the attention weight, which represents the degree of attention of each wavelength point to the global modal feature. Then, each intensity value of the original ultraviolet absorption spectrum is modulated: taking its original value and multiplying it by a modulation factor, which is calculated by the attention weight and the cross-response value. Specifically, the weight of each wavelength point is distributed to the infrared and Raman dimensions, a compression function is applied to the corresponding response value, and then weighted and summed. The modulated result forms a new two-dimensional data table with the dimension of the number of time slices multiplied by the number of ultraviolet wavelengths, in which the ultraviolet signals strongly correlated with the infrared and Raman features are enhanced, and the non-target signals are suppressed, thereby generating the separation spectrum of the target active ingredient.
[0031] The spectral band probability binning is performed on the separation spectrum of the active ingredient, the differential mapping relationship between the absorbance gradient field and the concentration is established, and the variable-weight local linear fitting is performed using an adaptive fern classifier. Finally, the hyperbolic curvature constraint term is fused to generate the concentration time course of the target component. The specific steps are as follows:
[0032] Starting from the separated ultraviolet absorption intensity data, the time axis is segmented according to a fixed time window. For the absorbance sequences at each wavelength within each window, kernel density estimation is used to generate a probability density distribution, and it is evenly divided into multiple intervals based on the cumulative probability. After statistically analyzing the sample proportion, a relationship model between the absorbance change rate and the concentration change rate is constructed. Subsequently, an adaptive fern classifier is introduced. Binary discrimination conditions are constructed by randomly selecting features from the sub-compartments to divide the regions. For the samples within each region, a local linear model is fitted by distance weighting, and low-precision ferns are eliminated according to the fitting error, and new ferns are added to optimize the modeling. Finally, the sub-compartment concentration estimation is mapped to the hyperbolic space, the geodesic distance is used to penalize the non-smoothness, and after back-projection, it is weighted and fused in combination with the sample proportion and the curvature loss to generate the final concentration time course.
[0033] The technical effects and advantages of a detection and analysis system for an oil-control and soothing extract of the present invention:
[0034] The detection and analysis system for an oil-control and soothing extract of the present invention improves the detection accuracy and selectivity of active ingredients in complex plant extracts through multi-spectral spatio-temporal synchronous acquisition and feature fusion, and overcomes the problem of signal cross-interference caused by spectral overlap of compounds such as flavonoids, phenolic acids, and triterpenoids and solvent effects in the traditional ultraviolet-visible spectrophotometry. First, the omnidirectional features of the sample are captured through multi-spectral synchronous acquisition - the ultraviolet spectrum reveals the absorbance change, the infrared spectrum locks the functional group vibration, and the Raman spectrum tracks the molecular skeleton displacement - and these information are fused into a high-dimensional data framework to avoid the limitations of a single spectrum. Then, noise is removed through precise calibration, and combined with sparse decomposition and cross-modal verification to ensure that the extracted component signals can truly reflect the chemical characteristics and are not masked by redundant interference. On this basis, the clear signals of the target active ingredients are further separated by using dynamic convolution and attention mechanism, and the influence of irrelevant backgrounds is eliminated. Finally, through piecewise modeling and geometric constraints, the absorbance change is accurately converted into concentration data, and it can remain stable even under background fluctuations or non-linear conditions. This progressive approach not only solves the problems of signal cross and poor selectivity, but also greatly improves the quantitative accuracy, making the detection results still credible in complex systems, and providing a practical and efficient solution for the component analysis of oil-control and soothing extracts. Brief Description of the Drawings
[0035] Figure 1 It is a structural schematic diagram of the present invention.
[0036] Figure 2 It is a schematic diagram of the steps of the robust fusion process with hyperbolic curvature constraint of the present invention. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] Embodiment 1: Figure 1 A detection and analysis system for an oil-control and soothing extract of the present invention is provided, including: a full-spectrum co-acquisition module, a matrix deconstruction module, a spectral correlation analysis module, and a curvature inversion module;
[0039] Full-spectrum co-acquisition module: Synchronously collect various spectral data of the sample using ultraviolet, infrared, and Raman combined spectroscopy, and generate a multi-dimensional spatio-temporal feature tensor through comprehensive weighting;
[0040] Matrix deconstruction module: After performing multivariate baseline calibration on the multi-dimensional spatio-temporal feature tensor, perform sparse constraint-based non-negative matrix factorization to obtain a preliminary screening component spectrum, and optimize the decomposition result through cross-modal absorption peak symmetry correlation verification and absorption continuity topological closure detection;
[0041] Matrix deconstruction module: Based on the optimized component spectrum, construct a correlation matrix of ultraviolet, infrared, and Raman spectra, extract multi-spectral feature cross-response values using deformable convolution kernels, and generate an active ingredient separation spectrum through attention probability density;
[0042] Curvature inversion module: Perform spectral band probability binning on the active ingredient separation spectrum, establish a differential mapping relationship between the absorbance gradient field and the concentration, use an adaptive fern classifier for variable-weight local linear fitting, and finally fuse the hyperbolic curvature constraint term to generate the concentration time course of the target component.
[0043] The process of obtaining the multi-dimensional spatio-temporal feature tensor is as follows:
[0044] In the spectral analysis of the oil-control and soothing extract, through multi-spectral spatio-temporal synchronous acquisition and feature fusion, high-dimensional data support is provided for the component analysis of complex solution systems. The full-spectrum co-acquisition module relies on triple spectroscopy combined technology to synchronously collect ultraviolet-visible spectroscopy (UV-Vis), Fourier transform infrared spectroscopy (FTIR), and Raman spectroscopy to form a high-dimensional data set, comprehensively capturing the molecular structure and chemical reaction characteristics of the sample.
[0045] To ensure the efficient combination of the three spectra, a triple optical path combination system is used to ensure synchronization and precision. The ultraviolet light source uses a deuterium-tungsten lamp, covering the wavelength range of 190 - 1100 nm, ensuring accurate acquisition of the absorbance of the sample within the full wavelength range. The infrared part is equipped with a mid-infrared ATR module (4000 - 400 cm-1, resolution 4 cm-1), and provides efficient signal transmission through a liquid ATR accessory (ZnSe crystal). The Raman spectrum is equipped with a 785 nm laser (power 50 mW), which is combined with a backscattering collection structure to ensure high-quality Raman signals. The three spectral channels are triggered for timing synchronization through an FPGA chip, ensuring the time-domain precision of each spectral data within a single measurement cycle (the acquisition error of the three spectra is less than 1 millisecond).
[0046] Data acquisition process:
[0047] (1) Ultraviolet-visible spectroscopy (UV-Vis): During the acquisition of the ultraviolet-visible spectrum, the sample is injected into a quartz cuvette (optical path 10 mm) for full wavelength scanning with a scanning step of 1 nm. In each scanning cycle, the ultraviolet spectrum records two-dimensional matrix data of the sample absorbance changing with time. Special attention is paid to the evolution of the characteristic absorption peaks of chromophores such as benzene rings (200 - 280 nm) and conjugated double bonds (250 - 300 nm). This data contains changes in time (rows) and wavelength (columns), and can reveal the dynamic absorption characteristics of chemical components at different wavelengths.
[0048] (2) Fourier transform infrared spectroscopy (FTIR): The acquisition of the FTIR spectrum relies on a liquid ATR accessory. The sample interacts with infrared radiation through a ZnSe crystal to capture the vibration characteristics of functional groups. To eliminate water vapor interference, a dynamic compensation algorithm is used to ensure the stability of the data. Sampling mainly focuses on key regions such as hydroxyl groups (3200 - 3600 cm-1) and carbonyl groups (1650 - 1750 cm-1). The average value is taken through 64 consecutive scans to improve the signal quality and reduce noise interference.
[0049] (3) Raman spectroscopy: The Raman spectrum is excited using a 785 nm laser, with the integration time set to 1 second and three repeated measurements are carried out. Inelastic scattered light is removed through a filter, and the displacement signal (300 - 1800 cm-1) region is collected. Attention is paid to characteristic regions such as sulfur groups (500 - 550 cm-1) and carbon chain backbone vibrations (1000 - 1200 cm-1). For these key regions, the Lorentz fitting method is applied to extract parameters such as peak position and full width at half maximum, and the time-course data is recorded. This process can reveal the changes and dynamic characteristics of the sample molecular structure.
[0050] Construction of the triple spectrum correlation database:
[0051] (1) Physical space-time alignment: To ensure the precise temporal alignment of ultraviolet, infrared, and Raman spectroscopy data, the acquisition moments of the three spectra are synchronized at the microsecond level using the FPGA clock signal. The time alignment of sampling points every 5 milliseconds ensures the consistency of data acquisition. In the wavelength dimension, a non-linear mapping table is established: for example, a quantum chemical correlation chain is established between the ultraviolet maximum absorption peak (257 nm) of the C=C bond, the infrared vibration wavenumber (1250 cm-1) of C-O, and the Raman shift (660 cm-1) of C-S. This correlation table provides the mapping of physical and chemical characteristics between different spectra.
[0052] (2) Feature parameterization: To further process these multi-dimensional data, first, the first derivative spectrum of the ultraviolet spectrum is extracted to identify the shoulder peak positions and eliminate the influence of noise. The infrared spectrum is processed by the second derivative to resolve overlapping peaks and extract more detailed functional group information. The Raman spectrum uses the full width at half maximum to calibrate the broadening effect of the instrument and extract important peak position information. After these processes, the time-varying characteristics of the three spectra are parameterized, and a structured parameter library is generated. The data structure of each time slice includes: ultraviolet absorbance (500×1 dimension), infrared second derivative value (1600×1 dimension), Raman shift correction gradient (1500×1 dimension). These data lay the foundation for subsequent fusion and analysis.
[0053] (3) Molecular fingerprint indexing: To facilitate the extraction of valuable information from large-scale data, a three-level retrieval hierarchy is constructed:
[0054] The first layer: The samples are binned according to the ultraviolet maximum absorption wavelength, with each interval being 10 nm (such as 200 - 210 nm, 210 - 220 nm, etc.) for preliminary classification.
[0055] The second layer: The infrared spectrum is split into 11 standard vibration band segments, and each segment is assigned an eight-bit binary vibration mode code to achieve finer-grained molecular fingerprint recognition.
[0056] The third layer: By correlating the centroid coordinates (represented by three-dimensional scale values) of the Raman shift characteristic peak groups, a composite index tree is formed to support multi-dimensional reverse queries. This hierarchical structure ensures that relevant component information can be quickly and efficiently queried in large-scale data sets.
[0057] Data fusion and feature weighted fusion:
[0058] Based on feature parameterization, the KL divergence algorithm is used to perform weighted fusion on ultraviolet, infrared, and Raman spectral features. KL divergence can measure the differences between different features, thereby allocating appropriate weights to different spectral information. Through this method, the three spectral data are dynamically weighted to generate the final multi-dimensional spatio-temporal feature tensor (dimension: time × spectral parameter × feature weight). This tensor reflects the dynamic chromophore concentration of ultraviolet absorbance, the evolution path of functional groups in infrared vibration modes, and the molecular backbone vibration trajectory revealed by Raman shift gradients, providing a comprehensive feature basis for subsequent intelligent component separation and analysis.
[0059] The full-spectrum co-acquisition module eliminates the spectral acquisition time-domain error through synchronous triggering technology, combines the multi-dimensional data fusion and feature weighted fusion of ultraviolet, infrared, and Raman spectra, and provides reliable data support for the component analysis of the oil-control and soothing extract. Through this innovative data acquisition and processing method, not only the selectivity of the spectrum is optimized, but also a high-dimensional and accurate data basis is provided for subsequent intelligent analysis and component separation.
[0060] After performing multivariate baseline calibration on the multi-dimensional spatio-temporal feature tensor, sparse constraint-based non-negative matrix factorization is performed to obtain the preliminary screening component spectrum. The specific process is as follows:
[0061] First, based on the multi-dimensional spatio-temporal feature tensor database, baseline correction is performed on each time slice in each mode to eliminate various noises and unnecessary background interferences.
[0062] For ultraviolet-visible spectra, infrared vibration intensities, and Raman shift gradients, calibration methods corresponding to their respective characteristics are used to remove baseline drift and interferences:
[0063] Ultraviolet-visible spectrum baseline correction: The moving smoothing and adaptive polynomial fitting methods are used. By setting the window width to 10 nm and the order of the polynomial to 3, the interference of scattered light is eliminated, thereby obtaining more accurate absorbance data.
[0064] Infrared vibration intensity baseline correction: The wavelet decomposition (Symlets-4 basis function, 4-layer decomposition) technology is used to remove low-frequency drift, thereby ensuring better resolution of the infrared spectrum in vibration characteristics.
[0065] Raman shift gradient baseline correction: The time-shifted median filtering method (the sliding window is set to 20 time slices) is used to suppress the fluorescence baseline and enhance the analyzability of the Raman spectrum.
[0066] These calibration methods are optimized for the spectral characteristics of different modes to ensure the quality of the input data for subsequent decomposition.
[0067] After completing the baseline correction, the three-dimensional spectral data matrix is expanded into a three-dimensional data matrix by time slice Reconstructed into a generalized two-dimensional matrix Use non-negative matrix factorization with sparse constraints to extract the preliminary candidate component spectra and concentration time courses. The objective function of sparse constraint non-negative matrix factorization NMF is:
[0068]
[0069] , where: is the sparse component spectrum matrix, k is the preset number of components, λ is the number of wavelengths, and v is the vibration band coding dimension; is the sparse concentration time course matrix, t represents the time dimension of the data, corresponding to different time points during data acquisition or the number of different samples; α and β are the L1 regularization weights, used to enforce the locality characteristics and temporal sparsity of the components.
[0070] Reconstruction error This term represents the reconstruction error of the three-dimensional data matrix X by W and H. The goal of optimization is to minimize this error, that is, to make the gap between WH and X as small as possible.
[0071] Sparsity constraint ||W|| 1 and ||H|| 1 : The L1 regularization term forces the elements of W and H to be as close to zero as possible, thereby achieving feature sparsification. This makes the decomposed W and H more interpretable. L1 regularization can prompt the solution to contain only a few important features, reduce redundancy, and enhance the interpretability and generalization ability of the model.
[0072] The goal of the optimization process is to solve the minimum value of the above objective function. Since the objective function is non-convex and there are multiple local optimal solutions, non-negative matrix factorization optimization usually uses the alternating least squares method or the gradient descent method to iteratively update W and H. In each iteration, by fixing one matrix (for example, fixing H and optimizing W), and then updating the other matrix until the objective function converges.
[0073] The objective function of non-negative matrix factorization finds the optimal sparse component spectrum matrix and sparse concentration time course matrix by balancing the reconstruction error of the data and the sparsity of the matrix, thereby decomposing the three-dimensional data matrix into a low-dimensional matrix that is easy to interpret. The introduction of the L1 regularization term makes the columns of each basis matrix and the rows of the coefficient matrix sparse, thereby enhancing the interpretability of the decomposition result and the generalization ability of the model.
[0074] After the preliminary decomposition, by calculating the absorption peak symmetry across modalities (ultraviolet-visible spectroscopy, infrared spectroscopy, and Raman spectroscopy), the cross-modal cross-spectrum monotonicity index is obtained to verify the physical rationality and spatio-temporal alignment of the decomposition. The specific steps are as follows:
[0075] Ultraviolet-Visible Spectral Feature Peak Gradient Field: For each candidate component spectrum, at the position of the maximum absorption peak λ p extend a predetermined length, such as 5 nm, to the left and right respectively, and calculate its left and right gradient fields:
[0076] Left gradient: g left (λ) = S(λ + 1) - S(λ), which represents the degree of change of the signal between the position (λ + 1) on the left side of the wavelength λ and the position of the wavelength λ, reflecting the rate of signal change.
[0077] Right gradient: g right (λ) = S(λ) - S(λ - 1), which represents the degree of change of the signal between the position of the wavelength λ and its position on the right side (λ - 1), reflecting the rate of signal change.
[0078] S(λ) represents the signal intensity or absorption intensity at the wavelength λ. S is a function of the signal, representing the absorption or other measurement values at a specific wavelength.
[0079] Infrared / Raman Spectral Feature Peak Gradient Field: At the position of the main peak encoded by the vibration band, extract the first derivative of the normalized intensity within the range of ±10 cm -1 as the gradient value of the vibration feature.
[0080] Manifold Similarity Calculation: By projecting the gradient fields of the ultraviolet, infrared, and Raman modalities into a shared embedding space, construct a cross-modal symmetric manifold matrix and calculate its co-monotonicity score:
[0081] Intra-modal Monotonicity Weight:
[0082] m = UV, IR, Raman
[0083] , where w m represents the monotonicity weight of modality m; n is the number of characteristic peaks of each modality; i represents the index of each gradient point participating in the calculation; g m,left and g m,right represent the left gradient and right gradient of modality m respectively; corr is the Pearson correlation coefficient.
[0084] UV, IR, and Raman represent the ultraviolet spectrum, infrared spectrum, and Raman spectrum respectively.
[0085] Cross-modal Cross-spectrum Monotonicity Index:
[0086]
[0087] , where is the cross-modal cross-spectrum monotonicity index; w m and w m′Denote the symmetry weights of modes m and m′, which reflect the strength of each mode in terms of gradient symmetry; JS is the Jensen-Shannon divergence, used to measure the similarity between the gradients (g m and g m′ ) of modes m and m′. The Jensen-Shannon divergence is a symmetric distance metric, and the smaller the value, the more similar the two distributions; overlap(g m ,g m′ ) calculates the overlapping region of the signals of modes m and m′ on the wavelength or frequency axis.
[0088] To further ensure the rationality of the decomposed residuals, by incorporating continuous topological closure detection, the gradient field dispersion index is obtained to identify potential redundancies or inappropriate decompositions. The specific steps are as follows:
[0089] Construction of the residual gradient field complex: For each time slice, calculate its second derivatives along the wavelength and vibration band and construct a discrete Morse function for further analysis of the topological structure of the residuals.
[0090] Calculation of persistent homology: Use the Vietoris-Rips complex to generate a persistence diagram and extract the birth-death radius pairs (b l ,d l ) of the residual structure. Among them, the Vietoris-Rips complex is a topological structure constructed using the neighborhood relationship of the point set, which helps to capture low-dimensional topological features in high-dimensional data.
[0091] Gradient field dispersion index:
[0092]
[0093] , the gradient field dispersion index:
[0094]
[0095] , where l is the index of the hole, traversing all residual holes; C is the gradient field dispersion index, measuring the complexity of the residual structure; is the gradient field dispersion index; δ is a smoothing factor to avoid numerical instability in the calculation; entropy(p(d l )) is the entropy value of the hole lifetime distribution, representing the complexity of the hole lifetime distribution; p(d l ) represents the probability distribution of the hole duration, describing the lifetime of each hole.
[0096] Finally, by comprehensively calculating the cross-spectrum monotonicity index and the gradient field dispersion index Construct the judgment coefficient Λ to determine the effectiveness of the decomposition result. Example of the calculation formula:
[0097]
[0098] , where and are the standard deviations of the cross-spectrum monotonicity index and the gradient field dispersion index in historical iterations, used for dynamic normalization; γ1 is the dispersion penalty weight (empirical value).
[0099] Judgment rule:
[0100] If the judgment coefficient is greater than the critical threshold, it is considered that the initially screened component spectrum satisfies cross-modal consistency and low residual interference, and the decomposition result is retained;
[0101] If the judgment coefficient is less than or equal to the critical threshold, non-convex optimization recalibration is triggered: when non-convex optimization recalibration is triggered, the joint sparsity and smoothness of the sparse component spectrum matrix and the sparse concentration matrix are constrained by non-convex regularization terms, and the proximal gradient descent method is used for optimization, and non-negativity is ensured through adaptive step size and hard threshold truncation; subsequently, the misaligned peaks of the optimized spectral lines are corrected, the overlapping spectral lines are merged and the concentration time course is redistributed; the Savitzky-Golay filter is applied to the concentration curve to eliminate high-frequency noise; iterative optimization is performed until the judgment coefficient is greater than the critical threshold or the maximum number of iterations is reached, and the finally converged sparse component spectrum matrix and sparse concentration time course matrix are output, the decomposition kernel matrix of the fusion tensor database is updated and the residual boundary is corrected. Specifically:
[0102] Add a non-convex regularization term ρ·||W⊙H|| to the objective function of sparse constraint-based non-negative matrix factorization 2,1 , forcing the column sparsity and row smoothness of the kernel matrix.
[0103] ρ·||W⊙H|| 2,1 is a non-convex regularization term, which means that by imposing an L 2,1 norm constraint on the element-wise product (Hadamard product W⊙H) of W and H, the joint sparsity and smoothness between the component spectrum and the concentration are forced, where represents calculating the sum of the L :j norms of the element-wise product of each column of the component spectrum W j: and the corresponding concentration row H 2 This structure encourages only a few components to be active in a specific time period, suppresses the splitting of redundant components, and at the same time maintains the continuity of the spectral line and the concentration curve; the parameter ρ is the weight coefficient of the regularization term, used to balance the relative importance of this constraint and the reconstruction error in the objective function. A smaller value (such as 0.3) indicates that the data fitting accuracy is prioritized and the sparsity is not overly penalized.
[0104] Update W using the proximal gradient descent method (1) and H (1) , calculate the gradient direction, the adaptive step size (initial η = 1e-3), and apply a hard threshold truncation to ensure non-negativity.
[0105] W (1) is the sparse component spectral matrix after non-convex optimization and recalibration; H (1) is the sparse concentration time course matrix after non-convex optimization and recalibration;
[0106] Spectral line rearrangement: Perform Gaussian fitting on the optimized W (1) . If the peak spacing of the same component in different modes exceeds the device resolution (ultraviolet ±0.5 nm, infrared ±1 cm -1 , Raman ±1.5 cm -1 ), merge the overlapping spectral lines and reallocate the concentration time course.
[0107] Perform Savitzky-Golay filtering (window width of 7 time slices) on the concentration curve of H (1) to eliminate high-frequency oscillation noise.
[0108] Termination condition: Iterate until the coefficient of determination is greater than the critical threshold or the maximum number of iterations is reached, and output the final W * and H * .
[0109] W * is the finally converged sparse component spectral matrix; H * is the finally converged sparse concentration time course matrix.
[0110] The construction process of the correlation matrix is as follows:
[0111] Align the characteristic data of the three modes (ultraviolet absorption phase-frequency characteristics, infrared vibration eigenmodes, Raman displacement gradient field) along the time axis, and integrate them into a high-dimensional space through splicing and linear transformation to form a correlation matrix. This provides a unified representation basis for subsequent cross-modal feature extraction and facilitates the exploration of non-linear relationships between different modes. Specific processing:
[0112] The ultraviolet absorption phase-frequency characteristics are a two-dimensional data table with the number of rows equal to the total number of time slices (resolution of 5 milliseconds, assuming 1000 time points) and the number of columns equal to the number of ultraviolet wavelength points (e.g., from 200 nm to 800 nm, one point every 0.5 nm, a total of 1200 columns).
[0113] The infrared vibration eigenmodes are also a two-dimensional data table with the same number of rows as the total number of time slices and the number of columns equal to the number of infrared vibration band points (e.g., from 400 to 4000 wavenumbers, one point every 1 wavenumber, a total of 3600 columns).
[0114] The Raman shift gradient field is also a two-dimensional data table, with the number of rows being the total number of time slices and the number of columns being the number of Raman shift points (for example, from 100 to 3500 wavenumbers, with one point every 1.5 wavenumbers, approximately 2333 columns).
[0115] These data have been baseline calibrated and reorganized in the array deconstruction module to ensure the consistency of the time axis.
[0116] High-dimensional correlation mapping process:
[0117] Concatenate the ultraviolet, infrared, and Raman eigenvectors of each time slice in sequence into an extremely long vector. For example, if there are 1200 points in the ultraviolet, 3600 points in the infrared, and 2333 points in the Raman, the length of the vector for each time slice after concatenation is 7133.
[0118] Apply a linear transformation to this concatenated vector and project it into a high-dimensional space through a set of learnable weight matrices (the output dimension of each set of weight matrices is 256, i.e., the latent space dimension). Each weight matrix is responsible for generating one dimension in the latent space, and finally 256 latent eigenvalues are formed.
[0119] After the linear transformation, use the GeLU activation function (a smooth non-linear function based on the Gaussian error distribution) to process the result and enhance the non-linear expression ability of the features.
[0120] To solve the problem of inconsistent dimensions between modalities (such as the lack of the low-frequency band in Raman), adjust the feature length through a one-dimensional convolutional layer. For example, fill the missing area with zeros and align the band ranges of infrared and Raman before concatenation. The size of the convolutional kernel is determined according to the missing range, usually 5 or 7 points.
[0121] The obtained correlation matrix is a three-dimensional data structure, with dimensions of the number of time slices multiplied by the total number of concatenated features multiplied by the latent space dimension, such as 1000 (time) × 7133 (features) × 256 (latent dimensions). This matrix provides a unified representation of cross-modal features for subsequent analysis.
[0122] Extract the cross-response values of multi-spectral features through a deformable convolutional kernel. The specific process is as follows:
[0123] Traditional fixed convolutional kernels cannot adapt to the dynamic changes between modalities in spectral data (such as the drift of absorption peaks over time). This step introduces deformable convolution to extract the correlation between the three modalities in time and space by dynamically adjusting the sampling positions, generating cross-response values, which provide a basis for subsequent attention calculation.
[0124] Design of the deformable convolution module:
[0125] Use a 3D convolutional kernel with a size of 3×3×3, covering the time axis (1 time slice before and after, a total of 3 points), the modality axis (local feature range), and the latent space axis (local latent dimension region) respectively.
[0126] Generate an offset field for each time slice, describing the adjustment amount of each sampling point within the convolutional kernel. This offset field is a five-dimensional data structure, with the dimension being the number of time slices multiplied by the 3×3×3 grid of the convolutional kernel multiplied by 3 directions (corresponding to the lateral, longitudinal, and depth adjustments of the wavelength or vibration band).
[0127] The offset field is calculated as follows: Extract the local windows before and after each time slice from the correlation matrix (the time range is 1 before and after the current slice, a total of 3 slices), and then predict the offset by a two-layer linear network. The first layer compresses the latent features (e.g., from 256 dimensions to 128 dimensions), and the second layer outputs 27 offset values (corresponding to the adjustments of each point in the 3×3×3 kernel in 3 directions).
[0128] Cross-response value generation process:
[0129] For each time slice and ultraviolet wavelength point, the deformable convolution adjusts the sampling position according to the offset field and extracts data from the correlation matrix. The sampling position is based on the reference coordinates (e.g., the original points of ultraviolet wavelength, infrared vibration band, and Raman shift), and after adding the offset, it may fall at non-integer positions.
[0130] Perform a weighted sum on the values of these adjusted sampling points, and the weights are provided by a 3×3×3 learnable convolutional kernel. The value of each sampling point is calculated by bilinear interpolation to ensure sub-pixel-level accuracy (e.g., if the sampling position is 300.2 nm, then interpolate according to the values at 300 and 301 nm).
[0131] The finally generated cross-response value reflects the coupling strength between the ultraviolet wavelength point and the infrared and Raman features within the local time window.
[0132] Generate the separation spectrum of the active ingredient through the attention probability density, and the specific process is as follows:
[0133] Multi-modal attention weight calculation:
[0134] Based on the cross-response value, calculate the correlation weights between the ultraviolet wavelength and other modal features through the attention mechanism, perform local enhancement and global suppression on the ultraviolet absorption spectrum, and finally generate the separation spectrum of the target active ingredient. This process highlights the ultraviolet signals strongly correlated with the infrared and Raman vibration features and sparsifies the influence of non-target components.
[0135] Extract the values of each time slice and UV wavelength point from the cross-response values, and generate corresponding feature representations through two linear transformations (referred to as query and key transformations respectively). The query transformation maps a single response value into a 256-dimensional vector, and the key transformation maps all response values of this time slice into a 256-dimensional matrix.
[0136] Calculate the dot product of the query vector and the key matrix to measure the correlation between the current wavelength point and the global features. The dot product result is scaled by dividing by the square root of the latent space dimension (i.e., 16) to avoid excessive numerical values, and then normalized by the softmax function to obtain the attention weights.
[0137] These attention weights represent the degree of attention of each UV wavelength point to all modal features within this time slice, and the sum of the weights is 1.
[0138] Probability density modulation process:
[0139] Modulate each intensity value of the original UV absorption spectrum: take its original value and multiply it by a modulation factor. The modulation factor is jointly determined by the attention weights and the cross-response values.
[0140] Specifically, for each UV wavelength point, calculate the sum of the attention weights with all infrared vibration bands and Raman shift points. Assume that the attention weights are decomposed into infrared and Raman dimensions through an additional mapping (e.g., assigned by a linear layer), apply the Sigmoid function (compress the values to the range of 0 to 1) to the cross-response values of each infrared and Raman point, and then multiply and sum with the corresponding weights.
[0141] The modulated result is a new UV absorption intensity value, which enhances the signals strongly correlated with infrared and Raman features and suppresses the irrelevant signals. The finally generated separated spectrum directly reflects the UV absorption time course of the target component.
[0142] The separated spectrum of the active ingredient is a two-dimensional data table with the dimension of the number of time slices multiplied by the number of UV wavelengths, such as 1000×1200. The non-target components are sparsified due to low correlation, and the signals of the target components are highlighted.
[0143] The process of spectral band probability binning and absorbance gradient field construction is as follows:
[0144] At the beginning of the process, the ultraviolet absorption intensity data is segmented along the time axis according to a fixed time window, and each window contains a number of time slices. For each absorbance sequence at each wavelength within each window, the kernel density estimation method is used to generate the probability density distribution. A Gaussian kernel is used, and the smoothing parameter is preset according to the data characteristics to capture the distribution characteristics of the absorbance. Based on the cumulative probability of this distribution, the absorbance range is evenly divided into multiple intervals, each interval covering an equal cumulative probability share, and the proportion of absorbance samples within each interval is counted. Then, a relationship model between the absorbance change rate and the concentration change rate is established within each interval: the absorbance change rate is expressed as the concentration change rate multiplied by a sensitivity coefficient (this coefficient varies piecewise linearly with the concentration and is fitted through experimental calibration points), and a noise term (calculated from the sample standard deviation within the interval) is superimposed. This step provides a local correlation basis for the subsequent analysis of the absorbance and concentration changes.
[0145] The variable-weight local fitting process of the adaptive fern classifier is as follows:
[0146] In this step, a classifier based on random feature combinations (called fern) is used. Several are randomly selected from multiple bins as features, and the features include the central value of the bin and the sensitivity value at zero concentration. Binary discrimination conditions are set according to these features. For example, whether the absorbance is greater than the central value, and the data space is divided into multiple regions (leaf nodes). Within each region, the corresponding absorbance and concentration samples are collected, and the sample weights are calculated. The weights decay exponentially with the distance of the sample from the region mean, and the decay rate is controlled by a dynamically adjusted bandwidth. The weighted least squares method is used to fit a local linear model to describe the relationship between the concentration and the absorbance deviation. The model includes an intercept and a term of the slope multiplied by the deviation. At the same time, the fitting error of each fern (represented by the root mean square of the residuals) is monitored. If the error exceeds the preset threshold, the fern is eliminated, and a new fern is generated from the bins that have not been fully modeled, and the model accuracy is iteratively optimized. This process improves the adaptability of the concentration estimation through dynamic weighting and local regression.
[0147] As Figure 2 shown, the robust fusion process with hyperbolic curvature constraints is as follows:
[0148] In the fusion stage, the concentration estimation values of each bin are first mapped to a high-dimensional hyperbolic space (using the Poincaré ball model). The estimation values are converted to hyperbolic coordinates through normalization and exponential mapping to capture the non-linear relationship between the data.
[0149] Subsequently, the geodesic distance between adjacent bins in the hyperbolic space is calculated, and a loss function is constructed to penalize the non-smoothness of the estimation values, ensuring the local continuity of the concentration change between adjacent bins.
[0150] Finally, the hyperbolic space coordinates are back-projected onto the original space. Combining the concentration estimates and confidence levels of each sub-compartment (determined jointly by the sample proportion and the exponential decay of the curvature loss), the final concentration time course is generated through weighted average fusion. The weight during weighting is negatively correlated with the curvature loss, and the penalty intensity parameter is preset according to experience.
[0151] This step utilizes the characteristics of hyperbolic geometry to suppress noise and non-linear perturbations, enhancing the robustness of the results.
[0152] Aiming at the problems of poor selectivity and severe signal cross-interference in the detection of complex plant extracts by ultraviolet-visible spectrophotometry, the curvature inversion module correlates absorbance with concentration changes through probability-based sub-compartmentalization, adaptively weights and models the sub-compartmental data using a fern classifier, and fuses local estimates with hyperbolic curvature constraints to generate highly robust concentration inversion results. This method effectively overcomes the influence of overlapping absorption spectra of compounds such as flavonoids, phenolic acids, and triterpenoids, as well as solvent effects, improving the selectivity and accuracy of quantitative analysis of active ingredients, and is applicable to complex systems with background fluctuations or non-linear responses.
[0153] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0154] Only some exemplary embodiments of the present invention have been described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0155] It should be noted that in this article, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
[0156] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.
Claims
1. A detection and analysis system for oil-controlling soothing extracts, characterized in that: include: Full spectrum cooperative acquisition module, matrix deconstruction module, pattern-linked spectrum analysis module and curvature inversion module; Full spectrum cooperative sampling module: It uses ultraviolet, infrared and Raman coordinated spectra to synchronously collect various spectral data of samples, and comprehensively weights them to generate multi-dimensional spatiotemporal feature tensors; Array deconstruction module: After implementing multivariate baseline calibration on the multi-dimensional spatiotemporal feature tensor, sparse constrained non-negative matrix decomposition is performed to obtain the initial screening component spectrum, and the decomposition result is optimized through cross-modal absorption peak symmetry correlation verification and absorption continuity topological closure detection; Matrix deconstruction module: constructs the correlation matrix of ultraviolet, infrared and Raman spectra based on the optimized component spectra, uses deformable convolution kernel to extract the cross-response value of multi-spectral features, and generates the active ingredient separation spectrum through attention probability density; Curvature inversion module: The active ingredient separation spectrum is divided into spectral bands by probability, the differential mapping relationship between the absorbance gradient field and the concentration is established, the adaptive fern classifier is used for variable weight local linear fitting, and finally the hyperbolic curvature constraint term is integrated to generate the concentration time course of the target component.
2. The oil-controlling and soothing extract detection and analysis system according to claim 1, characterized in that: The process of obtaining the multi-dimensional spatiotemporal feature tensor is as follows: The full-band absorbance of the UV-visible spectrum, the functional group vibration characteristics of the Fourier transform infrared spectrum, and the molecular skeleton displacement signal of the Raman spectrum are collected; after data collection, the time-varying matrix of the UV absorbance is recorded by a quartz cuvette, and the quality of the infrared signal is improved by a dynamic compensation algorithm and multiple scan averages, and the Raman peak position parameters are extracted by Lorentz fitting; then, the time axis is aligned based on the FPGA clock signal and a nonlinear mapping table is established to associate the three spectral features, and the parameterized features are calibrated by the first-order derivative, second-order derivative, and half-peak width, respectively, to generate a structured parameter library containing absorbance, vibration derivative values, and displacement gradients; on this basis, a three-level molecular fingerprint index is constructed, and the KL divergence algorithm is used to weightedly fuse the three spectral features to form a multi-dimensional spatiotemporal feature tensor composed of time, spectral parameters, and feature weights.
3. The oil-controlling and soothing extract detection and analysis system according to claim 2, characterized in that: After multivariate baseline calibration is performed on the multi-dimensional spatiotemporal feature tensor, sparse constrained non-negative matrix decomposition is performed to obtain the initial screening component spectrum. The specific process is as follows: Firstly, based on the multi-dimensional spatiotemporal feature tensor database, baseline correction is performed modally for each time slice. After completing the baseline correction, the three-dimensional spectral data matrix is expanded into a three-dimensional data matrix by time slices and reconstructed into a generalized two-dimensional matrix, and sparse constrained non-negative matrix factorization is used to extract preliminary candidate component spectra and concentration time courses.
4. The oil-controlling and soothing extract detection and analysis system according to claim 3, characterized in that: The specific steps for cross-modal absorption peak symmetry correlation verification are as follows: For each candidate component spectrum, the predetermined length is extended to the left and right of the maximum absorption peak position, and the left and right gradient fields are calculated; at the main peak position of the vibration band encoding, the ±10cm -1 The first-order derivative of the normalized intensity in the interval is used as the gradient value of the vibration feature; by projecting the gradient fields of the ultraviolet, infrared and Raman modes into a shared embedding space, a cross-modal symmetric manifold matrix is constructed And calculate its co-monotonicity score: Intramodal monotonicity weight: m=UV,IR,Ramam, where w m represents the monotonicity weight of mode m; n is the number of characteristic peaks of each mode; i represents the index of each gradient point involved in the calculation; g m,left and g m,right They represent the left gradient and right gradient of mode m respectively; corr is the Pearson correlation coefficient; UV, IR and Raman represent ultraviolet spectrum, infrared spectrum and Raman spectrum respectively; Cross-modal cross-spectral monotonicity index: in is the cross-modal cross-spectral monotonicity index; w m and w m′ represents the symmetry weight of mode m and mode m′, reflecting the strength of each mode in terms of gradient symmetry; JS is the Jensen-Shannon divergence; overlap(g m ,g m′ ) calculates the overlapping area of the signals of mode m and mode m′ on the wavelength or frequency axis.
5. The oil-controlling and soothing extract detection and analysis system according to claim 4, characterized in that: The specific steps of absorbing continuity topological closure detection are as follows: For each time slice, calculate the second-order derivative along the wavelength and vibration band The Vietoris-Rips complex is used to generate persistence graphs and extract the birth-death radius pairs of the residual structure (b l ,d l ); Gradient field dispersion index: Gradient field dispersion index: Where l is the index of the hole; C is the gradient field discreteness index; is the gradient field dispersion index; δ is the smoothing factor; entropy(p(d l )) is the entropy of the hole lifetime distribution; p(d l ) represents the probability distribution of hole duration, describing the lifetime of each hole.
6. The oil-controlling and soothing extract detection and analysis system according to claim 5, characterized in that: Finally, the cross-spectral monotonicity index and the gradient field dispersion index Construct the judgment coefficient Λ; If the judgment coefficient is greater than the critical threshold, it is considered that the spectrum of the initial screening group meets the cross-modal consistency and low residual interference, and the decomposition result is retained; If the judgment coefficient is less than or equal to the critical threshold, non-convex optimization recalibration is triggered.
7. The oil-controlling and soothing extract detection and analysis system according to claim 6, characterized in that: The process of constructing the correlation matrix is as follows: The ultraviolet, infrared and Raman feature vectors of each time slice are concatenated in sequence into a long vector. To solve the problem of inconsistent dimensions, the feature length is adjusted through a one-dimensional convolutional layer, and the missing parts are filled with zeros and aligned to the band range. Then, a set of learnable weight matrices are applied to the concatenated vectors for linear transformation. Each set of weight matrices generates a latent eigenvalue of a fixed dimension, which is processed using a smooth nonlinear function based on Gaussian error distribution to improve the feature expression capability. Finally, an association matrix is formed, which includes the number of time slices, the total number of concatenated features and the latent space dimension for subsequent cross-modal analysis.
8. The oil-controlling and soothing extract detection and analysis system according to claim 7, characterized in that: The multi-spectral feature cross-response value is extracted through the deformable convolution kernel. The specific process is as follows: To achieve dynamic sampling, the local windows before and after each time slice are first extracted from the association matrix, and the offset is predicted through a two-layer linear network. The first layer compresses the hidden feature dimension, and the second layer outputs the adjustment value of each convolution kernel sampling point in three directions to form a five-dimensional data structure that describes the offset field. Then, for each time slice and ultraviolet wavelength point, the sampling position is adjusted according to the offset field, and data is extracted from the association matrix. When the sampling point falls on a non-integer position, the exact value is calculated by bilinear interpolation, and then a set of learnable 3×3×3 weights are used to weighted sum the sampling values to generate a cross-response value. The cross-response value reflects the coupling strength between the UV wavelength point and the infrared and Raman features in the local time, and finally forms a two-dimensional data table with the dimension of the number of time slices multiplied by the number of UV wavelengths.
9. The oil-controlling and soothing extract detection and analysis system according to claim 8, characterized in that: The active ingredient separation spectrum is generated by attention probability density. The specific process is as follows: The values of each time slice and ultraviolet wavelength point are extracted from the cross-response values, and feature representation is generated through two linear transformations: one transformation maps a single response value into a high-dimensional vector, and the other transformation maps all response values of the corresponding time slice into a high-dimensional matrix. The dot product of the two is then calculated and scaled, and a normalization function is applied to obtain the attention weight, which represents the degree of attention paid to the global modal features by each wavelength point. Next, each intensity value of the original ultraviolet absorption spectrum is modulated: its original value is taken and multiplied by a modulation factor. The modulation factor is calculated through the attention weight and the cross-response value. Specifically, the weight of each wavelength point is assigned to the infrared and Raman dimensions, and a compression function is applied to the corresponding response values, followed by a weighted sum. The modulated result forms a new two-dimensional data table with the dimension of the number of time slices multiplied by the number of ultraviolet wavelengths, in which the ultraviolet signals that are strongly correlated with the infrared and Raman features are enhanced and the non-target signals are suppressed, thereby generating a separation spectrum of the target active ingredient.
10. The oil-controlling and soothing extract detection and analysis system according to claim 9, characterized in that: The active ingredient separation spectrum is divided into spectral band probability compartments, the differential mapping relationship between the absorbance gradient field and the concentration is established, and the adaptive fern classifier is used for variable weight local linear fitting. Finally, the hyperbolic curvature constraint term is integrated to generate the concentration time course of the target component. The specific steps are as follows: Starting from the ultraviolet absorption intensity data after separation, the time axis is divided into fixed time windows, and the kernel density estimation is used to generate the probability density distribution of the absorbance sequence of each wavelength in each window, and it is evenly divided into multiple intervals based on the cumulative probability. After counting the sample proportion, the relationship model between the absorbance change rate and the concentration change rate is constructed; then the adaptive fern classifier is introduced, and features are randomly selected from the compartments to construct binary discriminant conditions to divide the regions. The local linear model is weighted by distance for the samples in each region, and low-precision ferns are eliminated according to the fitting error, and new ferns are supplemented for optimization modeling; finally, the compartment concentration estimate is mapped to the hyperbolic space, and the geodesic distance is used to penalize the non-smoothness. After back-projection, the sample proportion and curvature loss are weightedly fused to generate the final concentration time course.
Citation Information
Cited By
Intelligent monitoring method for sleeve gas recovery
CN120847014A
Intelligent monitoring method for casing recovery
CN120847014B
Self-adaptive fluorescent immune layer quantitative detection feature extraction method and system
CN120877013A
Accurate detection system for medicinal liquor components
CN121612833A
Method and system for rapidly estimating residual nuclide concentration of radioactive waste liquid
CN121917482A