A method and application for identifying hyaluronic acid with different degradation pathways based on water spectromics.

By using water spectromics technology, combined with near-infrared spectral data processing and model building, the problem of distinguishing between acid-hydrolyzed and enzymatically hydrolyzed hyaluronic acid has been solved, enabling efficient quality monitoring of hyaluronic acid.

CN116124733BActive Publication Date: 2026-04-03SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately distinguish between hyaluronic acid prepared by acid hydrolysis and enzymatic hydrolysis, and traditional near-infrared spectroscopy classification methods have low accuracy.

Method used

Using water spectromics technology, near-infrared spectral data of hyaluronic acid were collected, and preprocessed with standard normal transformation and continuous wavelet transform. Twelve characteristic wavelengths were selected to establish PCA-DA and PLS-DA models, enabling accurate differentiation between acid-hydrolyzed and enzymatically hydrolyzed hyaluronic acid.

Benefits of technology

It enables rapid and accurate identification of hyaluronic acid, with a recognition rate and accuracy of 100%, reducing the professional requirements for testing personnel and effectively monitoring the quality of hyaluronic acid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116124733B_ABST
    Figure CN116124733B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of rapid qualitative identification and quality control technology using near-infrared spectroscopy, specifically relating to a method and application for identifying hyaluronic acid with different degradation pathways based on aqueous spectromics. The method includes: S1, acquiring the near-infrared spectrum of the sample to be tested and obtaining spectral data; S2, performing standard normal transformation and continuous wavelet transform preprocessing on the spectral data obtained in step S1; S3, selecting characteristic wavelengths based on the principles of aqueous spectromics, establishing a model for the spectral data at the characteristic wavelengths, and making a judgment. This invention uses aqueous spectromics as its theoretical basis, selecting characteristic wavelengths that characterize the classification information of hyaluronic acid with different degradation pathways, thereby establishing a discrimination model for identifying hyaluronic acid from acid hydrolysis, enzymatic hydrolysis, and mixtures of the two. This method can rapidly, simply, and accurately identify hyaluronic acid obtained through different degradation pathways, and reduces the professional requirements for testing personnel, thus having good practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rapid qualitative identification and quality control technology of near-infrared spectroscopy, specifically involving a method and application for identifying hyaluronic acid with different degradation modes based on water spectromics. Background Technology

[0002] The information disclosed in this background section is intended only to enhance understanding of the overall background of the invention and is not necessarily to be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

[0003] Hyaluronic acid (HA), also known as hyaluronic acid, is a high-molecular-weight polysaccharide that has not been sulfated. It possesses physiological functions such as joint lubrication, wound repair, anti-aging, and promoting growth and development, and is widely used in the pharmaceutical, cosmetic, and food industries due to its unique physiological functions. The production process of hyaluronic acid determines its quality. Currently, the most common degradation methods for low-molecular-weight hyaluronic acid in production are acid hydrolysis and enzymatic hydrolysis. Enzymatic hydrolysis has mild reaction conditions, a simple process, and is pollution-free. Acid degradation has low production costs. The price of enzymatically hydrolyzed hyaluronic acid with a molecular weight below 10,000 Da is approximately 3 to 4 times that of acid-hydrolyzed products, and the structure of enzymatically hydrolyzed hyaluronic acid is more uniform. However, the hyaluronic acid obtained by the two degradation methods is structurally very similar, and there is currently no clear method to distinguish between hyaluronic acid prepared by acid hydrolysis and enzymatic hydrolysis.

[0004] Near-infrared spectroscopy is a rapid and non-destructive detection technique that has emerged in recent years. It features convenient analytical instruments, requires no sample pretreatment, and allows for real-time monitoring, making it widely used in food, pharmaceutical, and other fields. By performing chemometric analysis and machine learning on near-infrared spectral data, research objects can be quickly compared and identified. Given the similarity in structure between the two degradation pathways of hyaluronic acid, their near-infrared spectra show little difference. Therefore, traditional classification methods are not ideal for classifying their near-infrared spectral data, and the accuracy of identification needs further improvement. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method and application for identifying hyaluronic acid obtained through different degradation pathways based on aqueous spectromics. This invention utilizes aqueous spectromics as its theoretical foundation, selecting characteristic wavelengths to characterize the classification information of hyaluronic acid obtained through different degradation pathways, thereby establishing a discrimination model for identifying hyaluronic acid from acid hydrolysis, enzymatic hydrolysis, and mixtures of both. This method can rapidly, simply, and accurately identify hyaluronic acid obtained through different degradation pathways, and reduces the professional requirements for testing personnel. This invention is based on the above research findings.

[0006] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:

[0007] A first aspect of the present invention provides a method for identifying hyaluronic acid with different degradation pathways based on aqueous spectromics, the method comprising:

[0008] S1. Collect the near-infrared spectrum of the sample to be tested and obtain spectral data;

[0009] S2. Perform standard normal variable transformation and continuous wavelet transform preprocessing on the spectral data obtained in step S1;

[0010] S3. Select characteristic wavelengths based on the principles of water spectromics and establish models for the spectral data at the characteristic wavelengths;

[0011] The 12 characteristic wavelengths are located at wavelengths of 1336-1348nm, 1360-1366nm, 1370-1376nm, 1380-1388nm, 1398-1418nm, 1421-1430nm, 1432-1444nm, 1448-1454nm, 1458-1468nm, 1472-1482nm, 1482-1495nm, and 1506-1516nm.

[0012] The sample to be tested may be hyaluronic acid from acid hydrolysis, hyaluronic acid from enzymatic hydrolysis, or a mixture of the two.

[0013] Therefore, in a second aspect, the present invention provides the application of the above-described method in the quality monitoring of hyaluronic acid.

[0014] The beneficial technical effects of one or more of the above technical solutions are as follows:

[0015] The above-mentioned technical solution, based on water spectromics technology, is a method for identifying hyaluronic acid with different degradation methods. It involves collecting near-infrared spectra of acid-hydrolyzed, enzymatically hydrolyzed, and mixed samples of hyaluronic acid. The spectral data undergoes SNV and CWT preprocessing, and calibration and validation sets are divided according to the KS method. Based on the principles of water spectromics, 12 characteristic wavelengths are selected, and PCA-DA and PLS-DA models are established for the spectral data at these 12 wavelengths. The classification results show that the distribution of enzymatically hydrolyzed, acid-hydrolyzed, and mixed hyaluronic acid is relatively concentrated, and the recognition rate and accuracy of the test set are both 100%. Compared to models established without wavelength selection or using commonly used wavelength selection methods, the above method can effectively and accurately distinguish hyaluronic acid with different degradation methods.

[0016] In summary, the above technical solution acquires near-infrared spectra of hyaluronic acid with different degradation methods and extracts characteristic wavelengths containing classification information based on aqueous spectromics technology to establish a discriminant analysis model. By employing aqueous spectromics technology combined with near-infrared spectroscopy to establish the model, the quality of hyaluronic acid with different degradation methods can be effectively monitored, and the professional requirements for testing personnel are reduced, thus demonstrating significant practical application value. Attached Figure Description

[0017] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0018] Figure 1 Examples of the present invention include (a) the original near-infrared spectrum of the sample and (b) the near-infrared spectrum after pretreatment.

[0019] Figure 2 Examples of the present invention are (a) the original near-infrared spectrum of the sample after outlier removal and (b) the near-infrared spectrum after preprocessing.

[0020] Figure 3 The continuous wavelet variation spectrum is shown in the embodiment of the present invention;

[0021] Figure 4 This refers to the standard deviation of each wavelength within the coordinate range of the water matrix in this embodiment of the invention.

[0022] Figure 5 This describes the selection of the first principal component load points for PCA in this embodiment of the invention.

[0023] Figure 6 This embodiment of the invention uses the PCA-DA model based on water spectromics band selection.

[0024] Figure 7 This is the prediction result of the PCA-DA model based on water spectromics band selection in an embodiment of the present invention;

[0025] Figure 8 This embodiment of the invention uses a PCA-DA model based on four different band selection methods.

[0026] Figure 9 This is the prediction result based on the four band selection methods in the embodiments of the present invention;

[0027] Figure 10 The accuracy of model discrimination in the embodiments of the present invention;

[0028] Figure 11 The PLS-DA model is selected based on water spectromics band selection in this embodiment of the invention;

[0029] Figure 12 This is the prediction result of the PLS-DA model based on the selection of water spectromics bands in an embodiment of the present invention;

[0030] Figure 13 This embodiment of the invention is based on the PLS-DA model with four different band selection methods;

[0031] Figure 14 This is the prediction result of the PLS-DA model based on four band selection methods in an embodiment of the present invention. Detailed Implementation

[0032] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0033] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0034] As mentioned earlier, hyaluronic acid prepared by acid hydrolysis and enzymatic hydrolysis is structurally very similar, and conventional methods are difficult to distinguish accurately.

[0035] In view of this, in a typical embodiment of the present invention, a method for identifying hyaluronic acid with different degradation pathways based on aqueous spectromics is proposed, the method comprising:

[0036] S1. Collect the near-infrared spectrum of the sample to be tested and obtain spectral data;

[0037] S2. Perform standard normal variable transformation and continuous wavelet transform preprocessing on the spectral data obtained in step S1;

[0038] S3. Select characteristic wavelengths based on the principles of water spectromics and establish models for the spectral data at the characteristic wavelengths;

[0039] The 12 characteristic wavelengths are located at wavelengths of 1336-1348nm, 1360-1366nm, 1370-1376nm, 1380-1388nm, 1398-1418nm, 1421-1430nm, 1432-1444nm, 1448-1454nm, 1458-1468nm, 1472-1482nm, 1482-1495nm, and 1506-1516nm.

[0040] In step S1, the sample to be tested can be hyaluronic acid from acid hydrolysis, hyaluronic acid from enzymatic hydrolysis, or a mixture of the two.

[0041] The near-infrared spectral acquisition was performed using a near-infrared spectrometer. Specifically, the sampling method was integrating sphere diffuse reflectance, and the spectral scanning range was 4000-10000 cm⁻¹. -1 The resolution is set to 8cm. -1 The number of scans was 32. The sample was sampled 3 times, and the average value was used as the data for subsequent model building.

[0042] In step S2, the data is first preprocessed using standard normal variable transformation to remove outliers, and then the data is denoised using continuous wavelet transform.

[0043] The specific method of step S3 includes: selecting the peak and valley values ​​of principal component analysis loads within the 12 water matrix coordinates specified by water spectromics according to the point selection method of water spectromics; calculating the standard deviation of the absorption values ​​of each variable after continuous wavelet transform of the 12 water matrix coordinates; selecting the variable with the largest standard deviation in each band as the characteristic wavelength; and combining principal component analysis to obtain 12 activated water absorption bands as the final modeling data.

[0044] The wavelength positions of the 12 activated water absorption bands are specifically 1344nm, 1364nm, 1376nm, 1388nm, 1411nm, 1421nm, 1438nm, 1448nm, 1467nm, 1475nm, 1491nm, and 1505nm.

[0045] In another specific embodiment of the present invention, the method further includes: using the absorption values ​​at the above-mentioned characteristic wavelengths, dividing the calibration set and the validation set according to the KS method, and establishing a qualitative model of principal component discriminant analysis and partial least squares discriminant analysis, thereby realizing the qualitative discrimination of hyaluronic acid and its mixed products with different degradation methods (acid hydrolysis and enzymatic hydrolysis).

[0046] The classification results obtained using the above method show that the distribution of enzymatically hydrolyzed, acidically hydrolyzed, and mixed hyaluronic acid is relatively concentrated, and the test set recognition rate and accuracy are both 100%. Compared with models established without band selection and those built using commonly used band selection methods, the identification method of this invention can effectively and accurately distinguish hyaluronic acid with different degradation methods.

[0047] In another specific embodiment of the present invention, the application of the above-described method in the quality monitoring of hyaluronic acid is provided. Specifically, the present invention, through the above-described identification method, can qualitatively determine the presence of acid-hydrolyzed and enzymatically hydrolyzed hyaluronic acid, thereby enabling quality monitoring of hyaluronic acid from different degradation sources.

[0048] The following examples further illustrate the present invention, but do not constitute a limitation thereof. It should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0049] Example

[0050] 1. Materials and Instruments

[0051] 1.1 Experimental Materials

[0052] 50 batches of acid-hydrolyzed hyaluronic acid A samples

[0053] 50 batches of enzymatically hydrolyzed hyaluronic acid B samples (provided by Bloomage Biotechnology)

[0054] 1.2 Instruments and Software

[0055] Antaris II Fourier Transform Near Infrared Spectrometer (Thermo Nicolet, USA)

[0056] XSR105 analytical balance (Mettler Toledo, Switzerland)

[0057] MATLAB R2020a (The MathWorks Inc., USA)

[0058] 2 Experimental Methods

[0059] 2.1 Sample preparation and spectral acquisition

[0060] Different batches of acid-hydrolyzed hyaluronic acid A were numbered A1, A2, A3...A40, and different batches of enzymatically hydrolyzed hyaluronic acid B were numbered B1, B2, B3...B40. Hyaluronic acid A from A1 to A40 and hyaluronic acid B from B1 to B40 were randomly matched and mixed in 2% increments (2%-98% by mass) to obtain a total of 98 mixed samples, numbered C1, C2, C3...C98. Hyaluronic acid A from A41, A42, A43...A50 and hyaluronic acid B from B41, B42, B43...B50 were also randomly mixed to obtain another 20 mixed samples, numbered C99, C100, C101...C118.

[0061] Near-infrared spectra of the above 218 samples were collected. The sampling method of the near-infrared spectrometer was integrating sphere diffuse reflectance, and the spectral scanning range was 4000-10000 cm⁻¹. -1 The resolution is set to 8cm. -1 The number of scans was 32. Each sample was sampled 3 times, and the average value was used as the experimental data for subsequent model building, resulting in a total of 218 spectra, with each spectral line consisting of 1557 data points.

[0062] 2.2 Spectral Preprocessing Methods

[0063] Preprocessing of near-infrared spectral data is crucial for handling background interference or noise unrelated to sample information. Choosing an appropriate preprocessing method is an important step before building a model. In this example, spectral preprocessing includes Standard Normal Variate Transformation (SNV), Derivative algorithm, and Continuous Wavelet Transform (CWT).

[0064] 2.2.1 Standard Normal Variable Transformation (SNV)

[0065] SNV (Surface Normalization) is primarily used to eliminate the effects of solid particle size, surface scattering, and optical path transformation on diffuse reflectance. The specific calculation principle is as follows: subtract the mean value of the original spectrum from the original spectrum, and then divide by the standard deviation. Essentially, this is the standard normalization of the original spectral data. The calculation formula is as follows: It is the average value of the spectrum of the i-th sample, k = 1, 2, 3, ..., m, where m is the number of wavelength points.

[0066]

[0067] 2.2.2 Continuous Wavelet Transform (CWT)

[0068] Derivative calculation is a powerful technique in chemometrics, used for resolving spectra, sharpening peaks, and performing qualitative and quantitative analysis. While widely applied, higher-order derivative calculations can increase noise levels. To improve the signal-to-noise ratio of higher-order derivatives, denoising is typically performed before calculating continuous derivatives, which complicates the calculation. Wavelet transform (WT) is widely used in signal analysis for derivative calculations. Compared to traditional methods, this technique offers advantages such as computational simplicity, flexible parameters, and good smoothing effects. CWT, in particular, is a common tool in near-infrared spectroscopy for background removal and noise reduction to improve models. The WT formula is as follows, where f(t) is the input spectral signal, and t is the time-domain signal, which can be considered as the wavenumber. Let be the wavelet basis function, a be the translation parameter, and C be the wavelet coefficients.

[0069]

[0070] The SymletsA function is an approximately symmetric, compactly supported, bioorthogonal wavelet basis function, which is generally expressed in the form of symN.

[0071] 2.3 Principal Component Analysis (PCA)

[0072] Principal component analysis (PCA) is a multivariate statistical method used to examine the correlation between multiple variables. It reveals the internal structure of these variables through a few principal components. The specific calculation process of PCA is as follows:

[0073] (1) Combine the autocorrelation spectra of all samples into a matrix X = {x1, x2, x3, x...} n If we need to reduce the dimension to k, we first need to decentralize, that is, subtract the average value from each feature:

[0074]

[0075] (2) Calculate the decentralized matrix The specific calculation process is as follows:

[0076]

[0077] (3) Use eigenvalue decomposition to find the eigenvalues ​​and eigenvectors of the covariance matrix C. If vector ν is an eigenvector of matrix C, then there exists a corresponding eigenvalue λ, as follows:

[0078] Cν=λν

[0079] (4) For matrix C, there is a set of eigenvectors. Orthogonalizing and normalizing this set of vectors yields a set of orthogonal unit vectors Q. Eigenvalues ​​are defined for the elements on the diagonal of Σ, and their specific mathematical expressions are as follows:

[0080] C=Q∑Q -1

[0081] (5) The eigenvalues ​​are sorted in descending order of numerical value. The first k eigenvalues ​​are selected to transform the dataset X into a new space constructed by the k eigenvectors, i.e., Y = PX, thus achieving data dimensionality reduction.

[0082] 2.4 Band Selection Method

[0083] Choosing different bands can optimize the predictive ability of the model. In this example, the variable selection methods include: Competitive Adaptive Reweighted Sampling (CARS), Variable Importance in the Projection (VIP) algorithm, and Correlation Coefficient (CC).

[0084] 2.4.1 Competitive Adaptive Reweighting (CARs)

[0085] CARS is a method for selecting characteristic wavelengths of regression coefficients in a PLSr model that combines Monte Carlo sampling. CARS utilizes a combination of an exponential decay function and adaptive reweighted sampling techniques to select the variable points with the largest absolute values ​​of regression coefficients in the PLS model. This sampling process is similar to the "survival of the fittest" principle in Darwinian evolution. Each sampling run of the CARS method includes four steps:

[0086] (1) Use Monte Carlo model sampling to ensure sample randomness;

[0087] (2) Calculate the correlation coefficient bi between each wavelength variable and the first-level data, and then calculate the contribution w corresponding to each wavelength variable. i :

[0088]

[0089] Where m is the number of variables, b i Let be the regression coefficient between the i-th variable and the first-level data, where n is the number of variables, i = 1, 2, ..., n;

[0090] (3) Starting with the full spectrum, the PLSR regression coefficients with smaller absolute values ​​are reduced using an exponentially decreasing function. i Gradually remove variables, then use adaptive reweighted sampling to further filter them;

[0091] (4) Repeat the above operation to obtain N variable combinations. Based on the interactive verification method, compare these combinations and finally select the subset with the lowest RMSECV value as the optimal variable combination, where N is the number of Monte Carlo random samplings required in the whole process.

[0092] 2.4.2 Variable Importance Projection Algorithm (VIP)

[0093] The VIP algorithm is based on parameters such as the scoring factor in the PLS model. This algorithm comprehensively considers the contribution of each variable to both the independent variable (X) and the dependent variable (Y), calculating the importance index V for each variable. The specific calculation formula is as follows:

[0094]

[0095] Where F represents the values ​​of the latent variable factors in the model, w jf SSY represents the importance of the j-th variable when the number of components is f. f SSY represents the variance of the Y values ​​explained by the f-th component of the PLS model, where J represents the number of variables. total This represents the sum of the variances of the Y values ​​explained by all components of the PLS model. The weights in the PLS model reflect the covariance between the independent variable X and the dependent variable Y. The formula calculates the V value of each variable based on its importance, and selects variables with V > 1 as characteristic wavelengths.

[0096] 2.4.3 Correlation Coefficient Method (CC)

[0097] The correlation coefficient method calculates the correlation coefficient between the absorbance values ​​in each column of the spectral matrix and the primary data by regressing them. Then, band selection is based on the absolute value of the correlation coefficient, |R|. The correlation coefficient spectrum reflects the relationship between absorbance and primary data; wavelength regions showing high correlation are selected, while regions showing low or no correlation are ignored. The calculation method is as follows: [Equation omitted for brevity]. Let be the mean of n spectral data points and the mean of the primary data points, i = 1, 2, 3, ..., n. The correlation coefficient r ranges between -1 and +1, i.e., |r| ≤ 1. The closer |r| is to 1, the stronger the linear correlation between x and y.

[0098]

[0099] 2.5 Water Spectromics

[0100] Water spectromics is a new omics discipline proposed in 2005 by Professor Roumiana Tsenkova of Kobe University, Japan. This discipline aims to use water-light interactions to extract structural information about water in the spectra of the research system, and indirectly analyze the impact of perturbation factors on the system by analyzing the structure and structural changes of water.

[0101] The general research method in water spectromics is as follows: spectral data of the research system are collected, and the raw spectra are preprocessed and analyzed using chemometrics to assign them to activated water absorption bands. This determines the Water Absorbance Spectral Pattern (WASP) of the system, which is then visualized as a radar chart showing the impact of perturbations on the water structure. The WASP describes the state of the entire system and contains a wealth of physical and chemical information about water. Determining the WASP of a system under specific perturbations is the core of water spectromics research. Establishing a WASP database for different research systems allows for rapid comparison and identification of the states of the studied systems. Based on extensive previous research, water spectromics has defined 12 water matrix coordinates (WAMACS) ranges in the first overtone absorption region of water within the near-infrared spectral range. Their naming and assignment are shown in the table below.

[0102]

[0103]

[0104] Where ν refers to the OH stretching vibration of hydrogen bonds in water molecules, such as ν1 for symmetric stretching fundamental vibration, ν2 for doubly degenerate bending fundamental vibration, and ν3 for antisymmetric stretching fundamental vibration; S refers to the number of hydrogen bonds in water, such as S0 for water molecules without hydrogen bonds (i.e., free water molecules), S1 for water molecules with one hydrogen bond, and so on, with S4 for water molecules with four hydrogen bonds; OH-(H2O) n The solvation layer of water; O2-(H2O) n It refers to superoxide in water.

[0105] 2.6 Qualitative Judgment Methods

[0106] 2.6.1 Principal Component Discriminant Analysis (PCA-DA)

[0107] PCA-DA is often used in pattern recognition. Based on PCA (see 2.2.2 for the principle), we obtain the principal components (PCs) we want. We can take any two of the principal components to make a two-dimensional discriminant map based on the principal components. The first principal component contains the largest percentage of data variation, and the subsequent principal components contain data variation that gradually decreases.

[0108] 2.6.2 Partial Least Square Discriminant Analysis (PLS-DA)

[0109] PLS-DA is a statistical analysis method for multivariate data that combines least squares regression attributes with classification techniques for discriminant analysis. PLS regression is primarily used to establish statistical relationships between multiple dependent and independent variables, and its principle is as follows:

[0110] Y0+β1X1+β2X2+…+β P X P

[0111] Where β0 is the intercept and β1 is the independent variable X. i The regression coefficients, the basic PLS model is:

[0112] Y = uQ + F

[0113] X = TP + E

[0114] Where X(n×p) is the independent variable matrix, representing the spectral absorbance values ​​at p wavelengths for n samples; Y(n×m) is the dependent variable matrix, representing the primary data of m components for n samples; T and U(n×f) are the latent variable score matrices of X and Y, respectively; P(p×f) and Q(m×f) are the orthogonal loading matrices; and E and F are the fitted residual matrices. The dependent variable matrix for PLS-DA is categorical data, and a binarized dependent variable matrix is ​​used for regression modeling.

[0115] 3 Results

[0116] 3.1 Data Preprocessing

[0117] 3.1.1 Standard Normal Variable Transform (SNV)

[0118] Near-infrared raw spectra of acid-hydrolyzed, enzymatically hydrolyzed, and mixed hyaluronic acid used for model building, such as... Figure 1 As shown in (a), after SNV processing, the transformed spectrum is as follows: Figure 1 (b)

[0119] 3.1.2 Removing outliers

[0120] To avoid outliers affecting the classification performance of the discrimination model, outlier removal was performed on the 178 preprocessed spectra. B1, C49, and C50 were identified as outliers and removed from the original data. The remaining 175 spectra were then reprocessed using SNV (Search and Variants) analysis. The results are as follows: Figure 2 As shown.

[0121] 3.1.3 Continuous Wavelet Transform (CWT)

[0122] To improve the resolution and reduce the signal-to-noise ratio of near-infrared spectra, CWT was performed on 175 near-infrared spectral data after SNV. The wavelet basis was Sym2, and the scale factor was 20. The transformed spectra are shown below. Figure 3 As shown. Subsequent band selection and model building all used data processed by CWT.

[0123] 3.2 Data Partitioning

[0124] A total of 175 samples were used as modeling data, including hyaluronic acid A (A1-A40), hyaluronic acid B (B2-B40), and mixed hyaluronic acid (C1-C48, C51-C98). The samples were partitioned according to the KS method with a calibration set:validation set ratio of 3:1, with 130 data points selected for the calibration set and 45 for the validation set. Hyaluronic acid samples A41-A50, B42-B50, and C99-C118 were used as the test set samples. The data partitioning is shown in Table 1.

[0125] Table 1 Data Division

[0126]

[0127] 3.3 Band Selection Method

[0128] 3.3.1 Water matrix coordinate method based on water spectromics

[0129] This example determines the locations of absorption peaks for different aquatic species based on the 12 water matrix coordinate ranges specified by water spectromics and the characteristic wavelength selection method. The absorption values ​​at the 12 characteristic locations are used as modeling data to perform discriminant analysis on different samples, and the discriminant effect is compared with that of models established by commonly used band selection methods.

[0130] 3.3.1.1 Wavelet Transform Extraction of Activated Water Absorption Bands

[0131] Calculate the standard deviation of the spectral data at each wavelength after CWT, and locate the position with the largest standard deviation within the 12 water matrix coordinate ranges, such as... Figure 4 As shown, the 12 wavelengths are summarized in Table 2.

[0132] Table 2 CWT Site Selection Status

[0133]

[0134] 3.3.1.2 Principal component analysis: Extraction of activated water absorption zone

[0135] PCA was performed on the modeling data. The first principal component explained 73.8% of the spectral variations, while the second principal component explained only 18.6%. Based on water spectromics methods, the peaks and troughs located within 12 water matrix coordinate ranges in the first principal component loading plot were selected as characteristic wavelengths. Figure 5 As shown, a total of 5 characteristic wavelengths were selected.

[0136] The characteristic wavelengths extracted by CWT and PCA were integrated to obtain 12 activated water absorption bands, which are summarized in Table 3. A model was established based on the absorbance values ​​at the 12 characteristic wavelengths.

[0137] Table 3. Selection of wavebands based on water spectroomics

[0138]

[0139]

[0140] 3.3.2 Other Band Selection Methods

[0141] The data were band-selected using the CARs method, VIP method, and CC method, respectively, with 68, 219, and 67 characteristic wavelengths selected. Their absorption values ​​were used as modeling data, and the calibration set and validation set were divided according to the KS method to establish the model and perform discriminant analysis.

[0142] 3.4 Establishment of the discriminant model

[0143] 3.4.1 PCADA

[0144] This example establishes a PCADA model, and the calibration set and validation results are as follows: Figure 6 As shown, solid shapes represent the calibration set, and hollow shapes represent the validation set. The model achieves 100% recognition rate and accuracy.

[0145] The reliability of the model was further verified by feeding 40 test set data into the model, and the prediction results are as follows: Figure 7 As shown, the solid shapes represent the calibration set, and the hollow shapes represent the test set. The recognition rate and accuracy both reach 100%.

[0146] The classification results of the discriminant model established without band selection, CARs method, VIP method, and CC method are as follows: Figure 8 As shown, solid shapes represent calibration set data, and hollow shapes represent validation set data. The accuracy of both the calibration and validation sets is 100%.

[0147] Further validation of the models built using the test set data revealed that their test set accuracy was significantly lower than that of the patented method. The prediction performance of the models built using the four different processing methods is shown below. Figure 9As shown in the figure, the test set prediction sensitivity, specificity, and accuracy of the patented method compared with four other methods are visualized and analyzed. The results are as follows. Figure 10 As shown, although the recognition performance of the five models on the calibration and validation sets is good, the recognition performance of the patented method on the test set is significantly better than the other four models, with an accuracy of 100%.

[0148] 3.4.2 PLSDA

[0149] A PLSDA model was built using a patented method. The calibration set and validation results are as follows: Figure 11 As shown, solid shapes represent the calibration set, and hollow shapes represent the validation set. Both the recognition rate and accuracy reach 100%.

[0150] The reliability of the model was further verified by feeding 40 test set data into the model, and the prediction results are as follows: Figure 12 As shown, the solid shapes represent the calibration set, and the hollow shapes represent the test set. The recognition rate and accuracy both reach 100%.

[0151] The classification results of the PLSDA discriminant model established without band selection, CARs method, VIP method, and CC method are as follows: Figure 13 As shown, solid shapes represent calibration set data, and hollow shapes represent validation set data. The accuracy of both the calibration and validation sets is 100%.

[0152] Further validation of the models built using the test set data revealed poor prediction accuracy. The prediction performance of the models built using the four different processing methods is shown below. Figure 14 As shown in Table 4, the test set recognition rate and accuracy of the patented method and four other methods are summarized.

[0153] Table 4. Model test set recognition rate and accuracy

[0154]

[0155] This invention presents a method for identifying hyaluronic acid based on different degradation methods using water spectromics. Near-infrared spectra of acid-hydrolyzed, enzymatically hydrolyzed hyaluronic acid and their mixtures are collected. The spectral data are preprocessed using SNV and CWT methods. A calibration and validation set are divided according to the KS method. Based on the principles of water spectromics, 12 characteristic wavelengths are selected, and PCA-DA and PLS-DA models are established for the spectral data at these 12 wavelengths. The classification results are as follows: Figure 6 and Figure 11 It can be seen that the distribution of hyaluronic acid from enzymatic hydrolysis, acid hydrolysis, and their mixtures is relatively concentrated, and the recognition rate and accuracy of the test set are both 100%. Compared with models built without band selection and those built using commonly used band selection methods, the patented method can effectively and accurately distinguish hyaluronic acid from different degradation methods.

[0156] In summary, this invention acquires near-infrared spectra of hyaluronic acid with different degradation pathways and, based on aqueous spectromics technology, extracts characteristic wavelengths containing classification information to establish a discriminant analysis model. By employing aqueous spectromics technology combined with near-infrared spectroscopy to build the model, the quality of hyaluronic acid with different degradation pathways can be effectively monitored.

[0157] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of them. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention. Although the specific embodiments of the present invention have been described above, they are not intended to limit the protection scope of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A method for identifying hyaluronic acid with different degradation pathways based on aqueous spectromics, the method comprising: S1. Collect the near-infrared spectrum of the sample to be tested and obtain spectral data; S2. Perform standard normal variable transformation and continuous wavelet transform preprocessing on the spectral data obtained in step S1; S3. Select characteristic wavelengths based on the principles of water spectromics, establish models for the spectral data at the characteristic wavelengths, and make judgments. The 12 characteristic wavelengths are located at wavelengths of 1336-1348nm, 1360-1366nm, 1370-1376nm, 1380-1388nm, 1398-1418nm, 1421-1430nm, 1432-1444nm, 1448-1454nm, 1458-1468nm, 1472-1482nm, 1482-1495nm, and 1506-1516nm.

2. The method as described in claim 1, characterized in that, In step S1, the sample to be tested is hyaluronic acid from acid hydrolysis, hyaluronic acid from enzymatic hydrolysis, or a mixture of the two.

3. The method as described in claim 1, characterized in that, In step S1, the near-infrared spectral acquisition is performed using a near-infrared spectrometer. Specifically, the sampling method is integrating sphere diffuse reflectance, and the spectral scanning range is 4000-10000 cm⁻¹. -1 The resolution is set to 8cm. -1 The number of scans was 32.

4. The method as described in claim 3, characterized in that, The sample was taken three times, and the average value was used as the data for subsequent model building.

5. The method as described in claim 1, characterized in that, In step S2, the data is first preprocessed using standard normal variable transformation to remove outliers, and then the data is denoised using continuous wavelet transform.

6. The method as described in claim 1, characterized in that, The specific method of step S3 includes: according to the water spectromics point selection method, selecting the peak and valley values ​​of principal component analysis loads within the 12 water matrix coordinates specified by water spectromics, and simultaneously calculating the standard deviation of the absorption values ​​of each variable after continuous wavelet transform of the 12 water matrix coordinates. Selecting the variable with the largest standard deviation in each band as the characteristic wavelength, and combining principal component analysis, summarizing 12 activated water absorption bands as the final modeling data.

7. The method as described in claim 6, characterized in that, The wavelength positions of the 12 activated water absorption bands are specifically 1344nm, 1364nm, 1376nm, 1388nm, 1411nm, 1421nm, 1438nm, 1448nm, 1467nm, 1475nm, 1491nm, and 1505nm.

8. The method as described in claim 7, characterized in that, The method further includes: using the absorption value at a characteristic wavelength, dividing the calibration set and validation set according to the KS method, and establishing qualitative models of principal component discriminant analysis and partial least squares discriminant analysis, thereby realizing the qualitative discrimination of hyaluronic acid and its mixed products with different degradation methods.

9. The application of the method according to any one of claims 1-8 in the quality monitoring of hyaluronic acid.

10. The application as described in claim 9, characterized in that, The specific application is to qualitatively determine the acid hydrolysis and enzymatic hydrolysis of hyaluronic acid.

Citation Information

Patent Citations

  • Preparation method of low molecular weight hyaluronic acid based on near infrared spectroscopy

    CN110358869A

  • Method and apparatus for biopsy spectroscopy

    CN115135240A