Establishment method of dimensionality reduction training model of high-dimensional raman spectrum in noninvasive detection
Through artificial intelligence supervision and learning methods, problems such as dimensionality reduction of high-dimensional data and inaccurate selection of feature peak parameters in Raman spectroscopy technology are solved, and more accurate and efficient quantitative analysis capabilities are achieved, and the equipment is continuously optimized in actual applications.
Patent Information
- Application Number
- CN202510493913.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-19
AI Technical Summary
In the quantitative analysis of the monitored substances, the problems of high-dimensional data reduction, inaccurate selection of characteristic peak parameters, uncertain equipment system error, superposition of characteristic peaks of quantitative measurement of mixed substances, and difficulty in supporting the accurate selection of dynamic characteristic peaks in the system.
Using artificial intelligence supervised learning method, the dimensionality reduction of high Viraman spectral data and the continuous optimization of feature peak parameters are achieved through pre-training models and dynamic optimization algorithms, reducing equipment system errors, and supporting dynamic feature peak selection.
It improves the accuracy and efficiency of Raman spectroscopic quantitative analysis, reduces measurement errors, enhances the ability to quantitatively measure mixed substances, and enables the equipment to be "smart" more and more used in practical applications.
Smart Images

Figure CN120045846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of chemoinformatics and biomedicine, specifically to the sub-field of artificial intelligence quantitative detection of Raman spectra of specific chemical substances. In particular, it is related to the research and development of quantitative detection equipment for specific substances in mixed substances, and provides an artificial intelligence solution for supervised learning to obtain the characteristic peaks of Raman spectra of designated chemical substances. Background Art
[0002] Using Raman spectroscopy for non-destructive and non-invasive measurement of chemical substances in the human body is an important research focus in the field of spectroscopic measurement.
[0003] As early as 1928, Indian physicist Mr. C.V. Raman discovered the Raman scattering phenomenon, that is, when an incident light with a wavelength of irradiates a substance, due to the action of the incident light on the electrons of the substance, energy level transitions occur, resulting in a batch of scattered light with wavelengths different from and forming a spectrum. People call this scattered light and spectrum Raman scattered light and Raman spectrum, and it is found that there is a "fingerprint" - like connection between the spectra of these scattered lights and the irradiated substances. Based on this, Mr. C.V. Raman won the Nobel Prize in Physics in 1930. With the breakthrough development of laser technology in the 1960s, the method of using Raman spectroscopy to measure the content of substances began to develop rapidly.
[0004] According to the inventor's search, in recent years, there has been much research on the qualitative monitoring of specific substances using Raman spectroscopy, especially for non-destructive and non-invasive detection, and the development has been relatively rapid. The existing main technologies related to Raman spectroscopy analysis and big data in the industrial and medical fields are as follows: I. Application of Raman Spectroscopy in Qualitative Analysis of Substances Laser Raman Spectroscopy Currently, Raman spectroscopy technology is widely used for qualitative analysis of the substances to be detected in the fields of chemistry, medicine and pharmacy, industry, life science, agriculture, food, and even criminal investigation. Especially with the development of laser technology and the maturity of technologies such as Raman spectrum acquisition equipment or Raman spectrum testing equipment, it objectively promotes the rapid application and development of Raman spectroscopy analysis in the above fields, especially as a non-destructive and non-invasive qualitative analysis of substances.
[0005] Raman Characteristic Peaks By analyzing the laser Raman spectrum of the monitored substance, according to the characteristic of the unique Raman shift characteristic peaks corresponding to different substances, the vibration and rotation energy level conditions of the substance can be identified, and the properties of the substance can be analyzed, so as to identify what kind of substance the detected substance is or what substances it contains, that is, to conduct qualitative analysis of the identification of the detected substance.
[0006] Identification of mixture components When performing Raman spectroscopy analysis on a mixture of multiple substances, the composition components of the mixture can be identified, which also belongs to the relatively popular qualitative analysis of mixture components using Raman spectroscopy.
[0007] Surface-enhanced Raman scattering and tip-enhanced Raman spectroscopy According to the inventor's search, in terms of the sensitivity and accuracy of Raman spectroscopy analysis, with the development of modern nanotechnology, it has promoted the birth of surface-enhanced Raman scattering (SERS: Surface-Enhanced Raman Scattering) and tip-enhanced Raman spectroscopy (TERS: Tip-Enhanced Raman Spectroscopy) technologies, which have increased the intensity of Raman scattered light of the monitored substance, improved the signal-to-noise ratio, and made great progress in ultra-high-sensitivity detection. This has also opened up a feasible technical direction for Raman laser detection technology in quantitative analysis of substance components.
[0008] II. Technical applications of Raman spectroscopy in quantitative analysis of substances Quantitative analysis of the concentration of substances in solution With the development of Raman spectroscopy technology, people not only need to perform qualitative analysis on the tested substances, but in more application fields, the demand for quantitative analysis of the content or concentration of unknown substances to be detected, or the content or concentration of substances with known properties in solid mixtures or solutions of substances is becoming increasingly urgent. For example, in the chemical field and industrial field, Raman spectroscopy is used to analyze the alcohol content concentration in a mixed solution. In the medical field, a Raman laser spectrometer is used to analyze the blood samples taken from patients for blood glucose concentration using Raman spectroscopy. Currently, for quantitative analysis based on Raman spectroscopy, it is generally carried out in a dedicated laboratory through complex sample sampling and detection. Factors such as complex sample feature recognition, heavy reliance on professional analysts for data processing and analysis, long analysis time, low processing efficiency, and large errors in test results will greatly affect the application of this technology in practice.
[0009] Basic calculations for Raman spectroscopy quantitative analysis In the application of Raman spectroscopy for quantitative analysis, first, interference problems such as fluorescence need to be addressed, and at the same time, the requirements for hardware device noise are also very high. Generally, only a professional laboratory can test the substances to be detected. The basic technical process is as follows: Establish the corresponding relationship between known different concentrations of a pure substance solution and the intensity of Raman characteristic peaks. This relationship can be in the form of a table or a fitted curve equation, forming a reference comparison standard.
[0010] The Raman spectroscopy equipment is used to detect the solution of the substance to be detected with unknown concentration to obtain the Raman characteristic peak spectrum.
[0011] Read the intensity value of the characteristic peak of the Raman spectrum of the solution of the substance to be tested, compare it with the known correspondence table between the concentration of the substance and the Raman peak intensity, and analyze the concentration of the substance to be tested manually or with the built-in software of the equipment based on the correspondence table or the fitted curve equation.
[0012] Non-invasive Raman blood glucose analysis At present, some scholars and inventors have proposed to realize non-invasive blood glucose concentration analysis by using Raman spectra generated by Raman laser irradiation of specific tissue parts of the human body. However, whether it is direct Raman quantitative analysis of the monitored substances in the industrial field or non-invasive blood glucose Raman analysis in the medical field, the characteristic peaks of the Raman spectrum are not obvious or even submerged due to the noise of the equipment system itself, interference of fluorescent components, interference of Raman light of substances such as human tissue proteins, and measurement errors. The quantitative test effect of the content or concentration of the measured substance is very pessimistic, and it can only be realized in laboratories with very strict conditions, which affects the progress of related equipment research and development and is very unfavorable to the market promotion of related Raman testing products.
[0013] 3. Application of artificial intelligence and big data technology in the medical field With the development of information technology and computer technology, people have ushered in the era of "artificial intelligence, big data and cloud". In the past, traditional embedded devices or local systems could not realize complex algorithms such as qualitative and quantitative judgment, estimation, diagnosis, and planning with artificial intelligence (AI) functions due to their own hardware conditions or data processing calculation limitations. Now, they can be realized one by one through cloud servers and big data technology.
[0014] The core of big data technology is essentially to perform various complex and efficient data operations on massive data, especially the operation algorithms based on statistical theory. For example, the current epidemiological big data analysis in the medical field and the remote self-service medical diagnosis expert system are typical applications of "Internet + big data + artificial intelligence".
[0015] If big data technology is applied to Raman spectroscopy analysis, it will undoubtedly promote the rapid development of current Raman spectroscopy monitoring in quantitative monitoring applications, break through the limitations of traditional hardware equipment in performing complex artificial intelligence algorithm computing capabilities, and especially in the medical field of blood glucose monitoring and analysis. Some traditional and expensive equipment such as Raman blood glucose analysis equipment that are fixed in laboratories can be gradually moved out of professional laboratories to achieve miniaturization, portability, and even wearable non-invasive products, and promoted to the daily blood glucose monitoring and management of the majority of users, especially diabetic users.
[0016] Disadvantages of existing technology methods Based on the above analysis, the inventor believes that the existing Raman spectroscopy technology has the following deficiencies in the technology or method for quantitative analysis of the monitored substances.
[0017] 1. The problem of significant dimensionality reduction of high-dimensional Raman spectra for low-dimensional characteristic peaks.
[0018] In a professional Raman spectrometer, there are usually thousands of spectral data points, such as 2048 points, while the number of characteristic peaks of a certain substance is usually less than two digits, such as 5. From a mathematical perspective, reducing 2048-dimensional data to 5 dimensions while ensuring a certain level of credibility and accuracy is a difficult problem.
[0019] 2. The problem of inaccurate selection of Raman spectrum characteristic peak parameters.
[0020] Currently, for the characteristic peak parameters of most substances, such as wavenumber shift values and amplitudes, their selection is based on previous papers and statistical data, and it is difficult to determine whether they are accurate.
[0021] 3. The problem of uncertainty in the systematic error of Raman measurement equipment.
[0022] As a Raman spectrum measurement device, it not only has its own systematic errors. For example, the spatial position error in the installation of the CCD sensor on the beam splitter and the spectrometer integration module will directly affect the accuracy of the Raman wavenumber shift. At the same time, the performance of the CCD and the temperature of the cooler will also directly affect the accuracy of the spectral measurement value.
[0023] 4. The problem of characteristic peak superposition in quantitative measurement of mixed substances.
[0024] When measuring mixed substances, due to the mixing of Raman spectra generated by the molecular bonds of various substances, for the characteristic peaks of pure substances, the waves of these characteristic peaks are superimposed on each other. It is very difficult to evaluate the concentration of a single measured substance under the condition of changing concentrations of various substances in the mixed substance.
[0025] 5. The problem that the system is difficult to support accurate selection of dynamic characteristic peaks.
[0026] As a measurement instrument, the internal measurement function of a Raman spectrometer is usually fixed and does not support dynamic adjustment, so the accuracy of characteristic peak selection is questionable. Summary of the Invention
[0027] In view of the deficiencies of the prior art, the present invention designs a method for establishing a dimensionality reduction training model of high-dimensional Raman spectroscopy in non-invasive detection, which supports dynamic and iterative self-learning. Artificial intelligence supervised learning is used to continuously optimize the characteristic peaks of Raman spectroscopy for a specified substance through continuous training, enabling this device to become "smarter" with use. The establishment of the pre-training model of the present invention provides a better platform for intelligent Raman spectroscopy for measuring blood glucose.
[0028] The objectives and intentions of the present invention are achieved through the following working steps of the technical solution.
[0029] 1. Basic solution steps
[0030] As a method for establishing a dimensionality reduction training model of high-dimensional Raman spectroscopy in non-invasive detection, the present invention at least includes the following steps: According to the patient's laboratory test report, while collecting the patient's venous blood, synchronously collect more than one set of Raman spectroscopy data of the patient, digitize the test results and input them into the Raman spectroscopy data to establish an original data set.
[0031] In the original data set, for the Raman spectroscopy data, use methods including but not limited to inspection methods to eliminate abnormal Raman spectroscopy data groups.
[0032] Use the asymmetric least squares method to perform baseline correction on the Raman spectroscopy data, select the regularization smoothing parameter from 10 3 to 10 5 levels, the initial weight is 1.0, and the updated weight starts from 0.001 to 0.1 until the baseline is flat.
[0033] Use the discrete wavelet transform (DWT) method to reduce noise in the Raman spectroscopy data, select the wavelet basis function as db4 - db6, the decomposition level is 7 layers, and the soft threshold adjustment factor is 0.5 - 2.0.
[0034] In the Raman spectroscopy data, select and mark pre-characteristic peaks, optimize the parameters and sort the pre-characteristic peaks, select the pre-characteristic peaks with the top reduced dimensionality as characteristic peaks, and use the established dimensionality reduction result as the pre-training data set.
[0035] 2. Steps for establishing the original data set On the basis of the foregoing basic solution, in the aspect of the steps for establishing the original data set, the present invention specifically includes, but is not limited to, the following steps or methods for establishing the original data set in one or more combinations: Mark the sampling group number and the Raman spectroscopy number, where the test results include but are not limited to the names and arrays of blood test items. The sampling group includes all the Raman spectroscopy data corresponding to the synchronous collection of one laboratory test report, and the Raman spectroscopy data includes but is not limited to the Raman spectroscopy array and the blood test item array.
[0036] Generate test results by means including but not limited to manual entry or automatic recognition of paper test report images, where the test results also include but are not limited to patient information and test report numbers.
[0037] The Raman spectrum array is established as including but not limited to a one-dimensional array, composed of the Raman shift points determined by a Raman spectrometer and the Raman spectrum amplitudes at those points. The sequence number of the Raman shift points starts from 1 and ends at the maximum Raman shift point.
[0038] The blood test item array is established as including but not limited to a one-dimensional array, composed of more than one blood test item name and more than one blood test item data obtained from a single blood draw. The sequence number starts from 1 and ends at the last blood test item name.
[0039] The Raman spectrum array also includes but is not limited to patient numbers, and the Raman spectrum data of all patients are established as the original data set.
[0040] The model for establishing the original data set includes but is not limited to: determining the Raman spectrum data of the i-th sampling group by formula 2.1, determining the blood test item array of the i-th sampling group by formula 2.2, determining the Raman spectrum array of the i-th sampling group by formula 2.3, and determining the original data set by formula 2.4: , 2.1 2.2 , 2.3 , 2.4 Where: The Raman spectrum data of the i-th sampling group, The Raman spectrum array of the i-th sampling group, The blood test item array of the i-th sampling group, The original data set, is the number of the sampling group, is the maximum number of the sampling group, , is the Raman spectrum number, is the maximum Raman spectrum number, , is the blood test item number, is the maximum blood test item number, , Number the Raman shift points, , is the Raman spectrum shift range, determined by the spectral sensor of the Raman spectroscopy device, is the name of the blood test item, is the blood test item data, is the Raman spectrum amplitude.
[0041] 3. T-test to eliminate abnormal Raman spectra Based on the above technical solutions, the present invention includes but is not limited to the step of using a T-test to eliminate abnormal Raman spectra, and specifically, one or more combinations of the following local improvement measures can also be adopted: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is below 30, use the T-test method to find abnormal Raman spectrum data and eliminate the Raman spectrum data. The elimination methods include but are not limited to deletion or marking.
[0042] According to the known Raman characteristic peaks of the substance to be detected, select the Raman shift points where the characteristic peaks are located and one or more Raman shift points before and after, and accumulate or calculate the average value of the Raman characteristic peaks at these shift points as , as T the difference vector for the test, calculate the overall average of the th Raman characteristic peak according to formula 3.1, calculate the test value of this Raman characteristic peak according to formula 3.2, calculate the standard deviation of this Raman characteristic peak according to formula 3.3, and find the built-in T-test table or call, but not limited to, Python functions to calculate the p-value. When it is determined that this Raman characteristic peak belongs to abnormal Raman spectrum data: 3.1 3.2 3.3 is the overall average, is the deviation statistic of the sample number from the overall average, is the total wavenumber points of each Raman spectrum or the shift range of the Raman spectrum, is the standard deviation, represents the difference between the th Raman spectrum and the mean at each wavenumber point, Number the sampling groups, is the maximum number of the sampling groups, , is the Raman spectrum number, is the maximum number of the Raman spectra, , is the overall average of the Raman characteristic peaks.
[0043] 4. Z - test to eliminate abnormal Raman spectra On the basis of the foregoing technical solution, the present invention includes, but is not limited to, using the Z - test to eliminate abnormal Raman spectra. Specifically, the following local improvement measures in one or more combinations can also be adopted: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is 30 or more, use the Z - test method to find abnormal Raman spectral data and eliminate the Raman spectral data. The elimination methods include, but are not limited to, deletion or marking.
[0044] According to the known Raman characteristic peak values of the substance to be detected, select the Raman shift points where the characteristic peaks are located and one or more Raman shift points before and after, and accumulate or calculate the average value of the Raman characteristic peaks at these shift points , as Z the difference vector for the test, and calculate the test value of the Raman characteristic peak of the 4.1 4.2 is the overall average of the Raman characteristic peaks, is the maximum number of the Raman spectra, , is the standard deviation.
[0045] 5. Baseline correction On the basis of the foregoing technical solution, the present invention includes, but is not limited to, the baseline correction step. Specifically, the following local improvement measures and steps in one or more combinations can be adopted: Decompose the Raman spectrum array into a baseline, characteristic peaks, and noise, as shown in Formula 5.1 and Formula 5.2, 5.1 5.2 Adopt, including but not limited to, polynomial fitting of the baseline, solve the coefficients by the least squares method, as shown in Equation 5.3, and solve the minimized sum of squared residuals according to Equation 5.4. 5.3 5.4 and is the function and independent variable of the Raman spectrum array of the i-th sampling group and is the baseline in the Raman spectrum array of the i-th sampling group is the characteristic peak in the Raman spectrum array of the i-th sampling group is the noise in the Raman spectrum array of the i-th sampling group is the least squares coefficient is the order of the regularization smoothing parameter, with a value range of 3 to 5 is the Raman shift point number , is the Raman spectrum displacement range, which is determined by the spectral sensor of the Raman spectrum device.
[0046] S330: For the Raman spectrum amplitudes of all the Raman shift points in all the Raman spectrum arrays, select the regularization smoothing parameter from 10 3 to 10 5 levels, the initial weight is 1.0, and the updated weight starts from 0.001 to 0.1 until the baseline is flat. Subtract the baseline and noise at this point to obtain the characteristic peak and complete the baseline correction.
[0047] 6. Noise elimination On the basis of the foregoing technical solutions, the present invention includes but not limited to the noise elimination step, which specifically includes: Select the wavelet basis function as db4 to db6, Symlets select sym6, the decomposition level is 0 - 7 layers, and the soft threshold adjustment factor is 0.5 - 2.0.
[0048] For the Raman spectrum array after baseline correction, define the low-pass filter and high-pass filter respectively as shown in Equation 6.1 and Equation 6.2. For the Raman spectrum amplitudes of the Raman shift points in the Raman spectrum array , perform the steps of obtaining the low-frequency component by decomposition using Equation 6.3 through low-pass filtering, the steps of obtaining the high-frequency component by decomposition using Equation 6.4 through high-pass filtering, and Equation 6.5 for reconstruction to obtain noise reduction and perform downsampling. 6.1 6.2 6.3 6.4 6.5 is the low-pass filter. is the high-pass filter. is the reconstruction low-pass filter. is the reconstruction high-pass filter. is the filter decomposition layer number, and for db4, the value range is 0 - 7. is the sampling group number. is the Raman shift point number. is the downsampling number. is the Raman spectrum amplitude in the Raman spectrum array of the i-th sampling group. of the function.
[0049] 7. Dimensionality reduction to generate characteristic peaks On the basis of the foregoing technical solutions, the present invention includes but is not limited to the steps of establishing characteristic peaks, specifically including: Based on one or more assay item names, set the dimensionality reduction number corresponding to the assay item name in the Raman spectrum array to be less than 30.
[0050] In the Raman spectrum data after noise elimination, select the Raman spectrum amplitude with a fluctuation amplitude greater than 2% and being the largest at both the front and back wavenumber positions as the pre-characteristic peak corresponding to the assay item name.
[0051] For each assay item name, sort the pre-characteristic peaks from largest to smallest, and select the top dimensionality reduction number as the characteristic peak. The characteristic peak includes but is not limited to the Raman spectrum amplitude, Raman shift value, and assay item name.
[0052] Optimally, based on the name of the test item, among the Raman spectral data after noise elimination, select the Raman spectral amplitudes with a fluctuation amplitude greater than 10% and being the maximum at both the front and back wavenumber positions as the pre-feature peaks corresponding to the name of the test item, calculate the number of pre-feature peaks as the dimensionality reduction number, and use all the pre-feature peaks as the feature peaks corresponding to the name of the test item, including the Raman spectral amplitude, the Raman shift value, and the name of the test item.
[0053] Establish all the feature peaks as the pre-training data set.
[0054] 8. Feature peaks generated by wavelet algorithm dimensionality reduction Based on the above technical solutions, the present invention includes but is not limited to the following steps for establishing feature peaks, specifically including: Based on one or more names of test items, set the dimensionality reduction number corresponding to the name of the test item in the Raman spectral array to be below 30.
[0055] Preferably, use the wavelet algorithm to process the pre-feature peaks to obtain feature peaks including the Raman spectral amplitude, the Raman shift value, and the name of the test item.
[0056] 9. Feature peaks generated by recursive algorithm dimensionality reduction Based on the above technical solutions for generating feature peaks by dimensionality reduction, the present invention includes but is not limited to the following steps for feature peaks using the recursive algorithm, specifically including: For the feature peaks, use the recursive algorithm to check and fine-tune the feature peaks corresponding to the name of one test item to obtain the optimized feature peaks corresponding to the name of one test item.
[0057] Preferably, check and fine-tune the feature peaks corresponding to all the names of test items to obtain the optimized feature peaks corresponding to all the names of test items.
[0058] 10. Feature peaks generated by recursive algorithm dimensionality reduction Preferably, based on the technical solutions for generating feature peaks by wavelet algorithm dimensionality reduction, use the recursive algorithm to check and fine-tune the feature peaks corresponding to the name of one test item to obtain the optimized feature peaks corresponding to the name of one test item.
[0059] Check and fine-tune the feature peaks corresponding to all the names of test items to obtain the optimized feature peaks corresponding to all the names of test items.
[0060] 11. Invention purpose and intention Through long-term research, observation, and experiments, the inventor proposes a method for establishing a dimensionality reduction training model for high-dimensional Raman spectroscopy in non-invasive detection, which is also a training method for artificial intelligence supervised learning.
[0061] The purpose and intention of the present invention is to establish a complete method for establishing a dimensionality reduction training model for high-dimensional Raman spectroscopy in non-invasive detection, and specifically include the following solutions to the deficiencies of the prior art.
[0062] 1. The high-dimensional Raman spectrum can significantly reduce the dimensionality of low-dimensional characteristic peaks.
[0063] In the present invention, the 2048-dimensional Raman spectrum is greatly reduced to single digits, such as 5 dimensions, through pre-characteristic peak learning, sorting method, wavelet dimensionality reduction method, and recursive optimization method.
[0064] 2. Inaccurate selection of Raman spectrum characteristic peak parameters.
[0065] The supervised learning method and the biochemical analyzer test results of the same period are used for continuous comparison and learning, so as to continuously iterate and optimize the characteristic peak parameters.
[0066] 3. Uncertainty of systematic errors in Raman measurement equipment.
[0067] Through supervised learning, the systematic error of the equipment is iteratively reduced and the measurement accuracy is continuously improved.
[0068] 4. The problem of characteristic peak superposition in quantitative measurement of mixed substances.
[0069] Within the visible range of mixed material ratios, for example, although the subcutaneous mixture of human skin is complex, its material types and mixing ratios are within a certain range after all. Through a large number of biochemical blood tests, the system can determine the relevant parameters. In addition, after the device is set to cloud mode, as all devices continue to learn, the system will quickly achieve the optimized result. Therefore, the present invention can make the device more and more "smart".
[0070] 5. The system has difficulty supporting accurate selection of dynamic characteristic peaks.
[0071] The internal measurement function of the traditional measurement equipment is static, while the present invention adopts dynamic supervised learning. With continuous iteration, the present invention can perfectly support the dynamic optimization of characteristic peaks.
[0072] 12. Beneficial Effects of the Invention 1. The invention realizes the purpose of the invention, solves the problem of large-scale dimension reduction from Raman spectrum to characteristic peaks, solves the problem of inaccurate selection of characteristic peak parameters of Raman spectrum, the problem of uncertainty of system error of Raman measurement equipment, the problem of superposition of characteristic peaks in quantitative measurement of mixed substances, and the problem that the system is difficult to support accurate selection of dynamic characteristic peaks; 2. Provides independence between algorithms and hardware devices; 3. Provide optional characteristic peak sequence; 4. Provide the preliminary innovative achievements and foundation for the quantitative measurement of specific substances in the mixture in the later stage. Description of the Drawings
[0073] List of the Drawings: Figure 1 : Schematic diagram of the system layer structure Figure 2 : Reference structure diagram of the hardware in the embodiment Figure 3 : Original Raman spectrogram Figure 4 : Baseline-corrected spectrogram Figure 5 : Application example of the pre-training set Detailed description of the drawings: See the embodiments for details. Specific Embodiments
[0074] The objectives and intentions of the present invention can be achieved by the design methods of the following embodiments. It should be particularly noted here that since the specific embodiments have specific uses and industrial applicability, the embodiments cannot include all the features and steps of the present invention. The description in the claims of the present invention is the summary of the invention.
[0075] The specific embodiments of the present invention are as follows: Embodiment: Method for establishing a pre-training model for non-invasive in vitro blood glucose Raman spectroscopy quantitative monitoring This example is a general example of the present invention, which is a general example of the method for establishing an AI pre-training model for glucose concentration when using Raman spectroscopy to non-invasively measure subcutaneous glucose through human skin.
[0076] It should be stated that: The content and views of this embodiment are not a limitation to the present invention, nor can they fully interpret the present invention. It is only one of the embodiments of the industrial application of the present invention.
[0077] 1. Explanation of the Drawings The content of this embodiment mainly includes but is not limited to the following main schematic drawings: Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 . It should be noted here that Figure 4 On the left side of , due to data overflow caused by the measuring instrument exceeding the range, it is marked with a red line.
[0078] 2. Explanation of the Implementation Steps The method steps of this embodiment mainly include Step 1 to Step 7. Among them, unless otherwise specified, the step numbers of these 7 parts do not have a sequential order, nor do all embodiments require the combination of these 7 parts. In addition, each of these 7 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and their sequential order is not necessary either. Instead, according to the requirements of some specific tasks, the patent implementer makes an optimized choice.
[0079] The specific working steps are described as follows.
[0080] 2.1 Description of the basic solution steps.
[0081] As a method for establishing a dimensionality reduction training model of high-dimensional Raman spectroscopy in non-invasive detection, the present invention at least includes the following steps: According to the patient's test report, while collecting the patient's venous blood, synchronously collect more than one set of Raman spectroscopy data of the patient, digitize the test results and input them into the Raman spectroscopy data to establish an original data set.
[0082] In the original data set, for the Raman spectroscopy data, adopt inspection methods including but not limited to, and eliminate abnormal Raman spectroscopy data groups.
[0083] Use the asymmetric least squares method to perform baseline correction on the Raman spectroscopy data, select the regularization smoothing parameter from 10 3 to 10 5 levels, the initial weight is 1.0, and the updated weight starts from 0.001 to 0.1 until the baseline is flat.
[0084] Use the discrete wavelet transform DWT method to reduce noise in the Raman spectroscopy data, select the wavelet basis function as db4-db6, the decomposition level is 7 layers, and the soft threshold adjustment factor is 0.5-2.0.
[0085] Preferably or further, in the Raman spectroscopy data, select and mark pre-feature peaks, optimize the parameters and sort the pre-feature peaks, select the pre-feature peaks with the top-dimensionality reduction numbers as feature peaks, and use the established dimensionality reduction result as the pre-training data set.
[0086] Preferably or further, as the hardware platform for running this method, from a functional perspective, it includes a Raman spectroscopy acquisition subsystem, a lens subsystem, a test report recognition subsystem, a computer subsystem, a communication subsystem, a display subsystem, and a housing. For test report collection, use sub-devices based on video or image automatic collection and character recognition, and realize linkage with Raman spectroscopy acquisition, so as to realize automatic character recognition of test report information and synchronous collection of Raman spectroscopy while the patient draws venous blood.
[0087] Figure 1It is a schematic diagram of the system layer structure of the present invention. Logically, the system is mainly divided into an original data set layer, an outlier removal layer, a baseline correction layer, a noise reduction layer, and a pre-training set layer.
[0088] Figure 2 It is an example of the system hardware platform solution adopted in this embodiment. Structurally, it mainly includes a skin detection point, an excitation light module and a laser control module, a semi-reflective semi-transmissive mirror, a condenser lens 1, a filter lens, a condenser lens 2, a slit, a condenser lens 3, a spectroscope, a spectrometer integration module, a cooler, and a master control module.
[0089] 2.2 Step description for establishing the original data set On the basis of the foregoing basic solution, in the steps of establishing the original data set, the present invention specifically includes, but is not limited to, the following steps or methods for establishing the original data set in one or more combinations: Mark the sampling group number and the Raman spectrum number, where the test results include, but are not limited to, the names of blood test items and the blood test item arrays. The sampling group includes the Raman spectrum data of all the Raman spectrum numbers synchronously collected corresponding to one test report, and the Raman spectrum data includes, but is not limited to, the Raman spectrum array and the blood test item array.
[0090] Preferably, the test results are generated by means including, but not limited to, manual entry or automatic recognition of the paper test report image, and the test results also include, but are not limited to, patient information and the test report number.
[0091] Further, the Raman spectrum array is established as including, but not limited to, a one-dimensional array, which is composed of the Raman shift points determined by the Raman spectrometer and the Raman spectrum amplitudes at these points. The sequence numbers of the Raman shift points start from 1 and end at the maximum Raman shift point.
[0092] Further, the blood test item array is established as including, but not limited to, a one-dimensional array, which is composed of more than one blood test item name and more than one blood test item data obtained from one blood draw test. The sequence numbers start from 1 and end at the last blood test item name.
[0093] Preferably, the Raman spectrum array also includes, but is not limited to, the patient number, and the Raman spectrum data of all patients are established as the original data set.
[0094] Further, the model for establishing the original data set includes, but is not limited to: the Raman spectrum data of the i-th sampling group is determined by formula 2.1, the blood test item array of the i-th sampling group is determined by formula 2.2, the Raman spectrum array of the i-th sampling group is determined by formula 2.3, and the original data set is determined by formula 2.4: 2.1 2.2 2.3 2.4 Wherein: Raman spectral data of the i-th sampling group, Raman spectral array of the i-th sampling group, Blood test item array of the i-th sampling group, Original data set, is the sampling group number, is the maximum sampling group number, , is the Raman spectrum number, is the maximum Raman spectrum number, , is the blood test item number, is the maximum blood test item number, , is the Raman shift point number, , is the Raman spectrum displacement range, which is determined by the spectral sensor of the Raman spectroscopy device, is the name of the blood test item, is the blood test item data, is the Raman spectrum amplitude.
[0095] Preferably or further, based on the definition of the above data structure, the data structure itself can be in the form of a two-dimensional table, a multi-dimensional array, or a structure form customized based on functions included in the software development language; the content of the data structure can be increased with other content according to the system requirements, such as patient personal information, hospital information, electronic medical record information, medication information, etc.; the operation relationships of the data structure are not limited to Formulas 2.1 to 2.4, and can also be extended or modified to other calculation relationships according to the requirements of system users.
[0096] Figure 3 is the planar expansion diagram of the measured Raman spectrum data in the original data set constructed by the present invention. Among them, the abscissa is the displacement of the Raman spectrum, and the ordinate is the Raman spectrum amplitude.
[0097] 2.3 Step description of T-test for eliminating abnormal Raman spectra Based on the foregoing technical solutions, the present invention includes, but is not limited to, the step of using a T-test to eliminate abnormal Raman spectra, and specifically, the following local improvement measures in one or more combinations can also be adopted: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is below 30, the T-test method is used to find abnormal Raman spectral data and eliminate the Raman spectral data. The elimination methods include, but are not limited to, deletion or marking.
[0098] Further, according to the known Raman characteristic peaks of the substance to be detected, the Raman shift points where the characteristic peaks are located and one or more Raman shift points before and after are selected, and the Raman characteristic peaks at these shift points are accumulated or averaged to obtain , as T the difference vector for testing, and the overall average of the th Raman characteristic peak is calculated according to Formula 3.1, the test value of this Raman characteristic peak is calculated according to Formula 3.2, the standard deviation of this Raman characteristic peak is calculated according to Formula 3.3, and the p-value is calculated by looking up the built-in T-test table or calling, but not limited to, Python functions. When it is determined that this Raman characteristic peak belongs to abnormal Raman spectral data, 3.1 3.2 3.3 is the overall average, is the deviation statistic between the sample size and the overall average, is the total wavenumber points of each Raman spectrum or the displacement range of the Raman spectrum, is the standard deviation, represents the difference between the th Raman spectrum and the mean at each wavenumber point, is the sampling group number, is the maximum number of the sampling group, , is the Raman spectrum number, is the maximum number of the Raman spectra, , is the overall average of the Raman characteristic peaks.
[0099] Preferably or further, based on a single Raman pre-characteristic peak, the T-test is used to further eliminate abnormal pre-characteristic peaks to further screen characteristic peaks.
[0100] Preferably or further, based on the entire Raman spectrum, the T-test is used to further eliminate abnormal entire Raman spectra to further screen Raman spectra.
[0101] 2.4 Description of the steps for eliminating abnormal Raman spectra by Z-test Based on the foregoing technical solutions, the present invention includes, but is not limited to, eliminating abnormal Raman spectra by Z-test, and specifically, the following local improvement measures and steps in one or more combinations can be adopted: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is 30 or more, the Z-test method is used to find abnormal Raman spectral data and eliminate the Raman spectral data. The elimination methods include, but are not limited to, deletion or marking.
[0102] Further, according to the known Raman characteristic peak values of the substance to be detected, the Raman shift points where the characteristic peak values are located and one or more Raman shift points before and after are selected, and the Raman characteristic peak values at these shift points are accumulated or averaged , as Z the differential vector for the test, and the test value of the th Raman characteristic peak value is calculated according to Formula 4.1, and it is determined whether the Raman characteristic peak value belongs to abnormal Raman spectral data according to Formula 4.2: , 4.1 4.2 is the overall average of the Raman characteristic peaks, is the maximum number of the Raman spectrum, , is the standard deviation.
[0103] Preferably or further, based on a single Raman pre-characteristic peak, the Z-test is used to further eliminate abnormal pre-characteristic peaks to further screen characteristic peaks.
[0104] Preferably or further, based on the entire Raman spectrum, the Z-test is used to further eliminate abnormal entire Raman spectra to further screen Raman spectra.
[0105] 2.5 Description of the baseline correction steps Based on the foregoing technical solutions, the present invention includes, but is not limited to, the baseline correction step, and specifically, the following local improvement measures and steps in one or more combinations can be adopted: Decompose the Raman spectrum array into a baseline, characteristic peaks, and noise, as shown in Equations 5.1 and 5.2, 5.1 5.2 Furthermore, adopt, including but not limited to, polynomial fitting of the baseline, and solve the coefficients by the least squares method, as shown in Equation 5.3, and solve for the minimized sum of squared residuals according to Equation 5.4, 5.3 5.4 and is the Raman spectrum array of the i-th sampling group the function and independent variable of, is the baseline in the Raman spectrum array of the i-th sampling group, is the characteristic peak in the Raman spectrum array of the i-th sampling group, is the noise in the Raman spectrum array of the i-th sampling group, is the least squares coefficient, is the order of the regularization smoothing parameter, with a value range of 3 to 5, is the Raman shift point number, , is the Raman spectrum shift range, determined by the spectral sensor of the Raman spectroscopy device.
[0106] Preferably or further, adjust the parameters in Equations 5.1 to 5.4 so that the baseline becomes a straight line and only the characteristic peaks are retained on the baseline.
[0107] Figure 4 The measured graph of the Raman spectrum after baseline correction is unfolded. Compared with Figure 3 compared, Figure 3 is similar to a downward slope, while Figure 4 zeroes the baseline to facilitate highlighting the pre-characteristic peaks of the Raman spectrum. Figure 4 There is a high probability correspondence between the pre-characteristic peaks in and the molecular bonds of the detected substance. The reason for calling it a high probability is because of the influence of wave superposition.
[0108] 2.6 Description of the noise elimination steps Based on the foregoing technical solutions, the present invention includes but is not limited to the noise elimination steps, specifically including: The wavelet basis function is selected as db4, the decomposition level is from 0 to 7 layers, and the soft threshold adjustment factor is from 0.5 to 2.0.
[0109] Furthermore, for the Raman spectrum array after baseline correction, steps including but not limited to obtaining the low-frequency component through low-pass filtering according to formula 6.1, obtaining the high-frequency component through high-pass filtering according to formula 6.2, and downsampling are performed: 6.1 6.2 6.3 6.4 6.5 is the low-pass filter, is the high-pass filter, is the reconstructed low-pass filter, is the reconstructed high-pass filter, is the filter decomposition layer number, and for db4, the value range is 0 - 7, is the sampling group number, is the Raman shift point number, is the number of downsampling, is the Raman spectrum amplitude in the Raman spectrum array of the i-th sampling group of the function.
[0110] Preferably or further, adjust the parameters of formulas 6.1 and 6.2. By viewing the low-frequency component, make it a flat line to eliminate low-frequency noise interference. By viewing the high-frequency component and comparing it with the characteristic peak, taking the width of the characteristic peak as the unit, determine that the fluctuation less than the width of the characteristic peak is high-frequency interference. Finally, eliminate all low-frequency and high-frequency interferences.
[0111] 2.7 Step description of dimensionality reduction to generate characteristic peaks Based on the foregoing technical solutions, the present invention includes but is not limited to the steps of establishing characteristic peaks, specifically including: Based on one or more assay item names, set the dimensionality reduction number corresponding to the assay item name in the Raman spectrum array to be less than 30.
[0112] It should be noted that a Raman spectrum is usually composed of Raman spectral amplitude data of more than a thousand Raman sites. For example, the Raman spectrum collected by a certain Raman spectrometer has a total of 2048 points, and in almost all calculations, each of them is regarded as a dimension of a high-dimensional space. In order to reduce the amount of calculation and save computing power, in the present invention, a substantial dimensionality reduction measure is adopted, so that the original 2048 points are reduced to less than 30 points. Because usually for a characteristic peak, 5 to 12 points at one point are sufficient.
[0113] Among the Raman spectral data after noise elimination, select the Raman spectral amplitude with a fluctuation amplitude greater than 2% and being the largest at both the front and back wavenumber positions as the pre-characteristic peak corresponding to the name of the test item. Here, the "Raman spectral amplitude being the largest at both the front and back wavenumber positions" is actually an extreme value on this Raman spectral line. From the perspective of the continuous spectrum, the derivative of this point is 0.
[0114] Furthermore, for each name of the test item, sort the pre-characteristic peaks from largest to smallest, and select the top-dimensionality reduction numbers as the characteristic peaks. The characteristic peaks include but are not limited to Raman spectral amplitude, Raman shift value, and the name of the test item.
[0115] Optimally, based on the name of the test item, among the Raman spectral data after noise elimination, select the Raman spectral amplitude with a fluctuation amplitude greater than 10% and being the largest at both the front and back wavenumber positions as the pre-characteristic peak corresponding to the name of the test item, calculate the number of pre-characteristic peaks as the dimensionality reduction number, and use all the pre-characteristic peaks as the characteristic peaks corresponding to the name of the test item, including Raman spectral amplitude, Raman shift value, and the name of the test item.
[0116] Use all the characteristic peaks as the pre-training data set.
[0117] Preferably or further, for a specific substance, although there are many characteristic peaks, due to the superposition of mixtures on the Raman spectrum, their permutations and combinations will present infinite possibilities on the Raman shift axis, as Figure 4 shown. Mathematically, the actual data dimension will be extremely large. As Figure 2 shown in the hardware structure, the data generated by the spectrometer integration module is 2048, that is, 2048 dimensions. This not only consumes a large amount of computing power in calculation, but also in actual measurement, using subsequent artificial intelligence statistical algorithms will generate meaningless redundancy. Therefore, the dimension must be reduced. According to the specific detected substance and the pre-characteristic peaks detected on the human skin, reduce the dimension, that is, reduce the number of characteristic peaks to 5 to 30, so as to stably and accurately distinguish the specific substance on the premise of ensuring the measurement accuracy, and then quantitatively measure the content of the specific substance.
[0118] It should be noted that the amplitude of fluctuation here refers to taking the pre-feature peak as the unit of examination, and the difference between the peak and the trough of the pre-feature peak is the fluctuation value. The fluctuation value of the largest pre-feature peak is used as the denominator for calculation. For example: Amplitude of fluctuation = (fluctuation value of a certain pre-feature peak / fluctuation value of the largest pre-feature peak) × 100%.
[0119] Figure 5 It is a schematic diagram of the feature peaks required for the detection of subcutaneous glucose content. In the figure, only 5 feature peaks are needed to complete the reliable detection of glucose content, and the measurement error does not exceed 5%.
[0120] 2.8 Step description of generating feature peaks by wavelet algorithm dimensionality reduction On the basis of the foregoing technical solutions, the present invention includes but is not limited to the following steps of establishing feature peaks, specifically including: Preferably, based on more than one test item name, the dimensionality reduction number corresponding to the test item name in the Raman spectrum array is set to be less than 30.
[0121] Preferably, the wavelet algorithm is used to process the pre-feature peak to obtain feature peaks including Raman spectrum amplitude, Raman shift value and test item name.
[0122] 2.9 Step description of generating feature peaks by recursive algorithm dimensionality reduction On the basis of the foregoing technical solutions of generating feature peaks by dimensionality reduction, the present invention includes but is not limited to the following steps of using the recursive algorithm for feature peaks, specifically including: Preferably, for the feature peaks, the recursive algorithm is used to check and fine-tune the feature peaks corresponding to one test item name to obtain the optimized feature peaks corresponding to one test item name.
[0123] Preferably, check and fine-tune the feature peaks corresponding to all test item names to obtain the optimized feature peaks corresponding to all test item names.
[0124] The adoption of the recursive algorithm here is a recursive process based on the dynamic accumulation of the original data set. For example, at a certain moment, 5000 test report samples are collected in the original data set. At this time, a recursive calculation is performed, and several feature peak sets are found and modified. When the test samples in the original data set increase dynamically, for example, to 6000 samples, and the recursive calculation is performed again. Since 1000 new samples are added, several new feature peaks are found to need to be modified at this time. The intention of the recursion here is a process of iteratively optimizing the feature peak data set, making the feature peak data set more and more accurate.
[0125] 2.10 Step description of generating feature peaks by recursive algorithm dimensionality reduction Preferably, on the basis of the technical solution of generating characteristic peaks by wavelet algorithm dimensionality reduction, a recursive algorithm is adopted to verify and fine-tune the characteristic peaks corresponding to a test item name, and the optimized characteristic peaks corresponding to a test item name are obtained.
[0126] Verify and fine-tune the characteristic peaks corresponding to all test item names, and obtain the optimized characteristic peaks corresponding to all test item names.
Claims
1. A method for establishing a dimensionality reduction training model for high-dimensional Raman spectroscopy in non-invasive detection, characterized in that: include: S100: According to the patient's test sheet, while collecting the patient's venous blood, synchronously collect one or more sets of Raman spectrum data of the patient, digitize the test results and enter them into the Raman spectrum data to establish an original data set; S200: In the original data set, for the Raman spectrum data, using a testing method to remove abnormal Raman spectrum data sets; S300: Using an asymmetric least squares method to perform baseline correction on the Raman spectrum data, selecting a regularization smoothing parameter of 10 3 Up to 10 5 Level, with an initial weight of 1.0, and updating the weight from 0.001 to 0.1 until the baseline is flat; S400: using a discrete wavelet transform (DWT) method to reduce noise from the Raman spectrum data, selecting a wavelet basis function of db4-db6, a decomposition layer number of 7 layers, and a soft threshold adjustment factor of 0.5-2.0; S500: Select and mark pre-characteristic peaks in the Raman spectrum data, optimize parameters and sort the pre-characteristic peaks, select the pre-characteristic peaks with the front dimensionality reduction number as the characteristic peaks, and use the established dimensionality reduction results as the pre-training data set.
2. The method according to claim 1, characterized in that The S100 specifically includes the following steps: S110: Marking the sampling group number and the Raman spectrum number, wherein the test result includes the blood test item name and the blood test item array, the sampling group includes the Raman spectrum data of all the Raman spectrum numbers collected synchronously corresponding to one test sheet, and the Raman spectrum data includes the Raman spectrum array and the blood test item array; S120: Generate the test result by manual entry or automatic recognition of a paper test report image, wherein the test result also includes the patient information and the test report number; S130: establishing the Raman spectrum array to include a one-dimensional array, which is composed of Raman shift points determined by the Raman spectrometer and the Raman spectrum amplitudes at the points, wherein the sequence numbers of the Raman shift points start from 1 and end at the maximum Raman shift point; S140: Establishing the blood test item array to include a one-dimensional array, which is composed of one or more blood test item names and one or more blood test item data obtained from a blood test, and the sequence number starts from 1 and ends at the last blood test item name; S150: The Raman spectrum array also includes the patient's serial number, and the Raman spectrum data of all the patients are established as the original data set; and / or, S160: Establishing the model of the original data set includes: determining the Raman spectrum data of the i-th sampling group by formula 2.1, determining the blood test item array of the i-th sampling group by formula 2.2, determining the Raman spectrum array of the i-th sampling group by formula 2.3, and determining the original data set by formula 2.4: 2.1 2.2 2.3 2.4 in: The Raman spectrum data of the i-th sampling group, The Raman spectrum array of the i-th sampling group, The array of blood test items for the i-th sampling group, The original dataset, is the sampling group number, is the maximum number of the sampling group, , is the Raman spectrum number, is the maximum number of the Raman spectrum, , is the blood test item number, is the maximum number of the blood test items, , is the number of the Raman shift point, , is the Raman spectrum shift range, determined by the spectrum sensor of the Raman spectrum device, is the name of the blood test item, is the blood test item data, is the Raman spectrum amplitude.
3. The method according to claim 2, characterized in that The S200 specifically includes the following steps: S210: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is less than 30, a T-test method is used to find abnormal Raman spectrum data and remove the Raman spectrum data, and the removal method includes deletion or marking; S211: According to the known Raman characteristic peak of the substance to be detected, the Raman shift point where the Raman characteristic peak is located and one or more Raman shift points before and after the Raman shift point are selected, and the Raman characteristic peak of the Raman shift point is accumulated or averaged. , as stated T The difference vector of the test is calculated according to formula 3.1 The overall average of the Raman characteristic peaks is calculated according to formula 3.
2. Test value, calculate the standard deviation of the Raman characteristic peak according to formula 3.3, obtain the T test table by looking up the built-in or call the Python function to calculate the p value, when Determining that the Raman characteristic peak value belongs to the abnormal Raman spectrum data, 3.1 3.2 3.3 is the overall mean, is the deviation statistic of the sample number from the population mean, is the total wave number points of each Raman spectrum or the displacement range of the Raman spectrum, is the standard deviation, To indicate the The difference between the Raman spectrum at each wave number point and the mean value, is the Raman spectrum number, is the maximum number of the Raman spectrum, , is the overall average of the Raman characteristic peaks.
4. The method according to claim 2, characterized in that The S200 specifically further includes the following steps: S220: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is 30 or above, a Z test method is used to find abnormal Raman spectrum data and remove the Raman spectrum data, and the removal method includes deletion or marking; S221: According to the known Raman characteristic peak of the substance to be detected, the Raman shift point where the Raman characteristic peak is located and one or more Raman shift points before and after the Raman shift point are selected, and the Raman characteristic peak of the Raman shift point is accumulated or averaged. , as stated Z The difference vector of the test is calculated according to formula 4.1 The Raman characteristic peak The test value is used to determine, according to formula 4.2, that the Raman characteristic peak value belongs to the abnormal Raman spectrum data. 4.1 4.2 is the overall average of the Raman characteristic peaks, is the maximum number of the Raman spectrum, , is the standard deviation.
5. The method according to claim 3 or 4, characterized in that The step S300 specifically includes: S310: Decomposing the Raman spectrum array into the baseline, the characteristic peak and the noise, as shown in Formula 5.1 and Formula 5.2, 5.1 5.2; S320: Fit the baseline using a polynomial, and solve the coefficients using the least square method, such as formula 5.3, and solve the minimized residual square according to formula 5.
4. 5.3 5.4 and The Raman spectrum array for the i-th sampling group The function and independent variable of is the baseline in the Raman spectrum array of the i-th sampling group, is the characteristic peak in the Raman spectrum array of the i-th sampling group, is the noise in the Raman spectrum array of the i-th sampling group, is the least squares coefficient, is the order of the regularized smoothing parameter, ranging from 3 to 5. is the number of the Raman shift point, , is the Raman spectrum shift range, which is determined by the spectrum sensor of the Raman spectrum device; S330: For the Raman spectrum amplitudes of all the Raman shift points in all the Raman spectrum arrays, select a regularized smoothing parameter 10 3 Up to 10 5 Level, the initial weight is 1.0, and the weight is updated from 0.001 to 0.1 until the baseline is flat. The baseline and noise at this point are deducted to obtain the characteristic peak and complete the baseline correction.
6. The method according to claim 5, characterized in that The step S400 specifically includes: S410: Select wavelet basis function from db4 to db6, select sym6 for Symlets, the number of decomposition layers from 0 to 7, and the soft threshold adjustment factor from 0.5 to 2.0; S420: For the Raman spectrum array after the baseline correction, as shown in Formula 6.1 and Formula 6.2 respectively define a low-pass filter and a high-pass filter, for the Raman spectrum amplitude of the Raman shift point in the Raman spectrum array , executing the steps including formula 6.3 to obtain low-frequency components by the low-pass filtering decomposition, formula 6.4 to obtain high-frequency components by the high-pass filtering decomposition, formula 6.5 to reconstruct to obtain noise reduction and perform downsampling, 6.1 6.2 6.3 6.4 6.5 is the low-pass filter, is the high-pass filter, is the reconstructed low-pass filter, is the reconstructed high-pass filter, is the filter decomposition layer number. For db4, the value range is 0-7. is the sampling group number, is the number of the Raman shift point, is the number of the down-sampling, is the Raman spectrum amplitude in the Raman spectrum array of the i-th sampling group function.
7. The method according to claim 6, characterized in that The step S500 specifically includes: S510: Based on one or more test item names, setting the dimension reduction number corresponding to the test item name in the Raman spectrum array to be less than 30; S520: In the Raman spectrum data after executing S400, the Raman spectrum amplitude with a fluctuation amplitude greater than 2% and the largest at both the front and rear wavenumber positions is selected as the pre-characteristic peak corresponding to the test item name; S530: for each of the test item names, sort the pre-characteristic peaks from large to small, and select the dimension reduction number arranged in front as the characteristic peak, wherein the characteristic peak includes the Raman spectrum amplitude, the Raman shift value and the test item name; S540: Establishing all the characteristic peaks as the pre-training data set; and / or, S550: Based on the test item name, in the Raman spectrum data after executing S400, select the Raman spectrum amplitude with a fluctuation amplitude greater than 10% and the largest at the front and rear wavenumber positions as the pre-characteristic peak corresponding to the test item name, calculate the number of the pre-characteristic peaks as the dimensionality reduction number, and take all the pre-characteristic peaks as the characteristic peaks corresponding to the test item name, including the Raman spectrum amplitude, the Raman shift value and the test item name.
8. The method according to claim 6, characterized in that Also includes S600 steps: S610: Based on one or more test item names, setting the dimension reduction number corresponding to the test item name in the Raman spectrum array to be less than 30; S620: Processing the pre-characteristic peak using a wavelet algorithm to obtain the characteristic peak including the Raman spectrum amplitude, the Raman shift value and the test item name.
9. The method according to claim 7, characterized in that Also includes S700, including: S710: For the characteristic peak, a recursive algorithm is used to verify and fine-tune the characteristic peak corresponding to the name of the test item, so as to obtain an optimized characteristic peak corresponding to the name of the test item; S720: Verify and fine-tune the characteristic peaks corresponding to all the test item names to obtain optimized characteristic peaks corresponding to all the test item names.
10. The method according to claim 8, characterized in that Also includes S800, specifically including: S810: For the characteristic peak, a recursive algorithm is used to verify and fine-tune the characteristic peak corresponding to the name of the test item, so as to obtain an optimized characteristic peak corresponding to the name of the test item; S820: Verify and fine-tune the characteristic peaks corresponding to all the test item names to obtain optimized characteristic peaks corresponding to all the test item names.
Citation Information
Patent Citations
Blended fabric component Raman spectra qualitative checking method
CN101285773A
Raman spectral preprocessing method
CN103217409A
Lossless method for quantitatively detecting blood glucose by utilizing Raman spectrum
CN103505221A
Thyroid dysfunction model and establishment method thereof
CN109346156A
Raman spectrum detection method and system with sensitivity and response speed
CN113624734A
Cited By
Noninvasive biochemical substance content detection method and system based on Raman spectrum
CN121101552A