Method for establishing high-dimensional raman spectrum in non-invasive detection dimension reduction training model
Through artificial intelligence supervised learning and dimensionality reduction training models, the selection of Raman spectral characteristic peaks is optimized, the problems of dimensionality reduction and systematic error of high-dimensional data are solved, the accuracy of quantitative measurement of mixed substances and the intelligence of equipment are achieved, and it is suitable for non-invasive detection equipment.
Patent Information
- Application Number
- CN202510493913.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-04-19
AI Technical Summary
Existing Raman spectroscopy technology has problems in the quantitative analysis of substances, such as difficulty in dimensionality reduction of high-dimensional data, inaccurate selection of characteristic peak parameters, large system errors of measurement equipment, superposition of characteristic peaks in quantitative measurement of mixed substances, and difficulty in the system supporting dynamic characteristic peak selection, resulting in poor quantitative detection results.
By adopting artificial intelligence supervised learning methods, through the high-dimensional Raman spectroscopy dimensionality reduction training model, combined with asymmetric least squares method, discrete wavelet transform and T/Z test technologies, the characteristic peak selection and equipment system error are optimized to achieve iterative optimization of dynamic characteristic peaks.
It achieves a significant dimensionality reduction of the Raman spectral characteristic peaks, improves the measurement precision and accuracy, solves the problem of characteristic peak superposition in the quantitative measurement of mixed substances, supports the accurate selection of dynamic characteristic peaks, and enhances the intelligence level of the equipment.
Smart Images

Figure CN120045846B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cheminformatics and biomedical, in particular to the sub-field of artificial intelligence quantitative detection of Raman spectrum of specific chemicals, especially the development of quantitative detection equipment for specific substances in mixed substances, and provides an artificial intelligence solution for obtaining Raman spectrum characteristic peaks of specified chemicals by supervised learning. BACKGROUND
[0002] Using Raman spectrum to non-destructively and non-invasively measure chemicals is an important hotspot in the field of spectroscopy measurement.
[0003] As early as 1928, Indian physicist C.V. Raman discovered the phenomenon of Raman scattering, that is, when an incident light with a wavelength of v0 irradiates a substance, due to the action of the incident light on the electrons of the substance, energy level transition occurs, thereby generating a batch of scattered light with a wavelength different from v0, and forming a spectrum. People call this scattered light and spectrum Raman scattered light and Raman spectrum, and find that the spectrum of these scattered light has a "fingerprint" relationship with the substance being irradiated. Accordingly, C.V. Raman won the Nobel Prize in Physics in 1930. With the breakthrough development of laser technology in the 1960s, the method of measuring the content of substances by Raman spectrum began to develop rapidly.
[0004] According to the inventor's search, in recent years, Raman spectrum has been more studied in the qualitative monitoring of specific substances, especially in non-destructive detection and non-invasive detection, and has developed rapidly. The existing main technologies related to Raman spectrum analysis and big data in industry, medical treatment are as follows:
[0005] I. Application of Raman spectrum in qualitative analysis of substances
[0006] Laser Raman spectrum
[0007] Currently, Raman spectrum technology is widely used for qualitative analysis of detected substances in the fields of chemistry, medicine, industry, life science, agriculture, food, and even criminal investigation. Especially with the development of laser technology and the maturity of Raman spectrum acquisition equipment or Raman spectrum testing equipment, the Raman spectrum analysis in the above fields is rapidly applied and developed, especially as a non-destructive and non-invasive qualitative analysis of substances.
[0008] Raman characteristic peak
[0009] The laser Raman spectrum analysis of the monitored substance can identify the vibration and rotation energy level of the substance according to the characteristic Raman shift peak of different substances, analyze the properties of the substance, and thus identify the substance or the substance contained in the detected substance, i.e. qualitative analysis of the substance.
[0010] Identification of mixture components
[0011] When the Raman spectrum analysis is performed on a mixture of multiple substances, the components of the mixture can be identified, which also belongs to the popular qualitative analysis of mixture components using Raman spectrum.
[0012] Surface-enhanced Raman scattering and tip-enhanced Raman spectrum
[0013] According to the inventor's search, in terms of sensitivity and accuracy of Raman spectrum analysis, with the development of modern nanotechnology, surface-enhanced Raman scattering (SERS) and tip-enhanced Raman spectrum (TERS) technology have been developed, which improves the intensity of the Raman scattering light of the monitored substance and improves the signal-to-noise ratio, making great progress in ultra-high sensitivity detection, which also opens up the quantitative analysis of substance components using Raman laser detection technology and provides a feasible technical direction.
[0014] II. Technical application of Raman spectrum in quantitative analysis of substances
[0015] Quantitative analysis of solution substance concentration
[0016] With the development of Raman spectrum technology, people not only need to perform qualitative analysis on the tested substance, but also have more application fields, such as the need for quantitative analysis of the content or concentration of the unknown substance detected or the content or concentration of the known substance in the solid mixture or solution of the substance. For example, in the chemical and industrial fields, Raman spectrum is used to analyze the concentration of alcohol in the mixed liquid, and in the medical field, Raman laser spectrum analyzer is used to analyze the blood sugar concentration of the patient's sampling blood sample using Raman spectrum. At present, the quantitative analysis based on Raman spectrum is generally performed in a special laboratory through complex sample sampling and detection, and the sample feature recognition is complex, the data processing and analysis largely depend on professional analysts, the analysis time is long, the processing efficiency is low, and the test result error is large, which will greatly affect the application of the technology in practice.
[0017] Basic calculation of Raman spectrum quantitative analysis
[0018] Raman spectrum in quantitative analysis applications, first to deal with the problem of interference such as fluorescence, while the hardware device noise requirements are very high, the general case can only be professional laboratory to be detected material testing, basic technical processes as follows:
[0019] Establish a certain pure substance solution known different concentrations and Raman peak intensity between the corresponding relationship, this relationship can be in the form of a table, but also can be fitted curve equation form, form a kind of reference comparison standard.
[0020] Use Raman spectroscopy equipment for unknown concentration of the detected substance solution detection, obtain Raman characteristic peak spectrum.
[0021] Read the Raman spectrum of the detected substance solution characteristic peak intensity value, known concentration of the substance and Raman peak intensity between the corresponding relationship table, artificial analysis or device embedded software according to the corresponding relationship table or fitting curve equation analysis of the concentration of the measured substance.
[0022] Noninvasive Raman blood glucose analysis
[0023] Now there are scholars and inventors put forward, through the Raman laser irradiation of human tissue parts generated by Raman spectroscopy, the realization of noninvasive blood glucose concentration analysis. However, whether it is industrial field directly on the monitored substance Raman quantitative analysis, or medical field noninvasive blood glucose Raman analysis, the device itself system noise, fluorescence interference, human tissue protein and other substances Raman light interference, and measurement error caused by Raman spectrum characteristic peak is not obvious even submerged, the quantitative test effect of the measured substance content or concentration is very not optimistic, only in the condition of very demanding laboratory can be realized, so as to affect the related equipment development progress, is not conducive to the market promotion of related Raman test products.
[0024] Three, artificial intelligence big data technology in the medical field
[0025] With the development of information technology and computer technology, people ushered in the era of "artificial intelligence big data and cloud", the traditional embedded devices or local system in the past due to the limitations of hardware conditions or data processing operation can not realize the qualitative, quantitative research and judgment, estimation, diagnosis, planning and other complex algorithms with artificial intelligence (AI) function, now through the cloud server, big data technology, one by one can be realized.
[0026] The core of big data technology is essentially a variety of complex and efficient data operation on massive data, especially the operation algorithm based on statistical theory. For example, the current epidemiological analysis of big data in the medical field, remote self-service mode of medical diagnosis expert system, etc. are typical "Internet + big data + artificial intelligence" typical application.
[0027] If big data technology is applied to Raman spectrum analysis, it can undoubtedly promote the development of current Raman spectrum monitoring in quantitative monitoring applications, break through the limitations of traditional hardware devices in complex artificial intelligence algorithm operation capability, and especially in the field of blood glucose monitoring analysis, some traditional and expensive fixed laboratory Raman blood glucose analysis equipment can gradually move out of the professional laboratory, realize miniaturization, portability, and even wearable non-invasive products, and be popularized to the daily blood glucose monitoring and management of users, especially diabetic users.
[0028] Disadvantages of prior art methods
[0029] According to the above analysis, the inventors believe that the existing Raman spectrum technology has the following disadvantages in the quantitative analysis of the monitored substance:
[0030] 1. High-dimensional Raman spectrum has a large dimension reduction problem for low-dimensional characteristic peaks.
[0031] In a professional Raman spectrometer, the spectral data are in the order of thousands, for example, 2048 points, and the number of characteristic peaks of a certain substance is usually less than two digits, for example, 5. From a mathematical point of view, it is a difficult problem to reduce 2048-dimensional data to 5-dimensional data while ensuring a certain reliability and accuracy.
[0032] 2. Inaccurate selection of Raman spectrum characteristic peak parameters.
[0033] At present, the selection of most substance characteristic peak parameters, such as wave number displacement value and amplitude, is based on previous papers and statistical data. It is difficult to determine whether they are accurate.
[0034] 3. Uncertainty of system error of Raman measurement equipment.
[0035] As a Raman spectrum measurement device, it not only has system errors itself, such as the spatial position error of the CCD sensor installed on the spectrometer integration module, which will directly affect the accuracy of the Raman wave number displacement, but also the performance of the CCD and the temperature of the cooler, which will directly affect the accuracy of the spectral measurement assignment.
[0036] 4. Characteristic peak superposition problem in quantitative measurement of mixed substances.
[0037] When measuring mixed substances, the Raman spectra of various molecular bonds are mixed. At this time, the characteristic peaks of pure substances are superimposed. It is very difficult to evaluate the concentration of a single measured substance in a mixed substance with varying concentrations of each substance.
[0038] 5. Difficulty in supporting accurate selection of dynamic characteristic peaks.
[0039] The measurement function inside the Raman spectrometer as a kind of measuring instrument is usually fixed, and does not support dynamic adjustment, so that the accuracy of characteristic peak selection is doubtful. SUMMARY
[0040] The present application is designed according to the deficiencies of the prior art, a "high-dimensional Raman spectrum in non-invasive detection dimensionality reduction training model establishment method" supporting dynamic and iterative self-learning, artificial intelligence supervised learning is used to continuously optimize the characteristic peak of Raman spectrum for the specified material, so that the device can be more "intelligent". The establishment of the pre-training model provides a better platform for intelligent Raman spectrum measurement of blood glucose.
[0041] The purpose and intention of the present application is realized by the working steps of the following technical solutions:
[0042] 1. Basic scheme step
[0043] The present application as a high-dimensional Raman spectrum in non-invasive detection dimensionality reduction training model establishment method, at least includes the following steps:
[0044] According to the patient's test sheet, while collecting the patient's venous blood, a group of more than one Raman spectrum data of the patient's skin is collected, the test result is digitized and recorded into the Raman spectrum data, and the original data set is established.
[0045] In the original data set, for the Raman spectrum data, the abnormal Raman spectrum data group is removed by using the test method including but not limited to.
[0046] The baseline correction is carried out on the Raman spectrum data by using the asymmetric least square method, the regularization smoothing parameter is selected to be 10 3 to 10 5 level, the initial weight is 1.0, the update weight starts from 0.001 to 0.1, until the baseline is flat.
[0047] The noise of the Raman spectrum data is reduced by using the discrete wavelet transform DWT method, the wavelet base function is selected to be db4-db6, the decomposition layer is selected to be 7 layers, and the soft threshold adjustment factor is 0.5-2.0.
[0048] In the Raman spectrum data, the pre-characteristic peak is selected and marked, the pre-characteristic peak tuning parameter and sorting are selected, the pre-characteristic peak arranged in the front of the dimensionality reduction number is selected as the characteristic peak, and the established dimensionality reduction result is used as the pre-training data set.
[0049] 2. Establishing the original data set step
[0050] On the basis of the foregoing basic scheme, the application, in the aspect of establishing the original data set step, specifically comprises but is not limited to one or more combinations of the following establishment of the original data set step or method:
[0051] The sample group number and the Raman spectrum number are marked, wherein the test result comprises but is not limited to a blood test item name and a blood test item array, the sample group comprises Raman spectrum data of all Raman spectrum numbers synchronously collected corresponding to one test sheet, and the Raman spectrum data comprises but is not limited to a Raman spectrum array and a blood test item array.
[0052] The test result is generated by a mode comprising but not limited to manual input or automatic recognition of a paper test sheet image, wherein the test result further comprises but is not limited to patient information and a test sheet number.
[0053] The Raman spectrum array is established to comprise but not limited to a one-dimensional array, and is composed of Raman shift points and Raman spectrum amplitudes of the Raman shift points determined by a Raman spectrometer, the order number of the Raman shift points starts from 1 and ends at a maximum Raman shift point.
[0054] The blood test item array is established to comprise but not limited to a one-dimensional array, and is composed of one or more blood test item names and one or more blood test item data obtained by one-time blood sampling, the order number starts from 1 and ends at the last blood test item name.
[0055] The Raman spectrum array further comprises but is not limited to a patient number, and Raman spectrum data of all patients is established into the original data set.
[0056] The model for establishing the original data set comprises but is not limited to: the Raman spectrum data of the i-th sample group is determined by formula 2.1, the blood test item array of the i-th sample group is determined by formula 2.2, the Raman spectrum array of the i-th sample group is determined by formula 2.3, and the original data set is determined by formula 2.4:
[0057] LM_DATA i =[LM_ARRAY i BT_ARRAY i ], 2.1
[0058] BT_ARRAY iQ =[(btn btd) i1 … (btn btd) iQ ], 2.2
[0059]
[0060] Wherein:
[0061] LM_DATA i The Raman spectrum data of the i-th sample group,
[0062] LM_ARRAY i the Raman spectrum array of the i-th sampling group,
[0063] BT_ARRAY i the blood test item array of the i-th sampling group,
[0064] ORG_DATA original data set,
[0065] i is the sampling group number, N is the maximum sampling group number, 1≤i≤N,
[0066] j is the Raman spectrum number, M is the maximum Raman spectrum number, 1≤j≤M, 1≤M≤500,
[0067] q is the blood test item number, Q is the maximum blood test item number, 1≤q≤Q,
[0068] k is the Raman shift point number, 1≤k≤K,
[0069] K is the Raman spectrum shift range, which is determined by the spectrum sensor of the Raman spectrum device,
[0070] btn is the blood test item name, btd is the blood test item data, and lm is the Raman spectrum amplitude.
[0071] 3. T-test to eliminate abnormal Raman spectrum
[0072] On the basis of the foregoing technical solutions, the application adopts, but is not limited to, the T-test to eliminate abnormal Raman spectrum step, and can also adopt one or more combinations of the following local improvement measures:
[0073] In the same sampling group of the original data set, when the maximum number of collected Raman spectra is less than 30, the T-test method is used to find abnormal Raman spectrum data, and the Raman spectrum data is eliminated, including but not limited to deletion or marking.
[0074] According to the known Raman characteristic peak value of the substance to be detected, the Raman shift point where the characteristic peak value is located and one or more Raman shift points before and after the Raman shift point are selected, and the Raman characteristic peak value of the shift point is accumulated or calculated as x j , as the difference vector of the T-test, the overall average of the i-th Raman characteristic peak value is calculated according to formula 3.1, the t-test value of the Raman characteristic peak value is calculated according to formula 3.2, and the standard deviation of the Raman characteristic peak value is calculated according to formula 3.3, and the p value is calculated by searching the built-in T-test table or calling but not limited to Python function, when p<0.05, it is determined that the Raman characteristic peak value belongs to abnormal Raman spectrum data.
[0075]
[0076]
[0077] μ is the population mean,
[0078] t is the sample number and the deviation of the population mean,
[0079] K is the total wave number point of each of the Raman spectra or the displacement range of the Raman spectra,
[0080] s is the standard deviation,
[0081] is the difference between the i-th Raman spectrum at each wave number point and the mean,
[0082] i is the sample group number, N is the maximum sample group number, 1≤i≤N,
[0083] j is the Raman spectrum number, M is the maximum Raman spectrum number, 1≤j≤M, 1≤M≤500,
[0084] is the population mean of the Raman characteristic peak.
[0085] 4. Z test to remove abnormal Raman spectra
[0086] On the basis of the foregoing technical solutions, the application includes but is not limited to using Z test to remove abnormal Raman spectra, and the following one or more combinations of local improvements can also be used:
[0087] In the same sample group of the original data set, when the maximum number of collected Raman spectra is 30 or more, the Z test method is used to find abnormal Raman spectrum data, and the Raman spectrum data is removed, including but not limited to deletion or marking.
[0088] According to the known Raman characteristic peak value of the substance to be detected, the Raman displacement point where the characteristic peak value is located and one or more Raman displacement points before and after it are selected, and the Raman characteristic peak value x of the displacement point is accumulated or calculated as a difference vector for Z test. j The Z test value of the i-th Raman characteristic peak is calculated according to formula 4.1, and it is determined according to formula 4.2 that the Raman characteristic peak belongs to abnormal Raman spectrum data.
[0089]
[0090] |z|≥2.58 4.2
[0091] is the population mean of the Raman characteristic peak.
[0092] M is the maximum number of Raman spectrum, 1≤j≤M, 1≤M≤500,
[0093] s is the standard deviation.
[0094] 5. Baseline correction
[0095] On the basis of the foregoing technical solutions, the present application includes but is not limited to the baseline correction step, and the following one or more combinations of local improvement measures and steps can be specifically used:
[0096] Decompose the Raman spectrum array into baseline, characteristic peak and noise, as shown in formula 5.1 and formula 5.2,
[0097] LM_ARRAY i (x)=LM_ARRAY_T i (x)+LM_ARRAY_B i (x)+LM_ARRAY_H i (x) 5.1
[0098] LM_ARRAY_B i (x)=LM_ARRAY i (x)-LM_ARRAY_T i (x)-LM_ARRAY_H i (x) 5.2
[0099] S320: adopt the baseline including polynomial fitting, solve the coefficients by least square method, as shown in formula 5.3, solve the minimum residual square according to formula 5.4,
[0100] LM_ARRAY_B i (x)=a0+a1x+a2x 2 +…+a n x n 5.3
[0101]
[0102] LM_ARRAY i (x) and x are functions and independent variables of the Raman spectrum array lm ik of the i-th sampling group,
[0103] LM_ARRAY_B i (x) is the baseline in the Raman spectrum array of the i-th sampling group,
[0104] LM_ARRAY_T i (x) is the characteristic peak in the Raman spectrum array of the i-th sampling group,
[0105] LM_ARRAY_H i (x) is noise in the Raman spectrum array of the i-th sampling group,
[0106] a is a least square coefficient,
[0107] n is an order of a regularization smoothing parameter, and the value range is 3 to 5,
[0108] k is the Raman shift point number, and 1≤k≤K,
[0109] K is the Raman spectrum shift range, which is determined by a spectrum sensor of the Raman spectrum device.
[0110] S330: for the Raman spectrum amplitude of all the Raman shift points in all the Raman spectrum arrays, a regularization smoothing parameter 10 3 to 10 5 levels, the initial weight is 1.0, the updated weight starts from 0.001 to 0.1, until the baseline is flat, the baseline and noise of the point are deducted, the characteristic peak is obtained, and the baseline correction is completed.
[0111] 6, eliminating noise
[0112] On the basis of the foregoing technical solutions, the present application includes but is not limited to the step of eliminating noise, specifically comprising:
[0113] The wavelet base function is selected as db4 to db6, sym6 is selected for Symlets, the decomposition layer number is 0-7 layers, and the soft threshold adjustment factor is 0.5-2.0.
[0114] For the Raman spectrum array after the baseline correction, the low-pass filter and the high-pass filter are respectively defined as formula 6.1 and formula 6.2, for the Raman spectrum amplitude lm of the Raman shift point in the Raman spectrum array, the steps of executing formula 6.3 through low-pass filter decomposition to obtain low-frequency components, formula 6.4 through high-pass filter decomposition to obtain high-frequency components, and formula 6.5 for reconstruction to obtain noise reduction and perform down-sampling,
[0115]
[0116] vA m [k]=∑ k [k-2m]·cA m-1 [k] 6.3
[0117] cD m [k]=∑ k [k-2m]·cA m-1 [k] 6.4
[0118]
[0119] h[m] is the low-pass filter,
[0120] g[m] is the high-pass filter,
[0121] is the reconstructed low-pass filter,
[0122] is the reconstructed high-pass filter,
[0123] m is the filter decomposition layer number, for db4, the value range is 0-7,
[0124] i is the sampling group number,
[0125] k is the Raman shift point number,
[0126] [k-2m] is the down-sampling number,
[0127] LM_ARRAY i is a function of the Raman spectrum amplitude lm in the Raman spectrum array of the i-th sampling group. ik
[0128] 7, Dimensionality reduction generates characteristic peaks
[0129] On the basis of the foregoing technical solutions, the application includes but is not limited to establishing a characteristic peak step, specifically including:
[0130] Based on one or more assay item names, the dimensionality reduction number corresponding to the assay item name in the Raman spectrum array is set to be 30 or less.
[0131] In the Raman spectrum data after executing noise elimination, the Raman spectrum amplitude with a fluctuation amplitude greater than 2% and the maximum at the front and rear wave number positions is selected as the pre-characteristic peak corresponding to the assay item name.
[0132] For each assay item name, the pre-characteristic peaks are sorted from large to small, and the dimensionality reduction number arranged in the front is selected as the characteristic peak, which includes but is not limited to the Raman spectrum amplitude, the Raman shift value and the assay item name.
[0133] Optimally, based on the assay item name, in the Raman spectrum data after executing noise elimination, the Raman spectrum amplitude with a fluctuation amplitude greater than 10% and the maximum at the front and rear wave number positions is selected as the pre-characteristic peak corresponding to the assay item name, the number of pre-characteristic peaks is calculated as the dimensionality reduction number, and all pre-characteristic peaks are taken as the characteristic peak corresponding to the assay item name, including the Raman spectrum amplitude, the Raman shift value and the assay item name.
[0134] All feature peaks are established as a pre-training data set.
[0135] 8. Wavelet algorithm dimension reduction to generate feature peaks
[0136] On the basis of the foregoing technical solutions, the present application includes but is not limited to the following feature peak establishment steps, specifically comprising:
[0137] Based on one or more assay item names, the dimension reduction number corresponding to the assay item name in the Raman spectrum array is set to be 30 or less.
[0138] Preferably, the wavelet algorithm is used to process the pre-feature peaks to obtain feature peaks including Raman spectrum amplitude, Raman shift value and assay item name.
[0139] 9. Recursive algorithm dimension reduction to generate feature peaks
[0140] On the basis of the foregoing technical solutions of dimension reduction to generate feature peaks, the present application includes but is not limited to the following recursive algorithm feature peak steps, specifically comprising:
[0141] For feature peaks, the recursive algorithm is used to calculate and fine-tune the feature peaks corresponding to one assay item name to obtain optimized feature peaks corresponding to one assay item name.
[0142] Preferably, the calculation and fine-tuning are performed on the feature peaks corresponding to all assay item names to obtain optimized feature peaks corresponding to all assay item names.
[0143] 10. Recursive algorithm dimension reduction to generate feature peaks
[0144] Preferably, on the basis of the technical solutions of wavelet algorithm dimension reduction to generate feature peaks, the recursive algorithm is used to calculate and fine-tune the feature peaks corresponding to one assay item name to obtain optimized feature peaks corresponding to one assay item name.
[0145] The calculation and fine-tuning are performed on the feature peaks corresponding to all assay item names to obtain optimized feature peaks corresponding to all assay item names.
[0146] 11. Purpose and intention of the application
[0147] The inventor proposes a method for establishing a dimension reduction training model of high-dimensional Raman spectrum in non-invasive detection through long-term research, observation and experiment, which is also a training method of artificial intelligence supervised learning.
[0148] The purpose and intention of the present application is to establish a complete method for establishing a dimension reduction training model of high-dimensional Raman spectrum in non-invasive detection, and to solve the problems of the prior art, specifically including the following solutions:
[0149] 1. High-dimensional Raman spectrum is greatly reduced in dimension for low-dimensional characteristic peaks.
[0150] In the present application, 2048-dimensional Raman spectrum is greatly reduced in dimension to single digits, for example, 5 dimensions, by pre-characteristic peak learning, ordering method, wavelet dimension reduction method, and recursive optimization method.
[0151] 2. Inaccurate selection of Raman spectrum characteristic peak parameters.
[0152] The method of supervised learning and the biochemical instrument test results at the same time are continuously compared and learned, so as to continuously iterate and optimize the characteristic peak parameters.
[0153] 3. System error uncertainty of Raman measurement device.
[0154] Through supervised learning, the system error of the device is iteratively reduced, and the measurement accuracy is continuously improved.
[0155] 4. Characteristic peak superposition problem of mixed substance quantitative measurement.
[0156] Within the visible range of the mixing ratio of the mixture, for example, although the subcutaneous mixture of human skin is complex, the substance types and mixing ratio are ultimately within a certain range, and through a large number of biochemical instrument blood tests, the system can determine the relevant parameters. In addition, after the device is set to cloud mode, as all devices continuously learn, the system will quickly reach an optimized result. Therefore, the present application can make the device more and more "intelligent".
[0157] 5. System difficulty in supporting dynamic characteristic peak accurate selection.
[0158] The internal measurement function of the traditional measurement device is static, while the present application adopts dynamic supervised learning, and with continuous iteration, the present application can perfectly support dynamic optimization of characteristic peaks.
[0159] 12. Beneficial effects of the present application
[0160] 1. The invention achieves the intended purpose, solves the problem of greatly reducing the dimension of Raman spectrum to characteristic peaks, and solves the problems of inaccurate selection of Raman spectrum characteristic peak parameters, system error uncertainty of Raman measurement device, characteristic peak superposition problem of mixed substance quantitative measurement, and system difficulty in supporting dynamic characteristic peak accurate selection.
[0161] 2. The present application provides the independence of algorithm and hardware device.
[0162] 3. The present application provides a selectable characteristic peak sequence.
[0163] 4. The present application provides the pre-innovation achievement and foundation for later quantitative measurement of specific substances in mixed substances. BRIEF DESCRIPTION OF DRAWINGS
[0164] List of drawings:
[0165] Figure 1 : System layer structure schematic diagram
[0166] Figure 2 : Embodiment hardware reference structure diagram
[0167] Figure 3 : Original Raman spectrum diagram
[0168] Figure 4 : Baseline correction spectrum diagram
[0169] Figure 5 : Pre-training set application example
[0170] Detailed description of the drawings:
[0171] See the embodiments for details. Specific embodiments
[0172] The purpose and intention of the present application can be realized by the design method of the following embodiments. It should be particularly noted that since the specific embodiments have specific purposes and industrial applicability, the embodiments cannot include all the features and steps of the present application, and the description of the claims of the present application is the summary of the invention.
[0173] The specific embodiments of the present application are as follows:
[0174] Pre-training model establishment method for non-invasive blood glucose Raman spectrum quantitative monitoring in vitro
[0175] This example is a general example of the present application, which is a general example of the AI pre-training model establishment method for glucose concentration when using Raman spectrum to measure subcutaneous glucose non-invasively through human skin.
[0176] It should be pointed out that the content and view of the present embodiment are not a limitation of the present application, nor can they fully interpret the present application, but only one of the embodiments of the industrial application of the present application.
[0177] 1. Diagram explanation
[0178] The content of the present embodiment mainly includes but is not limited to the following main schematic diagrams, which are: Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 It should be pointed out that, Figure 4 in the left side, the data overflow is marked with a red line due to the measurement instrument exceeding the range.
[0179] 2. Implementation step explanation
[0180] The method steps of the embodiment mainly include steps 1 to 7. Among them, unless otherwise specified, the step numbers of the 7 parts do not have a sequence, and the combination of the 7 parts is not required in any embodiment. In addition, each of the 7 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and their sequence is not necessary, but is optimized by the patent implementer according to the needs of some specific tasks.
[0181] The specific working steps are as follows:
[0182] 2.1, basic scheme step description
[0183] The method for establishing a dimension reduction training model of high-dimensional Raman spectrum in non-invasive detection, at least includes the following steps:
[0184] According to the patient's test sheet, while collecting the patient's venous blood, a group of more than one Raman spectrum data of the patient is collected synchronously, the test result is digitized and recorded into the Raman spectrum data, and an original data set is established.
[0185] In the original data set, for the Raman spectrum data, test methods including but not limited to are used to eliminate abnormal Raman spectrum data groups.
[0186] The baseline correction of the Raman spectrum data is carried out by using the asymmetric least squares method, the regularization smoothing parameter is selected to be 10 3 to 10 5 level, the initial weight is 1.0, the update weight starts from 0.001 to 0.1, and the baseline is flat.
[0187] The Raman spectrum data is denoised by using the discrete wavelet transform (DWT) method, the wavelet basis function is selected to be db4-db6, the decomposition layer is selected to be 7 layers, and the soft threshold adjustment factor is 0.5-2.0.
[0188] Preferably or further, in the Raman spectrum data, the pre-feature peaks are selected, the pre-feature peak tuning parameters and sorting are selected, the pre-feature peaks arranged in the front are selected as the feature peaks, and the established dimension reduction result is used as the pre-training data set.
[0189] Preferably or further, as the hardware platform running the method, from the functional point of view, it includes a Raman spectrum acquisition subsystem, a lens subsystem, a test sheet recognition subsystem, a computer subsystem, a communication subsystem, a display subsystem and a shell. In the test sheet acquisition, a sub-device based on video or image automatic acquisition and character recognition is used, and is linked with the Raman spectrum acquisition to facilitate automatic recognition of the test sheet information characters and synchronous acquisition of the Raman spectrum while the patient is drawing venous blood.
[0190] Figure 1 is a schematic diagram of the system layer structure of the present application. From a logical point of view, the system is mainly divided into an original data set layer, an outlier elimination layer, a baseline correction layer, a noise reduction layer and a pre-training set layer.
[0191] Figure 2 is an example of the system hardware platform scheme adopted in the present embodiment. From a structural point of view, it mainly includes a skin detection point, an excitation light module and a laser control module, a half-reflective half-transmissive lens, a condenser lens 1, a filter, a condenser lens 2, a slit, a condenser lens 3, a beam splitter, an integral module of a spectrometer, a refrigerator and a general control module.
[0192] 2.2, Step-by-step explanation of establishing an original data set
[0193] On the basis of the foregoing basic scheme, the present application, in terms of the step of establishing an original data set, specifically includes but is not limited to the following one or more combinations of steps or methods of establishing an original data set:
[0194] The sampling group number and the Raman spectrum number are marked, wherein the test results include but are not limited to blood test item names and blood test item arrays, the sampling group includes all Raman spectrum data of the Raman spectrum numbers synchronously collected corresponding to one test sheet, and the Raman spectrum data include but are not limited to Raman spectrum arrays and blood test item arrays.
[0195] Preferably, the test results are generated in a manner including but not limited to manual input or automatic recognition of paper test sheet images, wherein the test results further include but are not limited to patient information and test sheet numbers.
[0196] Further, the Raman spectrum array is established to include but not limited to a one-dimensional array, which is composed of Raman shift points determined by a Raman spectrometer and Raman spectrum amplitudes at the points, and the order number of the Raman shift points starts from 1 and ends at the maximum Raman shift point.
[0197] Further, the blood test item array is established to include but not limited to a one-dimensional array, which is composed of one or more blood test item names and one or more blood test item data obtained by one-time blood sampling, and the order number starts from 1 and ends at the last blood test item name.
[0198] Preferably, the Raman spectrum array further includes but is not limited to patient numbers, and the Raman spectrum data of all patients are established into an original data set.
[0199] Further, the model of establishing an original data set includes but is not limited to: the Raman spectrum data of the i-th sampling group is determined by formula 2.1, the blood test item array of the i-th sampling group is determined by formula 2.2, the Raman spectrum array of the i-th sampling group is determined by formula 2.3, and the original data set is determined by formula 2.4:
[0200] LM_DATA i = [LM_ARRAY i BT_ARRAY i ], 2.1
[0201] BT_ARRAY iQ = [(btn btd) i1 … (btn btd) iQ ], 2.2
[0202]
[0203]
[0204] wherein:
[0205] LM_DATA i Raman spectrum data of the i-th sampling group,
[0206] LM_ARRAY i Raman spectrum array of the i-th sampling group,
[0207] BT_ARRAY i Blood test item array of the i-th sampling group,
[0208] ORG_DATA original data set,
[0209] i is the sampling group number, N is the maximum number of sampling groups, 1≤i≤N,
[0210] j is the Raman spectrum number, M is the maximum number of Raman spectra, 1≤j≤M, 1≤M≤500,
[0211] q is the blood test item number, Q is the maximum number of blood test items, 1≤q≤Q,
[0212] k is the Raman shift point number, 1≤k≤K,
[0213] K is the Raman spectrum shift range, determined by the spectral sensor of the Raman spectrum device,
[0214] btn is the blood test item name, btd is the blood test item data, and lm is the Raman spectrum amplitude.
[0215] Preferably or further, in the definition of the above data structure, the data structure itself can adopt a two-dimensional table, a multi-dimensional array, or a structure defined based on functions included in a software development language; the content of the data structure can be increased with other content according to the needs of the system, such as patient personal information, hospital information, electronic medical record information, medication information, etc.; the operation relationship of the data structure is not limited to Formulas 2.1 to 2.4, and according to the needs of the system users, other calculation relationships can also be extended or modified.
[0216] Figure 3 is a plane expansion diagram of the Raman spectrum actually measured in the original data set constructed by the present application, wherein the abscissa is the displacement of the Raman spectrum, and the ordinate is the amplitude of the Raman spectrum.
[0217] 2.3, T-test step for removing abnormal Raman spectrum
[0218] On the basis of the foregoing technical solutions, the present application includes but is not limited to the step of using T-test to remove abnormal Raman spectrum, and the following one or more combinations of local improvement measures can also be used:
[0219] In the same sampling group of the original data set, when the maximum number of the collected Raman spectrum is below 30, the T-test method is used to find abnormal Raman spectrum data, and the Raman spectrum data is removed, including but not limited to deletion or marking.
[0220] Further, according to the known Raman characteristic peak value of the substance to be detected, the Raman displacement point at the characteristic peak value and one or more Raman displacement points before and after the characteristic peak value are selected, and the Raman characteristic peak value at the displacement point is accumulated or calculated as x j , as the difference vector of T-test, the population mean of the ith Raman characteristic peak value is calculated according to Formula 3.1, the T-test value of the Raman characteristic peak value is calculated according to Formula 3.2, the standard deviation of the Raman characteristic peak value is calculated according to Formula 3.3, the p value is calculated by looking up the built-in T-test table or calling but not limited to Python function, and when p < 0.05, it is determined that the Raman characteristic peak value belongs to abnormal Raman spectrum data,
[0221]
[0222] μ is the population mean,
[0223] t is the deviation statistic of the sample number and the population mean,
[0224] K is the total wave number point of each Raman spectrum or the displacement range of the Raman spectrum,
[0225] s is the standard deviation,
[0226] to represent the difference between the i-th Raman spectrum at each wave number point and the mean value,
[0227] i is the sample group number, N is the maximum sample group number, 1≤i≤N,
[0228] j is the Raman spectrum number, M is the maximum Raman spectrum number, 1≤j≤M, 1≤M≤500,
[0229] is the overall average of the Raman characteristic peaks.
[0230] Preferably or further, according to the single Raman pre-characteristic peak, the Z test is further used to eliminate abnormal pre-characteristic peaks, so as to further screen the characteristic peaks.
[0231] Preferably or further, according to the whole Raman spectrum, the Z test is further used to eliminate abnormal Raman spectra, so as to further screen the Raman spectra.
[0232] 2.4, Z test eliminates abnormal Raman spectrum step description
[0233] On the basis of the foregoing technical solutions, the application includes but is not limited to the Z test for eliminating abnormal Raman spectra, and the following one or more combinations of local improvements can also be used:
[0234] In the same sample group of the original data set, when the maximum number of collected Raman spectra is 30 or more, the Z test method is used to find abnormal Raman spectrum data, and the Raman spectrum data is eliminated, including but not limited to deletion or marking.
[0235] Further, according to the known Raman characteristic peak value of the substance to be detected, the Raman shift point where the characteristic peak value is located and one or more Raman shift points before and after the Raman shift point are selected, and the Raman characteristic peak value x j at the displacement point is accumulated or calculated as a difference vector for the Z test, the Z test value of the i-th Raman characteristic peak value is calculated according to formula 4.1, and it is determined that the Raman characteristic peak value belongs to abnormal Raman spectrum data according to formula 4.2.
[0236]
[0237] |z|≥2.58 4.2
[0238] is the overall average of the Raman characteristic peaks,
[0239] M is the maximum Raman spectrum number, 1≤j≤M, 1≤M≤500,
[0240] s is the standard deviation.
[0241] Preferably or further, according to the single Raman pre-characteristic peak, Z test is used to further eliminate abnormal pre-characteristic peak, so as to further screen the characteristic peak.
[0242] Preferably or further, according to the whole Raman spectrum, Z test is used to further eliminate abnormal whole Raman spectrum, so as to further screen the Raman spectrum.
[0243] 2.5, baseline correction step description
[0244] On the basis of the foregoing technical solutions, the present application includes but is not limited to the baseline correction step, and the following one or more combinations of local improvement measures and steps can be specifically used:
[0245] Decompose the Raman spectrum array into baseline, characteristic peak and noise, as shown in formula 5.1 and formula 5.2,
[0246] LM_ARRAY i (x)=LM_ARRAY_T i (x)+LM_ARRAY_B i (x)+LM_ARRAY_H i (x) 5.1
[0247] LM_ARRAY_B i (x)=LM_ARRAY i (x)-LM_ARRAY_T i (x)-LM_ARRAY_H i (x) 5.2
[0248] Further, the baseline is fitted by including but not limited to a polynomial, and the coefficients are solved by least square method, as shown in formula 5.3, and the minimum residual square is solved according to formula 5.4,
[0249] LM_ARRAY_B i (x)=a0+a1x+a2x 2 +…+a n x n 5.3
[0250]
[0251] LM_ARRAY i (x) and x are functions and independent variables of the i-th sampling group Raman spectrum array lm ik ,
[0252] LM_ARRAY_B i (x) is the baseline in the i-th sampling group Raman spectrum array,
[0253] LM_ARRAY_Ti (x) is a characteristic peak in the i-th sampling group Raman spectrum array,
[0254] LM_ARRAY_H i (x) is a noise in the i-th sampling group Raman spectrum array,
[0255] a is a least square coefficient,
[0256] n is an order of a regularization smoothing parameter, and the value range is 3 to 5,
[0257] k is a Raman shift point number, and 1≤k≤K,
[0258] K is a Raman spectrum shift range, which is determined by a spectrum sensor of a Raman spectrum device.
[0259] Preferably or further, the parameters in the formulae 5.1 to 5.4 are adjusted so that the baseline becomes a straight line, and only the characteristic peak is reserved on the baseline.
[0260] Figure 4 The measured Raman spectrum is unfolded after the baseline correction. Compared with Figure 3 , Figure 3 it is similar to a downward slope, and Figure 4 the baseline is zeroed so as to highlight the pre-characteristic peak of the Raman spectrum. Figure 4 The pre-characteristic peak in the formula 5.1 has a high probability of corresponding to the molecular bond of the detected substance. The reason for being called high probability is that there is a superposition of waves.
[0261] 2.6, elimination of noise step description
[0262] On the basis of the foregoing technical solutions, the present application includes but is not limited to the noise elimination step, and specifically includes:
[0263] The wavelet base function is selected as db4, the decomposition layer is 0-7 layers, and the soft threshold adjustment factor is 0.5-2.0.
[0264] Further, for the Raman spectrum array after the baseline correction, the steps including but not limited to the formula 6.1 through low-pass filtering to obtain a low-frequency component, the formula 6.2 through high-pass filtering to obtain a high-frequency component, and down-sampling are performed.
[0265]
[0266] cA m [k]=∑ k [k-2m]·cA m-1 [k] 6.3
[0267] cD m [k]=∑ k[k-2m] · cA m-1 [k] 6.4
[0268]
[0269] h[m] is a low-pass filter,
[0270] g[m] is a high-pass filter,
[0271] is a reconstruction low-pass filter,
[0272] is a reconstruction high-pass filter,
[0273] m is a filter decomposition layer number, for db4, the value range is 0-7,
[0274] i is a sampling group number,
[0275] k is a Raman shift point number,
[0276] [k-2m] is a down-sampling number,
[0277] LM_ARRAY i is the Raman spectrum amplitude lm in the i-th sampling group Raman spectrum array ik is a function.
[0278] Preferably or further, the parameters of formulas 6.1 and 6.2 are adjusted by looking at the low-frequency components, so that they become flat lines, to eliminate low-frequency noise interference. By looking at the high-frequency components, compared with the characteristic peaks, in units of the width of the characteristic peaks, fluctuations less than the width of the characteristic peaks are determined as high-frequency interference. Finally, all low-frequency and high-frequency interference is eliminated.
[0279] 2.7, Dimensionality reduction to generate characteristic peak step description
[0280] On the basis of the foregoing technical solutions, the present application includes but is not limited to establishing a characteristic peak step, specifically comprising:
[0281] Based on one or more assay item names, the dimensionality reduction number in the Raman spectrum array corresponding to the assay item name is set to 30 or less.
[0282] It should be noted that a Raman spectrum usually has more than 1,000 Raman spectrum amplitude data of Raman shift points. For example, a Raman spectrum collected by a certain Raman spectrometer has a total of 2048 points, which in almost all calculations, each is regarded as a dimension of a high-dimensional space. In order to reduce the amount of calculation and save computing power, in the present application, a large amount of dimensionality reduction measures are adopted, so that the original 2048 points are reduced to 30 points or less. Because usually for a characteristic peak, 5 to 12 points at a point are enough.
[0283] In the Raman spectrum data after noise elimination, the Raman spectrum amplitude with an amplitude of fluctuation greater than 2% and maximum at the front and rear wave number positions is selected as the pre-characteristic peak corresponding to the test item name. The "Raman spectrum amplitude with maximum at the front and rear wave number positions" is actually an extreme value on the Raman spectrum line, and the derivative of the point is 0 from the continuous spectrum.
[0284] Further, for each test item name, the pre-characteristic peaks are sorted from large to small, and the dimensionality reduction number arranged in the front is selected as the characteristic peak, which includes but is not limited to the Raman spectrum amplitude, the Raman shift value and the test item name.
[0285] Optimally, based on the test item name, in the Raman spectrum data after noise elimination, the Raman spectrum amplitude with an amplitude of fluctuation greater than 10% and maximum at the front and rear wave number positions is selected as the pre-characteristic peak corresponding to the test item name, the number of pre-characteristic peaks is calculated as the dimensionality reduction number, and all pre-characteristic peaks are taken as the characteristic peaks corresponding to the test item name, including the Raman spectrum amplitude, the Raman shift value and the test item name.
[0286] All characteristic peaks are taken as the pre-training data set.
[0287] Preferably or further, for a specific substance, although there are many characteristic peaks, based on the superposition of the mixture on the Raman spectrum, the arrangement and combination will show infinite possibilities on the Raman shift axis, as shown in Figure 4 Mathematically, the actual data dimension will be extremely large. As shown in the hardware structure of Figure 2 The data generated by the spectrometer integration module is 2048, that is, 2048 dimensions, which not only consumes a lot of computing power in calculation, but also produces meaningless redundancy in actual measurement using subsequent artificial intelligence algorithms. Therefore, the dimensionality must be reduced. According to the pre-characteristic peaks of the specific detection material and the human skin, the dimensionality is reduced, that is, the number of characteristic peaks is reduced to 5 to 30, which can stably and accurately distinguish the specific substance under the premise of ensuring the measurement accuracy, and then quantitatively measure the content of the specific substance.
[0288] It should be noted that the amplitude of fluctuation here refers to the difference between the peak and the trough of the pre-characteristic peak as the fluctuation value, and the maximum fluctuation value of the pre-characteristic peak is taken as the denominator for calculation, for example:
[0289]
[0290] Figure 5The figure is a characteristic peak diagram required for detecting the content of subcutaneous glucose, and only 5 characteristic peaks are required to complete reliable detection of the glucose content, with a measurement error of not more than 5%.
[0291] 2.8, Wavelet algorithm dimension reduction to generate characteristic peak step description
[0292] On the basis of the foregoing technical solutions, the present application establishes the characteristic peak step including but not limited to the following, specifically comprising:
[0293] Preferably, based on one or more test item names, the dimension reduction number corresponding to the test item name in the Raman spectrum array is set to be 30 or less.
[0294] Preferably, the wavelet algorithm is used to process the pre-characteristic peak to obtain the characteristic peak including the Raman spectrum amplitude, Raman shift value and test item name.
[0295] 2.9, Recursive algorithm dimension reduction to generate characteristic peak step description
[0296] On the basis of the foregoing dimension reduction to generate characteristic peak technical solutions, the present application adopts the recursive algorithm characteristic peak step including but not limited to the following, specifically comprising:
[0297] Preferably, for the characteristic peak, the recursive algorithm is used to calculate and fine-tune the characteristic peak corresponding to one test item name to obtain the optimized characteristic peak corresponding to one test item name.
[0298] Preferably, the calculation and fine-tuning are performed on the characteristic peaks corresponding to all test item names to obtain the optimized characteristic peaks corresponding to all test item names.
[0299] Here, the recursive algorithm is used according to the dynamic accumulation of the original data set, for example, at a certain moment, 5000 test sample samples are collected in the original data set, at this moment, the recursive calculation is performed once, and a plurality of characteristic peak sets are found and modified, and after the dynamic increase of the test samples in the original data set, for example, to 6000 samples, the recursive calculation is performed again, and since 1000 new samples are added, the recursive calculation finds that a plurality of new characteristic peaks need to be modified. The intention of the recursion here is a process of iteratively optimizing the characteristic peak data set, so that the characteristic peak data set becomes more and more accurate.
[0300] 2.10, Recursive algorithm dimension reduction to generate characteristic peak step description
[0301] Preferably, on the basis of the wavelet algorithm dimension reduction to generate characteristic peak technical solutions, the recursive algorithm is used to calculate and fine-tune the characteristic peak corresponding to one test item name to obtain the optimized characteristic peak corresponding to one test item name.
[0302] The characteristic peaks corresponding to all the test item names are checked and fine-tuned to obtain the optimized characteristic peaks corresponding to all the test item names.
Claims
1. A method for establishing a dimensionality reduction training model for high-dimensional Raman spectroscopy in non-invasive detection, characterized by: include: S100: According to the patient's test sheet, while collecting the patient's venous blood, simultaneously collect one or more sets of Raman spectral data of the patient's skin, digitize the test results and enter them into the Raman spectral data to establish an original data set; S200: using a test method for the Raman spectrum data in the original data set to eliminate abnormal Raman spectrum data sets; S300: Baseline correction is performed on the Raman spectrum data using an asymmetric least squares method, and a regularization smoothing parameter of 10 is selected. 3 to 10 5 Level, with an initial weight of 1.0, and updating weights starting from 0.001 to 0.1 until the baseline is flat; S400: using a discrete wavelet transform (DWT) method to reduce noise on the Raman spectrum data, selecting a wavelet basis function of db4-db6, a decomposition layer number of 7 layers, and a soft threshold adjustment factor of 0.5-2.0; S500: Select and mark pre-characteristic peaks in the Raman spectrum data, tune parameters and sort the pre-characteristic peaks, select the pre-characteristic peaks with the highest dimensionality reduction number as the characteristic peaks, and use the established dimensionality reduction results as a pre-training data set.
2. The method according to claim 1, characterized in that The S100 specifically includes the following steps: S110: Marking the sampling group number and the Raman spectrum number, wherein the test result includes the blood test item name and the blood test item array, the sampling group includes the Raman spectrum data of all the Raman spectrum numbers corresponding to the test sheet collected synchronously, and the Raman spectrum data includes the Raman spectrum array and the blood test item array; S120: Generate the test result by manual entry or automatic recognition of a paper test order image, wherein the test result also includes the patient information and the test order number; S130: Establishing the Raman spectrum array to include a one-dimensional array, which is composed of Raman shift points determined by a Raman spectrometer and Raman spectrum amplitudes of the Raman shift points, wherein the sequence number of the Raman shift points starts from 1 and ends at the maximum Raman shift point; S140: Establishing the blood test item array to include a one-dimensional array, which is composed of one or more blood test item names and one or more blood test item data obtained from a blood test, with the sequence number starting from 1 and ending with the last blood test item name; S150: The Raman spectrum array further includes the patient number, and the Raman spectrum data of all the patients are established as the original data set; and / or, S160: Establishing a model of the original data set includes: determining the Raman spectrum data of the i-th sampling group by formula 2.1, determining the blood test item array of the i-th sampling group by formula 2.2, determining the Raman spectrum array of the i-th sampling group by formula 2.3, and determining the original data set by formula 2.4: LM_DATA i =[LM_ARRAY i BT_ARRAY i ], 2.1 BT_ARRAY iQ =[(btn btd) i1 ... (btn btd) iQ ], 2.2 in: LM_DATA i The Raman spectrum data of the i-th sampling group, LM_ARRAY i The Raman spectrum array of the i-th sampling group, BT_ARRAY i The array of blood test items for the i-th sampling group, ORG_DATA original data set, i is the sampling group number, N is the maximum number of the sampling group, 1≤i≤N, j is the number of the Raman spectrum, M is the maximum number of the Raman spectrum, 1≤j≤M, 1≤M≤500, Q is the maximum number of blood test items, K is the Raman spectrum shift range, which is determined by the spectrum sensor of the Raman spectrum device. btn is the name of the blood test item, btd is the blood test item data, and lm is the Raman spectrum amplitude.
3. The method according to claim 2, characterized in that The S200 specifically includes the following steps: S210: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is less than 30, use a T-test method to find abnormal Raman spectrum data and eliminate the Raman spectrum data, wherein the elimination method includes deletion or marking; S211: Based on the known characteristic peak value of the substance to be detected, select the Raman shift point where the characteristic peak value is located and one or more Raman shift points before and after the characteristic peak value, and accumulate or calculate the average value as the characteristic peak value of the Raman shift point X. j , as the difference vector of the T test, calculate the overall mean of the characteristic peak of the i-th item according to formula 3.1, calculate the t test value of the characteristic peak according to formula 3.2, calculate the standard deviation of the characteristic peak according to formula 3.3, calculate the p value by looking up the built-in T test table or calling the Python function, and when p < 0.05, determine that the characteristic peak belongs to the abnormal Raman spectrum data, μ is the population mean, t is the deviation statistic between the sample number and the overall mean, K is the total wave number points of each Raman spectrum or the displacement range of the Raman spectrum, s is the standard deviation, To express the difference between the Raman spectrum in item i and the mean at each wave number point, i is the sampling group number, N is the maximum number of the sampling group, 1≤i≤N, j is the number of the Raman spectrum, M is the maximum number of the Raman spectrum, 1≤j≤M, 1≤M≤500, is the overall average of the characteristic peaks.
4. The method according to claim 2, characterized in that The S200 specifically further includes the following steps: S220: In the same sampling group of the original data set, when the maximum number of the collected Raman spectra is 30 or above, using a Z test method to find abnormal Raman spectrum data, and eliminating the Raman spectrum data, the elimination method includes deletion or marking; S221: Based on the known characteristic peak of the substance to be detected, select the Raman shift point where the Raman characteristic peak is located and one or more Raman shift points before and after it, and accumulate or calculate the average value as the characteristic peak value x of the Raman shift point. j , as the difference vector of the Z test, calculate the Z test value of the characteristic peak of the jth item according to formula 4.1, and determine that the characteristic peak belongs to the abnormal Raman spectrum data according to formula 4.2, |z|≥2.584.2 is the overall average of the characteristic peaks, M is the maximum number of the Raman spectrum, 1≤j≤M, 1≤M≤500, s is the standard deviation.
5. The method according to claim 3 or 4, characterized in that The step S300 specifically includes: S310: Decomposing the Raman spectrum array into the baseline, the characteristic peak and the noise, as shown in Formula 5.1 and Formula 5.2, LM_ARRAY i (x)=LM_ARRAY_T i (x)+LM_ARRAY_B i (x)+LM_ARRAY_H i (x),5.1 LM_ARRAY_B i (x)=LM_ARRAY i (x)-LM_ARRAY_T i (x)-LM_ARRAY_H i (x),5.2; S320: Fit the baseline using a polynomial, and solve the coefficients using the least squares method, such as Formula 5.3, and solve the minimized residual square according to Formula 5.
4. LM_ARRAY_B i (x)=a0+a1x+a2x 2 +…+a n x n , 5.3 LM_ARRAY i (x) and x are the Raman spectrum array lm of the i-th sampling group ik Functions and independent variables, LM_ARRAY_B i (x) is the baseline in the Raman spectrum array of the i-th sampling group, LM_ARRAY_T i (x) is the characteristic peak in the Raman spectrum array of the i-th sampling group, LM_ARRAY_H i (x) is the noise in the Raman spectrum array of the i-th sampling group, a is the least squares coefficient, n is the order of the regularized smoothing parameter, ranging from 3 to 5. k is the Raman shift point number, 1≤k≤K, K is the Raman spectrum shift range, which is determined by the spectrum sensor of the Raman spectrum device; S330: For the Raman spectrum amplitudes of all the Raman shift points in all the Raman spectrum arrays, select a regularized smoothing parameter 10 3 to 10 5 Level, the initial weight is 1.0, and the weight is updated from 0.001 to 0.1 until the baseline is straight. The baseline and noise at this point are deducted to obtain the characteristic peak and complete the baseline correction.
6. The method according to claim 5, characterized in that The step S400 specifically includes: S410: Select wavelet basis functions from db4 to db6, Symlets from sym6, decomposition levels from 0 to 7, and soft threshold adjustment factor from 0.5 to 2.0; S420: For the Raman spectrum array after the baseline correction, as shown in Formula 6.1 and Formula 6.2, which define a low-pass filter and a high-pass filter respectively, for the Raman spectrum amplitude lm of the Raman shift point in the Raman spectrum array, execute the steps including the steps of obtaining the low-frequency component by decomposing the Raman spectrum array through the low-pass filter in Formula 6.3 and obtaining the high-frequency component by decomposing the Raman spectrum array through the high-pass filter in Formula 6.
4. Formula 6.5 is the step of reconstructing the Raman spectrum to obtain noise reduction and performing downsampling. cA m [k]=∑ k h[k-2m]·cA m-1 [k],6.3 cont. m [k]=∑ k g[k-2m] cA m-1 [k],6.4 h[m] is the low-pass filter, g[m] is the high-pass filter, To reconstruct the low-pass filter, To reconstruct the high-pass filter, m is the filter decomposition layer number, for db4, the value range is 0-7, i is the sampling group number, k is the Raman shift point number, [k-2m] is the number of the downsampling, LM_ARRAY i is the Raman spectrum amplitude lm in the Raman spectrum array of the i-th sampling group ik function.
7. The method according to claim 6, characterized in that The step S500 specifically includes: S510: Based on one or more test item names, setting the dimensionality reduction number corresponding to the test item name in the Raman spectrum array to be less than 30; S520: selecting, from the Raman spectrum data after executing S400, the Raman spectrum amplitude having a fluctuation amplitude greater than 2% and being the largest at both the preceding and following wavenumber positions as the pre-characteristic peak corresponding to the test item name; S530: For each test item name, sort the pre-characteristic peaks from largest to smallest, and select the pre-characteristic peak with the highest dimensionality reduction value as the characteristic peak, wherein the characteristic peak includes the Raman spectrum amplitude, the Raman shift value, and the test item name; S540: Establishing the pre-training data set using all the characteristic peaks; and / or, S550: Based on the test item name, in the Raman spectrum data after executing S400, select the Raman spectrum amplitude with a fluctuation amplitude greater than 10% and the largest at the front and rear wavenumber positions as the pre-characteristic peak corresponding to the test item name, calculate the number of the pre-characteristic peaks as the dimensionality reduction number, and use all the pre-characteristic peaks as the characteristic peaks corresponding to the test item name, including the Raman spectrum amplitude, Raman shift value and the test item name.
8. The method according to claim 6, characterized in that Also includes S600 steps: S610: Based on one or more test item names, setting the dimensionality reduction number corresponding to the test item name in the Raman spectrum array to be less than 30; S620: Processing the pre-characteristic peak using a wavelet algorithm to obtain the characteristic peak including the Raman spectrum amplitude, the Raman shift value and the test item name.
9. The method according to claim 7, characterized in that Also includes S700, specifically including: S710: using a recursive algorithm to verify and fine-tune the characteristic peak corresponding to the test item name, thereby obtaining an optimized characteristic peak corresponding to the test item name; S720: Verify and fine-tune the characteristic peaks corresponding to all the test item names to obtain optimized characteristic peaks corresponding to all the test item names.
10. The method according to claim 8, characterized in that Also includes S800, specifically including: S810: using a recursive algorithm to verify and fine-tune the characteristic peak corresponding to the test item name, thereby obtaining an optimized characteristic peak corresponding to the test item name; S820: Verify and fine-tune the characteristic peaks corresponding to all the test item names to obtain optimized characteristic peaks corresponding to all the test item names.
Citation Information
Patent Citations
Raman spectral preprocessing method
CN103217409A
Thyroid dysfunction model and establishment method thereof
CN109346156A
Biological sample prediction method and device and storage medium
CN115308189A