A wavelength point screening method based on multi-sensor spectral data
By screening the spectral data of multi-sensor portable near-infrared spectrometers, the problems of large data volume and low analysis efficiency are solved, and the accuracy and stability of the spectral model are improved, making the data suitable for conventional analysis methods.
Patent Information
- Application Number
- CN202210328651.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Due to the increase in the number of sensors, the collected spectral data is large and redundant data is included in the multi-sensor portable near-infrared spectral data analysis method, the analysis efficiency is low, and the model accuracy and stability are not high.
The wavelength point screening method based on multi-sensor spectral data is adopted, and the partial least squares method modeling and root mean square error cross-validation are calculated, and the wavelength point is gradually regressed to screen, reducing the data volume and retaining the sample characterization.
Effectively reduce the amount of spectral data, improve analysis efficiency, improve the accuracy and stability of the spectral model, so that multi-sensor spectral data can be used for conventional spectral analysis methods.
Smart Images

Figure CN114689187B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a near-infrared spectrum analysis technology, and in particular to a wavelength point screening method based on multi-sensor spectrum data. Background Art
[0002] In recent years, near-infrared spectroscopy analysis technology has developed rapidly and has been applied in many fields such as chemical industry, pharmaceutical industry, military industry, food, etc. Near-infrared spectroscopy technology belongs to molecular spectroscopy technology, which can show the composition and properties of substances at the molecular level. It has achieved very high benefits in terms of both economic and social impact and has great development potential.
[0003] However, currently most material composition and property information detection is mainly carried out using large laboratory near-infrared spectroscopic instruments. Although these methods are quantitative, accurate and highly sensitive, the equipment required is bulky and expensive, the sample preparation time is long and the sample preparation method is strict. The detection equipment and sample preparation require professional operation, the detection environment is fixed, and the analysis time is long. They are not suitable for on-site detection and are not easy to promote and use.
[0004] With the development of portable near-infrared spectroscopy technology, the mainstream large-scale near-infrared spectrometers on the market are moving towards small size, low price and portable direction. However, the portable near-infrared spectrometers on the market are mainly single-sensor devices. Single-sensor devices are limited by sensor technology and cover a very limited band range. The collected spectral data has poor stability and is prone to deviation, which can easily lead to unstable prediction results and low accuracy.
[0005] In order to increase the sensor band coverage and improve the application scenarios of portable near-infrared spectrometers, multi-sensor portable near-infrared spectrometers came into being. Although multi-sensor portable near-infrared spectrometers can solve various problems of single-sensor devices, due to the increase in the number of sensors in the multi-sensor portable near-infrared spectrometer, the number of wavelength points of spectral data collected and acquired is increased exponentially compared with the original single-sensor spectral data wavelength points, and the collected data is also relatively redundant, containing too much data information with little correlation with the sample. Not only does it make conventional spectral data analysis methods unsuitable for multi-sensor spectral data analysis, but also due to the increase in the amount of spectral data, the analysis efficiency of portable near-infrared spectral analysis technology is greatly reduced. Summary of the invention
[0006] The technical problem to be solved by the present invention is to propose a wavelength point screening method based on multi-sensor spectral data, the purpose of which is to reduce the amount of spectral data, improve the efficiency of spectral analysis, and improve the accuracy and stability of the spectral model.
[0007] The technical solution adopted by the present invention to solve the above technical problems is:
[0008] A wavelength point screening method based on multi-sensor spectral data comprises the following steps:
[0009] S1, import the original multi-sensor spectral data, classify the spectral data according to the number of sensors, and obtain single-volume data;
[0010] S2. Perform partial least squares modeling on each group of single quantity data to obtain the root mean square error value of each model;
[0011] S3, calculating the characterization coefficient of each group of single quantity data based on the root mean square error value;
[0012] S4, screening the number of wavelength points of each group of single quantity data by combining the number of wavelength points of the single quantity data and the characterization coefficient;
[0013] S5. Filter the wavelength points based on the number of wavelength points in each group of single-volume data, and reorganize the filtered single-volume data to obtain multiple-volume data.
[0014] Furthermore, in step S2, the root mean square error is generated by cross-validation using the leave-one-out method, and the expression is:
[0015]
[0016] Among them, M is the number of original samples, y i For sample x i The calibration value of For sample x i The predicted value of .
[0017] Furthermore, in step S3, the calculation of the characterization coefficient of each group of single quantity data in combination with the root mean square error value specifically includes:
[0018] The smaller the root mean square error RMESCV value of the spectral model, the higher the characterization coefficient is. The calculation method is:
[0019]
[0020] Among them, α n is the characterization coefficient of the nth group of single quantity data, A n The root mean square error RMESCV value of the spectral model established for the nth set of single data.
[0021] Furthermore, in step S4, the wavelength point number screening of each group of single quantity data is performed in combination with the wavelength point number of the single quantity data and the characterization coefficient, specifically including:
[0022] The higher the characterization coefficient corresponding to each group of single quantity data, the stronger the sample characterization ability of the single quantity data is, and the more wavelength points should be retained when screening the number of wavelength points. The calculation method is:
[0023] X i =Round(α i *100%*m)
[0024] Among them, α i is the characterization coefficient of the i-th group of single quantity data, m is the total number of wavelength points that need to be screened out from all single quantity data, Round() is the rounding function, X i is the number of wavelength points that need to be retained in the i-th group of single quantity data.
[0025] Furthermore, in step S5, the wavelength points are screened in combination with the number of wavelength points of each group of single quantity data, specifically including:
[0026] The stepwise regression method was used to verify the wavelength points in each group of single-volume data one by one, and the wavelength points with weak characterization ability were eliminated one by one through the root mean square error value, and the wavelength points with strong characterization ability were retained.
[0027] Furthermore, step S5 specifically includes:
[0028] S51, randomly select X from the i-th group of single quantity data i wavelength points as the initial wavelength points, and the remaining (zX i ) wavelength points are used as iterative wavelength points, i=1; z is the number of wavelength points in the i-th group of single quantity data;
[0029] S52, perform spectral modeling on the initial wavelength point, and calculate the root mean square error RMSECV value T of the spectral model 0 ;
[0030] S53, select one wavelength point from the iterative wavelength points and introduce it into the initial wavelength points, and replace the first to the Xth wavelength points in the initial wavelength points one by one i wavelength points, always keep the number of initial wavelength points as X i indivual;
[0031] S54, perform spectral modeling on the replaced initial wavelength points, and calculate the root mean square error RMSECV value T of these spectral models j (j=1, 2, ... X i );
[0032] S55, Comparison T 0 With T j The size of , and perform corresponding operations:
[0033] If T jBoth are greater than T 0 , it means that after adding the iterative wavelength point, the root mean square error RMSECV value of the spectral model increases, and the corresponding spectral model sample characterization ability becomes weaker, so this iterative wavelength point is discarded, and the initial wavelength point remains unchanged;
[0034] If T j Equal to T 0 , it means that the introduction of the iterative wavelength point has no effect on the spectral model, so this iterative wavelength point is discarded and the initial wavelength point remains unchanged;
[0035] If T j A value in is less than T 0 , it means that after replacing the corresponding initial wavelength point and adding the iterative wavelength point, the root mean square error RMSECV value of the spectral model decreases, and the corresponding spectral model sample characterization ability becomes stronger. In this way, the iterative wavelength point is introduced. At the same time, in order to ensure that the number of initial wavelength points remains unchanged, T j A value in is less than T 0 In this case, the corresponding initial wavelength point is removed;
[0036] If T j There are multiple values less than T 0 , then introduce this iterative wavelength point and remove the one that makes T j When it is the minimum value, it corresponds to the initial wavelength point;
[0037] S56, repeating steps S53-S55, introducing the remaining iterative wavelength points one by one, until all iterative wavelengths are traversed, and the wavelength points finally screened out are the screened wavelength points of the i-th group of single quantity data;
[0038] S57, assign i=i+1, and return to step S51 until the wavelength point screening of all single quantity data is completed;
[0039] S58, reorganize the wavelength points screened out from all single-volume data to obtain multiple-volume data.
[0040] The beneficial effects of the present invention are:
[0041] This method screens the wavelength points of each group of single data through the spectral model parameters of the single data, so that the original multi-sensor spectral data can be simplified to the same number of wavelength points as the single data while retaining the sample representativeness. This not only makes the multi-sensor spectral data applicable to conventional spectral analysis methods, reduces the amount of spectral data, and improves the efficiency of spectral analysis, but also greatly improves the accuracy and stability of the spectral model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1is a flow chart of a wavelength point screening method based on multi-sensor spectral data in an embodiment of the present invention;
[0043] Figure 2 Schematic diagram of the arrangement of the sensors of the portable multi-sensor near-infrared spectrometer in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present invention aims to propose a wavelength point screening method based on multi-sensor spectral data, the purpose is to reduce the amount of spectral data, improve the efficiency of spectral analysis, and improve the accuracy and stability of spectral models. The method first imports the original multi-sensor spectral data, classifies the spectral data according to the number of sensors, obtains single-volume data, and then performs partial least squares modeling on each group of single-volume data, obtains the root mean square error value of each model, calculates the characterization coefficient of each group of single-volume data in combination with the root mean square error value, and then screens the number of wavelength points of each group of single-volume data in combination with the number of wavelength points of the single-volume data and the characterization coefficient, and finally screens the wavelength points in combination with the number of wavelength points of each group of single-volume data, and reorganizes each group of single-volume data after screening to obtain multi-volume data. The method screens the wavelength points of each group of single-volume data by the spectral model parameters of the single-volume data, which can not only reduce the amount of spectral data and improve the efficiency of spectral analysis, but also greatly improve the accuracy and stability of the spectral model.
[0045] Example:
[0046] like Figure 1 As shown, the wavelength point screening method based on multi-sensor spectral data in this embodiment includes the following steps:
[0047] Step 101: importing original multi-sensor spectral data, classifying the spectral data according to the number of sensors, and obtaining single-volume data;
[0048] In this step, the essence of the original multi-sensor spectral data is the collection of each single-sensor spectral data. The number of wavelength data points it covers is the sum of the number of wavelength points of each single sensor. After being quantized, it is decomposed into multiple single-sensor data. Conventional single-sensor spectral analysis methods can be used to model and analyze it, effectively improving the analysis efficiency.
[0049] As a specific example, Figure 2As shown in the figure, the portable multi-sensor near-infrared spectrometer used in this example contains 4 sensors, which are 4 different types of sensors. The band ranges of the 4 sensors are 1255nm~1450nm, 1455nm~1650nm, 1655nm~1850nm, 1855nm~2050nm, the resolution is 5nm, and the distribution of wavelength points is wavelength averaging. The band range of sensor 1 is 1255nm~1450nm, and the number of wavelength points included is M 1 =1+(1450-1255) / 5=40; the wavelength range of sensor 2 is 1455nm~1650nm, and the number of wavelength points included is M 2 =1+(1650-1455) / 5=40; the wavelength range of sensor 3 is 1655nm~1850nm, and the number of wavelength points included is M 3 =1+(1850-1655) / 5=40; the wavelength range of sensor 4 is 1855nm~2050nm, and the number of wavelength points included is M 4 =1+(2050-1855) / 5=40.
[0050] From the above, we can see that the original spectral data of the multi-sensor is composed of 4 segments of spectral data, with a specific wavelength range of 1255nm to 2050nm, including M 5 =M 1 +M 2 +M 3 +M 4 = 160 wavelength points. The multi-sensor spectral data are classified according to the number of sensors to obtain 4 groups of single-volume data, which are set as single-volume data P 1 , P 2 , P 3 , P 4 , corresponding to the spectral data of sensors 1, 2, 3, and 4.
[0051] Step 102: Perform partial least squares modeling on each group of single quantity data to obtain the root mean square error value of each model;
[0052] In this step, the partial least squares method is the most commonly used and effective linear fitting modeling method in spectral analysis; the root mean square error is the most commonly used model indicator in the spectral modeling and analysis process, which can directly reflect the quality of the spectral model.
[0053] by Figure 2Taking the example of as an example, the single data of sensor 1, sensor 2, sensor 3, and sensor 4 are modeled by partial least squares method, and the root mean square error values of the four models are further calculated. The root mean square error (RMESCV) actually reflects the deviation relationship between the predicted calibration value and the predicted value in the model. It is generated by cross-validation using the leave-one-out method. The expression is as follows:
[0054]
[0055] Assuming that there are M original spectral data, there are M corresponding original samples. Each time a sample X is taken from them i (i-th sample, i=1, 2, ..., M), use the remaining (M-1) spectral sample data to build a model, and use the built model to predict the sample X i of y i Value (calibration value), get the predicted value of the sample Then, the root mean square error of all the predicted values is calculated to obtain the RMSECV value. Since the leave-one-out method traverses all samples of the original spectral data, the RMSECV indicator has a higher accuracy in judging the quality of the model than the traditional MSE and MAE indicators, and is more stable and applicable.
[0056] Step 103: Calculate the characterization coefficient of each group of single quantity data in combination with the root mean square error value;
[0057] In this step, since the spectral data of different sensors cover different band ranges, their ability to characterize samples will also vary. If each group of single data is simply screened for the same number of wavelength points, it is easy to cause weak sample characterization capabilities and thus affect the spectral prediction and analysis capabilities. The model effect of each group of single data is judged by the root mean square error value, which further reflects the characterization ability of each group of single data for the sample, and then the characterization coefficient is calculated to perform corresponding processing on each group of single data, which can effectively solve this problem.
[0058] In this embodiment, the single quantity data P is set 1 The characterization coefficient is α 1 The RMESCV value of the spectral model established by spectral data is A 1 , single quantity data P 2 The characterization coefficient is α 2 The RMESCV value of the spectral model established by spectral data is A 2 , single quantity data P 3 The characterization coefficient is α 3 The RMESCV value of the spectral model established by spectral data is A 3 , single quantity data P 4 The characterization coefficient is α 4The RMESCV value of the spectral model established by spectral data is A 4 , because in the ability of the spectral model to characterize the sample, the smaller the RMESCV value is, the smaller the deviation between the calibration value and the predicted value is predicted in the spectral model, that is, the better the characterization ability of the spectral model is. It can be further known that the smaller the RMESCV value of the spectral model is, the higher the characterization coefficient is, and the specific calculation formula for obtaining the characterization coefficient is:
[0059]
[0060] Step 104: screening the number of wavelength points of each group of single quantity data by combining the number of wavelength points of the single quantity data and the characterization coefficient;
[0061] In this step, the higher the characterization coefficient corresponding to each group of single quantity data, the stronger the sample characterization ability of the single quantity data, and the more wavelength points should be retained when screening the number of wavelength points. By calculating the number of wavelength points that should be retained for each group of single quantity data through the characterization coefficient and the number of wavelength points, the spectral sample characterization can be retained to the greatest extent.
[0062] In this embodiment, each group of single quantity data contains 40 wavelength points, that is, each single sensor spectral data contains 40 wavelength points, and the multi-sensor spectral data contains 160 wavelength points. In order to simplify the multi-sensor spectral data and make it possible to use the conventional single sensor spectral analysis method, it is necessary to filter the wavelength points of the multi-sensor spectral data into 40. Due to the different sample characterization capabilities of the spectral data of each sensor, the specific number of wavelength points that each group of single quantity data should contribute can be calculated in combination with the characterization coefficient. Suppose the single quantity data P 1 , P 2 , P 3 , P 4 The specific number of wavelength points that should be contributed is X 1 , X 2 , X 3 , X 4 , combined with the characterization coefficient, the specific number of wavelength points is:
[0063]
[0064] Among them, Round() is a rounding function. Usually (X 1 +X 2 +X 3 +X 4 )=40, if (X 1 +X 2 +X 3 +X 4 )=41 or 42, in X 1 , X 2 , X3 , X 4 One or two wavelength points are randomly selected and discarded, so that (X 1 +X 2 +X 3 +X 4 )=40 is always true.
[0065] Step 105: screening the wavelength points based on the number of wavelength points of each group of single quantity data, and reorganizing the single quantity data of each group after screening to obtain multiple quantity data;
[0066] In this step, after determining the number of wavelength points in each group of single quantity data, it is necessary to further determine the specific wavelength points in each group of single quantity data, and use the stepwise regression method to verify the wavelength points in each group of single quantity data one by one, and eliminate the wavelength points with weak characterization ability one by one through the root mean square error value, and retain the wavelength points with strong characterization ability.
[0067] In this embodiment, the single amount data P 1 For example, the number of wavelength points after screening is X 1 Then in the single quantity data P 1 Randomly select X 1 wavelength points as the initial wavelength points, and the remaining (40-X 1 ) wavelength points are used as iterative wavelength points, and the specific stepwise regression steps are as follows:
[0068] (1) Spectral modeling is performed on the initial wavelength point, and the RMSECV value T of the spectral model is calculated. 0 ;
[0069] (2) Select one wavelength point from the iterative wavelength points and introduce it into the initial wavelength point, and replace the 1st to the xth wavelength points in the initial wavelength point one by one 1 wavelength points, always keep the number of initial wavelength points as X 1 indivual;
[0070] (3) Perform spectral modeling on the replaced initial wavelength points and calculate the RMSECV values of these spectral models. j (j=1, 2, ... X 1 );
[0071] (4) Compare T 0 With T j The size of T j Greater than T 0 , it means that the RMSECV value of the spectral model increases after adding the iterative wavelength point, and the corresponding spectral model sample characterization ability becomes weaker, so this iterative wavelength point is discarded; if T j Always equal to T 0, it means that the introduction of the iterative wavelength point has no effect on the spectral model, so this iterative wavelength point is discarded; if T j A value in is less than T 0 , it means that after removing this initial wavelength point and adding the iterative wavelength point, the RMSECV value of the spectral model decreases, and the corresponding spectral model sample characterization ability becomes stronger. This iterative wavelength point is introduced, and in order to ensure that the number of initial wavelength points remains unchanged, T j A value in is less than T 0 In this case, the corresponding initial wavelength point is removed; if T j There are multiple values less than T 0 , this iterative wavelength point is introduced, when T j When it is the minimum value, the corresponding initial wavelength point is eliminated;
[0072] (5) Repeat steps (2), (3), and (4) to introduce the remaining iterative wavelength points one by one until all iterative wavelength points have completed the introduction and comparison operation. The spectral wavelength points selected after stepwise regression are the single quantity data P 1 The screening wavelength point.
[0073] Similarly, for single data P 2 , P 3 , P 4 The same stepwise regression process is performed to obtain the corresponding screening wavelength points, and each group of single-volume data after the screening is completed is reorganized to obtain multiple-volume data, which is the multi-sensor spectral data after the screening is completed.
[0074] Finally, it should be noted that the above embodiments are only preferred implementations and are not intended to limit the present invention. It should be pointed out that for those skilled in the art, several modifications, equivalent replacements, improvements, etc. can be made without departing from the scope of the present invention and the scope of protection of the claims, and all of these should be included in the protection scope of the present invention.
Claims
1. A wavelength point screening method based on multi-sensor spectral data, characterized in that: The following steps are involved: S1, import the original multi-sensor spectral data, classify the spectral data according to the number of sensors, and obtain single-volume data; S2. Perform partial least squares modeling on each group of single quantity data to obtain the root mean square error value of each model; S3. Calculate the characterization coefficient of each group of single quantity data in combination with the root mean square error value, including: The smaller the root mean square error RMESCV value of the spectral model, the higher the characterization coefficient is. The calculation method is: Among them, α1, α2, α3, and α4 are the characterization coefficients of the 1st, 2nd, 3rd, and 4th groups of single-volume data, respectively; A1, A2, A3, and A4 are the root mean square error RMESCV values of the spectral model established by the 1st, 2nd, 3rd, and 4th groups of single-volume data, respectively; S4. The number of wavelength points of each group of single quantity data is screened by combining the number of wavelength points of the single quantity data and the characterization coefficient, including: The higher the characterization coefficient corresponding to each group of single quantity data, the stronger the sample characterization ability of the single quantity data is, and the more wavelength points should be retained when screening the number of wavelength points. The calculation method is: X i =Round(α i *100%*m) Among them, α i is the characterization coefficient of the i-th group of single quantity data, m is the total number of wavelength points that need to be screened out from all single quantity data, Round() is the rounding function, X i is the number of wavelength points that need to be retained in the i-th group of single quantity data; S5. Filter the wavelength points based on the number of wavelength points in each group of single-volume data, and reorganize the filtered single-volume data to obtain multiple-volume data.
2. A wavelength point screening method based on multi-sensor spectral data as claimed in claim 1, characterized in that: In step S2, the root mean square error is generated by cross-validation using the leave-one-out method, and the expression is: Among them, M is the number of original samples, y i For sample x i The calibration value of For sample x i The predicted value of .
3. The wavelength point screening method based on multi-sensor spectral data according to claim 1, characterized in that: In step S5, the wavelength points are screened in combination with the number of wavelength points of each group of single quantity data, specifically including: The stepwise regression method was used to verify the wavelength points in each group of single-volume data one by one, and the wavelength points with weak characterization ability were eliminated one by one through the root mean square error value, and the wavelength points with strong characterization ability were retained.
4. A wavelength point screening method based on multi-sensor spectral data as claimed in claim 3, characterized in that: Step S5 specifically includes: S51, randomly select X from the i-th group of single quantity data i wavelength points as the initial wavelength points, and the remaining (zX i ) wavelength points are used as iterative wavelength points, i=1; z is the number of wavelength points in the i-th group of single quantity data; S52, performing spectral modeling on the initial wavelength point, and calculating a root mean square error RMSECV value T0 of the spectral model; S53, select one wavelength point from the iterative wavelength points and introduce it into the initial wavelength points, and replace the first to the Xth wavelength points in the initial wavelength points one by one i wavelength points, always keep the number of initial wavelength points as X i indivual; S54, perform spectral modeling on the replaced initial wavelength points, and calculate the root mean square error RMSECV value T of these spectral models j (j=1, 2, ... X i ); S55, compare T0 and T j , and perform corresponding operations: If T j If both are greater than T0, it means that the root mean square error RMSECV value of the spectral model increases after adding the iterative wavelength point, and the corresponding spectral model sample characterization ability becomes weaker, then this iterative wavelength point is discarded, and the initial wavelength point remains unchanged; If T j If they are all equal to T0, it means that the introduction of the iterative wavelength point has no effect on the spectral model, so the iterative wavelength point is discarded and the initial wavelength point remains unchanged; If T j If a value in is less than T0, it means that after replacing the corresponding initial wavelength point and adding the iterative wavelength point, the root mean square error RMSECV value of the spectral model decreases, and the corresponding spectral model sample characterization ability becomes stronger. In this case, the iterative wavelength point is introduced. At the same time, in order to ensure that the number of initial wavelength points remains unchanged, T j When a certain value in is less than T0, the corresponding initial wavelength point is removed; If T j If there are multiple values less than T0, then this iterative wavelength point is introduced and the value that makes T j When it is the minimum value, it corresponds to the initial wavelength point; S56, repeating steps S53-S55, introducing the remaining iterative wavelength points one by one, until all iterative wavelengths are traversed, and the wavelength points finally screened out are the screened wavelength points of the i-th group of single quantity data; S57, assign i=i+1, and return to step S51 until the wavelength point screening of all single quantity data is completed; S58, reorganize the wavelength points screened out from all single-volume data to obtain multiple-volume data.
Citation Information
Patent Citations
Method for establishing multiple models of near infrared spectrums
CN103528990A
Near infrared spectrum noninvasive blood glucose detecting method and detecting network model training method thereof
CN107192690A