Direct standardization method for segmenting infrared spectra of animal milk after agglomerative clustering between different brands of instruments

The ACPDS method solves the problem of spectral differences between infrared spectrometers from different brands, achieves data standardization and model calibration across brands, and improves prediction accuracy and repeatability.

CN118782187BActive Publication Date: 2026-08-25HUAZHONG AGRI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410918840.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-08-25
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

Differences in spectral equipment and configurations among infrared spectrometers from different brands lead to incompatibility in chemometric methods. Existing calibration methods are either costly or have significant limitations, making it difficult to achieve model calibration and transfer between instruments from different brands.

Method used

The Agglomerative Clustering and Segmented Direct Standardization (ACPDS) method was adopted. By flexibly selecting the response window for each spectral point and combining clustering algorithms and multivariate regression analysis, a spectral standardization model for cross-brand instruments was established.

Benefits of technology

It improves the repeatability and accuracy of model prediction results across different brands of instruments, reduces root mean square error, and achieves data standardization and unification across brands of instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present application belongs to the field of engineering technology, cow performance measurement and milk detection, and particularly relates to a method for segmenting and directly standardizing animal milk by using different brand instruments after condensation clustering of mid-infrared spectra. The method provided by the present application is more flexible in the selection of windows for the response of each spectral wave point between different instruments, and is superior to the single and limited fixed window selection of the past method. The model prediction result repeatability level based on the mid-infrared spectrum between different instruments is superior to the present technology, the determination coefficient (R 2 ) after standardization is higher, and the root mean square error (RMSE) is smaller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of engineering technology, dairy cow performance testing and milk detection, and specifically relates to a method for direct standardization of segments after aggregation and clustering among different brands of infrared spectrometers in animal milk. Background Technology

[0002] Fourier Transform Mid-Infrared Spectroscopy (FT-MIRS) identifies specific chemical bonds in milk samples by scanning them, ultimately generating a spectral curve composed of individual infrared wavenumber transmittance values. FT-MIRS wavenumbers, combined with functions and hyperparameters generated by machine learning algorithms, predict milk traits, enabling rapid, non-destructive, and high-throughput phenotypic analysis of milk, such as protein composition, fatty acid composition, lactose, and minerals. Furthermore, bovine physiological traits are influenced by various genetic and environmental factors; studies have found that FT-MIRS can also predict bovine pregnancy diagnosis, energy status, nitrogen output, and methane emissions. FT-MIRS is now incorporated into Dairy Herd Improvement (DHI) programs in several countries, providing important insights for improving dairy cow performance.

[0003] Due to differences in light sources or optoelectronic components, FT-MIRS applications suffer from severe limitations. Currently, the three most commonly used FT-MIRS brands internationally are Foss, Bentley, and Perkin Elmer. Given the diversity of spectroscopic equipment and configurations, these three instruments exhibit significant differences in spectral range and resolution. This means that chemometric methods can typically only be matched to spectra measured by the same instrument, rendering the established models ineffective for data collected from instruments of different brands. Therefore, spectral calibration between different brands of FT-MIRS instruments is necessary for model calibration and transfer. Calibration methods are mainly divided into three categories: transferring and calibrating models through the standardization of predicted values, model coefficients, and spectral responses. Pre-calibration is primarily used in the early stages of spectral analysis. Lynch et al. used chemical analysis and coefficient evaluation of reference samples, employing these as standardized coefficients to adjust the phenotypic predictions of the reference samples. However, this method is only applicable to reference samples with existing phenotypes. Model coefficient calibration requires establishing prediction models separately for different brands of instruments to obtain the corresponding instrumental conversion coefficients. While this method offers high accuracy, it is extremely costly. Therefore, the standardization of spectral responses between instruments is receiving increasing attention.

[0004] Current methods standardize the wavenumber absorbance values ​​of slave instruments using the master instrument to match the spectral response of the master instrument. Grelet et al. proposed a Piecewise Direct Standardization (PDS) method based on Wang's description, while Bonfatti et al. proposed a traced percentile standardization method, mapping the absorbance percentile of each wavenumber of the master instrument to the corresponding percentile of the slave instrument, and using a linear regression formula to calculate and plot the standardization coefficients between the instruments. KMTiplady et al. evaluated and compared these two spectroscopic methods, showing that using PDS for standardization is the best method because it is insensitive to changes in milk composition and instrument errors. Furthermore, Grelet et al. used the results of a milk fat content prediction model to evaluate the effectiveness of spectral PDS, finding that the prediction bias and root mean square error between master and slave instruments were significantly reduced after standardization compared to the original spectra. Grelet et al. also tested the prediction model results for methane emissions, fatty acid content, and cheese yield before and after standardization, and calculated that the Mahalanobis distance from the slave instrument spectrum to the master instrument was significantly reduced after standardization. KMTiplady et al. compared the root mean square error of predictions for three common milk components after using PDS and RPS methods, further demonstrating the effectiveness of standardization and the superiority of the PDS method.

[0005] The study by Toledo-Alvarado et al. (2018) showed that using a single wavenumber to associate with a related phenotype may be affected by external factors such as management strategies and husbandry environment. Therefore, the core idea of ​​PDS is based on the fact that the spectral responses of adjacent wavenumbers in FT-MIRS are usually highly correlated and confined to a small region, which is low-dimensional compared to single nucleotide polymorphism (SNP) markers in the genome. However, PDS generally uses a fixed number of wavenumbers as a window for spectral response, resulting in a lack of diversity in window selection.

[0006] Therefore, this invention proposes a band segmentation method based on Agglomerative Clustering (AC). Before PDS, a flexible window selection is performed on the spectral response of each wavenumber across different brands of instruments based on the samples, defined as ACPDS. Simultaneously, the optimized method is compared with Single Wave Point Direct Standardization (SDS), PDS, CCA, and SBC, providing a reference method for further improving the standardization level of FT-MIRS across different brands of instruments. This opens up new avenues for the widespread application of FT-MIRS-based prediction models and provides a new perspective for the establishment and optimization of dairy cow performance phenotypic databases in China and even internationally, holding significant value in dairy herd improvement (DHI). Summary of the Invention

[0007] This invention provides a method for direct standardization of infrared spectra in animal milk after clustering and segmentation from different brands of instruments. It can standardize infrared data measured by Bentley FTS NexGen-600 and / or Perkin Elmer Lactoscope-300 instruments into Foss FT+ data, and can directly output data using the prediction model of Foss FT+.

[0008] To achieve the above objectives, the present invention adopts the following technical measures:

[0009] A method for direct standardization of segments after aggregation and clustering among different brands of infrared spectrometers for animal milk includes the following steps:

[0010] Step 1: Collect infrared spectral data of milk measured on two instruments: Bentley FTS NexGen-600 and / or Perkin Elmer Lactoscope-300. Use the 956.78 cm⁻¹ value. -1 -1597.21cm -1 2824.06cm -1 -2974.52cm -1 The data consists of 205 wave points;

[0011] Step 2: Perform PDS normalization on the spectral data of the 205 wave points output by Bentley FTS NexGen-600 according to the window shown in the table below;

[0012]

[0013]

[0014] The spectral data from the 205 spectral points output by the Perkin Elmer Lactoscope 300 were normalized using the window method described in the table below;

[0015]

[0016]

[0017] Step 3:

[0018] By substituting the standardized MIR data into the prediction model prepared using the Foss FT+ instrument, the prediction results for the target object can be obtained.

[0019] The animal milk mentioned in the above-described method is cow's milk, buffalo milk, or camel milk.

[0020] Preferably, the prediction model described above is:

[0021] The predictive model for A2 type β-casein in milk is: diff1+PLSR(n_component=6);

[0022] The predictive model for α-lactalbumin in milk is: none + PLS-DA (n_component = 31);

[0023] The predictive model for αs1-casein in milk is: SG1+CARS+SVR;

[0024] The predictive model for κ-casein in milk is: SG(w=21,p=3)+PLS-DA(n_component=38);

[0025] The predictive model for total casein in milk is: SG (w=21, p=3) + standardization + ridge regression (α=0.2245);

[0026] The predictive model for protein content in buffalo milk is: SG(w=7,p=4)+Ridge(α=0.12);

[0027] The predictive model for β-casein content in milk is: none (no pretreatment) + PLS-DA (n_component = 33);

[0028] The predictive model for lactoferrin content in milk is: SNV+PLS-DA(n_component=16);

[0029] The predictive model for the content of free amino acids in milk is: none + PLSR (n_component = 24);

[0030] The predictive model for the content of free essential amino acids in milk is: none + PLSR (n_component = 24);

[0031] The predictive model for the content of free methionine in milk is: SG(w=5,p=2)+PLSR(n_component=24);

[0032] The predictive model for the content of free lysine in milk is: SG(w=5,p=3)+PLSR(n_component=24);

[0033] The predictive model for the content of free taurine in milk is: SG(w=5,p=2)+PLSR(n_component=29);

[0034] The predictive model for the content of free valine in milk is: SG(w=15,p=3)+PLSR(n_component=35);

[0035] The predictive model for the content of free isoleucine in milk is: SG(w=9,p=4)+PLSR(n_component=29);

[0036] The predictive model for furosine content in buffalo milk is: none + PLSR (n_component = 25);

[0037] The predictive model for fat content in buffalo milk is: none + Ridge (α = 1.0);

[0038] The prediction model for total solids content in buffalo milk is: none + PLSR (n_component = 19);

[0039] The prediction model for magnesium content in camel milk is: D2 + SG(W = 13, P = 3) + PLSR(n_component = 13);

[0040] The predictive model for calcium content in camel milk is: D2 + SG(W = 13, P = 3) + PLSR(n_component = 6);

[0041] The predictive model for the protein content in camel milk is: D2+D2+PLSR(n_component=6);

[0042] The predictive model for milk fat content in camel milk is: D2+D2+PLSR(n_component=7);

[0043] The predictive model for immunoglobulin G content in bovine colostrum is: standardized + 1D + PLSR (n_component = 10);

[0044] The predictive model for the calcium content in milk is: D2+MMS+PLSR(n_component=9).

[0045] Compared with the prior art, the advantages of this invention are:

[0046] 1. The window selection for the response of each spectral point is more flexible across different instruments, which is superior to the single and limited fixed window selection of the past methods.

[0047] 2. The repeatability of prediction results based on mid-infrared spectroscopy across different instruments is superior to current technology, with a standardized coefficient of determination (R²). 2 It is higher and the root mean square error (RMSE) is smaller. Attached Figure Description

[0048] Figure 1 To address the spectral differences, band unification, and modeling band selection for the three brands of instruments;

[0049] (a) shows the original spectral absorbance of the three instruments. (b) shows the spectrum of the three instruments after selecting the common band and after sample interpolation, which includes two water wave absorption regions. (c) and (d) show the selected characteristic bands used for the prediction model establishment of the main instrument.

[0050] Figure 2 Six calibration transfer methods for subordinate instruments;

[0051] (a) Figure is a schematic diagram of PDS, (b) Figure is a schematic diagram of SWS, (c) Figure is a schematic diagram of ACPDS, (d) Figure is a schematic diagram of CCA, and (e) Figure is a schematic diagram of SBC. Detailed Implementation

[0052] Unless otherwise specified, the technical solutions described in this invention are all conventional solutions in the field; unless otherwise specified, the reagents or materials described are all from commercial sources.

[0053] 1. Data Acquisition and Measurement

[0054] This study analyzed the total solids content in milk as an example. From March to May 2023, milk samples from Holstein cows in four provinces across three regions of China (North China, Central China, and Northwest China) were collected and divided into two identical portions. Under refrigeration at 4°C for 24 hours, one portion was sent to the nearest dairy cow performance testing (DHI) laboratory for FT-MIRS data acquisition using three different brands of spectrometers (Foss FT+, Bentley FTS NexGen-600, and Perkin Elmer Lactoscope-300). The other portion was sent to a chemical laboratory for determination of the total solids content using national standard testing methods. The valid data structure is shown in Table 1, with the measurement unit being a percentage of milk concentration.

[0055] In addition, following the method proposed in the International Dairy Federation's Guide to the Application of Mid-Infrared Spectroscopy (ISO 9622: IDF 141), 12 milk samples with orthogonal gradient variations were prepared in the Chinese Standard Sample Laboratory, and FT-MIRS was measured on three different brands of instruments to cover as much of the variable value region of each wavenumber as possible, in order to study the standardization of the spectra and predicted values ​​of the three brands of instruments.

[0056] Table 1. Composition and statistical description of the research dataset

[0057]

[0058]

[0059]

[0060] 2.2 Standardization Methods:

[0061] 2.2.1 Segmented Direct Standardization (PDS)

[0062] like Figure 2 As shown in Figure a, the PDS method is a spectral transformation after establishing a multivariate model with spectral variables limited to a small region. In PDS, the spectral intensity response r measured by the master instrument at wavelength j is related to the wavelengths within a small window surrounding j measured by the slave instrument.

[0063] X m =X s b j

[0064] In the formula, X s =[X s-i ,..,X s ,…,X s+i [ ] represents the small window response matrix of the subordinate instrument samples. j Let F be the transform coefficient vector for the j-th wavelength. The regression vector can be calculated using PCR or PLSR, and the regression coefficients can be combined into a strip diagonal transform matrix F.

[0065] F = diag(b1) T b2 T ,…b j T ,…,b k T )

[0066] k is the number of wavelengths. Predictions can then be made by multiplying these transmitted spectra by regression coefficients established on the main instrument.

[0067] However, this method has certain limitations. For example, if the critical point in the spectral normalization band is less than half the number of spectral windows, it cannot be accurately normalized using this method (e.g., if the number of windows is 8, the critical points 1, 2, and 3, and the last 1, 2, and 3 cannot be accurately normalized).

[0068] 2.2.2 Single Wavenumber Normalization (SWS)

[0069] like Figure 2 As shown in b, in SWS, each wavelength is modeled separately from the others, allowing for different intensity variations within the wavelength range. Then, the conversion coefficients b for each wavenumber are used. i Combined into a transformation matrix F = [b1, b2, ..., b n ].

[0070] X m(i) =b i *X s(i)

[0071] X tran =F*X s,new

[0072] 2.2.3 Agglomerative Clustering and Segmented Direct Standardization (ACPDS)

[0073] like Figure 2 As shown in Figure c, this method is an improvement on the PDS method. The response measured by wavenumber on the master instrument is related to the response of a small window near the same wavenumber measured on each slave instrument, and is usually highly correlated with the absorbance value of the wavenumber. This invention uses agglomerative clustering to cluster highly correlated and adjacent wavenumbers, creating a window of highly correlated wavenumbers. After determining the window, each wavepoint is PDS standardized using all wavepoints in the corresponding window, no longer restricted by the zero point. Specifically, it is implemented using the AgglomerativeClustering package in the sklearn library of Python 3.10. This study integrates the window selection and standardization calculation process and defines it as Agglomerative Clustering and Piecewise Direct Standardization (ACPDS). Agglomerative clustering iteratively merges the most similar clusters based on the distance function D calculated according to the linking criterion. Therefore, this method specifically adopts the widely used Ward association linking on the basis of PDS. It defines clustering objects according to Euclidean distance and can intuitively express the variance within / between classes. D(C m C n )=ED(C m ∪C n )-ED(C m )-ED(C n ), xi is the response vector.

[0074] 2.2.4 Canonical Correlation Analysis (CCA)

[0075] like Figure 2 As shown in Figure d, the standardized set X of the same sample is measured using master and slave instruments. m and X s After centering the mean of the data, X m and X s Canonical correlation analysis yields the regularized vector W. m and W s .

[0076]

[0077]

[0078]

[0079] From the constraints of the Lagrangian function calculation, we know that p = λ = v, where p is the correlation coefficient between the two datasets and also a characteristic value of canonical correlation. m and W s These are canonical eigenvectors. For X m and X s Canonical correlation analysis can yield W m and W s Then, using the following equation, we can obtain the typical variable L. m and L s .

[0080] L m =X m ×W m

[0081] L s =X s ×W s

[0082] The transformation matrix F1 is obtained by PLSR calculation:

[0083] L m =L s ×F1

[0084] X trans =X s,new ×W s ×F1×W m

[0085] 2.2.5 Slope and Bias Correction (SBC)

[0086] like Figure 2 As shown in e, X m and X s The given spectra are subsets of samples collected by two instruments (master and slave). Where y m To simultaneously use the concentration values ​​of a subset of samples determined by the spectrum and calibration model on the main instrument, y s To determine the concentration values ​​of the same subset of samples using the spectra obtained from the instruments, without considering the different instrument responses of the two instruments.

[0087] y m The predicted value and y sThe predicted values ​​correspond to the calculated values, and a fitted univariate linear model is calculated. Under the assumption of equal variances between master and slave instrument measurements, the bias and slope of this linear model are given by the formula.

[0088]

[0089]

[0090]

[0091]

[0092]

[0093] The combination of the regression coefficient b and the slope / bias correction term allows us to directly calculate the corrected concentration value y of the spectrum collected from the slave instrument. tran .

[0094] y tran =slope*(X s,new *b)+bias

[0095] 2.3 Construction Environment and Assessment Standards

[0096] Based on the sample size and spectral variables, a Linux system with an Intel(R) Core(TM) i9-10850K 3.60GHz processor and 32GB of memory was selected for computation. The model was developed using the sklearn package in Python version 3.10.

[0097] Model evaluation criteria include the training set determination coefficient (Training R-value). 2 ), Training Root Mean Square Error (RMSE), Test Set Coefficient of Decision (RTD) 2 The calibration transfer method is determined by the root mean square error (RMSE) of the test set and the relative prediction error ratio (RPD) of the test set. The criteria for determining the calibration transfer method are the coefficient of determination and the root mean square error of the test data of the two dependent instruments. The specific formula is shown below:

[0098]

[0099]

[0100]

[0101] Where y n and The reference and predicted values ​​represent total solids. It is the average of the y-values, N represents the sample size, and STDEV The standard deviation of the sample.

[0102] Example 1:

[0103] Selection of infrared spectral bands for milk in the main instrument (Foss FT+):

[0104] Of the total solids sample dataset from the main instrument (Foss FT+) (i.e., the 201 milk samples in Table 1), 75% was used for model training and cross-validation, and 25% was used for model testing. Based on previous research, the coefficient of variation and variance in the water absorption noise region were both large, making it difficult to extract sufficient effective information; therefore, it was removed and not used in model building. In other absorption regions, 956.78 cm⁻¹... -1 -1597.21cm -1 This is the most commonly used fingerprint region for modeling, containing the main characteristic bands of functional groups of milk proteins, milk fats, and milk solids, at 2824.06 cm⁻¹. -1 -2974.52cm -1 These are the main absorption regions for milk fat and milk solids, so we selected 205 spectral points in these two segments to establish the main instrument model.

[0105] In this embodiment, the applicant uses the previously constructed total solids prediction model (CN114166786B) as an example for prediction. The prediction model is: none + PLSR (n_component = 19), and 956.78 cm is selected. -1 -1597.21cm -1 and 2824.06cm -1 -2974.52cm -1 The MIR data in the band are shown in Table 3.

[0106] like Figure 1 As shown in c, no feature preprocessing was performed on the spectral variables before training.

[0107] Table 3. Optimal Model Performance of Total Solids Model Based on PLSR

[0108]

[0109] The table above shows that the RPD of the test set is greater than 2, which has reached the application level and can be used for standardized performance verification.

[0110] Example 2:

[0111] Selection of PDS normalization method and clustering window distribution:

[0112] like Figure 2As shown, the applicant performed data segmentation and calibration transfer analysis on the data measured by the master instrument and slave instruments. Specifically, spectral data of 12 standard milk samples were measured on both the master and slave instruments. Four standardization methods—PDS, SWS, ACPDS, and CCA—were used to calculate the spectral response between the master and slave instruments and obtain the spectral transformation matrix. For PDS, the optimal window width was selected between 1 and 20; the regression vector was calculated using PLSR (with latent variables optimally selected between 1 and 3); and for ACPDS, the optimal clustering threshold was selected between 0 and 2. Simultaneously, a model constructed using the master instrument was used to predict the spectra of the 12 standard milk samples, and the SBC standardization method was used to calculate the transformation matrix of the predicted values ​​from the master and slave instruments.

[0113] Tables 4 and 5 show the window distribution of the 205 wave points modeled by the main instrument for Bentley FTS NexGen-600 and Perkin Elmer Lactoscope 300 after agglomeration clustering. The clustering result for Bentley FTS NexGen-600 is 36 windows, and the clustering result for Perkin Elmer Lactoscope 300 is 9 windows. The window size varies from 1 to 166 wave points, which further proves that there are significant differences in the correlation between different spectral absorption regions and different wave point ranges. Therefore, the selection of windows should be more flexible.

[0114] Table 4 shows the 36 windows (numbered from 1 to 36) resulting from clustering 205 points from the master instrument (Foss FT+) to the slave instrument (Bentley FTS NexGen-600).

[0115]

[0116]

[0117] Notes: Window: window; Upper: upper critical point of the window; Lower: lower critical point of the window; NW: number of points. Table 5 shows the 36 windows (numbered from 1 to 36) resulting from clustering 205 points from the master instrument (Foss FT+) for the slave instrument (Perkin Elmer Lactoscope 300).

[0118]

[0119] Notes: Window: window; Upper: upper threshold of window; Lower: lower threshold of window; NW: number of wave points. Example 3:

[0120] Standardization method usage process and validation results:

[0121] Step 1: Collect infrared spectral data of milk measured using two instruments: a Bentley FTS NexGen-600 (185 milk samples in Table 1) and a Perkin Elmer Lactoscope-300 (189 milk samples in Table 1). The 956.78 cm⁻¹ value was used. -1 -1597.21cm -1 2824.06cm -1 -2974.52cm -1 The data consists of 205 wave points;

[0122] Step Two:

[0123] The 205 wave point data output by the Bentley FTS NexGen-600 were PDS normalized according to the window described in Table 4.

[0124] The spectral data of 205 spectral points output by the Perkin Elmer Lactoscope 300 were normalized to PDS according to the window described in Table 5;

[0125] Control group:

[0126] The 205 wave point data obtained in the first step were substituted into PDS (when using data from Bentley FTS NexGen-600, select window 10; when using data from Perkin Elmer Lactoscope-300, select window 12), SWS, CCA, and SBC for normal standardization.

[0127] Step 3:

[0128] The standardized dot data were substituted into the milk total solids prediction model none+PLSR(n_component=19) to make predictions, obtaining the predicted value of total solids. This predicted value was then compared with the actual value measured by the national standard to calculate R. 2 The final results, compared with R MSE, are shown in Table 6. The validation results show that ACPDS achieves R MSE performance on both types of slave instruments. 2 The ACPDS method outperforms the other four standardization methods in both RMSE and RMSE, demonstrating the superiority of the ACPDS method and the rationality of the window selection.

[0129] Table 6 shows the verification results of the two instruments based on five standardized methods.

[0130]

[0131] Note: "-" represents the R-value of the calibration transfer method performance. 2 ≤0.

Claims

1. A method for direct standardization of segments after aggregation and clustering among different brands of infrared spectrometers for animal milk, comprising the following steps: Step 1: Collect infrared spectral data of milk measured on a Bentley FTS NexGen-600 and / or Perkin Elmer Lactoscope-300 instrument, using the 956.78 cm⁻¹ value. -1 -1597.21cm -1 2824.06cm -1 -2974.52cm -1 The data consists of 205 wave points; Step 2: Perform PDS normalization on the spectral data of the 205 wave points output by Bentley FTS NexGen-600 according to the window shown in the table below; ; The spectral data from the 205 spectral points output by the Perkin Elmer Lactoscope 300 were normalized using the window method described in the table below; ; The window is created by clustering highly correlated and adjacent wavenumbers using agglomerative clustering. After the window is determined, each wavenumber is normalized using all wavenumbers in the corresponding window using PDS. Step 3: By substituting the standardized MIR data into the prediction model prepared using the Foss FT+ instrument, the prediction results for the target object can be obtained.

Citation Information

Patent Citations

  • Mid-infrared spectroscopy rapid batch detection method and application of total solid content in buffalo milk

    CN114166786B