Method for establishing mid-infrared spectrum prediction model based on concentration of anti-Mullerian hormone (AMH) in cow serum

By constructing a mid-infrared spectral random forest model based on bovine serum AMH concentration, the gap in the prediction of bovine reproductive lifespan in existing technologies has been filled, achieving efficient and accurate prediction of reproductive potential and productive lifespan, and providing an economical method for assessing reproductive potential.

CN122016703APending Publication Date: 2026-05-12HUAZHONG AGRI UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG AGRI UNIV
Filing Date
2026-01-09
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Current technologies have failed to effectively predict the reproductive lifespan of dairy cows and lack a stable prediction system. Existing methods such as ELISA have limitations and suffer from problems such as low fertility and short milk production lifespan. A more economical and efficient alternative is needed.

Method used

Based on the serum anti-Müllerian hormone (AMH) concentration in dairy cows, a random forest model was constructed using mid-infrared spectral data. By measuring the characteristic wavelength infrared spectral data of milk, the productive lifespan of dairy cows was predicted, and they were divided into low, medium, and high reproductive potential groups. An RF model was then established for prediction.

Benefits of technology

It achieves efficient and accurate prediction of dairy cow reproductive potential and productive lifespan. The model has an average AUC of 0.806, an accuracy of 0.656, a precision of 0.674, a recall of 0.656, and an F1 score of 0.664. It can distinguish between samples with high and low reproductive potential and provides a fast and convenient means of assessing reproductive potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122016703A_ABST
    Figure CN122016703A_ABST
Patent Text Reader

Abstract

The invention provides a method for establishing an intermediate infrared spectrum prediction model based on the concentration of anti-Mullerian hormone (AMH) in dairy cow serum, which comprises the following steps of: correspondingly dividing a dairy cow group into a low-concentration group, a middle-concentration group and a high-concentration group according to the concentration of AMH in the dairy cow serum, taking the low-concentration group, the middle-concentration group and the high-concentration group as expressions of low, middle and high reproductive potentials of dairy cows, and measuring full-wave band infrared spectrums of milk of corresponding dairy cows, selecting a characteristic wavelength; a random forest model is constructed based on the AMH low concentration range, the AMH middle concentration range and the AMH high concentration range and the corresponding infrared spectrums of the AMH low concentration range, the AMH middle concentration range and the AMH high concentration range, the RF model constructed by introducing the measured milk infrared spectrums can be obtained, and the production life of the dairy cow can be judged The method has the advantages that the model can accurately divide the breeding potential of a dairy cow group by using the serum AMH concentration, and then predict the AMH concentration of the dairy cow according to the difference of different AMH concentrations in the 1500-2000cm <-1 > section of the infrared spectrum in the milk, so that the breeding potential of the dairy cow, namely the production life, is predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a method for establishing a mid-infrared spectral prediction model based on the concentration of anti-Müllerian hormone (AMH) in bovine serum. Background Technology

[0002] Reproductive traits are key factors influencing the longevity and productive lifespan of dairy cows. Approximately one-third of dairy cows are prematurely culled due to declining fertility or reproductive problems. In countries with high milk production, the average productive lifespan of dairy cows is only 3-4 years. Currently, many developed countries have incorporated longevity into their breeding programs and conducted related selection work based on the importance of dairy production (Miglior F et al., 2005; De Vries A et al., 2020). Extending the productive lifespan of dairy cows has become an urgent need for industry development. However, due to the complexity of this trait, longevity breeding has not yet been included in national selection indices in most developing countries. Taking my country as an example, the latest revised China Dairy Cattle Performance Index (CPI) in 2020 includes milk protein content, milk fat content, somatic cell count, body type, lactation system, and limbs, but does not cover longevity (Hu H et al., 2021). Against this backdrop, identifying biomarkers that can effectively predict the longevity of dairy cows is crucial.

[0003] For a long time, members of the transforming growth factor-β (TGF-β) superfamily have been considered key players in the recruitment, selection, and development of primordial follicles, as well as in the formation of primary follicles through ovulation and luteinization (Knight PG et al., 2006; Trombly DJ et al., 2009; Myers M et al., 2010). Anti-Mullerian hormone (AMH), a member of this family, has been shown to play an important regulatory role in follicle growth and development; its concentration can indirectly reflect the total number of healthy follicles in the ovary, thus reflecting ovarian reserve. In animals, research on AMH is relatively limited. It is known that in cattle, AMH concentration changes very little during the estrous cycle and is closely related to bovine oocyte development or follicle development. In addition, other family members such as BMP-15 and GDF-9 can synergistically promote the transformation of primordial follicles into primary follicles, activate granulosa cell proliferation and inhibit apoptosis (e.g., upregulating the anti-apoptotic protein Bcl-2), and can also regulate oocyte development capacity (Yan C et al., 2001; Peng J et al., 2013). Mutations in these two genes can lead to premature ovarian failure (POI), ovulation disorders and decreased fertility (Di Pasquale E et al., 2004).

[0004] Milk composition is an indicator of a cow's health and physical condition. As a cow's health changes, its milk composition also changes (DU C et al., 2020). Therefore, investigating the changes in various components of milk (such as milk protein, milk fat, and lactose) is crucial. Fourier transform infrared spectroscopy is a widely used technique for identifying the characteristics of dairy products, with a spectral range from 5010.15 to 925.66 cm⁻¹. -1 The absorbance at 1060 wavenumbers within the range corresponds to the short-wave infrared (SWIR), mid-wave infrared (MWIR), and long-wave infrared (LWIR) regions. The absorbance at each wavenumber, generated by the interaction between infrared light and molecules, helps characterize the chemical composition of milk (NAN L et al., 2023), thus enabling the prediction of major milk components (such as milk fat, milk protein, lactose, urea nitrogen, and total solids), refined components (fatty acids, amino acids, etc.), and the physiological state of dairy cows (reproduction, metabolism, and health) (DE MARCHI M et al., 2014). Currently, mid-infrared spectroscopy has been used to establish detection techniques and prediction models for estrus and pregnancy in dairy cows, revealing significant differences in the spectral characteristics of milk during proestrus and estrus (DU C et al., 2021).

[0005] The importance of reproductive lifespan in dairy cows is widely recognized, and the use of multiple factor indicators and mid-infrared spectroscopy for prediction has shown great potential. However, there is currently a significant research gap in the field of reproductive lifespan prediction, and a stable prediction system has not yet been established. Therefore, this invention will comprehensively explore the correlation between cytokine levels such as AMH or mid-infrared spectroscopy and reproductive lifespan, screening reliable indicators that can be used to predict high reproductive performance and long productive lifespan in dairy cows, thus filling this research gap.

[0006] Over the past 100 years, dairy cattle improvement programs have made significant progress in assessing the phenotypic traits of milk composition. Technological advancements have shifted from farm-based sample collection and processing to large-scale laboratory-based sample processing, enabling the assessment of not only milk yield and fat levels but also animal characteristics and reproductive traits (Tiplady KM et al., 2020). However, assessing the bovine energy status index (BCS) is costly in terms of resources and accuracy. McParland et al. first proposed an alternative using only mid-infrared spectral information from milk samples to estimate the ΔBCS of Irish cows. Frizzarin et al. further refined this method, demonstrating an improvement of up to 7% in the accuracy of predicting ΔBCS in early lactation by using age in days (DIM) as a predictor variable and replacing the traditional partial least squares regression used by McParland et al. with a neural network.

[0007] Besides using spectral data as a predictor to forecast milk composition, animal characteristics, and reproductive status, spectral bands can also be considered as traits and analyzed from a genetic perspective (Lainé A et al., 2017). Previous studies (Soyeurt H et al., 2010; Bittante G et al., 2013; Wang Q et al., 2016; Zaalberg RM et al., 2019) have revealed genetic analyses of milk spectra, showing different ranges of heritability, from low to high.

[0008] A woman's reproductive potential depends on her ovarian reserve, which is measured by the quantity and quality of follicles in the ovary at a specific time point (Gleicher N et al., 2011). However, since the number of follicles in the body cannot be directly measured, ovarian reserve can be indirectly assessed through biochemical indicators and ultrasound markers (Broekmans FJ et al., 2006). Increasing evidence suggests that AMH is currently the most sensitive and specific detection method compared to other parameters (such as FSH, E2, and inhibin B concentrations) (van Rooij IA et al., 2002; Kaya C et al., 2010; Fanchin R et al., 2003), and can be used as an indicator for diagnosing the reproductive potential of dairy cows.

[0009] Serum AMH is a reliable predictor of reproductive potential in dairy cows. However, this method mainly relies on ELISA assays, which has limitations in large-scale applications. Finding a more economical and efficient alternative is crucial. Current dairy farming suffers from industry problems such as low fertility, short milk-producing lifespan, and low heritability of reproductive and milk-producing lifespan traits. Therefore, establishing a selection technology system for dairy cows with high reproductive performance and long productive lifespan is extremely important. Summary of the Invention

[0010] To address the aforementioned technical problems, this invention provides a method for establishing a mid-infrared spectral prediction model based on the concentration of anti-Müllerian hormone (AMH) in dairy cow serum. Mid-infrared spectral data is routinely collected in farms and is inexpensive, providing support for the preliminary establishment of a screening method for dairy cows with high reproductive potential. This method has certain potential application value for screening high-fertility herds in actual production.

[0011] To achieve the above objectives, the present invention adopts the following technical solution:

[0012] A method for predicting the productive lifespan of dairy cows using infrared spectroscopy in milk based on serum anti-Müllerian hormone (AMH) concentration, comprising: AMH concentration in bovine serum was detected using an AMH quantitative detection kit. Bovine populations with AMH concentrations below 458.3 pg / mL were classified as having low reproductive potential, while those above 814.5 pg / mL were classified as having high reproductive potential. The bovine population was further divided into low, medium, and high concentration groups to represent low, medium, and high reproductive potential, respectively. The characteristic wavelengths of the corresponding milk samples were then selected based on the full-band infrared spectra. A random forest (RF) model was constructed based on the infrared spectral data of the three AMH concentration ranges and their corresponding characteristic wavelengths. This model predicts the productive lifespan of bovines by measuring the characteristic wavelengths of the milk. The characteristic wavelength infrared spectral data of the milk samples from the tested bovine populations were then imported into the constructed RF model to predict the productive lifespan of the bovines.

[0013] According to the RF model, when the infrared spectral data of milk from dairy cows corresponds to a high AMH concentration, it indicates that the cows have high reproductive potential and long productive lifespan.

[0014] As described above, preferably, the division into low, medium, and high concentration groups is based on AMH concentration ≤ 458.3 pg / mL, 458.3 pg / mL < AMH concentration < 814.5 pg / mL, and AMH concentration ≥ 814.5 pg / mL.

[0015] The method described above preferably uses a characteristic wavelength range of 926-5012 cm⁻¹ for the infrared spectrum of milk. −1 The study included five characteristic ranges: SWIR: 5010~3673 cm⁻¹, SWIR-MWIR: 3669~3052 cm⁻¹, MWIR-1: 3048~1701 cm⁻¹, MWIR-2: 1698~1585 cm⁻¹, and MWIR-LWIR: 1582~925 cm⁻¹. The study compared the mid-infrared spectra of dairy products from cows with low, medium, and high AMH concentrations, and correlated serum AMH concentrations with mid-infrared spectral characteristics. The results showed a correlation between AMH levels in dairy cows and mid-infrared spectral data, particularly in the 1500~2000 cm⁻¹ range. −1 The most significant difference is between the segments, namely the wavelengths of the infrared spectrum of milk preferably in the MWIR-2 and MWIR-LWIR segments.

[0016] In the method described above, preferably, the infrared spectral data of the characteristic wavelength is the absorbance value within the characteristic wavelength range.

[0017] As described above, preferably, when the predicted AMH concentration is ≤458.3 pg / mL, the predicted parity is ≤2, indicating low-parity culling of cattle, which indicates poor reproductive performance; When the predicted AMH concentration is ≥814.5 pg / mL, the predicted parity is ≥3, indicating high-parity breeding cattle and good reproductive performance.

[0018] Milk infrared spectroscopy is used to predict the reproductive performance of dairy cows, wherein the characteristic wavelength of the milk infrared spectrum is in the range of 1500–2000 cm⁻¹. −1 Section.

[0019] More preferably, the infrared spectrum is the absorbance value with characteristic wavelengths in the range of 1698~1585 cm⁻¹ and 1582~925 cm⁻¹.

[0020] This invention uses AMH concentration as an indicator of high reproductive potential (where AMH concentrations below 458.3 pg / mL indicate poor reproductive potential, while concentrations above 814.5 pg / mL indicate high reproductive potential in dairy cows), and establishes a random forest (RF) mid-infrared spectral prediction model for dairy cow reproductive potential. The trained random forest model is encapsulated as a standalone, reusable R language script module. In practical applications, after obtaining new dairy cow milk spectral data, users only need to import the data into this script to automatically run the preprocessing and prediction process, directly outputting the corresponding dairy cow AMH concentration prediction classification result (low, medium, or high concentration), achieving rapid and convenient reproductive potential assessment.

[0021] This invention relates to a method for predicting the productive lifespan of dairy cows based on full-band infrared spectral data of milk. Specifically, it uses the concentration of anti-Müllerian hormone (AMH) in dairy cow serum as a key biological indicator, classifying different AMH concentrations into low, medium, and high reproductive potential groups. A random forest model is then used to establish a mapping relationship between spectral data and AMH concentration classifications, enabling accurate assessment of the reproductive potential of dairy cows. This method, through steps such as data acquisition and preprocessing, model construction and optimization, performance evaluation, and system deployment, achieves automated and high-precision analysis from spectral data to productive lifespan prediction, providing a reliable technical means for farm management.

[0022] The beneficial effects of this invention are as follows: This invention provides a method for establishing a superior model for predicting the reproductive potential of dairy cows using spectral analysis, namely the random forest model. Studies have found a positive correlation between actual AMH concentration and the probability of being predicted as having high reproductive potential. The model has an average AUC of 0.806, accuracy of 0.656, precision of 0.674, recall of 0.656, and F1 score of 0.664. Furthermore, the distribution of low and high reproductive potential in the probability histogram is relatively ideal, indicating that the model has a certain degree of discriminative power and can be used to predict the reproductive potential of dairy cows.

[0023] To select the optimal model from SVM and RF, two superior models were used to predict milk samples not involved in the modeling process as an external validation set. This set included low-parity culled cows (≤2 parities, poor reproductive performance) and high-parity breeding cows (≥3 parities, good reproductive performance). The probability of these samples being predicted as having high or low AMH concentrations was observed. Validating the RF model for predicting AMH concentration revealed that high-parity cows were predicted as having high AMH concentrations in 70.0% of cases, while low-parity cows were predicted as having low AMH concentrations in 99.1% of cases. This further demonstrates that AMH can serve as a biomarker for predicting dairy cow reproductive potential and optimizing herd management.

[0024] Bovine serum anti-Müllerian hormone (AMH) has shown good results as an indicator for predicting the longevity and reproductive potential of dairy cows. It can be used as an indirect indicator of the reproductive potential of dairy cows by utilizing infrared spectral data from milk to achieve the purpose of predicting the reproductive potential of dairy cows. Attached Figure Description

[0025] Figure 1 The correlation between different wavelengths of the milk spectrum and serum AMH levels;

[0026] Figure 2 ROC curves for multiple categories of the SVM model;

[0027] Figure 3 A probability plot showing AMH levels and predicted high AMH concentration groups;

[0028] Figure 4 The probability distribution diagram of each category predicted by the SVM model;

[0029] Figure 5 For optimal cross-validation performance of the PLSR model;

[0030] Figure 6 The coefficients of determination for the training set of the PLSR model;

[0031] Figure 7 The coefficients of determination for the model on the test set;

[0032] Figure 8 ROC curves for multiple categories in the RF model;

[0033] Figure 9 A probability plot showing AMH levels and predicted high AMH concentration groups;

[0034] Figure 10 Predicted proportions of AMH concentration in different parity groups using the RF model;

[0035] Figure 11 Predicted proportions of AMH concentration for different parity groups using the SVM model;

[0036] Figure 12 The predicted proportion of AMH concentration in different parity groups using the RF model. Detailed Implementation

[0037] Compared to other biochemical and biophysical markers, AMH (Amyotrophic Lateral Breast Milk) has significant advantages as a highly stable marker, particularly valuable in assessing ovarian reserve function and menopausal time. This study selected dairy cows with different AMH levels as a phenotypic trait for high reproductive potential. Based on previous research showing that AMH concentrations below 458.3 pg / mL indicate poor reproductive potential, while concentrations above 814.5 pg / mL indicate high reproductive potential, the dairy cow population was divided into low, medium, and high concentration groups. Comparison of milk spectral characteristics revealed a correlation between AMH levels and mid-infrared spectral data, with the 1500–2000 cm⁻¹ range showing the highest correlation. −1 The differences between segments are the most significant. Further consideration is to use the obtained AMH thresholds to divide dairy cows into three groups based on their AMH concentration: low, medium, and high reproductive potential (i.e., cows with high AMH concentration are considered to have high reproductive potential and long productive lifespan), and to use three modeling methods (SVM, PLSR, RF) to predict the productive lifespan of dairy cows.

[0038] First, an SVM model for predicting the reproductive lifespan of dairy cows was constructed. It was found that the probability of being predicted as having high reproductive potential increased with increasing actual AMH concentration. The model had an average AUC of 0.849, accuracy of 0.748, precision of 0.687, recall of 0.712, and F1 score of 0.685. The AUC for predicting low, medium, and high reproductive potential were all above 0.8, indicating that the model performs well in distinguishing between samples with high and low reproductive potential. However, only the distribution of high reproductive potential in the probability histogram conformed to the ideal distribution, indicating that the model's prediction effect for high reproductive potential was best. Next, a PLSR model for predicting the reproductive lifespan of dairy cows was constructed. The results were unsatisfactory; the R² values ​​for the training and test sets were 0.125 and 0.121, respectively, far less than 1, making it unsuitable for predicting dairy cow populations with different reproductive potentials. Finally, the RF model for predicting dairy cow reproductive lifespan was constructed. The study found a high degree of overlap between the actual and predicted values ​​of the high reproductive potential group in the confusion matrix of the validation set, indicating that the model effectively distinguishes high reproductive potential. The actual AMH concentration was positively correlated with the probability of being predicted as having high reproductive potential. The model's mean AUC was 0.806, accuracy was 0.656, precision was 0.674, recall was 0.656, and F1 score was 0.664. Furthermore, the distribution of low and high reproductive potential in the probability histogram was relatively ideal, indicating that the model has a certain degree of discriminative power.

[0039] To select the best model from SVM and RF, these two models were used to predict milk samples not included in the modeling, including low-parity culled cows (parity ≤2, poor reproductive performance) and high-parity breeding cows (parity ≥3, good reproductive performance). The probability of these samples being predicted as having high or low AMH concentrations was observed. The results showed that in the SVM model, high-parity samples were predicted as having high AMH concentrations in 48.8% of cases, while low-parity samples were predicted as having low AMH concentrations in only 17.4%, indicating poor model validation. In contrast, in the RF model, high-parity samples were predicted as having high AMH concentrations in 70% of cases, while low-parity samples were predicted as having low AMH concentrations in 99.1%. Previous research has shown that the RF model has higher prediction accuracy for samples with low and high reproductive potential, consistent with the model's prediction patterns, and the model validation results are more satisfactory. This result not only proves that cows with different AMH levels have different infrared spectral characteristics in their milk, which can be used to predict the reproductive potential and productive lifespan of cows, but also reaffirms that Holstein cows, which are long-lived and have high reproductive potential, possess the trait of high AMH levels.

[0040] In conclusion, AMH has proven effective as an indicator for predicting the longevity and reproductive potential of dairy cows. It can be used as a direct indicator correlated with parity, or as an indirect indicator, utilizing mid-infrared spectroscopy to predict the reproductive potential of dairy cows, thereby achieving the goal of predicting their productive lifespan.

[0041] The following embodiments are used to further illustrate the present invention, but should not be construed as limiting the present invention. Any modifications or substitutions made to the present invention without departing from its spirit and essence are within the scope of the present invention.

[0042] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art, and unless otherwise specified, all reagents used in the embodiments are analytical grade or higher.

[0043] Example 1

[0044] 1. Experimental Materials

[0045] 1.1 Experimental Animals

[0046] Milk samples used for modeling: From June 2023 to March 2025, milk samples were collected from 1438 healthy Holstein cows with complete breeding records from two Holstein dairy farms (Pastoral A and Pasture B) in Central China. Milk samples were collected using a fully automated rotary milking machine, with approximately 100 mL collected per sample. The samples were aliquoted into sampling bottles and numbered sequentially. Preservatives were immediately added to the sampling tubes, and the samples were then transported to the laboratory at 4°C for further analysis. Upon arrival at the DHI laboratory, the samples were immediately subjected to spectral analysis and data collection.

[0047] Serum samples were collected from 1,438 dairy cows at the same location and from June 2023 to March 2025. Blood was drawn from the tail vein of these 1,438 Holstein cows. To analyze the relationship between AMH concentration and age, all cows were grouped according to their parity: low parity group (<3 parities) and high parity group (≥3 parities).

[0048] 1.2 Collection and determination of MIRS and conventional milk components in milk

[0049] At the Hubei Provincial DHI Testing Center, the standard "NYT1450-2007 China MIRS Spectroscopy and Routine Milk Composition" was referenced. All milk sample spectra were acquired using a Danish FOSS milk composition analyzer, with a range of 926-5012 cm⁻¹. -1 1060 individual wave points within.

[0050] The MIRS determination process for milk samples includes: warming the milk samples and measuring the samples. Fresh milk samples stored at 2-4℃ are taken out and placed on a rack, then preheated in a 45℃ water bath for 30 minutes. After the preheated milk samples are shaken well, they are placed on the MilkoScan™ FT+ detection conveyor belt. The cap of each milk sample bottle is opened, and the milk samples are tested using the Danish FOSS milk composition analyzer.

[0051] 1.3 Collection of dairy cow serum samples

[0052] Blood was collected from the tail vein of dairy cows using dry vacuum tubes free of pyrogens, endotoxins, and additives. 5 mL of fresh blood was collected from each blood collection tube. After collection, the blood in the tubes was allowed to stand at room temperature for 2 hours (avoiding direct sunlight) until it coagulated. The tubes were then centrifuged at 3000 r / min for 10 min. The supernatant, i.e., serum, was collected and aliquoted into sterile 2 mL EP tubes and stored at -80°C.

[0053] 1.4 Test Equipment

[0054] Fourier transform mid-infrared spectroscopy milk component analyzer (Milko Scan™ FT+, FOSS, Denmark); Dumatec™ 8000 combustion nitrogen analyzer (FOSS, Denmark); Funke Gerber Super Vario-N milk fat centrifuge (Funke Gerber, Germany); Geller milk fat meter (Funke Gerber 3156, Funke Geber, Germany); electric heating drying oven (Shanghai Boxun Industrial Co., Ltd. Medical Equipment Factory); glass weighing dishes (60x30 mm); microplate reader; 37℃ constant temperature incubator.

[0055] 1.5 Test Reagents

[0056] Sulfuric acid (H2SO4, AR, 98%); Isoamyl alcohol (C5H) 12 O, AR, 98.5%); ethylenediaminetetraacetic acid (EDTA, 99.9% purity); distilled water.

[0057] 2. Test Methods

[0058] 2.1 Determination of serum AMH concentration

[0059] Serum AMH levels were measured using the Bovine Müllerian Hormone (AMH) Quantitative Detection Kit (ELISA) from Shanghai Enzyme Linked Products Co., Ltd., catalog number ml944588V.

[0060] This kit uses an enzyme-linked immunosorbent assay (ELISA) competitive assay to detect the level of anti-Müllerian hormone (AMH) in serum. The specific procedure is as follows: First, a solid-phase secondary antibody is prepared by coating a microplate with goat anti-rabbit antibody. Then, the test serum, horseradish peroxidase-labeled anti-Müllerian hormone (AMH), and anti-Müllerian hormone (AMH) antibody are added to form a coated secondary antibody-anti-Müllerian hormone (AMH) antibody-anti-Müllerian hormone (AMH) (HRP) complex. The amount of labeled anti-Müllerian hormone (AMH) bound is inversely proportional to the amount of anti-Müllerian hormone (AMH) in the serum. After color development, the absorbance (OD value) is measured using an ELISA reader. The concentration-absorbance curve is fitted by computer or by plotting to calculate the anti-Müllerian hormone (AMH) level in the test serum.

[0061] The specific operating steps for ELISA testing are as follows:

[0062] (1) Move all reagents to room temperature for two hours to equilibrate. Take the concentrated washing solution and dilute it with distilled water at a ratio of 1:20 according to the batch testing quantity. Mix well and set aside.

[0063] (2) Remove the pre-coated plate from the sealed bag and set up a blank control well without adding any liquid; set up two wells for each calibrator and add 50 μl of the corresponding calibrator; add 50 μl of the sample to be tested to the sample well; add 50 μl of enzyme-labeled antigen to each well except the blank well;

[0064] (3) Then add 50 μl of antibody to each well (except for the blank control well), mix thoroughly, attach the sealing film, and incubate at 37°C for 1 hour.

[0065] (4) Manual washing: Discard the liquid in the holes, fill each hole with washing liquid, let stand for 10 seconds and spin dry, repeat 3 times and then pat dry. Washing machine washing: Select the 3-wash program to wash the plate and then pat dry.

[0066] (5) Add 50 μl of chromogenic reagent A and 50 μl of chromogenic reagent B to each well, shake to mix, and place at 37°C in the dark for 15 minutes for color development. Use an ELISA reader to read the absorbance at a wavelength of 450 nm. For single-wavelength ELISA readers, the zero point needs to be adjusted first with a blank control well, and then the absorbance of each well is measured.

[0067] Result calculation:

[0068] After the test is completed, the concentration of the standard is used as the x-axis and the corresponding absorbance (OD value) is used as the y-axis. Using computer software, a standard curve equation is created by fitting a four-parameter Logistic curve (4-p1). The concentration of the sample is calculated using the sample absorbance (OD value) through the equation.

[0069] If the sample is diluted, the concentration value measured by the above method should be multiplied by the dilution factor to obtain the final concentration of the sample.

[0070] Following the sample collection and ELISA assay methods described above, serum AMH levels were measured in 1591 serum samples. After optimization, the AMH ELISA kit achieved a detection range of 156.25 pg / mL to 5000 pg / mL, a sensitivity of less than 10 pg / mL, an intra-plate coefficient of variation of less than 10%, and an inter-plate coefficient of variation of less than 15%. Standard curves were established by subtracting the blank well readings from the standard well readings in the ELISA kit, using the corresponding AMH standard concentrations (5000, 2500, 1250, 625, 312.5, and 156.25 pg / mL) and their corresponding OD values.

[0071] 2.2 Spectral Data Acquisition

[0072] All milk samples were vortexed in a 40℃ water bath for 5 min. The milk composition was analyzed using a FOSS FT-6000 Fourier transform mid-infrared spectrometer, and the infrared spectrum of each sample was acquired. Each sample was scanned twice, and the average value was taken. The FT-MIR spectral range was 926-5012 cm⁻¹. -1 Resolution 3.652 cm -1 There are a total of 1060 wave points.

[0073] Example 2: Establishment of an AMH-based model for predicting the reproductive lifespan of dairy cows using mid-infrared spectroscopy

[0074] 1. Correlation between different wavelengths of milk spectrum and AMH concentration

[0075] This study obtained AMH levels and their corresponding spectral data using the method described in Example 1. The preprocessed FT-MIR spectral data were standardized to ensure that the mean and total variance of the spectral intensity for each wavenumber were zero. Standardization was performed using the SCALE function in R. A portion of the 1591 serum samples, comprising 601 dairy cow spectral datasets, was randomly divided into a training set (80%) for model building and a test set (20%) for performance evaluation. Models were built using the training set, and their performance was evaluated on the test set. To evaluate the optimal algorithm for building the best model, SVM, PLS, and RF algorithms were used to construct models. The model with the best overall performance was selected based on parameters such as accuracy, precision, recall, and F1 score. All machine learning models were built using the CARET package (versions 6.0-93).

[0076] Sixty-one dairy cows with different AMH levels were selected as the phenotypic trait for high reproductive potential. Serum AMH concentration and mid-infrared spectral characteristics were correlated and analyzed.

[0077] Based on previous research showing that cows with AMH concentrations below 458.3 pg / mL have low reproductive potential, while those above 814.5 pg / mL have high reproductive potential, dairy cow serum AMH concentrations were grouped into low, medium, and high groups according to their concentrations: ≤458.3 pg / mL, between 458.3 pg / mL and 814.5 pg / mL, and ≥814.5 pg / mL. Correlation analysis was then performed between these groups and the corresponding milk spectra, and correlation coefficients were calculated.

[0078] Step 1: Sample Grouping

[0079] The samples were divided into three categories based on quantiles, using dairy cow reproductive potential as a classification label, as defined below: "Low": AMH≤458.3 pg / mL; "Medium": 458.3 pg / mL <AMH<814.5 pg / mL; “High”: AMH≥814.5 pg / mL;

[0080] Step 2: Spectral-AMH Correlation Analysis

[0081] 1. Extract the spectral matrix X of the high AMH group. high and AMH concentration vector Y high ; 2. Regarding X high For each column in the matrix (i.e., the spectral intensity corresponding to each wavelength or wavenumber), calculate its relationship with Y. high Pearson correlation coefficient r j :

[0082] Where r j Let x represent the correlation coefficient at the j-th wavelength. ij y represents the spectral intensity of the i-th sample at the j-th wavelength. i The corresponding AMH concentration is given, and n is the number of samples in the high AMH group. 3. Iterate through all wavelengths to obtain the correlation coefficient vector r = [r1, r2, ..., r m ], where m is the total number of wavelengths.

[0083] Step 3: Plotting the Spectrum-AMH Correlation Diagram ( Figure 1 )

[0084] 1. Wavelength extraction and formatting: Extract wavelength values ​​from the column names of the AMH group spectral data and convert them into numerical vectors; verify that the number of wavelengths matches the number of columns in the spectral data to ensure a one-to-one correspondence.

[0085] 2. Construct a correlation coefficient dataset: Combine the calculated correlation coefficients for each wavelength with the wavelength values ​​to create a data frame; remove missing values ​​(NA) to ensure the plotting data is complete.

[0086] 3. Graphical plotting: Plot a scatter plot with wavelength as the x-axis and correlation coefficient as the y-axis; add a local weighted regression smoothing curve (LOESS, span=0.1) and use a solid red line to show the overall correlation trend; add a gray dashed horizontal reference line (y=0) to visually distinguish between positive and negative correlation areas.

[0087] Step 4: Model Building

[0088] 1. Based on the absolute value of the correlation coefficient |r j | Set a threshold (preferably 0.2-0.3) to screen out characteristic wavelengths that are significantly correlated with AMH concentration; 2. Based on the spectral intensity at the selected characteristic wavelength, establish an AMH concentration prediction model, such as the Support Vector Machine (SVM) model, the Partial Least Regression Squares (PLSR) model, or the Random Forest (RF) model.

[0089] Based on previous research indicating that cows with AMH concentrations below 458.3 pg / mL have low reproductive potential, while those above 814.5 pg / mL have high reproductive potential, the dairy herd was divided into three groups: low, medium, and high AMH concentrations. Comparison of the mid-infrared spectra of milk products from cows with low, medium, and high AMH concentrations revealed a correlation between AMH levels and mid-infrared spectral data: analysis of the association between milk spectra and AMH showed correlation coefficients ranging from -0.32 to 0.30. Characteristic wavelengths with |r|>0.2 were selected, including five characteristic intervals: SWIR: 5010~3673 cm⁻¹, SWIR-MWIR: 3669~3052 cm⁻¹, MWIR-1: 3048~1701 cm⁻¹, MWIR-2: 1698~1585 cm⁻¹, and MWIR-LWIR: 1582~925 cm⁻¹. A PLSR model based on these wavelengths was used to predict the coefficient of determination R of AMH concentration. 2 The accuracy reached 0.86, with an average relative error of <8%, particularly in the 1500–2000 cm range. −1 The most significant differences were observed in the wavelength ranges of MWIR-2 (1698–1585 cm⁻¹) and MWIR-LWIR (1582–925 cm⁻¹). These results indicate a correlation between spectral characteristics and bovine serum AMH levels, with differences in spectra observed in cows of different parities. Therefore, AMH concentration and milk spectra can be used to predict the productive lifespan of dairy cows.

[0090] Subsequently, 601 sample data points with reference values ​​were used for modeling (80% of the sample size was used as the training set and 20% as the test set), while the remaining 763 data points (without AMH concentration measurement) were used as the external validation set for the model. Three algorithms, namely Support Vector Machine (SVM), Partial Least Squares Regression (PLSR) and Random Forest (RF), were used to build the model, and the performance of the models built by the three algorithms was comprehensively analyzed.

[0091] Specifically, the following is true: (1) Using specific AMH concentrations in dairy cows and specific spectral bands in milk (1500-2000 cm⁻¹) -1 This study employs a Support Vector Machine (SVM) classification model construction and analysis method to provide a non-invasive detection method for assessing the reproductive potential and productive lifespan of dairy cows. The steps include:

[0092] (1) Variable definition and model form Independent variable (explanatory variable): Cow milk spectral data, a continuous matrix variable containing absorbance or reflectance values ​​across the entire spectrum, preferably in the wavelength range of 1500-2000 cm⁻¹. -1 ; Dependent variable (response variable): Cow reproductive potential classification label, a three-category variable, defined as follows: "Low": AMH concentration ≤ 458.3 pg / mL "Medium": 458.3 pg / mL < AMH concentration < 814.5 pg / mL "High": AMH concentration ≥ 814.5 pg / mL Model Form: A Support Vector Machine (SVM) classification model is adopted, with the radial basis function (RBF) kernel being the preferred choice. Its decision function is expressed as:

[0093] Where: x is the standardized spectral feature vector. ∈{−1,0,1} corresponds to low, medium, and high AMH category labels. Here, γ represents the support vector weight coefficients, γ represents the kernel function parameters, and b represents the bias term.

[0094] (2) Data partitioning and preprocessing

[0095] To evaluate the model's generalization ability, the complete dataset was randomly divided into: Training set: accounting for 80%, used to build SVM models and optimize parameters; Test set: accounting for 20%, used only for the final evaluation of the model's predictive performance and to simulate new data prediction scenarios.

[0096] Preferably, a stratified sampling method is used to ensure that the proportion of different AMH concentration categories in the training and test sets is basically consistent with that in the original dataset.

[0097] Data preprocessing includes the following steps: ① Missing value handling: Identify missing values ​​in the spectral data and fill them in using the median of the corresponding feature column.

[0098] Identification: Traverse all feature columns of the spectral data matrix and detect whether there are missing values ​​(Not aNumber, NA, or NaN).

[0099] Imputation: For any feature column with missing values, calculate the median of all valid values ​​in the column and assign this median to the positions of all missing values ​​in the column. This method ensures the robustness of the imputed values ​​and avoids the influence of extreme values. ② Outlier Handling: Identify and replace infinite values ​​(Inf / -Inf) with the median of the finite values ​​in the corresponding feature column.

[0100] Identification: Traverse the spectral data matrix and detect whether there are any outlier values ​​(Inf or -Inf) that represent infinity or infinity.

[0101] Correction: For any feature column with infinite values, first extract all finite values ​​in the column, calculate the median, and then use this median to replace all infinite values ​​in the column. This operation eliminates numerical distortion caused by instrument signal saturation or measurement errors. ③ Feature Filtering: Remove wavelength features with zero or no variance in the training set to eliminate invalid spectral bands.

[0102] Variance calculation: For the training set spectral data, calculate the variance of all sample data under each wavelength feature column by column.

[0103] Feature filtering: Identify and remove feature columns with zero variance or missing variance (NA). This step removes spectral bands that are invariant or have very low information content throughout the training set, thereby reducing data dimensionality and noise interference. ④ Data standardization: Center and standardize the spectral matrix so that the mean of each wavelength feature is 0 and the standard deviation is 1.

[0104] Centering: The spectral data matrix after the above cleaning is processed column-wise to make the mean of the data in each wavelength characteristic column zero. The calculation method is to subtract the arithmetic mean of all values ​​in the column from each original value in the column.

[0105] Scaling (Standardization): Building upon centralization, each feature column is further scaled to ensure its standard deviation is 1. This is calculated by dividing each centered value by the standard deviation of all values ​​in that column.

[0106] Technical effect: After centering and standardization, all wavelength features are transformed to the same scale (mean is 0, standard deviation is 1), which effectively eliminates the adverse effects of differences in the baseline or intensity dimension of the spectral signal on the calculation of the kernel function of the SVM model, and ensures the convergence speed and generalization performance of the model.

[0107] (3) Model training and parameter optimization

[0108] The following steps are used to construct an SVM classification model: ① Set up the training framework: Use 5-fold cross-validation as the model selection and evaluation strategy to ensure the robustness of parameter optimization; ② Parameter search: Optimize the cost parameter C and kernel parameter γ through grid search, with the preferred search range being C∈[2]. −5 ,2 15 ],γ∈[2 −15

[23] , with classification accuracy as the optimization objective; ③ Model training: Use the optimized parameters to train the final SVM model on the training set, and record the support vectors and corresponding weights.

[0109] (4) Model performance evaluation

[0110] The predictive performance of the final model is evaluated on an independent test set, primarily using the following metrics: ① Accuracy: the proportion of correctly predicted samples out of the total number of samples in the test set; ② Precision: the proportion of samples predicted as belonging to a certain reproductive potential category that actually belong to that category, calculated separately for low, medium, and high categories; ③ Recall: the proportion of samples that actually belong to a certain reproductive potential category that are correctly predicted by the model to belong to that category; ④ F1 score: the harmonic mean of precision and recall, comprehensively reflecting the model's ability to identify each category.

[0111] Preferably, the ROC curve and AUC value are reported: a one-vs-rest strategy is adopted to plot the ROC curve and calculate the AUC value for each AMH category. The closer the AUC value is to 1, the stronger the model's discrimination ability.

[0112] (5) Implementation Examples and Results

[0113] A total of 601 spectral samples of milk from healthy dairy cows and corresponding serum AMH concentration data were collected. The AMH concentration data were classified and labeled according to the aforementioned threshold. The preprocessed spectral data were used as input features, and the AMH classification labels were used as outputs. A training set (481 samples) and a test set (120 samples) were created to enable the SVM model to learn the correlation between spectral features and reproductive potential categories.

[0114] To verify the performance of the SVM model, ROC curves were plotted (e.g., Figure 2 The ROC curve, plotted with the false positive rate (FPR) on the horizontal axis and the true positive rate (TPR) on the vertical axis, showed an average AUC of 0.849, indicating that the model has good discriminative ability and can effectively distinguish between low, medium, and high reproductive potential to a certain extent. It also found that... Figure 3 As the actual AMH concentration increases, the probability of being predicted as having high reproductive potential also increases, while when the serum AMH concentration is below 458.3 pg / mL, the probability of being predicted as having high reproductive potential is far less than 50%. Figure 4 The histograms of predicted probability distributions for each category show that when there is a peak (approaching 1) on the right side of the high reproductive potential group, a peak (approaching 0) on the left side of the low reproductive potential group, and the medium reproductive potential distribution is between 0.3 and 0.7, the model performance is considered relatively ideal. However, in the probability histogram of this model, the distributions of the low and medium reproductive potential groups do not conform to the ideal pattern, while the classification of the high reproductive potential group conforms to this pattern, indicating that the model can accurately predict the high reproductive potential group based on the spectrum.

[0115] Because the sample size in the low and medium reproductive potential groups of this model is smaller than that in the high reproductive potential group, there may be a data imbalance problem. Therefore, a comprehensive evaluation of model performance is necessary, especially when class imbalance occurs. AUC is one of the best options. Thus, for this model, the performance indicators in Table 1 should be considered as an auxiliary indicator of AUC. Further evaluation using accuracy (0.748), precision (0.687), recall (0.712), and F1 score (0.685) shows that the model correctly predicted approximately 74.8% of the samples. Of the samples predicted to have an AMH concentration ≥814.5 pg / mL, approximately 68.7% actually had this condition. The proportion of samples with an actual AMH concentration ≥814.5 pg / mL that were correctly identified was approximately 71.2%. The overall performance index F1 reached 0.685. These results demonstrate that the model has a relatively ideal effect in predicting the productive lifespan of dairy cows.

[0116] Table 1 Evaluation Indicators of the SVM Model for Predicting Dairy Cow Lifespan Based on Milk Spectrum Group accuracy Accuracy Recall rate F1 average 0.748 0.687 0.712 0.685 low concentration \ 0.379 0.500 0.431 medium concentration \ 0.364 0.129 0.190 high concentration \ 0.824 0.915 0.867

[0117] (2) PLSR model for predicting dairy cow productive lifespan based on serum anti-Müllerian hormone (AMH) concentration through milk spectroscopy

[0118] All figures in the accompanying drawings of this invention were generated using statistical calculations and visualization scripts in R language. Specifically, they relate to a method for predicting the productive lifespan of dairy cows based on the concentration of anti-Müllerian hormone (AMH) and specific milk spectral data (absorbance values ​​at 1698~1585 cm⁻¹ and 1582~925 cm⁻¹), employing a partial least squares regression (PLSR) model. This method aims to provide a technical means for the accurate assessment of the reproductive potential of dairy cows. The steps include: First, dairy cows are classified according to their specific AMH concentrations. Cows with high AMH concentrations have high reproductive potential, while those with low AMH concentrations do not. Then, based on mid-infrared milk spectral data, a predictive model for dairy cow productive lifespan is constructed using a PLSR model. This method includes the following steps: preprocessing the milk spectral data; grouping and analyzing the data according to AMH concentration; training and validating the PLSR model to establish the relationship between spectral characteristics and AMH concentration; and visualizing the analysis process and results using accompanying figures.

[0119] (1) Data preparation and preprocessing

[0120] 1.1 Data Sources and Structure Spectral data of milk from healthy dairy cows and corresponding serum AMH concentration data were collected from large-scale dairy farms.

[0121] The data consists of two main parts: AMH concentration data: continuous numerical variable, unit is pg / mL; ② Spectral data: including a specific wavelength range (500-2000 cm⁻¹) -1 The absorbance (transmittance values ​​converted to absorbance values ​​using a formula) within the specified wavelength range is preferably MWIR-2: 1698~1585 cm⁻¹, MWIR-LWIR: 1582~925 cm⁻¹, stored in matrix form to construct a spectral data matrix M, the mathematical representation of which is: M ∈ ℝ^(n×m)

[0122] ℝ is the set of real numbers; n is the number of samples; m represents the number of wavelength points, corresponding to the number of measurement points within the range of 1500-2000cm⁻¹; The matrix element M(i,j) represents the absorbance value of the i-th sample at the j-th wavenumber point.

[0123] The transmittance value (0~1) converted to absorbance is expressed by the following formula:

[0124] Where A is absorbance and %T is percentage transmittance.

[0125] 1.2 Sample Grouping Strategy

[0126] Based on a fixed threshold set according to AMH concentration, and using this as a criterion for reproductive potential, the samples were divided into three groups: "Low": AMH concentration ≤ 458.3 pg / mL; "Medium": 458.3 pg / mL < AMH concentration < 814.5 pg / mL; "High": AMH concentration ≥ 814.5 pg / mL.

[0127] 1.3 Dataset Partitioning

[0128] A special grouping strategy is used to construct the training and validation sets: Training set: Contains samples from low-fertility and high-fertility groups, used to build the PLSR model. Validation set: Contains only samples from the medium reproductive potential group, used to evaluate the model's predictive performance in the intermediate concentration range.

[0129] This partitioning method is based on the following principle: the extreme groups (low and high) have more obvious differences in spectral characteristics, which is beneficial for the model to learn the correlation between AMH concentration and spectrum; the intermediate group serves as an independent validation set to evaluate the model's generalization ability.

[0130] 1.4 Data Preprocessing

[0131] ① Outlier handling: Identify and remove outliers from spectral data.

[0132] Identification: The spectral data matrix is ​​scanned row by row (sample) and column by column (wavelength). Based on statistical distribution principles (such as Tukey's rule based on interquartile range) or preset physical thresholds (such as reasonable absorbance range), abnormal values ​​that significantly deviate from the main data distribution or do not conform to the physical meaning of the spectrum are identified.

[0133] Elimination: Identified outlier values ​​are marked as missing values, or the entire sample containing them is excluded from subsequent modeling analysis, in order to eliminate their distorting effect on the estimation of regression coefficients.

[0134] ② Missing value handling: The median imputation method is used to handle missing spectral data.

[0135] Imputation strategy: For missing values ​​that have been removed from outliers or that exist in the original data, the median imputation method is used.

[0136] Specific operation: For any wavelength feature column with missing values, calculate the median of all valid observations in that column, and use this median to fill in all missing data points in that column. This method effectively avoids the bias that may be introduced by using the mean for imputation, ensuring the robustness of the imputed values ​​to the overall data distribution.

[0137] ③ Data type conversion: Ensure that the spectral data is a numerical matrix.

[0138] Type validation and conversion: Confirm the storage format of the spectral data and perform forced type conversion to ensure that every element of the entire data matrix is ​​numerical data.

[0139] Matrixing: Organizing data into a standard two-dimensional matrix form, where rows represent samples and columns represent wavelength variables, providing a structural basis for subsequent matrix operations.

[0140] ④ Standardization: The spectral matrix is ​​centered and standardized to eliminate baseline drift and dimensional effects.

[0141] Centering: Centering is performed on each column of the spectral data matrix (i.e., each wavelength variable). Specifically, the arithmetic mean of all observations in that column is subtracted from each original observation in that column. This step aims to eliminate spectral baseline drift and make the data distribution of each wavelength variable symmetrical around the zero point.

[0142] Standardization (scaling): Building upon the centering, further standardization is performed on each column. Specifically, the centered value is divided by the standard deviation of all observations in that column. This step aims to unify the dimensions and dispersion of the wavelength variables, ensuring all features have comparable importance and preventing bias in regression coefficients due to differences in variable scaling.

[0143] (2) PLSR model construction and parameter optimization

[0144] 2.1 Training Data Preparation Independent variable (X): The spectral data matrix of the training set, with dimensions n×p (n is the number of samples, p is the number of wavelengths). Dependent variable (Y): AMH concentration vector of the training set, with dimension n×1.

[0145] 2.2 Cross-validation settings The model parameters were optimized using a 5-fold cross-validation strategy. The training set is randomly divided into 5 mutually exclusive subsets. Four of these subsets are used as training subsets and one subset is used as validation subset in turn. This process is repeated 5 times to ensure that each subset is used as a validation set once.

[0146] 2.3 Principal Component Optimization

[0147] Determine the optimal number of principal components (ncomp) through grid search: Set the range of the number of principal components to be tested (preferably 1-10), calculate the root mean square error (RMSE) of cross-validation for each candidate number of principal components, and select the number of principal components that minimizes the cross-validation RMSE as the optimal parameter.

[0148] 2.4 Model Training

[0149] Train the PLSR model on the full training set using the optimal principal component number: Algorithm: Partial Least Squares Regression Preprocessing: Spectral data centralization and normalization, Output: Model parameters such as regression coefficient matrix, score matrix, and loading matrix.

[0150] 2.5 Visualization of Parameter Optimization

[0151] Generate PLSR model parameter tuning chart (as attached) Figure 5 (As shown), the specific steps include: ① Extract the cross-validation results, including the RMSE values ​​corresponding to different principal component numbers. ② Draw a line graph showing the change in RMSE as a function of the number of principal components. ③ Mark the points corresponding to the optimal principal components, and add vertical lines and text labels. ④ Set appropriate titles, axis labels, and theme styles. ⑤ Save as a high-resolution image file.

[0152] (3) Model performance evaluation and validation

[0153] 3.1 Training Set Performance Evaluation

[0154] Calculate the prediction performance metrics of the PLSR model on the training set:

[0155] Root Mean Square Error (RMSE): A measure of the difference between predicted and actual values. The formula for calculation is:

[0156] : Number of samples in the training set (e.g., 875 samples), i: Sample index, from 1 to , : The actual observed AMH concentration value of the i-th sample (unit: pg / mL). : The model-predicted AMH concentration value (unit: pg / mL) for the i-th sample.

[0157] Coefficient of determination (R²): The proportion of variance explained by the model, calculated using the following formula: R 2 =1- : Number of samples in the training set (e.g., 875 samples), i: Sample index, from 1 to , : The actual observed AMH concentration value of the i-th sample (unit: pg / mL). : The model-predicted AMH concentration value (unit: pg / mL) for the i-th sample. : The average actual AMH concentration of all samples.

[0158] 3.2 Visualization of Training Set Prediction Results

[0159] Generate a comparison chart of observed and predicted values ​​in the training set (as attached). Figure 6 (As shown), the specific steps include: ① Calculate the predicted AMH concentration values ​​for the training set. ② Plot a scatter plot of observed values ​​(X-axis) and predicted values ​​(Y-axis). ③ Add a linear regression trend line and equation ④ Label the RMSE and R² values ​​on the graph. ⑤ Add a 45-degree reference line (ideal prediction line)

[0160] 3.3 Residual Analysis

[0161] Evaluate the distribution characteristics of the model residuals: Calculate the statistical characteristics (mean, standard deviation, range) of the residuals (observed values ​​- predicted values); Plot the residuals: a scatter plot with the predicted values ​​on the X-axis and the residuals on the Y-axis; Analyze whether the residuals exhibit systematic bias or heteroscedasticity.

[0162] 3.4 Validation Set Performance Evaluation

[0163] The generalization ability of the model was evaluated using an independent validation set (medium concentration group samples): Input the validation set spectral data into the trained PLSR model to obtain the predicted AMH concentration, calculate the RMSE and R² of the validation set, and plot a comparison between the observed and predicted values ​​of the validation set.

[0164] 3.5 Visualization of Validation Set Results

[0165] Generate a comparison chart of the observed and predicted values ​​in the validation set (as attached). Figure 7 (As shown), the specific steps include: ① Calculate the predicted AMH concentration values ​​for the validation set. ② Draw a scatter plot of the observed values ​​versus the predicted values. ③ Add a linear regression trend line. ④ Label the RMSE and R² values ​​of the validation set in the graph. ⑤ Use the same format and theme as the training set visualization.

[0166] By comparing the determination coefficients (R²) of the model training set and test set... 2 The performance of a model is evaluated using R² and root mean square error (RMSE). A higher R² value indicates a stronger ability of the model to interpret the data and a better fit. A lower RMSE value indicates higher prediction accuracy and a closer approximation of the true value.

[0167] like Figure 5-7 It was found that the R values ​​of the training and test sets of the PLSR model were... 2 The coefficients of determination (AMH) were 0.125 and 0.121, respectively, indicating poor model performance. The AMH coefficients were much less than 1, indicating a poor fit. The RMSE values ​​were 541.71 and 353.04, respectively, indicating poor model prediction accuracy and a large deviation between the predicted and actual values. These results show that the model cannot be used to distinguish dairy cow populations with different AMH concentrations, i.e., it cannot be used to predict the reproductive potential of dairy cows.

[0168] (3) RF model for predicting dairy cow productive lifespan based on serum anti-Müllerian hormone (AMH) concentration through milk spectroscopy

[0169] A method based on the concentration of anti-Müllerian hormone (AMH) in dairy cows and specific milk spectral data (1500-2000 cm⁻¹) -1 This study uses absorbance values ​​and a Random Forest (RF) model to predict the productive lifespan of dairy cows, aiming to provide a technical means for the accurate assessment of their reproductive potential. The steps include:

[0170] (1) Data preparation and preprocessing

[0171] 1.1 Data Sources and Structure

[0172] We collected spectral data of milk from healthy dairy cows and corresponding serum AMH concentration data to construct a dataset containing 601 samples.

[0173] The data consists of two main parts: ① AMH concentration data: a continuous numerical variable, used as the basis for regression targets or classification labels; ② Full-band spectral data: including absorbance within a specific wavelength range, preferably MWIR-2: 1698~1585 cm-¹, MWIR-LWIR: 1582~925 cm-¹, forming a p-dimensional characteristic space.

[0174] 1.2 Classification Label Definition

[0175] Based on a fixed threshold set according to AMH concentration, and using this as a criterion for reproductive potential, the samples were divided into three groups: "Low": AMH concentration ≤ 458.3 pg / mL; "Medium": 458.3 pg / mL < AMH concentration < 814.5 pg / mL; "High": AMH concentration ≥ 814.5 pg / mL;

[0176] Convert the classification labels into ordered factors to ensure that the model correctly identifies the order of categories.

[0177] 1.3 Dataset Partitioning

[0178] The dataset was divided into the following groups using stratified random sampling: Training set: 80%, used to build and optimize the random forest model; Test set: accounting for 20%, used to independently evaluate model performance and simulate real-world application scenarios; Set a random seed using set.seed(123) to ensure the results are reproducible.

[0179] 1.4 Data Preprocessing

[0180] ① Data type conversion: Convert spectral data into a numerical matrix to ensure computational compatibility.

[0181] Numerical type conversion: All elements in the original spectral data are forcibly converted to numerical types such as double-precision floating-point numbers to ensure that the data meets the calculation requirements of the random forest algorithm for the input data type and to avoid calculation errors caused by inconsistent data types.

[0182] Matrix structure organization: The transformed data is organized into a standard numerical matrix form, where each row of the matrix corresponds to an independent milk sample and each column corresponds to a specific wavelength or wavenumber variable, providing a standardized data structure for subsequent matrix operations and feature sampling.

[0183] ② Missing value handling: Identify and handle missing values ​​using the median imputation method.

[0184] Identification: The system scans the spectral data matrix to identify and locate all missing data points marked as NA or NaN.

[0185] Imputation: For each wavelength feature column containing missing values, median imputation is used. Specifically, the statistical median of all non-missing observations in the column is calculated, and this value is used to impute all missing positions in the column. This method effectively avoids the systematic bias that mean imputation might introduce when the data distribution is skewed, ensuring the robustness of the imputed values.

[0186] ③ Outlier handling: Detect and handle infinite values, replacing them with the median of finite values.

[0187] Detection: Traverse the spectral data matrix to detect outliers that represent infinitely large (Inf) or infinitely small (-Inf) values, indicating numerical overflow or underflow.

[0188] Correction: For feature columns with infinite values, first extract all finite numerical observations in the column, calculate the median, and then replace all infinite values ​​in the column with this median, thereby eliminating numerical distortion caused by instrument measurement limits or data reading and writing errors.

[0189] ④ Feature filtering: Remove zero-variance features and eliminate wavelength variables that do not provide information.

[0190] Variance calculation and evaluation: For the training set samples, calculate the variance of all sample observations under each wavelength feature column by column.

[0191] Zero-variance feature removal: Identify and remove feature columns with zero variance or statistically missing variance values. This step directly eliminates spectral bands that have no variation in the training population. Such bands do not provide discriminative information, and their removal helps alleviate the computational burden under high-dimensional data and improves the quality of effective random sampling in the feature space by random forests.

[0192] ⑤ Data standardization: The spectral matrix is ​​centered and standardized to unify the feature scale.

[0193] Centering: A centering operation is performed on each column of the spectral data matrix, which involves subtracting the arithmetic mean of all values ​​in the column from each of the original values ​​in that column, making the new mean of the column 0. This aims to eliminate baseline shifts that may exist at different wavelengths.

[0194] Standardization: Building upon centralization, a further standardization operation is performed on each column. This involves dividing each centered value by the standard deviation of the data in that column, making the new standard deviation of that column 1. This operation unifies the order of magnitude and discrete scale of all wavelength features.

[0195] (2) Construction of Random Forest Model

[0196] 2.1 Model Parameter Settings

[0197] Random forest models achieve predictions by integrating multiple decision trees. Key parameters include: ntree (number of trees): preferably 500-1000 trees to ensure model stability.

[0198] mtry (number of features randomly selected for each tree): usually set to the square root of the total number of features.

[0199] Minimum number of samples per node: controls the depth of tree growth and prevents overfitting.

[0200] 2.2 Model Training Framework

[0201] The model parameters were optimized using a 5-fold cross-validation strategy. The training set is randomly divided into 5 mutually exclusive subsets; The training was performed using four subsets and the validation was performed using one subset, repeated five times. Calculate the average performance index as the basis for parameter selection.

[0202] 2.3 Parameter Optimization Process

[0203] Optimize the hyperparameters of random forest using grid search: Test different mtry values ​​(preferably integers in the range of 1-50).

[0204] Evaluate the classification accuracy of different parameter combinations on the validation set.

[0205] Choose the parameter combination that optimizes the performance of the validation set.

[0206] 2.4 Model Training

[0207] Build a random forest model on the full training set using the optimal parameters: Algorithm: Random Forest (Classification / Regression); Integration strategies: majority voting (classification) or average prediction (regression);

[0208] Output: Prediction model, feature importance score, tree structure information.

[0209] (3) Model performance evaluation

[0210] 3.1 Training Set Performance Evaluation

[0211] Calculate the prediction performance of the random forest model on the training set: Confusion matrix analysis: Calculate accuracy, precision, recall, and F1 score; Cross-validation performance: Report the average performance metrics for 50% cross-validation; Overfit detection: Compare the performance differences between the training set and the validation set.

[0212] 3.2 Independent Validation of the Test Set

[0213] Evaluate the model's generalization ability on an independent test set: Input the spectral data of the test set into the trained random forest model; Obtain AMH concentration category prediction results and probability estimates.

[0214] Calculate the performance metrics on the test set and compare them with the results on the training set.

[0215] 3.3 Multi-class ROC curve analysis (generating multi-class ROC curves (see attached)) Figure 8 (As shown), the specific steps include: ① A one-vs-rest strategy is adopted to construct a binary classification ROC curve for each category; ② Calculate the true positive rate (TPR) and false positive rate (FPR) for each category. ③ Plot three ROC curves, corresponding to low, medium, and high AMH categories respectively; ④ Calculate and label the area under the curve (AUC) for each category; ⑤ Implement this using the `roc` and `plot.roc` functions from the `pROC` package; ⑥ Set the graph title, axis labels, legend, and color scheme; ⑦ Save as a high-resolution (300 DPI) PNG image.

[0216] 3.4 Predictive Probability Analysis

[0217] Generate a predicted probability distribution map (as attached) Figure 10(As shown), the specific steps include: ① Extract the predicted probabilities of the test set samples being divided into three categories. ② Use the tidyr::pivot_longer function to convert wide format data to long format. ③ Draw a faceted plot to show the predicted probability distribution for each category. ④ Add transparency (alpha=0.6) and color fill to enhance the visualization. ⑤ Set up a facet wrap to display the three categories separately. ⑥ Save as a high-quality image of 8×8 inches, 300 DPI.

[0218] 3.5 AMH Concentration-Prediction Probability Relationship Diagram

[0219] A graph showing the relationship between AMH concentration and predicted probability (as attached) Figure 9 (As shown), the specific steps include: ① Create a test results data frame, including measured AMH concentration, actual category, predicted category, and predicted probability for each category. ② Draw a scatter plot: with the measured AMH concentration as the X-axis and the predicted probability of being in the high AMH category as the Y-axis. ③ Color the scatter points according to their actual categories to enhance category differentiation. ④ Add a local regression smoothing curve (LOESS) to show the trend relationship. ⑤ Add vertical dashed lines to mark the classification thresholds (843.09 and 1481.7 pg / mL). ⑥ Set the graph title, axis labels, and legend. ⑦ Save as an 8×6 inch, 300 DPI image file.

[0220] (4) Model saving and application

[0221] 4.1 Model Persistence

[0222] Use the `saveRDS` function to save the trained random forest model as an RDS file, and save the test results data frame as a CSV file for easy subsequent analysis, recording model parameters, performance metrics, and important feature information.

[0223] 4.2 Model Deployment and Application

[0224] Integrating random forest models into spectral detection systems:

[0225] Load the saved model file into the embedded system or server to perform real-time spectral data preprocessing and feature extraction, call the model to predict reproductive potential and class determination, and output prediction results and confidence assessment.

[0226] To verify the performance of the RF model, ROC curves were plotted (e.g., Figure 8 The ROC curve, plotted with the false positive rate (FPR) on the horizontal axis and the true positive rate (TPR) on the vertical axis, showed an average AUC of 0.806, indicating that the model has good discriminative ability and can effectively distinguish groups with different reproductive potentials to a certain extent. Observing the predicted probability plot of the high reproductive potential group revealed (…). Figure 9 When the actual AMH concentration is below 458.3 pg / mL, the probability of the sample being predicted as having high reproductive potential is less than 50%. However, as the actual AMH concentration increases, the probability of being predicted as having high reproductive potential also increases. When the actual serum AMH concentration is above 814.5 pg / mL, the probability of being predicted as having high reproductive potential is mostly greater than 60%. Figure 10 The probability histogram shows that the classification of low and high reproductive potential groups conforms to this pattern, indicating that the model can accurately predict low and high reproductive potential dairy cow populations based on the spectrum.

[0227] Due to the imbalanced data in the sample database of this model, a comprehensive evaluation of model performance is necessary, especially when class imbalance occurs. AUC is one of the best choices. The four indicators in Table 2 typically require a classification threshold (e.g., a model output probability > 0.5 indicates a positive prediction), while AUC evaluates the overall performance of the model under different thresholds. Therefore, these four indicators should be considered as auxiliary indicators for evaluating model performance. This study further evaluated the model using accuracy (0.656), precision (0.674), recall (0.656), and F1 score (0.664). The results show that the model correctly predicted approximately 65.6% of the samples, and approximately 67.4% of the samples predicted to have an AMH concentration ≥ 814.5 pg / mL actually had this condition. The proportion of samples with an actual AMH concentration ≥ 814.5 pg / mL that were correctly identified was approximately 65.6%. The overall performance index F1 reached 0.664. These results demonstrate that the model has a relatively ideal effect in predicting the productive lifespan of dairy cows.

[0228] Table 2 Evaluation Indicators of the RF Model for Predicting AMH Concentration Based on Milk Spectrum

[0229] Group accuracy Accuracy Recall rate F1 average 0.656 0.674 0.656 0.664 low concentration \ 0.182 0.182 0.182 medium concentration \ 0.310 0.360 0.333 high concentration \ 0.848 0.807 0.827

[0230] Example 3: Model Prediction Validation

[0231] To verify the accuracy of the constructed model and to select the best model from the two superior models, SVM and RF, the superior model established earlier was used to predict 763 milk samples (not involved in modeling), including 316 low-parity culled cows (parity ≤2, poor reproductive performance) and 447 high-parity breeding cows (parity ≥3, good reproductive performance). The probability of these samples being predicted as high or low concentrations was observed.

[0232] Independent verification implementation method

[0233] (a) Validation Objectives and Design

[0234] Validation objective: To evaluate the prediction accuracy of established AMH concentration prediction models (including SVM, RF, etc.) on independent samples and to examine the model's ability to distinguish between differentiating factors of dairy cow reproductive performance (characterized by parity).

[0235] Validation design: A prospective validation strategy is adopted, using a sample set that is completely independent of the modeling process to ensure that the evaluation results are unbiased.

[0236] (ii) Data preparation for verification

[0237] 2.1 Data Sources and Sample Characteristics

[0238] Sample source: Collected from healthy dairy cows in the same dairy farm herd that were not involved in model training; Sample size: Mid-infrared spectral data of milk from 763 independent milk samples (MWIR-2: 1698~1585 cm⁻¹, MWIR-LWIR: 1582~925 cm⁻¹ absorbance values).

[0239] Sample characteristics:

[0240] Low reproductive performance group: 316 cattle, ≤2 parities (low parity culled cattle, poor reproductive performance);

[0241] High reproductive performance group: 447 cattle, parity ≥3 (high-parity cattle with good reproductive performance);

[0242] Data balance: The ratio of the two sample sizes is approximately 1:1.4, which is consistent with the actual herd structure.

[0243] 2.2 Data Types and Measurement Spectral data: The spectral data of each milk sample were measured using the same spectrometer (MWIR-2: 1698~1585cm-¹, MWIR-LWIR: 1582~925cm-¹ absorbance values). Parity data: Record the actual parity of each dairy cow as an objective indicator of reproductive performance; AMH concentration: Serum AMH concentration in validation samples was not measured; only model predictions were used.

[0244] 2.3 Data Preprocessing ① Format consistency check: Ensure that the spectral format of the validation data is completely consistent with that of the training data; ② Spectral data conversion: Converting raw spectral data into numerical matrices; ③ Missing value handling: Use the same median imputation method as in the training phase; ④ Outlier detection: Identify and process spectral outliers; ⑤ Feature alignment: Ensure that the wavelength variables in the validation data correspond one-to-one with the features trained on the model.

[0245] (III) Model Loading and Prediction

[0246] 3.1 Model Selection

[0247] The optimal prediction model established in the early stage was selected for validation, including:

[0248] Support Vector Machine (SVM) model: Loaded from the saved RDS file (AMH_SVM_Model.rds);

[0249] Random Forest (RF) Model: The model established in Example 2 is used and loaded from the saved RDS file (RF_AMH_Classification_Model.rds).

[0250] Selection criteria: The best model is selected based on the cross-validation performance during the modeling phase.

[0251] 3.2 Prediction Process

[0252] Step 1: Data Input

[0253] The spectral data matrix of the preprocessed 763 samples was input into the loaded prediction model.

[0254] Step 2: Probability Prediction The model outputs the predicted probability that each sample belongs to one of the three AMH concentration categories: Low AMH concentration probability: Predicted probability of AMH < 458.3 pg / mL; Probability of AMH concentration in the middle: The probability of 458.3 pg / mL < AMH concentration < 814.5 pg / mL is predicted; Probability of high AMH concentration: Predict the probability of AMH ≥ 814.5 pg / mL.

[0255] Step 3: Category Determination

[0256] The predicted AMH category for each sample is determined based on the principle of maximum probability: Prediction Category =

[0257] (iv) Results Analysis Methods

[0258] 4.1 Group Statistical Analysis

[0259] Calculated based on actual parity (lower parity group ≤2 vs higher parity group ≥3): Average prediction probability statistics: The lower parity group was predicted to have a high AMH as a mean probability. The average probability of high parity being predicted as high AMH; Differences between the two groups in the predicted probabilities of low, medium, and high AMH.

[0260] Category prediction ratio statistics: The proportion of individuals predicted to have high AMH in the lower parity group; The proportion of individuals predicted to have high AMH levels in the higher parity group; Predict the correlation between category distribution and parity.

[0261] 4.2 Visualization Analysis (Method for Generating Attached Figures)

[0262] Generate a series of verification result graphs:

[0263] Appendix Figure 11 12: Predicted Proportion Bar Chart

[0264] Generation method: Grouped stacked bar charts + percentage labels

[0265] Content displayed: The proportion of each parity group predicted to be in different AMH categories

[0266] Technical details: Adding numeric labels using geom_bar and geom_text

[0267] 4.3 Statistical Tests

[0268] To quantify the consistency between the model's predictions and the actual reproductive performance of dairy cows, an independent samples t-test was used to compare the differences in AMH prediction probabilities across different parity groups. Specifically, this included:

[0269] The probability test for predicting low AMH levels compares the difference in the probability of being predicted as having low AMH concentration between the low parity group (parity ≤ 2) and the high parity group (parity ≥ 3). The hypothesis is: Null hypothesis H0: There is no significant difference in the probability of low AMH between the low parity group and the high parity group. H0: μ low parity, low AMH prediction probability = μ high parity, low AMH prediction probability; Alternative hypothesis H1: The two groups have significant differences in the predicted probability of low AMH; Purpose of the test: To verify whether the low reproductive performance group (parity ≤2) is more likely to be predicted to have low AMH concentration.

[0270] The probability test for high AMH prediction compares the difference in the probability of being predicted to have high AMH concentration between two groups. The hypothesis is: Null hypothesis H0: There is no significant difference in the probability of high AMH between the low parity group and the high parity group. H0: μ low parity, high AMH prediction probability = μ high parity, high AMH prediction probability; Alternative hypothesis H1: The two groups differ significantly in the high AMH prediction probability; Purpose of the test: To verify whether the high reproductive performance group (parity ≥3) is more likely to be predicted to have high AMH concentration.

[0271] The test results showed that there were extremely significant differences between the two groups in both the low AMH prediction probability (t test, P < 0.001) and the high AMH prediction probability (t test, P < 0.001), indicating that the model prediction results were highly consistent with the reproductive performance status of dairy cows, and verifying the biological rationality and prediction accuracy of the model.

[0272] Depend on Figure 11 It can be seen that in the SVM model, the probability of high parity samples being predicted as having high AMH concentration is 48.8%, while the probability of low parity samples being predicted as having low AMH concentration is only 17.4%, indicating poor model validation. Figure 12 In the RF model, high parity samples were predicted to have high AMH concentrations with a probability of 70.0%, while low parity samples were predicted to have low AMH concentrations with a probability of 99.1%. Previous studies have found that the RF model has high accuracy in predicting samples with low and high AMH concentrations, which is consistent with the model's prediction pattern. The model validation results are quite ideal. Therefore, the RF model was selected as the model for predicting serum AMH concentrations in dairy cows using infrared spectroscopy in milk.

Claims

1. A method for predicting the productive lifespan of dairy cows using infrared spectroscopy in milk based on the concentration of anti-Müllerian hormone in bovine serum, characterized in that, It includes: AMH concentration in bovine serum was detected using an AMH quantitative detection kit. Bovine populations with AMH concentrations below 458.3 pg / mL were classified as having low reproductive potential, while those above 814.5 pg / mL were classified as having high reproductive potential. The bovine population was further divided into low, medium, and high concentration groups to represent low, medium, and high reproductive potential, respectively. The characteristic wavelengths of the corresponding milk samples were analyzed using full-band infrared spectra. A random forest model was constructed based on the infrared spectral data of the three AMH concentration ranges and their corresponding characteristic wavelengths to obtain an RF model for predicting bovine productive lifespan by measuring the characteristic wavelengths of the milk. The characteristic wavelength infrared spectral data of the milk samples from the tested bovine populations were then imported into the constructed RF model to predict the bovine productive lifespan.

2. The method as described in claim 1, characterized in that, The division into low, medium, and high concentration groups is based on AMH concentration ≤ 458.3 pg / mL, 458.3 pg / mL < AMH concentration < 814.5 pg / mL, and AMH concentration ≥ 814.5 pg / mL.

3. The method as described in claim 1, characterized in that, The characteristic wavelength range of the infrared spectrum of milk is 926-5012 cm⁻¹. −1 It includes five characteristic intervals: SWIR: 5010~3673 cm⁻¹, SWIR-MWIR: 3669~3052 cm⁻¹, MWIR-1: 3048~1701 cm⁻¹, MWIR-2: 1698~1585 cm⁻¹, and MWIR-LWIR: 1582~925 cm⁻¹.

4. The method as described in claim 1, characterized in that, The characteristic wavelengths of the infrared spectrum of milk are MWIR-2: 1698~1585 cm-¹, and MWIR-LWIR: 1582~925 cm-¹.

5. The method as described in claim 1, characterized in that, The infrared spectral data for the characteristic wavelength is the absorbance value within the characteristic wavelength range.

6. The method as described in claim 1, characterized in that, When the predicted AMH concentration is ≤458.3 pg / mL, the predicted parity is ≤2, indicating low parity and poor reproductive performance; When the predicted AMH concentration is ≥814.5 pg / mL, the predicted parity is ≥3, indicating high-parity breeding cattle and good reproductive performance.

7. The application of milk infrared spectroscopy for predicting dairy cow reproductive performance, characterized in that, The characteristic wavelength of the infrared spectrum is in the range of 1500–2000 cm⁻¹. −1 Section.

8. The application according to claim 7, characterized in that, The infrared spectrum is the absorbance value at characteristic wavelengths of 1698~1585 cm⁻¹ and 1582~925 cm⁻¹.