Method and system for evaluating nutrient composition of breast milk by using multivariate statistical method and application
Through multi-statistical methods and factor analysis of principal components, the evaluation of nutritional components of breast milk was simplified, and the problem of difficult to describe the dynamic changes of nutritional components of breast milk was solved. An accurate nutritional component evaluation system was established to be applied to the evaluation of breast milk and infant formula foods.
Patent Information
- Application Number
- CN202510237834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to describe the dynamic changes in breast milk nutritional components in comprehensively and accurately, especially when evaluating breast milk as a formula for infants and young children, the influencing factors are complex and the correlation increases the difficulty of analysis.
The macronutrients and micronutrients of breast milk were simplified by the multi-statistic method of principal components. Through factor analysis and principal component analysis, an evaluation model was constructed, and the main indicators were selected to evaluate the nutritional components of breast milk, including 22 nutrients such as phosphorus and solid substances, and data processing and analysis were used using SPSS software.
On the basis of retaining the original information, the evaluation of breast milk nutritional ingredients is simplified, and a more accurate evaluation system for breast milk nutritional ingredients is provided, providing a scientific basis for the design of infant formula foods, and improving the accuracy and efficiency of evaluation.
Smart Images

Figure BDA0005294248670000101 
Figure BDA0005294248670000111 
Figure BDA0005294248670000112
Abstract
Description
Technical Field
[0001] The present invention relates to the field of nutritional component evaluation, and particularly to a method, a system and an application for evaluating the nutritional component composition of breast milk by using multivariate statistical methods. Background Art
[0002] Breast milk, also known as human milk, is the juice produced by the breasts of postpartum women and is used to feed infants. It is known as the "white blood". Breast milk contains carbohydrates, proteins, fats, vitamins, minerals, bioactive substances, immune factors, etc. Breast milk can not only provide a nutritional source for the growth and development of infants and young children, but also play an important role in the construction of their own immune systems by infants and young children. Nutritionally, the main components of breast milk are macronutrients and micronutrients. Among them, macronutrients include three energy-producing nutrients: protein, fat, and carbohydrates; micronutrients include minerals and vitamins. For infants and young children, breast milk is the best food with comprehensive nutrition. For example, protein is the material basis for the growth and development of infants. Every cell and all important active substances in the infant's body require the participation of protein. Fat is the main source of energy for infants, provides essential fatty acids for growth and development, and is an important substance for promoting the development of the nervous system. Carbohydrates can provide energy and immune substances for newborns. Vitamins participate in the composition of holoenzymes and catalyze and regulate the metabolic activities of the body.
[0003] The World Health Organization and the United Nations Children's Fund have proposed that exclusive breastfeeding for the first 6 months after birth is the best way to feed newborns, and then continue breastfeeding and add appropriate complementary foods until the age of 2 or longer. As the most ideal and preferred food for newborns, the quality of breast milk has a direct impact on the growth and development of newborns. Therefore, domestic and foreign pediatricians and child health experts have never stopped researching the nutritional components of breast milk. As is well known, the composition of breast milk nutrients has differences in genetics, race, region, and different lactation stages. In addition to being affected by the above factors, it is significantly related to certain characteristics of lactating mothers, such as body mass index, mode of delivery, etc. Referring to domestic and foreign research, the relevant factors affecting the changes in breast milk nutrients are divided into the following four aspects: physiological factors, general demographic factors, exogenous factors, and endogenous factors. These factors act together, and some factors are also correlated, resulting in the complexity of the composition of breast milk nutrients and bringing inconvenience to the analysis and evaluation of breast milk components.
[0004] Currently, domestic and foreign researchers mainly collect breast milk samples through cross-sectional or longitudinal cohorts, and use detection equipment such as infrared breast milk composition analyzers and liquid chromatography-mass spectrometry to detect the macronutrients and micronutrients in breast milk, and evaluate breast milk through the mean (or median) of a single component. However, breast milk is a biological fluid with a complex structure and diverse compositions. The traditional nutritional evaluation of a single component cannot comprehensively and accurately describe the dynamic changes of breast milk nutrients, and is even less conducive to the development of infant formula foods with breast milk as the "gold standard".
[0005] Especially when the mother is unable to breastfeed sufficiently, newborns need to use infant formula powder designed based on the composition of breast milk as food. Therefore, the evaluation of breast milk nutrient components is particularly important for formula design. However, there are many factors affecting breast milk nutrient components, and some of these factors are correlated to a certain extent, increasing the complexity of breast milk nutrient component analysis and bringing inconvenience to breast milk nutritional evaluation. Summary of the Invention
[0006] Purpose of the Invention
[0007] To overcome the above deficiencies, the purpose of the present invention is to provide a method, system and application for evaluating the composition of breast milk nutrients using multivariate statistical methods. The present invention simplifies the macronutrients and micronutrients in breast milk by using the principal component multivariate statistical method, and selects the main indicators from among many quality evaluation indicators to evaluate the nutrient components of breast milk, laying a certain foundation for the establishment of a breast milk nutrient component evaluation system.
[0008] Solution
[0009] To achieve the purpose of the present invention, the technical solution adopted by the present invention is as follows:
[0010] In the first aspect, the present invention provides a method for evaluating the composition of breast milk nutrients using multivariate statistical methods. The method includes the following steps:
[0011] Obtain the content of nutrient components in breast milk or formula food to be evaluated as the original data variables; perform normalization processing on the original data variables to eliminate the dimension difference, and perform centering processing;
[0012] Input the centered data into a pre-constructed evaluation model to obtain the score output by the evaluation model, and compare the score with a preset median. The smaller the difference, the closer it is to breast milk;
[0013] Among them, the evaluation model includes four common factors obtained through principal component analysis and a scoring model obtained through the load and influence degree in factor analysis.
[0014] Further, each nutrient component includes one or several of phosphorus, solids, carbohydrates, potassium, magnesium, sodium, iron, riboflavin, pantothenic acid, fat, copper, niacin, zinc, vitamin B6, thiamine, vitamin E, beta-carotene, vitamin A, true protein, protein, calcium, and ascorbic acid.
[0015] Further, the construction method of the evaluation model is based on the factor analysis method in mathematical statistics. Optionally, it specifically includes:
[0016] Suppose there are n samples, and each sample observes p indicators with correlations. There is a strong correlation among these p indicators (the reason for requiring a strong correlation among the p indicators is clear. Only with a strong correlation can "common" factors be extracted from the original variables). To eliminate the influence caused by the differences in the observed variable dimensions and different orders of magnitude, the sample observation data is standardized so that the mean of the standardized variable is 0 and the variance is 1;
[0017] Optionally, the data standardization method includes: eliminating the dimension differences between the original samples through normalization; using the Pareto scaling method to centralize the normalized data. The Pareto scaling method includes centering on the mean and dividing by the square root of the standard deviation of each variable;
[0018] Check whether there is a correlation between the variables in the standardized data. Optionally, the methods used include: sampling adequacy test ((KMO test)) and Bartlett test (Bartlett's Test). Optionally, if the KMO statistic is greater than 0.6 and the probability value of the Bartlett sphericity test is less than the significance level, it is considered to have a correlation. Optionally, select 22 nutrient component indicators;
[0019] Perform factor analysis on the standardized and correlated data to obtain common factors, and obtain the factor load matrix, component matrix table, and data table for factor weight analysis of the indicators of each nutrient component;
[0020] For convenience, both the original variable and the standardized variable vector are represented by X, and use F1, F2,..., F m(m < p) represents the standardized common factor. Factor analysis is a multivariate statistical analysis method that reduces multivariate variables to a few common factors based on the correlation between variables and then analyzes and processes them. Its basic idea is to decompose the original variables into two parts: one part is the linear combination of common factors, which concentratedly represents most of the information in the original variables; the other part is the special factor unrelated to the common factors, which reflects the gap between the linear combination of common factors and the original variables. Let both the original variables and the standardized variable vector be represented by X. The factor analysis model for the p-dimensional variable X = [X1,..., Xi,..., Xp] is as follows:
[0021]
[0022] Among them, the vector F is the extracted common factor vector, representing m (m < p) mutually independent common influencing factors that are objectively existent but not directly observable in the original variables; the vector A is the factor loading matrix, indicating the loading of the variable X matrix on the common factor matrix F, reflecting the correlation coefficient between the two. The larger the absolute value, the higher the correlation;
[0023] Therefore, the key to establishing a factor analysis model for the multivariate variable X vector lies in solving the factor loading matrix A and the common factor vector F. In the factor model, the number of common factors is less than the number of original variables, and the common factors are unobservable latent variables. The loading matrix A is irreversible, so the common factors cannot be directly obtained. One way to solve this problem is to use the idea of regression to find the estimated values of the linear combination coefficients, that is, to establish the following regression equation with the common factor F i as the dependent variable and the standardized original variable X i as the independent variable. Optionally, all variables X can be expressed as a linear combination of the common factor F. The single-factor score calculation formula is as follows:
[0024] F i = a 1i X1 + a 2i X2 + … + a pi X p , p = 1, …, 22 (1)
[0025] In formula (1), the vector F is the extracted common factor vector, representing m mutually independent common influencing factors that are objectively existent but not directly observable in the original variables. Optionally, m < 22; a 1i , a 2i… a pi is the factor loading matrix of the vector A, indicating the loading of the variable X matrix on the common factor matrix F; then m principal components with eigenvalues greater than 1 are extracted from the factor weight analysis matrix, and according to the variance contribution rate w of the principal components, the calculation formula for the comprehensive score is obtained;
[0026] F总 = w1F1 + w2F2 + … + w m F m (2)
[0027] In formula (2), w1 …… w m represent the variance contribution rates of each common factor, and F1, F2, ……, F m represent the standardized common factors. Optionally, m < p;
[0028] Evaluate the nutritional components of breast milk or formula foods according to the calculation formula of the comprehensive score.
[0029] Perform factor analysis on the centralized data. Factor analysis is based on the idea of dimensionality reduction. Without losing or losing as little as possible the original data information, numerous complex variables are aggregated into a few independent common factors. These common factors can reflect the main information of the original numerous variables. While reducing the number of variables, it also reflects the internal relationship between the variables. Usually, factor analysis has three functions: one is for factor dimensionality reduction, the second is to calculate factor weights, and the third is to calculate the weighted comprehensive score of the factors. Quantitatively evaluate the rankings of multiple breast milk nutritional levels or the weights of each index according to 22 nutritional component indicators including protein, true protein, fat, etc.
[0030] An operable method is as follows: Through the data statistical software spss25, select the [Factor Analysis] method; among them, [Factor Analysis] requires the input data to be the [quantitative] independent variable X (the number of variables ≥ 2). Then select the number of principal components and the factor rotation method (Note: In factor analysis, it is inclined to describe the correlation relationship between the original variables. Therefore, generally, the number of principal components selected in factor analysis is the same as the number of independent variables X, and the eigenvalue selection is based on a set threshold. The number of principal components corresponding to the value greater than this threshold is selected as the number of principal components, with the default being 1). Click [Start Analysis], and three data tables of the factor load matrix, component matrix table, and factor weight analysis of 22 breast milk nutritional components will be obtained.
[0031] In a second aspect, a method for evaluating the composition of breast milk nutritional components using multivariate statistics is provided. The method includes the following steps:
[0032] Obtain the content of the nutritional components in the breast milk or formula food to be evaluated as the original data variables; perform normalization processing on the original data variables to eliminate the dimension difference, and perform centralized processing;
[0033] Input the centralized data into a pre-constructed evaluation model to obtain the score output by the evaluation model, and compare the score with a preset median. The smaller the difference, the closer it is to breast milk;
[0034] Among them, the evaluation model includes the following scoring formula:
[0035] F1 = 0.025×Z 脂肪 - 0.103×Z 蛋白质 + 0.087×Z 碳水化合物 + 0.061×Z 固形物 - 0.106×Z 真蛋白 + 0.014×Z 硫氨酸 + 0.142×Z 核黄素 + 0.148×Z 抗坏血酸 + 0.157×Z 烟酸 + 0.111×Z 泛酸 + 0.103×Z 维生素B6 - 0.064×Z 维生素A - 0.096×Z 维生素E - 0.143×Z β-类胡罗卜 + 0.106×Z 钠 + 0.066×Z 镁 + 0.075×Z 磷 + 0.099×Z 钾 + 0.094×Z 钙 + 0.044×Z 铁 + 0.025×Z 铜 - 0.18×Z 锌 (I)
[0036] F2 = 0.127×Z 脂肪 - 0.056×Z 蛋白质 - 0.015×Z 碳水化合物 + 0.017×Z 固形物 - 0.057×Z 真蛋白 + 0.101×Z 硫氨酸 - 0.094×Z 核黄素 - 0.063×Z 抗坏血酸 - 0.114×Z 烟酸 - 0.023×Z 泛酸 - 0.055×Z 维生素B6 + 0.304×Z 维生素A + 0.407×Z 维生素E + 0.395×Z β-类胡罗卜 - 0.02×Z 钠 + 0.014×Z 镁 + 0.006×Z 磷 - 0.024×Z 钾 - 0.123×Z 钙 + 0.027Z 铁 + 0.039×Z 铜 + 0.065×Z 锌(II)
[0037] F3 = -0.031×Z 脂肪 +0.481×Z 蛋白质 +0.032×Z 碳水化合物 +0.058×Z 固形物 +0.487×Z 真蛋白 -0.008×Z 硫氨酸 -0.045×Z 核黄素 -0.162×Z 抗坏血酸 -0.082×Z 烟酸 -0.063×Z 泛酸 -0.048×Z 维生素B6 -0.083×Z 维生素A -0.09×Z 维生素E +0.029×Z β-类胡罗卜 -0.011×Z 钠 +0.046×Z 镁 +0.035×Z 磷 +0.008×Z 钾 -0.128×Z 钙 +0.054×Z 铁 +0.058×Z 铜 +0.197×Z 锌 (III)
[0038] F4 = 0.083×Z 脂肪 -0.044×Z 蛋白质 +0.013×Z 碳水化合物 +0.035×Z 固形物 -0.042×Z 真蛋白 -0.078×Z 硫氨酸 -0.001×Z 核黄素 -0.532×Z 抗坏血酸 -0.089×Z 烟酸 +0.024×Z 泛酸 +0.11×Z 维生素B6 +0.066×Z 维生素A -0.095×Z 维生素E -0.163×Z β-类胡罗卜 -0.041×Z 钠 +0.011×Z 镁 +0.029×Z 磷 +0.013×Z 钾 +0.663×Z 钙 +0.143×Z 铁 +0.117×Z 铜 +0.145×Z 锌 (IV)
[0039] The four factors are weighted and added according to their respective variance contribution rates to obtain a comprehensive score, and its calculation formula is:
[0040] F 总 = 0.559×F1 + 0.206×F2 + 0.163×F3 + 0.072×F4 (V)
[0041] Among them, Z x1 ~Z x22 respectively represent the variable data after standardization of phosphorus, solids, carbohydrates, potassium, magnesium, sodium, iron, riboflavin, pantothenic acid, fat, copper, niacin, zinc, vitamin B6, thiamine, vitamin E, β-carotene, vitamin A, true protein, protein, calcium, and ascorbic acid.
[0042] In the first aspect or the second aspect, the standardization processing method includes: using logarithmic transformation to eliminate the dimensional differences between the original samples for the original data, and using the Pareto scaling method to centralize the data.
[0043] Optionally, the centralization processing uses the Pareto scaling method. Optionally, the Pareto scaling method includes centering on the mean and dividing by the square root of the standard deviation of each variable.
[0044] In the first aspect or the second aspect, the preset factor score range is (-1.06, 1.72), where the median of the factor scores is (-0.84, 0.4). Optionally, the median is -0.218. Optionally, the upper quartile is (-0.01, 1.23), and the upper quartile is 0.61. Optionally, the lower quartile is (-1.13, 0.12), and the lower quartile is -0.503; optionally, the preset median is determined according to the comprehensive scores of each data sample by factor analysis.
[0045] In the first aspect or the second aspect, the breast milk components are evaluated according to the preset median. If it is higher than the median, it is considered that the breast milk nutrient components are sufficient; if it is lower than the median, it is considered that the breast milk nutrient components are lacking, and the mother should take appropriate dietary measures for adjustment.
[0046] In the third aspect, a system for evaluating the nutrient components in breast milk or formula food is provided, which is applied to the method described in the first aspect, and includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, when the processor executes the program, it implements the steps of the method described in the first aspect.
[0047] In the fourth aspect, an application of the method described in the first aspect or the system described in the second aspect in the preparation of products for evaluating the nutrient components in breast milk or formula food is provided.
[0048] Advantageous Effects
[0049] The present invention simplifies the macronutrients and micronutrients in breast milk by using the principal component multivariate statistical method, selects the main indicators from numerous quality evaluation indicators to evaluate the nutritional components of breast milk or formula food, and lays a certain foundation for the establishment of the breast milk nutritional component evaluation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. The special term "exemplary" herein means "serving as an example, an embodiment or an illustration". Any embodiment illustrated as "exemplary" herein need not be construed as being superior to or better than other embodiments.
[0051] Figure 1 It is a heat map of the correlation of breast milk nutritional components drawn according to the original data in the principal component analysis of the present invention.
[0052] Figure 2 It is an analysis diagram of breast milk nutritional components after data centering processing in the principal component analysis of the present invention.
[0053] Figure 3 It is a heat map of the correlation of breast milk nutritional components after data centering processing in the principal component analysis of the present invention.
[0054] Figure 4 It is the explained variance of total variance after rotation in the factor analysis of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0056] In addition, to better illustrate the present invention, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present invention can be implemented without some specific details. In some embodiments, details of raw materials, components, methods, means, etc. well-known to those skilled in the art are not described in detail in order to highlight the gist of the present invention.
[0057] Unless otherwise clearly stated, in the whole specification and claims, the term "comprise" or its variations such as "comprises" or "including" etc. will be understood to include the stated elements or components, without excluding other elements or other components.
[0058] Factor analysis is a statistical method that uses a few factors to reflect most of the information of numerous original indicators. Through factor analysis, a few factors can be found to replace the original variables for regression analysis, clustering analysis, discriminant analysis, etc. Principal component analysis is a commonly used data reduction method, that is, to find several linear combination functions (i.e., principal components) that can explain the original variables. On the one hand, these principal components must be able to independently and non-overlappingly retain the information of the original variables. More importantly, they can replace the original more variables with fewer principal components to achieve the purpose of simplification. That is to say, principal component analysis transforms the original variables into a linear relationship to explain the variability of the original variables and determines the principal components to be retained to achieve the purpose of streamlining. The mathematical models of two commonly used dimensionality reduction statistical methods, factor analysis and principal component analysis, also differ: Principal component analysis represents the principal components as linear combinations of the original variables:
[0059] F i =a 1i X1+a 2i X2+…+a pi X p ,p=1,…,22 (1)
[0060] In formula (1), the vector F is the extracted common factor vector, representing m (m < 22) mutually independent common influencing factors that are objectively existent but not directly observable in the original variables; the vector A(a 1i ,a 2i ...a pi ) is the factor loading matrix, indicating the loading of the variable X matrix on the common factor matrix F and reflecting the correlation coefficient between the two. The larger its absolute value, the higher the correlation; both the original variables and the standardized variable vectors are represented by X. Then, m principal components with eigenvalues greater than 1 are extracted from the factor weight analysis matrix. According to the variance contribution rate w of the principal components, the calculation formula for the comprehensive score is obtained:
[0061] F 总 =w1F1+w2F2+…+w m F m (2)
[0062] In formula (2), w1 represents the variance contribution rate of each common factor, and F1, F2, …, F m (m < p) represents the standardized common factors.
[0063] Before performing factor analysis, it is necessary to first conduct principal component analysis on the samples. The eigenvector matrix is obtained through the unrotated factor loading matrix, and then the calculation formula for the principal components is derived. Through factor analysis and principal component analysis, on the basis of maximizing the retention of the original information, multiple original variables can be transformed into a few uncorrelated comprehensive indicators, thus achieving dimensionality reduction.
[0064] Example 1
[0065] A total of 198 breast milk samples were selected from four cities in China, namely Beijing, Tangshan, Luoyang, and Liuyang. These included colostrum (n = 11) from 0 - 5 days, transitional milk (n = 21) from 12 - 14 days, mature milk at 1 month (n = 61), mature milk at 3 months (n = 73), and mature milk at 6 months (n = 33). The samples used in this study have been approved by the Ethics Committee of Beijing Ditan Hospital, Capital Medical University (#2015 - 027 - 01).
[0066] 1. The detection of macronutrients in breast milk (fat, protein, true protein, carbohydrates, solids, and energy) was determined by infrared transmission spectroscopy. The instrument used was the MIRIS HMA breast milk analyzer (Miris AB, Sweden). After taking the breast milk samples out of the -80 °C refrigerator, they were homogenized using an ultrasonic process before testing. After preheating at 37 °C, 5 mL was injected into the breast milk analyzer for detection.
[0067] 2. For the detection of water-soluble vitamins, first, the samples were thawed at room temperature. Secondly, 4 mL of human milk was placed in a 40 mL plastic centrifuge tube and then weighed to an accuracy of 0.0001 g. The pH value was adjusted to 1.7 with 1 M HCl and then kept in the dark for 2 minutes. Then the pH value was adjusted to 4.7 with 1 M NaOH and left in the dark for 2 minutes. After making up the volume to 10 g with water, the samples were centrifuged at 10000 g for 20 minutes at 4 °C. Finally, the supernatant was removed and filtered through a 0.22 μm syringe filter. All samples were prepared in triplicate. In all the following steps, aluminum foil was used to avoid exposing the samples to light. Six water-soluble vitamins were analyzed using an UltiMate 3000 (Accela)-Q Exactive mass spectrometer (Thermo Fisher Scientific, Waltham, MA, USA) and an ACQUITY UPLC HSS T3 chromatographic column (HSS T3, 50 mm × 2.1 mm inner diameter, 1.8 μm) for separation. The mobile phase consisted of 10 mM ammonium formate and 0.1% formic acid as solvent A and acetonitrile as solvent B. The total running time for each injection was 8 minutes.
[0068] 3. Detection of fat-soluble vitamins: Accurately weigh 5 g (accurate to 0.0001 g) of breast milk sample and place it in a 50 mL round-bottom centrifuge tube. Add 5 mL of ethanol aqueous solution of vitamin C and mix well. Then add 10 mL of potassium hydroxide solution, mix well, place it in a water bath at 55 °C, saponify for 45 min, and then take it out and cool to room temperature. Add 10 ml of ethanol to the saponified solution, shake well and centrifuge for 5 min (6000 r / min). First, add 5 ml of methanol and 5 ml of water to the special solid-phase extraction column for vitamins for activation. After activation, add the supernatant of the saponified solution. After the supernatant completely flows out, then wash it with 10 ml (ethanol +
[0069] water, 1:1) solution, discard all the effluent, and dry it by suction for 20 min. Finally, elute with 2 ml of acetone and 8 ml of ethyl acetate, collect all the effluent, and blow the effluent to dryness with nitrogen at 40 °C. Dilute to volume with 1 ml of methanol, pass the collected solution through a 0.22 μm microporous filter membrane for high performance liquid chromatography detection and analysis.
[0070] 4. Detection of minerals: Digest the breast milk sample (250 - 500 mg) with 4 mL of 14 M nitric acid + 2 mL of 30% hydrogen peroxide. Determine the mineral concentration by inductively coupled plasma mass spectrometry (ICP-MS, ELANDRC II, Perkin Elmer Sciex, Shelton). The sample is subjected to microwave-assisted acid digestion in a closed system. Analyze using an ICP spectrometer equipped with a reaction cell (DRC-ICP-MS, ELANDRCII, PerkinElmer, SCIEX, Norwalk, CT, USA), and the reaction cell operates using ultra-pure argon (99.999%). The sample introduction system consists of a quartz cyclone spray chamber and a nebulizer connected to a continuous pump (adjusted to 20 rpm) through a tube to the ICP-MS. The ICP-MS operates using a sampling cone and a Pt separator from PerkinElmer.
[0071] 5. Data processing. For the large number of physical and chemical indicators of breast milk components in the queue, these indicators are either closely related or relatively independent. Through multivariate statistical methods, these indicators can be simplified while retaining most of the original information. Principal component analysis and factor analysis were performed on 22 physical and chemical indicators of breast milk at different lactation stages to simplify the quality evaluation indicators of breast milk components on the basis of maximizing the retention of the original indicator information, providing a certain theoretical basis for the formulation of the breast milk component quality evaluation system and the improvement research of simulated breast milk infant formula. Using the 22 physical and chemical indicators of the measured breast milk components as the original data, principal component analysis and factor analysis were performed using SPSS 25.0 software. Then, the original data was normalized using the Metaboanalyst online analysis platform, and principal component analysis and factor analysis were performed again using SPSS 25.0.
[0072] 1. Principal Component Analysis
[0073] The three most important statistics in principal component analysis are the eigenvalue, the variance contribution rate of the principal component, and the cumulative contribution rate. The eigenvalue (greater than 1) is usually considered an indicator of the explanatory power of the principal component, representing the number of average original variables that can be explained after introducing this principal component. The variance contribution rate of the principal component represents that the larger the original information content of the X1, X2,..., X it carries. The cumulative contribution rate. Generally speaking, when this indicator reaches 85%, it indicates that these principal components contain the main information possessed by all measurement indicators. The main role of the principal component is to not only reduce the number of variables but also facilitate the analysis and research of practical problems. The principal component is applicable to data with strong correlations between variables, that is, most correlation coefficients are greater than 0.3. p The larger the original information content. The cumulative contribution rate. Generally speaking, when this indicator reaches 85%, it indicates that these principal components contain the main information possessed by all measurement indicators. The main role of the principal component is to not only reduce the number of variables but also facilitate the analysis and research of practical problems. The principal component is applicable to data with strong correlations between variables, that is, most correlation coefficients are greater than 0.3.
[0074] The correlation heat map of breast milk nutritional components drawn based on the original data is as Figure 1 shown. The correlation coefficients of most variables in the original data are less than 0.3, that is, the correlation of the original data is weak. After applying principal component analysis, it cannot play a good role in dimensionality reduction, and the ability of each obtained principal component to condense the original variables is not much different, and the effect obtained by applying principal component analysis is not ideal.
[0075] Based on the above analysis, there are differences in the original dimensions of the original data. Usually, the multiple difference between data samples is above 10 3 Therefore, logarithmic transformation and data scaling were performed on the original data. Sample-specific standardization makes a general adjustment to the differences between samples through row-by-row normalization; data transformation uses logarithmic transformation to eliminate the dimensional differences between the original samples; Pareto scaling (centered on the mean and divided by the square root of the standard deviation of each variable) is used to make the data eigenvectors comparable.
[0076] There are a total of 198 breast milk samples, and 22 indicators are observed for each sample. There is a strong correlation among these 22 indicators. At the same time, in order to eliminate the influence caused by the differences in observation dimensions and orders of magnitude, the sample observation data is standardized so that the mean of the standardized variables is 0 and the variance is 1.
[0077] The analysis of the nutritional components of breast milk after being processed by the data center is as Figure 2 shown. After centralizing the 22 breast milk nutritional component indicators, it can be seen that the data distribution is approximately a normal distribution, eliminating the differences in the original data dimensions.
[0078] Perform a sufficiency verification on the centralized data, that is, test whether there is a correlation between variables, so as to judge whether it is suitable for factor analysis. The methods used are: sampling adequacy test ((KMO test)) and Bartlett test (Bartlett's Test). The main things to look at are the KMO statistic and the probability P value of Bartlett. If the KMO statistic is greater than 0.6 and the probability value of the Bartlett sphericity test is less than the significance level (0.01), it means that there is a correlation between the analyzed variables, and if there is a correlation, it is suitable for factor analysis. The results of the correlation test are as Figure 3 shown. After the data is centralized, the correlation coefficients of most variables are greater than 0.3, that is, the data has a strong correlation. After applying principal component analysis, it can play a good role in dimension reduction. The abilities of the obtained principal components to condense the original variables vary greatly. Therefore, all 22 breast milk nutritional components in this analysis are applicable to principal component analysis.
[0079] The eigenvalues of the 22 breast milk nutritional components are obtained through the factor analysis method of Spss software (the matrix of eigenvectors, original vector = eigenvalue · eigenmatrix). As shown in Table 1, four principal components are extracted according to the principle that the eigenvalue is greater than 1. These four principal components together explain 75.797% of the total variance of the original variables (the total explanatory degree of the cumulative variance). Generally speaking, these four principal components can reflect 75.797% of the information of the original variables. The eigenvalues of these four factors are: 12.221, 1.842, 1.489, and 1.124 respectively, and the variance contribution rates of each factor are 55.551%, 8.373%, 6.766%, and 5.107%. Therefore, in this analysis, it is selected that extracting the first four principal components can reflect 75.797% of the information of the original variables.
[0080] Table 1 Total variance explanation
[0081]
[0082]
[0083] In Table 1, 1 - 22 represent the matrix of eigenvectors, original vector = eigenvalue · eigenmatrix.
[0084] 2. Factor Analysis
[0085] Since factor analysis studies the internal dependence relationships of the correlation matrix (or covariance matrix) among multiple variables, aiming to find a few random variables that can comprehensively represent the main information of all variables, and these random variables cannot be directly measured, factor analysis can be regarded as a continuation of principal component analysis. The factors in factor analysis are mutually uncorrelated, and all variables can be expressed as linear combinations of common factors. Factor analysis requires a sufficient sample size (greater than 100), and in this case, there are 198 samples in total. At the same time, it is necessary to satisfy the correlation among variables. Generally, the KMO test and Bartlett test are used, that is, the KMO test statistic is greater than 0.5, indicating that there is information overlap among variables, and the significance of the Bartlett test is less than 0.05, that is, the hypothesis of independence of each variable is rejected. Thus, factor analysis can be carried out with this data. The model obtained by factor analysis has two characteristics. One is that the model is not affected by the dimension; the other is that the factor loadings are not unique. By rotating the factor axes, a new factor loading matrix can be obtained to make its meaning more obvious.
[0086] Before performing factor analysis, it is necessary to first conduct principal component analysis on the samples to obtain four components with eigenvalues greater than 1 (components 1, 2, 3, and 4 in Table 1) as four common factors, obtain the eigenvector matrix through the unrotated factor loading matrix, and then obtain the calculation formula for the principal components.
[0087] The component matrix in Table 2 is obtained through SPSS software. In factor analysis, each row represents the loadings and influence degrees of each factor on each variable.
[0088] Table 2 Component Matrix
[0089]
[0090]
[0091] Note: The solids in Table 2 refer to the substances remaining after removing water from dairy products, which is a proprietary term for nutritional components. Components 1 to 4 represent the four eigenvectors greater than 1 in Table 1.
[0092] According to Table 2, the original observed variables are decomposed into two parts: common factors and specific factors. Write the expressions of each variable composed of each factor and its loadings according to the rows (original data matrix = eigenvalue · eigenvector + ε), where Z represents the standardized data of the variable, and the expressions are as follows:
[0093] ZV1 = 0.971×F1 - 0.035×F2 + 0.005×F3 + 0.036×F4 + ε1 (3)
[0094] ZV2 = 0.971×F1 - 0.019×F2 + 0.047×F3 + 0.039×F4 + ε2 (4)
[0095] …
[0096] ZV 22 = 0.308×F1 - 0.381×F2 - 0.3713×F3 - 0.539×F4 + ε 22 (24)
[0097] Each variable in the expression is composed of four common factors and special factors. The special factors are factors that affect the variable other than the common factors. Factor analysis must have practical meanings for the extracted common factors. To make the coefficients in the factor loading matrix more significant, it is necessary to rotate the initial factor loading matrix, re - distribute the relationship between the factors and the original variables, and then it is easier to interpret the variables. Since principal component analysis only extracts four common factors, the rotation will be based on the four extracted common factors. The total variance explanation degree of the four common factors after rotation is as Figure 4 shown. The total variance explanation degree of the variables after rotation is adjusted compared with Table 1, but the total variance explanation degree remains unchanged, that is, the total explanation degree of the four common factors extracted by factor analysis for the variables reaches 75.8%.
[0098] Further analyze the situation of each factor after rotation. Arrange the loading matrix according to size and remove the coefficients less than 0.5, as shown in Table 3. Arrange the loading matrix according to size and remove the coefficients less than 0.5. From the rotated loading common factors, it can be seen that the first common factor has relatively large loadings in reflecting breast milk nutrients such as sodium, potassium, carbohydrates, riboflavin, phosphorus, solids, magnesium, zinc, niacin, pantothenic acid, iron, vitamin B6, fat, copper, and thiamine, and can be named the nutrition factor; the second common factor has relatively large loadings in fat - soluble vitamins. Since fat - soluble vitamins play a key role in the visual development of infants, it can be named the visual factor; the third common factor has relatively large loadings in total protein and true protein in breast milk immunity. Since the protein in breast milk can not only provide energy for infants, but more importantly, form immune protection for the infant body, it can be named the immune factor; the fourth common factor has relatively large loadings in calcium and ascorbic acid. Since the contents of both are affected by the maternal diet to a certain extent, it can be named the dietary factor. These factors are more reasonable in interpretation compared with before rotation.
[0099] Table 3 Component matrix before and after rotation
[0100]
[0101]
[0102] As shown in Table 4, the coefficients of the component score function calculated according to the multiple linear regression algorithm, and the factor score expression can be obtained from this table.
[0103] Table 4 Score Matrix
[0104]
[0105]
[0106] F1 = 0.025×Z 脂肪 - 0.103×Z 蛋白质 + 0.087×Z 碳水化合物 + 0.061×Z 固形物 - 0.106×Z 真蛋白 + 0.014×Z 硫氨酸 + 0.142×Z 核黄素 + 0.148×Z 抗坏血酸 + 0.157×Z 烟酸 + 0.111×Z 泛酸 + 0.103×Z 维生素B6 - 0.064×Z 维生素A - 0.096×Z 维生素E - 0.143×Z β-类胡罗卜 + 0.106×Z 钠 + 0.066×Z 镁 + 0.075×Z 磷 + 0.099×Z 钾 + 0.094×Z 钙 + 0.044×Z 铁 + 0.025×Z 铜 - 0.18×Z 锌 (I)
[0107] F2 = 0.127×Z 脂肪 - 0.056×Z 蛋白质 - 0.015×Z 碳水化合物 + 0.017×Z 固形物 - 0.057×Z 真蛋白 + 0.101×Z 硫氨酸 - 0.094×Z 核黄素 - 0.063×Z 抗坏血酸 - 0.114×Z 烟酸 - 0.023×Z 泛酸 - 0.055×Z 维生素B6 + 0.304×Z 维生素A + 0.407×Z 维生素E + 0.395×Z β-类胡罗卜 - 0.02×Z 钠 + 0.014×Z镁 +0.006×Z 磷 -0.024×Z 钾 -0.123×Z 钙 +0.027Z 铁 +0.039×Z 铜 +0.065×Z 锌 (II)
[0108] F3 = -0.031×Z 脂肪 +0.481×Z 蛋白质 +0.032×Z 碳水化合物 +0.058×Z 固形物 +0.487×Z 真蛋白 -0.008×Z 硫氨酸 -0.045×Z 核黄素 -0.162×Z 抗坏血酸 -0.082×Z 烟酸 -0.063×Z 泛酸 -0.048×Z 维生素B6 -0.083×Z 维生素A -0.09×Z 维生素E +0.029×Z β-类胡罗卜 -0.011×Z 钠 +0.046×Z 镁 +0.035×Z 磷 +0.008×Z 钾 -0.128×Z 钙 +0.054×Z 铁 +0.058×Z 铜 +0.197×Z 锌 (III)
[0109] F4 = 0.083×Z 脂肪 -0.044×Z 蛋白质 +0.013×Z 碳水化合物 +0.035×Z 固形物 -0.042×Z 真蛋白 -0.078×Z 硫氨酸 -0.001×Z 核黄素 -0.532×Z 抗坏血酸 -0.089×Z 烟酸 +0.024×Z 泛酸 +0.11×Z 维生素B6 +0.066×Z 维生素A -0.095×Z 维生素E -0.163×Z β-类胡罗卜 -0.041×Z 钠 +0.011×Z 镁+0.029×Z 磷 +0.013×Z 钾 +0.663×Z 钙 +0.143×Z 铁 +0.117×Z 铜 +0.145×Z 锌 (IV)
[0110] The four factors are weighted and added according to their respective variance contribution rates to obtain a comprehensive score. The calculation formula is as follows:
[0111] F 总 =0.559×F1 + 0.206×F2 + 0.163×F3 + 0.072×F4 (V)
[0112] For 198 columns of breast milk, calculate according to the calculation formula V of the comprehensive score of the above factor analysis. The median is -0.218, the upper quartile is 0.61, and the lower quartile is -0.503. Therefore, according to the median of the comprehensive score, evaluate the breast milk components. If it is higher than the median, it is considered that the breast milk nutrient components are sufficient; if it is lower than the median, it is considered that the breast milk nutrient components are lacking, and the mother should take appropriate dietary measures for adjustment.
[0113] Application Example 1: Formula Powder Evaluation
[0114] Breast milk is the gold standard for infant formula foods, especially in the creation of infant formula foods, which is of great significance. Currently, most infant formula foods adopt a configuration method that simulates the components of breast milk, but there is currently a lack of evaluation methods for formula foods. Based on the factor analysis of breast milk components, use the common factor and comprehensive scoring method as the standard for judging the formula.
[0115] Standardize the content of each component, referring to the standardization method of the principal component analysis in Example 1: Universally adjust the differences between samples through row-by-row normalization; since there are differences of more than 10 3 times in some data, logarithmic transformation is used for data conversion to eliminate the dimensional differences between the original samples; adopt Pareto scaling (centered on the mean and divided by the square root of the standard deviation of each variable) to make the data eigenvectors comparable. The scoring situations of the standardized processing of each component of the new and old formulas are shown in Table 5.
[0116] Table 5 Scoring Situations of Formula Milk Powder
[0117]
[0118]
[0119] The common factor scores and comprehensive factor scores of the two formulations can be calculated according to formulas (I) to (IV) and (V), and the results are shown in Table 6.
[0120] Table 6 Scoring of the new and old formulations
[0121] Component <![CDATA[F1]]> <![CDATA[F2]]> <![CDATA[F3]]> <![CDATA[F4]]> <![CDATA[F 总 > New Formula -0.919 0.697 0.397 0.279 -0.286 Old Formula -14.343 -1.208 0.477 -0.177 -8.207
[0122] According to the median of the comprehensive score in factor analysis being -0.218 (almost the same as the median of breast milk), it can be known that the composition of the new formulation is closer to breast milk.
[0123] The present invention simplifies the macronutrients and micronutrients of breast milk by using the principal component multivariate statistical method, selects the main indicators from numerous quality evaluation indicators to evaluate the nutritional components of breast milk, and lays a certain foundation for the establishment of the breast milk nutritional component evaluation system.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for evaluating the nutritional component composition of breast milk or formula foods by using multivariate statistical methods, characterized in that, The method includes the following steps: Obtain the content of nutritional components in breast milk or formula food to be evaluated as the original data variable; perform normalization processing on the original data variable to eliminate the dimension difference, and perform centering processing; Input the centered data into the pre-constructed evaluation model to obtain the score output by the evaluation model, and compare the score with the preset median. The smaller the difference, the closer it is to breast milk; Among them, the evaluation model includes four common factors obtained by principal component analysis and a scoring model obtained by load and influence degree in factor analysis.
2. The method according to claim 1, wherein Each nutritional component includes one or more of phosphorus, solids, carbohydrates, potassium, magnesium, sodium, iron, riboflavin, pantothenic acid, fat, copper, niacin, zinc, vitamin B6, thiamine, vitamin E, β-carotene, vitamin A, true protein, protein, calcium, ascorbic acid.
3. The method according to claim 1 or 2, characterized in that, The construction method of the evaluation model is constructed based on the factor analysis method in mathematical statistics. Optionally, it specifically includes: Suppose there are n samples, and each sample observes p nutritional component indicators with correlation. Standardize the sample observation data so that the mean of the standardized variables is 0 and the variance is 1; optionally, the data standardization method includes: eliminating the dimension difference between the original samples through normalization processing; using the Pareto scaling method to perform centering processing on the normalized data. The Pareto scaling method includes centering on the mean and dividing by the square root of the standard deviation of each variable; Check whether there is a correlation between the variables in the standardized data. Optionally, the methods used include: sampling adequacy test ((KMO test)) and Bartlett test. Optionally, if the KMO statistic is greater than 0.6 and the probability value of the Bartlett sphericity test is less than the significance level, it is considered to have a correlation. Optionally, select 22 nutritional component indicators; Perform factor analysis on the standardized and correlated data to obtain common factors, and obtain the factor load matrix, component matrix table, and data table for factor weight analysis of each nutritional component indicator; Let both the original variable and the standardized variable vector be represented by X. The factor analysis model of the p-dimensional variable X = [X1,..., Xi,..., Xp] is: Among them, the vector F is the extracted common factor vector, representing m (m < p) mutually independent common influencing factors that are objectively present but cannot be directly observed in the original variables; the vector A is the factor load matrix, indicating the load of the variable X matrix on the common factor matrix F, reflecting the correlation coefficient between the two. The larger the absolute value, the higher the correlation; Optionally, represent all variables X as a linear combination of the common factor F. The single-factor score calculation formula is as follows: F i = a 1i X1 + a 2i X2 + … + a pi X p , p = 1, …, 22 (1) In formula (1), vector F is the extracted common factor vector, representing m mutually independent common influencing factors that are objectively present but not directly observable in the original variables. Optionally, m < 22; a 1i , a 2i … a pi is the factor loading matrix of vector A, indicating the loading of the variable X matrix on the common factor matrix F; then m principal components with eigenvalues greater than 1 are extracted from the factor weight analysis matrix, and according to the principal component variance contribution rate w, the calculation formula for the comprehensive score is obtained; F 总 = w1F1 + w2F2 + … + w m F m (2) In formula (2), w1……w m represent the variance contribution rates of the respective common factors, F1, F2, ……, F m represent the standardized common factors, and optionally m < p; Evaluate the nutritional components of breast milk or formula food according to the comprehensive score calculation formula.
4. A method for evaluating the nutritional component composition of breast milk or formula food by using multivariate statistics, characterized in that, The method includes the following steps: Obtain the content of nutritional components in breast milk or formula food to be evaluated as the original data variable; perform normalization processing on the original data variable to eliminate the dimension difference, and perform centering processing; Input the centralized data into a pre-constructed evaluation model to obtain the score output by the evaluation model, and compare the score with a preset median. The smaller the difference, the closer it is to breast milk; Among them, the evaluation model includes the following scoring formula: F1 = 0.025×Z 脂肪 -0.103×Z 蛋白质 +0.087×Z 碳水化合物 +0.061×Z 固形物 -0.106×Z 真蛋白 +0.014×Z 硫氨酸 +0.142×Z 核黄素 +0.148×Z 抗坏血酸 +0.157×Z 烟酸 +0.111×Z 泛酸 +0.103×Z 维生素B6 -0.064×Z 维生素A -0.096×Z 维生素E -0.143×Z β-类胡罗卜 +0.106×Z 钠 +0.066×Z 镁 +0.075×Z 磷 +0.099×Z 钾 +0.094×Z 钙 +0.044×Z 铁 +0.025×Z 铜 -0.18×Z 锌 (I) F2 = 0.127×Z 脂肪 - 0.056×Z 蛋白质 - 0.015×Z 碳水化合物 + 0.017×Z 固形物 - 0.057×Z 真蛋白 + 0.101×Z 硫氨酸 - 0.094×Z 核黄素 - 0.063×Z 抗坏血酸 - 0.114×Z 烟酸 - 0.023×Z 泛酸 - 0.055×Z 维生素B6 + 0.304×Z 维生素A + 0.407×Z 维生素E + 0.395×Z β-类胡罗卜 - 0.02×Z 钠 + 0.014×Z 镁 + 0.006×Z 磷 - 0.024×Z 钾 - 0.123×Z 钙 + 0.027Z 铁 + 0.039×Z 铜 + 0.065×Z 锌 (II) F3 = -0.031×Z 脂肪 +0.481×Z 蛋白质 +0.032×Z 碳水化合物 +0.058×Z 固形物 +0.487×Z 真蛋白 -0.008×Z 硫氨酸 -0.045×Z 核黄素 -0.162×Z 抗坏血酸 -0.082×Z 烟酸 -0.063×Z 泛酸 -0.048×Z 维生素B6 -0.083×Z 维生素A -0.09×Z 维生素E +0.029×Z β-类胡罗卜 -0.011×Z 钠 +0.046×Z 镁 +0.035×Z 磷 +0.008×Z 钾 -0.128×Z 钙 +0.054×Z 铁 +0.058×Z 铜 +0.197×Z 锌 (III) F4 = 0.083×Z 脂肪 - 0.044×Z 蛋白质 + 0.013×Z 碳水化合物 + 0.035×Z 固形物 - 0.042×Z 真蛋白 - 0.078×Z 硫氨酸 - 0.001×Z 核黄素 - 0.532×Z 抗坏血酸 - 0.089×Z 烟酸 + 0.024×Z 泛酸 + 0.11×Z 维生素B6 + 0.066×Z 维生素A - 0.095×Z 维生素E - 0.163×Z β-类胡罗卜 - 0.041×Z 钠 + 0.011×Z 镁 + 0.029×Z 磷 + 0.013×Z 钾 + 0.663×Z 钙 + 0.143×Z 铁 + 0.117×Z 铜 + 0.145×Z 锌 (IV) The four factors are weighted and added together according to their respective variance contribution rates to obtain a comprehensive score, and its calculation formula is: F 总 = 0.559×F1 + 0.206×F2 + 0.163×F3 + 0.072×F4 (V) Among them, phosphorus, solids, carbohydrates, potassium, magnesium, sodium, iron, riboflavin, pantothenic acid, fat, copper, niacin, zinc, vitamin B6, thiamine, vitamin E, β-carotene, vitamin A, true protein, protein, calcium, and ascorbic acid are the standardized variable data.
5. The method according to any one of claims 1 to 4, characterized in that The standardization method includes: using logarithmic transformation to eliminate the dimensional difference between the original samples of the original data, and using the Pareto scaling method to centralize the data.
6. The method according to claim 5, wherein The centralization process uses the Pareto scaling method. Optionally, the Pareto scaling method includes centering on the mean and dividing by the square root of the standard deviation of each variable.
7. According to the method described in any one of claims 1 to 6, characterized in that The preset factor score range is (-1.06, 1.72), where the median of the factor scores is (-0.84, 0.4). Optionally, the median is -0.
218. Optionally, the upper quartile is (-0.01, 1.23). Optionally, the upper quartile is 0.
61. The lower quartile is (-1.13, 0.12). Optionally, the lower quartile is -0.503; Optionally, the preset median is determined according to the comprehensive scores of each data sample by factor analysis.
8. The method according to any one of claims 1 to 7, characterized in that Evaluate the breast milk components according to the preset median. If it is higher than the median, it is considered that the breast milk nutritional components are sufficient; if it is lower than the median, it is considered that the breast milk nutritional components are lacking, and the mother should take appropriate dietary measures for adjustment.
9. A system for evaluating nutritional components in breast milk or formula food, applied to the method according to any one of claims 1-8, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, when the processor executes the program, it implements the steps of any one of claims 1-8.
10. Use of the method according to any one of claims 1-8, or the system according to claim 9, in the preparation of products for evaluating the nutritional components in breast milk or formula foods.
Citation Information
Patent Citations
Urban livability evaluation model based on principal component analysis
CN107545380A
Evaluation index acquisition method and system
CN110717687A
Application of serum metabolic marker in preparation of diabetic kidney lesion early diagnosis reagent and kit
CN111289638A
Cited By
Breast milk fat globule membrane protein simulation degree evaluation method based on multi-dimensional similarity weighted fusion
CN121938450A