A Method for Evaluating the Simulation Quality of Breast Milk Fat Globule Membrane Proteins Based on Multidimensional Similarity Weighted Fusion
By constructing a multidimensional similarity-weighted fusion evaluation method for the simulation degree of breast milk MFGM protein, the problem of the gap between infant formula and breast milk protein composition was solved, enabling scientific evaluation of infant formula and guidance for formula optimization, and improving the scientificity and interpretability of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SANYUAN FOOD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-02
AI Technical Summary
Existing infant formula milk powders differ significantly from breast milk in terms of protein composition. There is a lack of an objective and quantifiable MFGM protein similarity evaluation system. The high-dimensional sparsity of omics data leads to information distortion. The evaluation perspective is singular, lacks a large-sample statistical basis, and cannot scientifically assess the breast milk similarity of different MFGM raw materials. There is also a lack of feedback to guide formula optimization.
A breast milk MFGM proteomics database covering different regions and lactation stages was constructed. The Z-score matrix and deletion mask matrix were calculated using a multidimensional similarity weighted fusion method. Combining abundance distribution and deletion pattern similarity, the Top-K score and overlap feature shrinkage coefficient were used to output key defective proteins and formulation optimization suggestions.
It enables precise scoring of the mimicry of MFGM protein in infant formula and breast milk, providing scientific evidence to guide formula optimization, improving the scientific nature and interpretability of the evaluation, and ensuring the stability and reliability of the evaluation results.
Smart Images

Figure CN121938450B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics, food nutritionomics, and quality evaluation of formulated foods, specifically to a method and apparatus for evaluating the protein mimicry of breast milk fat globule membrane (MFGM) based on multidimensional similarity weighted fusion. Background Technology
[0002] As is well known, breast milk is the ideal natural food for infants and young children. The milk fat globule membrane (MFGM) in breast milk is rich in bioactive components such as lipids and proteins, playing a vital role in the immune and cognitive development of newborns. The MFGM is composed of a unique and complex three-layered membrane structure surrounding the milk fat globules. Although MFGM proteins account for only 1% to 4% of the total protein content in breast milk, they contain a large number of diverse proteins. Proteomics analysis has identified over 800 different MFGM-related proteins. Some studies indicate that MFGM proteins account for approximately 90% of the protein types in breast milk. These key MFGM proteins include mucin-1 (MUC1), butyrophilin (BTN), lactadherin (MFG-E8), xanthine dehydrogenase / oxidase (XDH / XO), platelet glycoprotein 4 (CD36), fatty acid-binding protein (FABP), and perilipin. These proteins play a central role in infant immune development, intestinal maturation, brain development, and lipid absorption. Due to the presence of these key active components, the milk fat globule membrane is considered an important medium for the mother to transfer immune protection and nutrients necessary for growth and development to the infant.
[0003] Compared to breast milk, existing infant formula differs significantly in protein composition, particularly in MFGM protein components. Traditional formula fat is primarily derived from vegetable oils or skimmed milk, resulting in significant loss of milk fat globule membrane components during processing. This leads to a lack of MFGM active ingredients naturally present in breast milk. The composition of milk fat globule membranes in cow's milk and other animal milk also differs from breast milk: not only is the protein complexity lower, but the content of certain key proteins (such as MUC1 and lactamin) is lower or their structures differ. Furthermore, cow's milk contains some proteins not found in breast milk (such as whey proteins like β-lactoglobulin). This means that formula produced using traditional cow's milk as a raw material suffers from issues such as missing protein types and the introduction of "extra proteins" compared to breast milk. Moreover, the composition of MFGM proteins in breast milk varies from person to person; factors such as different mothers, lactation period, and regional diet all influence the spectrum and abundance distribution of milk fat globule membrane proteins.
[0004] In recent years, in an effort to narrow the nutritional differences between formula and breast milk, the industry has generally attempted to increase the "humanization" of infant formula by adding MFGM (milk-like compound). However:
[0005] 1) There is currently a lack of an objective and quantifiable evaluation system for the similarity of breast milk MFGM proteins;
[0006] 2) The high-dimensional sparsity of omics data: Due to the complexity of biological samples and detection limitations, omics data often contain a large number of missing values (NaN). Traditional methods forcibly filling or removing samples will lead to serious information distortion.
[0007] 3) The evaluation perspective is too narrow. Existing evaluations focus on the total amount or comparison of single proteins, lack a large sample statistical basis, and ignore factors such as protein abundance distribution, interference from non-target proteins, and differences in the importance of key functional proteins.
[0008] 4) Different MFGM raw materials (such as fortified whey powder, cow's milk, and MFGM isolate) lack scientifically comparable indicators of breast milk similarity;
[0009] 5) Lack of feedback and guidance; existing evaluations cannot provide targeted guidance for optimizing infant formula.
[0010] Given the above background, there is an urgent need for an evaluation system that can be based on a real breast milk protein database, has biological screening significance, and can generate a comprehensive report on "breadth-depth-purity" to assess the degree to which infant formula or milk-based raw materials simulate breast milk MFGM protein, and provide a scientific basis for formula development and quality improvement. Summary of the Invention
[0011] Therefore, embodiments of the present invention provide a method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion, in order to solve at least one technical problem existing in the prior art.
[0012] This invention provides a method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion, the method comprising:
[0013] We constructed a breast milk MFGM proteomics database covering different regions and different lactation stages, and obtained proteomics data of the samples to be evaluated.
[0014] Based on the breast milk MFGM proteomics database, a reference matrix was obtained through row and column filtering. The reference matrix was then logarithmically transformed, and the mean and standard deviation of each feature in the reference matrix were calculated to construct the Z-score matrix and the missing mask matrix corresponding to the reference matrix.
[0015] Based on the proteomics data of the samples to be evaluated, the sample matrix to be evaluated is obtained through row and column filtering. Logarithmic transformation is performed on the sample matrix to be evaluated, and the mean and standard deviation of each feature in the sample matrix to be evaluated are calculated. The Z-score matrix and missing mask matrix corresponding to the sample matrix to be evaluated are constructed.
[0016] Calculate the abundance distribution similarity and missing pattern similarity of the samples to be evaluated;
[0017] The shrinkage coefficient is set based on the actual number of overlapping features, and the fusion similarity is obtained through weighted fusion. The comprehensive score of each sample to be scored is obtained through Top-K scoring. The core protein set is screened according to the set threshold, and the types of missing MFGM proteins in the sample to be tested are obtained by comparison.
[0018] Construct a Top-K neighbor center, calculate the feature deviation contribution and abundance difference of the sample to be evaluated, and output key defective proteins and formulation optimization suggestions based on the preset feature importance index.
[0019] In some embodiments, a breast milk MFGM proteomics database covering different regions and different lactation stages is constructed, specifically including:
[0020] Raw breast milk samples were collected from different regions and at different stages of lactation, and MFGM protein was extracted from each of the breast milk samples.
[0021] The extracted MFGM protein was pretreated, and the MFGM protein obtained after pretreatment was detected by liquid chromatography-mass spectrometry to obtain the original sample data file.
[0022] The original data file of the sample is input into a pre-stored original database for comparison to obtain the comparison results;
[0023] The original breast milk sample, the original sample data file corresponding to the original breast milk sample, and the comparison results are stored to obtain the breast milk MFGM proteomics database.
[0024] In some embodiments, a reference matrix is obtained based on the breast milk MFGM proteomics database through row and column filtering, specifically including:
[0025] Construct the original reference dataset ,in Let n be the set of real numbers, p be the number of reference samples (rows), and NaN be the number of features (columns). The original reference dataset represents the set of real numbers. Sample-protein matrix;
[0026] The number of reference samples is filtered, and the proportion of non-missing features in the i-th sample is counted. When the number of valid features in the sample reaches a threshold... It shall be retained at that time;
[0027] The number of features is filtered, and the detection rate of the j-th feature in all samples is calculated. Features with a frequency lower than a threshold are removed. Features;
[0028] The reference matrix obtained after filtering by the number of reference samples and features is: .
[0029] In some embodiments, a logarithmic transformation is performed on the reference matrix to calculate the mean and standard deviation of each feature in the reference matrix, and the Z-score matrix and missing mask matrix corresponding to the reference matrix are constructed, specifically including:
[0030] Perform a logarithmic transformation on the reference matrix to make the data closer to a normal distribution:
[0031] ,in, This is the result after logarithmic transformation;
[0032] For each feature The mean is calculated based only on non-missing values. with standard deviation :
[0033] ;
[0034] in, Let j be the value of the i-th sample and the j-th protein feature after logarithmic transformation. Let be the original abundance value of the feature of the i-th sample and the j-th protein. ( ) is the mean calculation function. ( ) is the standard deviation calculation function;
[0035] Z-score matrix corresponding to the reference matrix The expression is:
[0036] ;
[0037] Missing mask matrix corresponding to the reference matrix The expression is:
[0038] .
[0039] In some embodiments, the Z-score matrix corresponding to the sample matrix to be evaluated The expression is:
[0040] ;
[0041] in, is the logarithmic transformation value of the j-th protein feature in the sample to be evaluated;
[0042] Missing mask matrix corresponding to the sample matrix to be evaluated The expression is:
[0043] ;
[0044] in, () is an indicator function.
[0045] In some embodiments, abundance distribution similarity The expression is:
[0046] ;
[0047] in, For any reference sample, For a set of common comparable features, For feature weights, For the j-th protein feature, Let be the Z-score normalized value of the j-th protein feature in the sample to be evaluated. The Z-score normalized value of the i-th sample and j-th protein feature in the reference matrix;
[0048] Missing pattern similarity The expression is:
[0049] ;
[0050] in, This represents the total number of valid features retained after filtering. Let be the value of the j-th feature in the missing mask matrix of the sample to be evaluated. This represents the missing mask matrix value for the i-th sample and the j-th feature in the reference matrix.
[0051] In some embodiments, a shrinkage coefficient is set based on the actual number of overlapping features, specifically including:
[0052] Let the actual number of overlapping features be... for:
[0053] ;
[0054] Shrinkage coefficient The expression is:
[0055] ;
[0056] in, These are the preset hyperparameters.
[0057] In some embodiments, a shrinkage coefficient is set based on the actual number of overlapping features, and a fusion similarity is obtained through weighted fusion. A comprehensive score is obtained through Top-K scoring, specifically including:
[0058] The numerical similarity and the missing pattern similarity are weighted and fused to obtain the fused similarity. :
[0059] ;
[0060] in, This refers to the similarity weighting coefficient;
[0061] Select the one with the highest similarity A set of reference samples Define the comprehensive score of the new sample to be tested. for:
[0062] .
[0063] The present invention also provides a device for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion, the device comprising:
[0064] The sample collection module is used to construct a breast milk MFGM proteomics database covering different regions and different lactation stages, and to acquire proteomics data of samples to be evaluated.
[0065] The reference data processing module is used to obtain a reference matrix based on the breast milk MFGM proteomics database through row and column filtering, perform logarithmic transformation on the reference matrix, calculate the mean and standard deviation of each feature in the reference matrix, and construct the Z-score matrix and missing mask matrix corresponding to the reference matrix.
[0066] The data processing module is used to obtain the sample matrix to be evaluated based on the proteomics data of the sample to be evaluated through row and column filtering, perform logarithmic transformation on the sample matrix to be evaluated, calculate the mean and standard deviation of each feature in the sample matrix to be evaluated, and construct the Z-score matrix and missing mask matrix corresponding to the sample matrix to be evaluated.
[0067] The similarity calculation module is used to calculate the abundance distribution similarity and missing pattern similarity of the samples to be evaluated.
[0068] The comprehensive scoring calculation module is used to set the shrinkage coefficient based on the actual number of overlapping features, obtain the fusion similarity through weighted fusion, obtain the comprehensive score of each sample to be scored through Top-K scoring, filter the core protein set according to the set threshold, and compare to find the types of missing MFGM proteins in the sample to be tested.
[0069] The evaluation result generation module is used to construct Top-K neighbor centers, calculate the feature deviation contribution and abundance difference of the samples to be evaluated, and output key defective proteins and formulation optimization suggestions based on preset feature importance indicators.
[0070] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described above.
[0071] The present invention provides a method for evaluating the mimicry of breast milk fat globule membrane proteins (MFGMs) based on multi-perspective similarity weighted fusion. This method utilizes NanoLC-Orbitrap MS liquid chromatography-mass spectrometry to detect breast milk samples from different regions and lactation stages, constructing a local breast milk MFGM proteomics database covering protein types and their relative abundance. Simultaneously, proteomics data of samples to be evaluated (such as formula powder, fortified whey powder, and milk) are acquired. A multi-perspective similarity evaluation mechanism is employed, using missing value-aware normalization to process high-dimensional sparse data, fusing numerical pattern similarity and missing value pattern similarity, and introducing an adaptive shrinkage mechanism for overlapping dimensions. This effectively addresses evaluation bias caused by missing values (NaN) in omics detection, achieving accurate scoring of sample mimicry. By constructing local reference centers, core deviation features leading to reduced mimicry scores are identified, providing quantitative evidence for precise auxiliary design and formula improvement of infant formula. Attached Figure Description
[0072] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0073] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0074] Figure 1 A flowchart illustrating the method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion provided by this invention;
[0075] Figure 2 This is a structural block diagram of the breast milk fat globule membrane protein simulation evaluation device based on multidimensional similarity weighted fusion provided by the present invention.
[0076] Figure 3 This is a structural block diagram of a computer device provided by the present invention. Detailed Implementation
[0077] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0078] The main objective of this invention is to build a breast milk MFGM protein database based on other breast milk samples from different regions and stages, and to provide a method for evaluating the similarity between infant formula and its raw materials and breast milk from multiple perspectives based on the constructed database. This method overcomes the limitations of traditional evaluation methods in handling incomplete omics datasets and significantly improves the scientific rigor and research guidance value of MFGM protein simulation evaluation.
[0079] In one specific implementation, such as Figure 1 As shown, the method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion provided by this invention includes the following steps:
[0080] S110: Construct a breast milk MFGM proteomics database covering different regions and different lactation stages, and obtain proteomics data of the samples to be evaluated; specifically, the method provided by this invention first establishes an MFGM protein database covering different regions and different lactation stages; the database construction and sample acquisition include the following steps:
[0081] (1) By collecting breast milk samples from different regions and different lactation stages, extracting MFGM protein and processing it, the NanoLC-Orbitrap MS liquid chromatography-mass spectrometry technology was used to obtain an omics database including protein types and their relative abundance, and a breast milk MFGM protein database, i.e., a reference matrix, was constructed.
[0082] (2) Data acquisition of samples to be evaluated: Proteomics data of infant formula, fortified whey powder, milk samples and other breast milk samples were acquired using the same technology.
[0083] S120: Based on the breast milk MFGM proteomics database, a reference matrix is obtained through row and column filtering. The reference matrix is then logarithmically transformed, and the mean and standard deviation of each feature in the reference matrix are calculated to construct the Z-score matrix and missing mask matrix corresponding to the reference matrix.
[0084] S130: Based on the proteomics data of the samples to be evaluated, a sample matrix is obtained through row and column filtering. A logarithmic transformation is performed on the sample matrix, and the mean and standard deviation of each feature in the matrix are calculated. This constructs the Z-score matrix and missing value mask matrix corresponding to the sample matrix. In other words, this invention employs a missing value-aware standardization mechanism in data processing when establishing the evaluation system. Traditional methods typically involve directly deleting samples with missing values or performing data imputation. Directly deleting samples with missing values wastes information, while data imputation easily introduces artificial noise. This invention introduces a missing value-preserving Z-score transformation, using only non-missing values to calculate statistics during the standardization process, and using a mask matrix to track the data presence in real time, ensuring the authenticity of the evaluation results.
[0085] S140: Calculate the abundance distribution similarity and missing pattern similarity of the sample to be evaluated; In the process of multi-perspective fusion evaluation of numerical and structural data, traditional methods only focus on the absolute difference of component content or only on whether it is detected. This invention proposes a dual evaluation of numerical similarity and missing pattern similarity, defining the degree of simulation from two perspectives: abundance distribution similarity and feature overlap similarity.
[0086] S150: A shrinkage coefficient is set based on the actual number of overlapping features, and a fusion similarity is obtained through weighted fusion. A comprehensive score for each sample to be scored is obtained through Top-K scoring and mapped to the [0,100] interval. A core protein set is selected based on a set threshold, and the types of MFGM proteins missing in the sample to be tested are compared. It should be understood that in this step, the comprehensive score of each sample to be scored is used to evaluate whether it is close to the reference breast milk. The score is mapped to the 0-100 interval, which facilitates the observation and determination of the maximum and minimum score ranges under the same evaluation system. Regarding the confidence control dimension of adaptive shrinkage, traditional methods treat all comparison samples equally, which can easily lead to false high scores due to a small number of accidentally overlapping features. This invention innovatively proposes an adaptive shrinkage mechanism based on the overlap dimension, through hyperparameters... The similarity weight of a sample is automatically adjusted when the sample to be tested and the reference sample share too few common features, ensuring that the evaluation results are based on a solid chain of evidence (sufficiently overlapping features). Furthermore, the multi-perspective similarity weighted fusion specifically evaluates the similarity of the MFGM proteome between the sample to be tested and breast milk from different perspectives. This is achieved through data preprocessing, missing-aware standardization, numerical pattern simulation, missing pattern similarity, and an adaptive shrinkage mechanism to comprehensively evaluate the sample simulation. In addition, referencing the obtained core protein set coverage, the screening methods used in the original data preprocessing for sample and feature data are applied to select core protein sets with set thresholds, thereby identifying protein sets that are prevalent and important in breast milk, and comparing them to identify the types of missing MFGM proteins in infant formula, etc.
[0087] S160: Construct Top-K neighbor centers, calculate the feature deviation contribution and abundance difference of the sample to be evaluated, and output key defective proteins and formulation optimization suggestions based on preset feature importance indicators, thereby shifting from isolated scoring to closed-loop diagnosis. After obtaining the comprehensive similarity score, this invention identifies the core features that lead to the score reduction by constructing local reference centers, providing a quantitative basis for formulation improvement. Traditional methods only output a final score, leaving researchers unable to know how to improve. This invention constructs a dynamic benchmark through Top-K local weighted centers, calculates feature-level deviation contributions, and can output key differential proteins that lead to a decrease in similarity and their contribution, realizing the interpretability of breast milk similarity evaluation. This improvement upgrades the "evaluation system" to a "decision support system," which can directly transform the evaluation results into priority directions for formulation and process optimization, improving improvement efficiency.
[0088] Thus, the method provided by this invention starts from real breast milk samples from multiple regions and lactation stages, establishes a standardized and large-scale MFGM protein reference system, and avoids evaluation bias caused by single similarity by using dual-dimensional quantitative evaluation of abundance distribution similarity and deletion pattern similarity. It improves the stability and reliability of evaluation results under different feature coverage by using shrinkage coefficient and weighted fusion, normalizes the simulation degree into intuitive scores, and automatically locates missing proteins and key defective proteins, realizing objective, quantitative, and interpretable evaluation of breast milk MFGM simulated formulas, directly supporting formula optimization.
[0089] In step S110 above, a breast milk MFGM proteomics database covering different regions and different lactation stages is constructed, specifically including the following steps:
[0090] Raw breast milk samples were collected from different regions and at different stages of lactation, and MFGM protein was extracted from each of the breast milk samples.
[0091] The extracted MFGM protein was pretreated, and the MFGM protein obtained after pretreatment was detected by liquid chromatography-mass spectrometry to obtain the original sample data file.
[0092] The original data file of the sample is input into a pre-stored original database for comparison to obtain the comparison results;
[0093] The original breast milk sample, the original sample data file corresponding to the original breast milk sample, and the comparison results are stored to obtain the breast milk MFGM proteomics database.
[0094] In a specific implementation scenario, breast milk samples can be selected from 85 healthy mothers from Beijing, Tangshan, Liuyang, Luoyang, Tibet and other regions. All volunteers have normal physical indicators and the babies are all full-term. A total of 232 breast milk samples from 0 to 6 months of age are collected from the above 85 healthy mothers, and the development status of the corresponding infants is recorded at the same time.
[0095] The breast milk collection method was as follows: all lactating mother volunteers had normal physical indicators, their babies were delivered at full term (38-42 weeks of gestation), and they had no congenital or hereditary diseases. All volunteers were required to empty one breast between 6:00 and 7:00 in the morning, and then collect whole milk from one breast (which had been emptied) between 9:00 and 11:00 in the morning. The whole milk was mixed and aliquoted into 1mL sterile cryovials and stored in an ultra-low temperature freezer at -80℃.
[0096] The infant formula powder was derived from a commercially available brand of stage 1 milk powder with added MFGM (hereinafter referred to as infant formula powder), the 6 cow milk samples were derived from a commercially available brand of sterilized milk (hereinafter referred to as cow milk), and the 5 MFGM samples were derived from a brand of fortified whey powder (hereinafter referred to as whey powder).
[0097] In the extraction of MFGM protein, 3 mL of each sample was taken, and sucrose was added at a ratio of 5 g / 100 mL. After mixing thoroughly, the samples were divided into two 1.5 mL portions and centrifuged at 4000×g for 20 min at 4℃. The fat layer was collected and washed 2-3 times with PBS buffer to obtain milk fat. An appropriate amount of ultrapure water was added to the obtained milk fat, and after mixing, it was centrifuged at 5000×g for 20 min. The fat layer was collected, and chloroform and methanol were added at a ratio of 1:1 (v / v). After mixing thoroughly, it was centrifuged at 11000×g for 20 min. The intermediate protein layer was collected. This process was repeated twice, followed by nitrogen blowing to obtain milk fat globule membrane protein.
[0098] Pretreatment for MFGM protein liquid chromatography-mass spectrometry analysis was performed according to the Waters RapiGest SF reagent instructions:
[0099] 1) Add an appropriate amount (50-100 uL) of dissolved 0.1% RapiGest SF to the protein solid, mix thoroughly and sonicate for 30 min;
[0100] 2) Add 10 mM DTT to make the final concentration 5 mM, and then place the sample in a water bath at 56°C for 1 h.
[0101] 3) After returning to room temperature, add iodoacetamide to a final concentration of 15 mM, and react in the dark for 40 min;
[0102] 4) Add trypsin at a ratio of 1:50 and perform enzymatic digestion overnight;
[0103] 5) Add an appropriate amount of 0.5 M HCl to terminate the reaction for 30 min, allowing the RapiGest SF reagent to precipitate (precipitation occurs when the pH is less than 2).
[0104] 6) Desalt, pass through an HLB column, freeze dry, reconstitute with 0.1% FA, centrifuge at 11000g for 10 min, and load the sample.
[0105] During chromatographic analysis, an UltiMate 3000 nano-high-performance liquid chromatograph was used; the analytical column was a C18 Acclaim Pep Map RSLC column (15 cm × 50 µm, 2 µm, 100 Å); the enrichment column was a C18 Acclaim Pep Map 100 column (2 cm × 75 µm, 5 µm, 100 Å); the column temperature was 35℃; the mobile phase was: A was 0.1% formic acid aqueous solution, and B was 0.1% formic acid acetonitrile solution; the gradient elution program was: 4% B held for 5 min, the concentration of B increased from 4% to 50% within 70 min and held for 3 min, and the concentration of B decreased to 4% within 2 min and held for 10 min; the flow rate was 0.25 µL / min, and the injection volume was 1 µL.
[0106] During the mass spectrometry testing, the mass spectrometer was a Thermo Fisher Q Exactive Orbitrap high-resolution mass spectrometer; the ion source was a Nanospray Flex nano-electrospray source; the electrospray voltage was 3.2 kV; the ion transmission tube temperature was 320℃; the RF lens was 50%; the scanning mode was data-dependent acquisition mode; the primary mass spectrometry was acquired using Orbitrap, with a scan range of 300-1600 m / z and a resolution of 70,000; secondary fragmentation was performed on the 15 most abundant ion peaks, using high-energy fragmentation (HCD) mode with a collision energy of 28%.
[0107] During data processing and statistical analysis, the original sample data files were compared against a database using Thermo ProtermeDiscoverer 1.4 software. The database search software used was SEQUEST HT and the Uniprot database (http: / / www.uniprot.org / ). Specific parameters were set as follows: enzyme: trypsin (full); maximum missed cleavage site: 2; precursor ion mass deviation: 10 ppm; daughter ion mass deviation: 0.02 Da; dynamic modification settings: oxidation (M+15.995 Da), deamination (N,Q+0.984 Da); fixed modification: urea methylation (C+57.021 Da). To control the reliability of the results, a false positive rate (FDR) was used to filter the results. The q-value was calculated using Percolator, with the following parameters: Peptide Confidence: High; FDR ≤ 0.01.
[0108] In the process of constructing the breast milk MFGM protein database and acquiring sample data, the processed protein samples were reconstituted and analyzed by NanoLC-Orbitrap MS. Two biological replicates were performed for each breast milk sample. The obtained RAW source files were compared using SEQUEST HT, yielding data for 442 MFGM proteins, containing a total of 9737 proteins. This resulted in an MFGM protein database covering multiple regions and lactation stages.
[0109] The same method was used to obtain data from infant formula, whey fortified powder, and cow's milk samples. Through homologous protein mapping technology, the non-human milk protein data were uniformly transformed into the coordinate system of the human proteome database to achieve quantitative evaluation of the similarity between samples from multiple species and breast milk.
[0110] Thus, the method provided by this invention establishes a breast milk MFGM proteomics database with authentic sources, wide coverage, complete lactation stages, and unified data format by standardizing the entire process of sample collection, protein extraction, mass spectrometry detection, data comparison, and unified storage. This ensures the authenticity, representativeness, and traceability of subsequent evaluation benchmarks and provides underlying data support for high-confidence similarity calculations.
[0111] In step S120 above, a reference matrix is obtained based on the breast milk MFGM proteomics database through row and column filtering, specifically including the following steps:
[0112] Construct the original reference dataset ,in Let n be the set of real numbers, p be the number of reference samples (rows), and NaN be the number of features (columns). The original reference dataset represents the set of real numbers. Sample-protein matrix;
[0113] The number of reference samples is filtered, and the proportion of non-missing features in the i-th sample is counted. When the number of valid features in the sample reaches a threshold... It shall be retained at that time;
[0114] The number of features is filtered, and the detection rate of the j-th feature in all samples is calculated. Features with a frequency lower than a threshold are removed. Features;
[0115] The reference matrix obtained after filtering by the number of reference samples and features is: .
[0116] In this way, by filtering out low-quality samples by row and removing features with low detection rate and low stability by column, the interference of noisy data and missing values on subsequent similarity calculations is reduced, the robustness and comparability of the reference matrix are improved, and the evaluation focuses on MFGM protein features with high confidence and high representativeness.
[0117] In step S120 above, a logarithmic transformation is performed on the reference matrix, the mean and standard deviation of each feature in the reference matrix are calculated, and the Z-score matrix and missing mask matrix corresponding to the reference matrix are constructed, specifically including:
[0118] Perform a logarithmic transformation on the reference matrix to make the data closer to a normal distribution:
[0119] ,in, This is the result after logarithmic transformation;
[0120] For each feature The mean is calculated based only on non-missing values. with standard deviation :
[0121] ;
[0122] in, Let j be the value of the i-th sample and the j-th protein feature after logarithmic transformation. Let be the original abundance value of the feature of the i-th sample and the j-th protein. ( ) is the mean calculation function. ( ) is the standard deviation calculation function;
[0123] Z-score matrix corresponding to the reference matrix The expression is:
[0124] ;
[0125] Missing mask matrix corresponding to the reference matrix The expression is:
[0126] .
[0127] In this way, logarithmic transformation compresses the order of magnitude of abundance, making protein expression data more consistent with a normal distribution and improving the rationality of subsequent correlation calculations; robust statistics using non-missing values avoid the contamination of the mean and standard deviation by missing values; Z-score standardization eliminates differences in the dimensions and orders of magnitude of different protein features; and the missing mask matrix explicitly marks the location of valid data, providing a unified basis for calculating missing pattern similarity.
[0128] In step S130 above, the Z-score matrix corresponding to the sample matrix to be evaluated The expression is:
[0129] ;
[0130] in, Let be the logarithmic transformation value of the j-th protein feature in the sample to be evaluated, i.e., the one-dimensional feature vector of the logarithmic transformation;
[0131] Missing mask matrix corresponding to the sample matrix to be evaluated The expression is:
[0132] ;
[0133] in, () is an indicator function; it is 1 if the condition in parentheses is met, and 0 if the condition is not met.
[0134] In this way, the test sample adopts the same standardization rules as the reference matrix, ensuring that the test sample and the reference sample are comparable in the same standardized space, avoiding systematic errors introduced by preprocessing differences; the missing mask in the form of an indicator function is simple, computable, and convenient for batch processing.
[0135] In step S140 above, the abundance distribution similarity The expression is:
[0136] ;
[0137] in, For any reference sample, For a set of common comparable features, For feature weights, For the j-th protein feature, Let be the Z-score normalized value of the j-th protein feature in the sample to be evaluated. The Z-score normalized value of the i-th sample and j-th protein feature in the reference matrix;
[0138] Missing pattern similarity The expression is:
[0139] ;
[0140] in, This represents the total number of valid features retained after filtering. Let be the value of the j-th feature in the missing mask matrix of the sample to be evaluated. This represents the missing mask matrix value for the i-th sample and the j-th feature in the reference matrix.
[0141] Abundance distribution similarity uses weighted cosine similarity, which can introduce feature importance weights to accurately quantify the overall similarity of protein expression levels; deletion pattern similarity uses Jaccard similarity coefficient, which specifically evaluates the consistency of protein detection or deletion patterns. The combination of the two achieves a dual similarity evaluation of expression level and presence pattern.
[0142] In step S150 above, the shrinkage coefficient is set based on the actual number of overlapping features, specifically including:
[0143] Let the actual number of overlapping features be... for:
[0144] ;
[0145] Shrinkage coefficient The expression is:
[0146] ;
[0147] in, These are the preset hyperparameters.
[0148] It is understandable that the fewer overlapping features, the lower the shrinkage coefficient. The smaller the value, the lower the reliability of similarity scores under low overlap conditions, avoiding inflated scores due to a few common features; this is achieved through hyperparameters. The shrinkage intensity can be flexibly adjusted to improve the robustness of evaluation results under different feature coverage scenarios.
[0149] Specifically, the above-mentioned multi-perspective similarity evaluation method includes the following steps:
[0150] 1. Initial parameter data matrix preprocessing: First, construct the original reference dataset. ,in Let n be the number of reference samples (rows), p be the number of features (columns), and NaN be the missing value. The dataset is a set of real numbers. It can be a sample-protein matrix or a combination thereof.
[0151] Row (sample) filtering:
[0152] ;
[0153] Meaning: Calculates the proportion of non-missing features in the i-th sample. Only when the number of valid features in a sample reaches a threshold... It will only be retained at certain times.
[0154] Column (feature) filtering:
[0155] ;
[0156] Meaning: Calculate the detection rate of the j-th feature across all samples. Remove features that occur with extremely low frequencies in the reference set (e.g., below a certain threshold). ).
[0157] The final reference matrix is as follows .
[0158] During row and column filtering:
[0159] This represents the detection value of the i-th sample in the original data matrix on the j-th feature.
[0160] i is the sample index, i = 1, 2, ..., n;
[0161] j is the feature index, j=1,2,...,n;
[0162] () is an indicator function that takes the value 1 when the condition inside the parentheses is true, and takes the value 0 otherwise;
[0163] This represents the proportion of non-missing features in the i-th sample;
[0164] This represents the proportion of the j-th feature that is not missing in all samples;
[0165] This is the row filtering threshold, used to remove samples with too many missing or insufficient information;
[0166] The column filtering threshold is used to limit the minimum effective detection ratio of a feature in a sample.
[0167] This is the filtered reference matrix;
[0168] Reference sample (number of rows) to be retained;
[0169] The number of valid features (columns) to be retained.
[0170] Log-normalization and missing value-preserving Z-score transformation of the reference set: In order to eliminate the influence of dimensions and handle skewed distributions, the reference set is normalized.
[0171] Perform a logarithmic transformation on the reference set to make the data closer to a normal distribution:
[0172] ;
[0173] For each feature The mean is calculated based only on non-missing values. with standard deviation :
[0174] ;
[0175] Define the Z-score matrix (missing values are kept as NaN):
[0176] ;
[0177] Simultaneously define a reference sample missing mask matrix:
[0178] ;
[0179] In the missing-aware standardization process for new samples to be tested, for any new input sample:
[0180] ;
[0181] Perform logarithmic transformation and standardization in the same manner:
[0182] ;
[0183] And define its missing mask:
[0184] ;
[0185] In this embodiment, the method provided by the present invention evaluates similarity from two perspectives: numerical similarity (abundance distribution similarity) and missing pattern similarity calculation (overlap similarity).
[0186] Abundance distribution similarity is calculated using missing-aware weighted cosine similarity, comparing the new sample to be tested with any reference sample. Define a set of common comparable features. :
[0187] ;
[0188] Let the feature weights be... Then numerical similarity Defined as:
[0189] ;
[0190] Abundance distribution similarity mainly measures whether the distribution patterns of components are consistent across data points that are available on both sides.
[0191] Missing pattern similarity (overlapping similarity) uses Jaccard similarity, and the Jaccard similarity of missing patterns is defined as follows:
[0192] ;
[0193] The missing pattern similarity assessment evaluates the degree of overlap in the pattern of "which components are present and which components are missing". If the new sample to be tested is missing a key component that is commonly found in breast milk, this value will be significantly reduced.
[0194] The adaptive shrinkage mechanism for overlapping dimensions (reliability correction) prevents false high similarity caused by too few overlapping features (e.g., only 1-2 overlapping features). A shrinkage coefficient is introduced, assuming the actual number of overlapping features is:
[0195] ;
[0196] Define the shrinkage coefficient:
[0197] ;
[0198] in This is a preset hyperparameter (usually set to 30% of the number of effective features after filtering) used to suppress low-overlapping samples.
[0199] when When smaller, The similarity is rapidly reduced, which has a punitive effect; only when there are enough overlapping dimensions is the similarity fully trusted.
[0200] The comprehensive score generation includes fusion similarity and Top-K score; where, fusion similarity is obtained by weighted fusion of numerical similarity and missing pattern similarity.
[0201] ;
[0202] in This is the similarity weighting coefficient. Generally, β is suggested to be in the range of 0.80-0.95, indicating a high emphasis on numerical similarity while also taking into account overlapping similarity.
[0203] When performing Top-K scoring, the neighbor samples are selected based on their similarity to the input sample scores, choosing the samples with the highest similarity. A set of reference samples The overall score of the new sample to be tested is defined as:
[0204] ;
[0205] And linearly mapped to the interval [0, 100].
[0206] For ease of understanding, the method provided by the present invention will be described below through Examples 1-3.
[0207] Example 1:
[0208] Based on this evaluation method, in the preprocessing of the original parameter data matrix, let when ≥1%, 419 data points were selected from 442 data points that met the requirements; the remaining 404 data points were used to construct a database (excluding 15 randomly selected breast milk samples). When the percentage is ≥80%, 97 characteristic proteins were identified.
[0209] The average scores of the tested infant formula (IF, n=2), whey fortified powder (MFGM, n=5), and cow's milk (Milk, n=6) samples were calculated and then scored. A randomly selected breast milk sample (MR, B110551) was also scored. The scoring results are shown in Table 1 below. From the table, it can be seen that MR scored 85.57, while the scores of IF, MFGM, and Milk were all below 80, showing a pattern of MR > Milk > IF > MFGM. The category coverage was 100% for MR, 64.95% for IF and Milk, and 54.64% for MFGM. When the coverage rate is ≥80%, IF and Milk cover 64.95% of the types of breast milk, and MFGM covers 54.64%. In terms of types, IF, Milk, and MFGM are still missing. To improve the simulation accuracy, it is necessary to supplement the missing proteins.
[0210] Table 1. Overall Scores of the Samples to be Tested
[0211] .
[0212] Example 2:
[0213] Based on this evaluation method, 15 breast milk samples were randomly selected from 419 data points (the first letter of their numbers represents B-Beijing, T-Tangshan, R-Tibet, H-Liuyang, DL-Dalian, L-Luoyang; the last digit of the number represents different lactation stages: 1 represents colostrum within 7 days, 2 represents transitional milk within 15 days, 3 represents mature milk within 1 month, 4 represents mature milk within 2 months, 5 represents mature milk within 3 months, and 6 represents mature milk within 4 months). MFGM protein was analyzed from these samples. The scoring results are shown in Table 2. From the scoring results, it can be seen that among the selected independent samples: 1) The breast milk from the two mothers in Beijing had different scores at different stages, and the trend changes were different. In each stage from B110551 to B1105555, the scores were relatively high and relatively stable; in each stage from B110601 to B110605, the scores of colostrum (B110601) and transitional milk (B110602) were relatively low, below 80 points. Mature milk was above 80 points; (2) Among the breast milk from different regions, the scores of the breast milk samples selected from Beijing, Tibet, Tangshan, and Liuyang were relatively high, above 80 points, and the coverage of types was basically around 90%. The scores of the breast milk samples selected from Dalian and Luoyang were lower, below 80 points, and the coverage of types was relatively low, indicating that there was a lack of MFGM protein types. A single sample cannot represent an entire region, but this scoring system can be used sequentially for: evaluating the MFGM protein in each breast milk sample obtained after evaluation; evaluating the MFGM protein in breast milk from different regions; and evaluating the MFGM protein in breast milk samples from different stages.
[0214] Table 2. Overall scores of different breast milk samples
[0215] .
[0216] Example 3:
[0217] According to this evaluation method, in the initial stage of upgrading infant formula to simulate breast milk, the variety coverage rate is an important indicator. It can guide researchers to measure the degree to which infant formula simulates the variety of breast milk. The initial goal is to supplement the core protein set of breast milk in terms of variety.
[0218] Specifically, when the species coverage is relatively low, this can be achieved by setting [specific parameters] in the first step of the original parameter data matrix preprocessing. The value of is used to construct the core breast milk MFGM protein set. When it is a certain value, it is set by The initial goal is to construct a core breast milk MFGM protein set for different values, thereby completing the core breast milk protein set in terms of variety. For example:
[0219] when When the threshold was 1%, 419 breast milk samples were screened, and 15 breast milk samples were randomly selected for scoring, resulting in 404 data points. Under this premise, the number of characteristic proteins in infant formula and fortified whey powder was counted, as shown in Table 3. When the content is ≥50%, the variety coverage of infant formula is 34.09%, the breast milk MFGM protein set contains 352 characteristic proteins, and 120 characteristic proteins are also present in infant formula; when When the content is ≥80%, the variety coverage of infant formula is 64.95%, and the breast milk MFGM protein set contains 97 characteristic proteins, of which 63 are also present in infant formula; when When the coverage rate is ≥95%, the infant formula type coverage rate is 75.00%. There are 36 characteristic proteins in the breast milk MFGM protein set, and 27 characteristic proteins also appear in infant formula. The characteristic proteins in fortified whey powder were counted in the same way.
[0220] In the initial stage of formula upgrades, the goal is to supplement the core protein set of breast milk in terms of variety. This can be set as follows: When the frequency is ≥95%, 36 characteristic proteins with a frequency of over 95% in breast milk are defined as the core protein set, as shown in Table 4. From Tables 3 and 4, it can be seen that: 1) 8 of these 36 characteristic proteins have a frequency of 100% in breast milk; 2) These proteins play a very important role in infant development, including: typical MFGM core proteins, such as XDH / XO, MFG-E8, BTN1A1, PLIN2, PLIN3, ACSL, CIDE-A, etc.; proteins detectable in MFGM, possibly exosome or glandular cell proteins, such as Clusterin, etc. FASN, CDC42, CNX, GRP78 / 94, Rab10 / 18, etc.; and background proteins such as β-CN, κ-CN, etc.; 3) 27 were detected in infant formula and 24 in whey powder (milk-based raw material). As shown in Table 4, typical MFGM core proteins such as XDH / XO, PLIN2, PLIN3, BTN1A1, BTN1A1, and CIDE-A were detected in both infant formula and whey powder. BAL and ACSL3, which are equally important, were not detected in infant formula, while ACSL3 was detected in whey powder, but BAL was not. Both ACSL3 and BAL play important roles in infant lipid digestion and brain development. ACSL3 affects PUFA metabolism in the brain, and BAL affects fat absorption in newborns; both are essential proteins for infant lipid digestion. However, they were not detected in infant formula with added MFGM proteins. This can be used to trace the reasons for their absence and how to improve the formula.
[0221] Furthermore, proteins such as pIgR, CNP, IgA1, and FOLR1, detected in breast milk, while not core components of MFGM protein, are equally important as they affect the development of the infant's immune system, brain, and nervous system, but were not detected in infant formula. This comparative approach can provide a data foundation for the development of infant formula.
[0222] Table 3. Coverage and Quantity of Characteristic Proteins in Infant Formula and Fortified Whey Powder
[0223] ;
[0224] Table 4. Information on over 95% protein content in breast milk and its presence in infant formula and fortified whey powder.
[0225]
[0226] In step S150 above, a shrinkage coefficient is set based on the actual number of overlapping features, and a fusion similarity is obtained through weighted fusion. A comprehensive score is obtained through Top-K scoring, specifically including:
[0227] The numerical similarity and the missing pattern similarity are weighted and fused to obtain the fused similarity. :
[0228] ;
[0229] in, This refers to the similarity weighting coefficient;
[0230] Select the one with the highest similarity A set of reference samples Define the comprehensive score of the new sample to be tested. for:
[0231] .
[0232] In this way, the contribution of abundance similarity and missing pattern similarity can be flexibly allocated through the weighting coefficient β, combined with the shrinkage coefficient α. i We obtain stable and reliable single-sample fusion similarity; we use Top-K average scoring to weaken the interference of extreme abnormal samples and output a more robust and representative overall simulation score, so as to achieve a quantitative evaluation of the overall similarity between the test sample and real breast milk.
[0233] Specifically, after obtaining the comprehensive simulation score, this invention identifies the core features that lead to a decrease in the score by constructing a local reference center, thus providing a quantitative basis for formula improvement.
[0234] The construction of the Top-K neighbor center (finding a local ideal benchmark) includes the following steps:
[0235] In Z-space, using the filtered Calculate the weighted center vector of the most similar reference samples (Top-K neighbors). :
[0236] For the Top-K neighbors, construct a weighted center in Z-space:
[0237] ;
[0238] Interpreted as the target value of the j-th feature at the local reference center, where the weights are... This is used to establish a tailored breast milk benchmark for the current sample.
[0239] In the definition of feature deviation contribution (localization optimization objective), the degree of deviation of each observed feature of the new sample relative to the reference feature is calculated:
[0240] ;
[0241] in Assign the feature deviation contribution and output the top feature with the largest deviation contribution. One characteristic.
[0242] ;
[0243] A positive value indicates that the nutrient content is higher than the baseline in breast milk; a negative value indicates that it is lower. This directly indicates the priority of formula improvement.
[0244] The feature importance metric (the frequency of each feature in the reference samples) incorporates the frequency of feature occurrence. Weights as an explanation:
[0245] This indicates the number of samples in the reference set that have non-missing values for the feature, used to measure interpretability. If a feature is prevalent in breast milk samples ( If the deviation is large, then its impact on the score is more significant, and improving this feature is of great scientific importance.
[0246] As can be seen from the above, this invention not only provides a comprehensive simulation score for the product, but also, through local weighted center calculation, eliminates background noise and accurately locates the "critical defective protein" that causes impaired simulation performance. According to and Values, taking the top 30 critical defective proteins, for negative and For researchers, a higher percentage of defective proteins is recommended. The content is supplemented according to the numerical value.
[0247] For ease of understanding, the method provided by the present invention will be described below through Example 4.
[0248] Example 4:
[0249] Based on the constructed multi-perspective similarity evaluation system, a systematic evaluation of the MFGM protein composition similarity between infant formula, fortified whey powder, and cow's milk and breast milk was conducted. The system not only outputs similarity scores but also accurately identifies "key defective proteins" that impair simulation accuracy and quantifies their contribution to deviations. Taking infant formula as an example, the Top 30 differential features identified by the evaluation system (Table 5) show that the overall simulation accuracy is not determined by a single protein but is driven by key features that contribute significantly to the score.
[0250] In the evaluation metrics, feature deviation contribution measures the strength of a single protein's influence on the overall similarity score; a larger value indicates a greater impact of the feature on the score. Abundance difference reflects its directional shift between infant formula and breast milk, allowing observation of which proteins are elevated and by how much (positive values), and which proteins are deficient and by how much (negative values). Furthermore, feature detection frequency represents the number of times it was detected in 404 breast milk samples, corresponding to the output feature detection frequency, which can be used to measure explanatory importance. Combining deviation contribution and abundance difference, and introducing the feature detection frequency in breast milk samples as an explanatory weight, helps identify key limiting factors that are prevalent in breast milk and decisive for simulation accuracy, thus providing data support for targeted optimization of infant formula.
[0251] Specifically, the ranking of the contribution of the differential features (Top 30, Table 5) shows:
[0252] 1) Ten proteins showed negative abundance differences: β-CN, LPL, RPN1, GFAT1, KRT10, HSPB1, PDI, MYL12A, CNX, and SDR1. These proteins are more abundant in breast milk MFGM than in infant formula, primarily originating from the natural secretion processes of mammary epithelial cells. Functionally, they encompass key biological processes such as lipid digestion and absorption (e.g., LPL), glycosylation modification (e.g., RPN1, GFAT1), protein folding and conformation maintenance (e.g., PDI, CNX), membrane structural stability (e.g., KRT10, MYL12A), and stress and immune regulation (e.g., HSPB1). Their absence or significant reduction in infant formula reflects that current infant formula simulations of MFGM mainly focus on component-level supplementation and have not yet effectively reconstructed the unique structural integrity, functional synergy, and mammary cell origin characteristics of breast milk MFGM.
[0253] It is worth noting that RPN1, PDI, HSPB1, and LPL are all highly heat-sensitive. The harsh heat treatment processes in infant formula production, such as high-temperature sterilization and spray drying, can easily cause irreversible denaturation, aggregation, or precipitation of these molecular chaperones and enzymes with complex tertiary structures, resulting in a significant decrease in their abundance in the final product. RPN1, PDI, and CNX are endoplasmic reticulum-related background signaling proteins. Due to the industrial high-pressure homogenization process used in infant formula, these tiny vesicle structures are often destroyed. Furthermore, purification processes may remove these non-core components as impurities, leading to a lack of accuracy in the simulation score.
[0254] 2) A total of 20 proteins showed positive abundance differences. Proteins with significantly higher levels in infant formula than in breast milk are mainly concentrated in three areas: substance transport, intracellular transport signaling, and the core MFGM scaffold structure. This excess is often due to the selection of raw materials, fortification, and residual cell debris during processing. Among these, PLIN2, BTN1A1, and ANXA5 are core components of MFGM. ApoE and ABCG2 synergistically participate in transmembrane lipid transport. Currently, high-end formulas enhance activity by adding concentrated MFGM powder derived from cow's milk. However, this industrial addition often prioritizes total quantity, resulting in absolute levels of these components in the formula far exceeding the natural proportions in breast milk. Furthermore, due to the lack of a native lipid environment, they cannot achieve precise membrane localization. Additionally, the release of the Rab family and vesicle transport proteins, which are responsible for vesicle transport within cells, occurs; in natural breast milk secretion, they exist only as trace signals. However, in infant formula production, high-pressure homogenization and intense heat treatment can cause the physical breakdown of the intracellular membrane system in the milk-based raw materials. This releases and enriches large amounts of proteins that were originally encapsulated within organelles, resulting in a negative simulation signal. By optimizing the homogenization pressure and employing low-temperature membrane filtration technology, the enrichment of non-characteristic proteins can be reduced, allowing these indicators to return to the golden ratio of breast milk, thus achieving a qualitative leap from "component simulation" to "structural twinning."
[0255] In summary, this embodiment demonstrates that the simulation of breast milk MFGM is not a simple accumulation of components, but rather relies on the fine regulation of lipid ratios, directional anchoring of core structural proteins, and synergistic effects of functional proteins. In the native breast milk system, lipid composition forms the physical basis. PLIN2, through its hydrophobic bundle structure, embeds itself in the lipid core as an "inner scaffold," BTN1A1 is responsible for membrane encapsulation, and proteins such as ANXA5 participate in membrane fusion and structural stability through a calcium-dependent mode. This highly coupled "structure-function" model determines the unique antiviral, immunomodulatory, and lipid metabolism properties of MFGM. Although existing infant formulas have improved the total amount of MFGM, they exhibit significant structural deviations. Therefore, the development of next-generation infant formulas should not be limited to the apparent approximation of nutrient abundance, but should focus on improving the complete replication of bioactive structural units in breast milk, thereby truly restoring the biological functions of breast milk at the molecular level.
[0256] Table 5. Top 30 Differences Between Infant Formula and Breast Milk
[0257]
[0258] In the above specific embodiments, the breast milk fat globule membrane protein (MFGM) simulation evaluation method based on multi-view similarity weighted fusion provided by the present invention utilizes NanoLC-Orbitrap MS liquid chromatography-mass spectrometry to detect breast milk samples from different regions and lactation stages, constructing a local breast milk MFGM proteomics database covering protein types and their relative abundance; simultaneously acquiring proteomics data of samples to be evaluated (such as formula powder, fortified whey powder, milk, etc.); employing a multi-view similarity evaluation mechanism, it uses missing-aware standardization to process high-dimensional sparse data, fuses numerical pattern similarity and missing pattern similarity, and introduces an adaptive shrinkage mechanism for overlapping dimensions to effectively solve the evaluation bias caused by missing values (NaN) in omics detection, achieving accurate scoring of sample simulation; by constructing a local reference center, it identifies the core deviation features that lead to a decrease in simulation score, providing quantitative basis for the precise auxiliary design and formula improvement of infant formula.
[0259] Specifically, the beneficial technical effects achieved by the present invention through the above technical solution are as follows:
[0260] This invention is the first to construct an MFGM protein database based on ≥400 real breast milk data. Based on this database, the similarity between infant formula and its raw materials and breast milk MFGM protein, as well as the quality of breast milk in different regions and at different lactation stages, are evaluated.
[0261] The scientific rigor and objectivity of the evaluation results are significantly improved by this invention. By processing skewed distributions through logarithmic transformation and combining missing-perceived similarity, this method can effectively shield against interference caused by the high-dimensional sparsity of biological samples, making the simulation score more consistent with biological meaning and avoiding evaluation bias caused by dimensional differences.
[0262] The evaluation method of this invention has extremely high industry applicability and robustness. This method does not require complete experimental data and can be compatible with incomplete datasets generated by different categories of nutrients, different laboratories, different test batches, and even different detection platforms (such as mass spectrometry and chromatography), which greatly reduces the cost and threshold of data cleaning.
[0263] This evaluation system can significantly shorten the formulation development cycle. This is mainly reflected in its precise guidance and reliable evidence. Researchers no longer need to blindly adjust formulations; by analyzing the system's output of the first N deviation features and their directions, they can intuitively identify the key differential proteins causing the decrease in similarity and their contribution. Furthermore, based on the core protein set coverage value, the system optimizes the types of missing MFGM proteins in the evaluated samples. Finally, combined with the feature importance index (nj), it helps researchers distinguish between "statistical error" and "true nutritional differences," prioritizing the optimization of key components commonly found in breast milk, thereby enabling targeted adjustments to raw materials and processes.
[0264] To enhance the market persuasiveness of products, the multi-perspective evaluation system proposed in this invention allows companies to obtain more rigorous and quantifiable evaluation criteria for breast milk simulation, providing data support for product development and technical specifications. This evaluation method based on Top-K neighbors (i.e., the breast milk group closest to the individual) has greater scientific communication value than simply comparing average breast milk values.
[0265] In summary, the method of this invention provides an objective measurement tool for the optimization of infant formula and has significant application value.
[0266] This invention also provides a device for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion, such as... Figure 2 As shown, the device includes:
[0267] The sample acquisition module 210 is used to construct a breast milk MFGM proteomics database covering different regions and different lactation stages, and to acquire proteomics data of the samples to be evaluated.
[0268] The reference data processing module 220 is used to obtain a reference matrix based on the breast milk MFGM proteomics database through row and column filtering, perform logarithmic transformation on the reference matrix, calculate the mean and standard deviation of each feature in the reference matrix, and construct the Z-score matrix and missing mask matrix corresponding to the reference matrix.
[0269] The data processing module 230 is used to obtain the sample matrix to be evaluated based on the proteomics data of the sample to be evaluated through row and column filtering, perform logarithmic transformation on the sample matrix to be evaluated, calculate the mean and standard deviation of each feature in the sample matrix to be evaluated, and construct the Z-score matrix and missing mask matrix corresponding to the sample matrix to be evaluated.
[0270] The similarity calculation module 240 is used to calculate the abundance distribution similarity and missing pattern similarity of the samples to be evaluated.
[0271] The comprehensive scoring calculation module 250 is used to set the shrinkage coefficient based on the actual number of overlapping features, obtain the fusion similarity through weighted fusion, obtain the comprehensive score through Top-K scoring and map it to the [0,100] interval, filter the core protein set according to the set threshold, and compare to obtain the types of missing MFGM proteins in the test sample.
[0272] The evaluation result generation module 260 is used to construct Top-K neighbor centers, calculate the feature deviation contribution and abundance difference of the sample to be evaluated, and output key defective proteins and formulation optimization suggestions based on preset feature importance indicators.
[0273] In some embodiments, a breast milk MFGM proteomics database covering different regions and different lactation stages is constructed, specifically including:
[0274] Raw breast milk samples were collected from different regions and at different stages of lactation, and MFGM protein was extracted from each of the breast milk samples.
[0275] The extracted MFGM protein was pretreated, and the MFGM protein obtained after pretreatment was detected by liquid chromatography-mass spectrometry to obtain the original sample data file.
[0276] The original data file of the sample is input into a pre-stored original database for comparison to obtain the comparison results;
[0277] The original breast milk sample, the original sample data file corresponding to the original breast milk sample, and the comparison results are stored to obtain the breast milk MFGM proteomics database.
[0278] In some embodiments, a reference matrix is obtained based on the breast milk MFGM proteomics database through row and column filtering, specifically including:
[0279] Construct the original reference dataset ,in Let n be the set of real numbers, p be the number of reference samples (rows), and NaN be the number of features (columns). The original reference dataset represents the set of real numbers. Sample-protein matrix;
[0280] The number of reference samples is filtered, and the proportion of non-missing features in the i-th sample is counted. When the number of valid features in the sample reaches a threshold... It shall be retained at that time;
[0281] The number of features is filtered, and the detection rate of the j-th feature in all samples is calculated. Features with a frequency lower than a threshold are removed. Features;
[0282] The reference matrix obtained after filtering by the number of reference samples and features is: .
[0283] In some embodiments, a logarithmic transformation is performed on the reference matrix to calculate the mean and standard deviation of each feature in the reference matrix, and the Z-score matrix and missing mask matrix corresponding to the reference matrix are constructed, specifically including:
[0284] Perform a logarithmic transformation on the reference matrix to make the data closer to a normal distribution:
[0285] ,in, This is the result after logarithmic transformation;
[0286] For each feature The mean is calculated based only on non-missing values. with standard deviation :
[0287] ;
[0288] in, Let j be the value of the i-th sample and the j-th protein feature after logarithmic transformation. Let be the original abundance value of the feature of the i-th sample and the j-th protein. ( ) is the mean calculation function. ( ) is the standard deviation calculation function;
[0289] Z-score matrix corresponding to the reference matrix The expression is:
[0290] ;
[0291] Missing mask matrix corresponding to the reference matrix The expression is:
[0292] .
[0293] In some embodiments, the Z-score matrix corresponding to the sample matrix to be evaluated The expression is:
[0294] ;
[0295] in, is the logarithmic transformation value of the j-th protein feature in the sample to be evaluated;
[0296] Missing mask matrix corresponding to the sample matrix to be evaluated The expression is:
[0297] ;
[0298] in, () is an indicator function.
[0299] In some embodiments, abundance distribution similarity The expression is:
[0300] ;
[0301] in, For any reference sample, For a set of common comparable features, For feature weights, For the j-th protein feature, Let be the Z-score normalized value of the j-th protein feature in the sample to be evaluated. The Z-score normalized value of the i-th sample and j-th protein feature in the reference matrix;
[0302] Missing pattern similarity The expression is:
[0303] ;
[0304] in, This represents the total number of valid features retained after filtering. Let be the value of the j-th feature in the missing mask matrix of the sample to be evaluated. This represents the missing mask matrix value for the i-th sample and the j-th feature in the reference matrix.
[0305] In some embodiments, a shrinkage coefficient is set based on the actual number of overlapping features, specifically including:
[0306] Let the actual number of overlapping features be... for:
[0307] ;
[0308] Shrinkage coefficient The expression is:
[0309] ;
[0310] in, These are the preset hyperparameters.
[0311] In some embodiments, a shrinkage coefficient is set based on the actual number of overlapping features, and a fusion similarity is obtained through weighted fusion. A comprehensive score is obtained through Top-K scoring, specifically including:
[0312] The numerical similarity and the missing pattern similarity are weighted and fused to obtain the fused similarity. :
[0313] ;
[0314] in, This refers to the similarity weighting coefficient;
[0315] Select the one with the highest similarity A set of reference samples Define the comprehensive score of the new sample to be tested. for:
[0316] .
[0317] In the above specific embodiments, the breast milk fat globule membrane protein (MFGM) simulation evaluation device based on multi-view similarity weighted fusion provided by the present invention utilizes NanoLC-Orbitrap MS liquid chromatography-mass spectrometry to detect breast milk samples from different regions and lactation stages, constructing a local breast milk MFGM proteomics database covering protein types and their relative abundance; simultaneously acquiring proteomics data of samples to be evaluated (such as formula powder, fortified whey powder, milk, etc.); employing a multi-view similarity evaluation mechanism, it uses missing-aware standardization to process high-dimensional sparse data, fuses numerical pattern similarity and missing pattern similarity, and introduces an adaptive shrinkage mechanism for overlapping dimensions to effectively solve the evaluation bias caused by missing values (NaN) in omics detection, achieving accurate scoring of sample simulation; by constructing a local reference center, it identifies core deviation features that lead to a decrease in simulation score, providing quantitative basis for precise auxiliary design and formula improvement of infant formula.
[0318] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and model predictions. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The model predictions of the computer device store static and dynamic information data. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0319] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0320] Corresponding to the above embodiments, this invention also provides a computer storage medium containing one or more program instructions. These one or more program instructions are used to execute the method described above.
[0321] The present invention also provides a computer program product, the computer program product including a computer program, the computer program being stored on a non-transitory computer-readable storage medium, and the computer being able to perform the above-described method when the computer program is executed by a processor.
[0322] In this embodiment of the invention, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0323] The various methods, steps, and logic diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor reads information from the storage medium and, in conjunction with its hardware, completes the steps of the above methods.
[0324] The storage medium can be memory, such as volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
[0325] Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.
[0326] Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0327] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0328] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using a combination of hardware and software. When applied as software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of computer programs from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0329] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for evaluating the degree of simulation of human milk fat globule membrane proteins based on multi-dimensional similarity weighted fusion, characterized in that, The method includes: We constructed a breast milk MFGM proteomics database covering different regions and different lactation stages, and obtained proteomics data of the samples to be evaluated. Based on the breast milk MFGM proteomics database, a reference matrix was obtained through row and column filtering. The reference matrix was then logarithmically transformed, and the mean and standard deviation of each feature in the reference matrix were calculated to construct the Z-score matrix and the missing mask matrix corresponding to the reference matrix. Based on the proteomics data of the samples to be evaluated, the sample matrix to be evaluated is obtained through row and column filtering. Logarithmic transformation is performed on the sample matrix to be evaluated, and the mean and standard deviation of each feature in the sample matrix to be evaluated are calculated. The Z-score matrix and missing mask matrix corresponding to the sample matrix to be evaluated are constructed. Calculate the abundance distribution similarity and missing pattern similarity of the samples to be evaluated; The shrinkage coefficient is set based on the actual number of overlapping features, and the fusion similarity is obtained through weighted fusion. The comprehensive score of each sample to be scored is obtained through Top-K scoring. The core protein set is screened according to the set threshold, and the types of missing MFGM proteins in the sample to be tested are obtained by comparison. Construct a Top-K neighbor center, calculate the feature deviation contribution and abundance difference of the sample to be evaluated, and output key defective proteins and formulation optimization suggestions based on the preset feature importance index.
2. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 1, characterized in that, A breast milk MFGM proteomics database covering different regions and different lactation stages was constructed, specifically including: Raw breast milk samples were collected from different regions and at different stages of lactation, and MFGM protein was extracted from each of the breast milk samples. The extracted MFGM protein was pretreated, and the MFGM protein obtained after pretreatment was detected by liquid chromatography-mass spectrometry to obtain the original sample data file. The original data file of the sample is input into a pre-stored original database for comparison to obtain the comparison results; The original breast milk sample, the original sample data file corresponding to the original breast milk sample, and the comparison results are stored to obtain the breast milk MFGM proteomics database.
3. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 1, characterized in that, Based on the breast milk MFGM proteomics database, a reference matrix was obtained through row and column filtering, specifically including: Construct the original reference dataset ,in Let n be the set of real numbers, p be the number of reference samples, and NaN be the number of missing values. The original reference dataset is... This is a sample-protein matrix; The number of reference samples is filtered, and the proportion of non-missing features in the i-th sample is counted. When the number of valid features in the sample reaches a threshold... It shall be retained at that time; The number of features is filtered, and the detection rate of the j-th feature in all samples is calculated. Features with a frequency lower than a threshold are removed. Features; The reference matrix obtained after filtering by the number of reference samples and features is: .
4. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 3, characterized in that, The reference matrix is logarithmically transformed, and the mean and standard deviation of each feature in the reference matrix are calculated. The Z-score matrix and missing mask matrix corresponding to the reference matrix are then constructed, specifically including: Perform a logarithmic transformation on the reference matrix to make the data closer to a normal distribution: ,in, This is the result after logarithmic transformation; For each feature The mean is calculated based only on non-missing values. with standard deviation : ; in, Let j be the value of the i-th sample and the j-th protein feature after logarithmic transformation. Let be the original abundance value of the feature of the i-th sample and the j-th protein. ( ) is the mean calculation function. ( ) is the standard deviation calculation function; Z-score matrix corresponding to the reference matrix The expression is: ; Missing mask matrix corresponding to the reference matrix The expression is: 。 5. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 1, characterized in that, Z-score matrix corresponding to the sample matrix to be evaluated The expression is: ; in, is the logarithmic transformation value of the j-th protein feature in the sample to be evaluated; Missing mask matrix corresponding to the sample matrix to be evaluated The expression is: ; in, () is an indicator function.
6. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 1, characterized in that, similarity of abundance distribution The expression is: ; in, For any reference sample, For a set of common comparable features, For feature weights, For the j-th protein feature, Let be the Z-score normalized value of the j-th protein feature in the sample to be evaluated. The Z-score normalized value of the i-th sample and j-th protein feature in the reference matrix; Missing pattern similarity The expression is: ; in, This represents the total number of valid features retained after filtering. Let be the value of the j-th feature in the missing mask matrix of the sample to be evaluated. This represents the missing mask matrix value for the i-th sample and the j-th feature in the reference matrix.
7. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 6, characterized in that, The shrinkage coefficient is set based on the actual number of overlapping features, specifically including: Let the actual number of overlapping features be... for: ; Shrinkage coefficient The expression is: ; in, These are the preset hyperparameters.
8. The method for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion according to claim 7, characterized in that, A shrinkage coefficient is set based on the actual number of overlapping features, and a fusion similarity is obtained through weighted fusion. A comprehensive score is obtained through Top-K scoring, specifically including: The numerical similarity and the missing pattern similarity are weighted and fused to obtain the fused similarity. : ; in, This refers to the similarity weighting coefficient; Select the one with the highest similarity A set of reference samples Define the comprehensive score of the new sample to be tested. for: 。 9. A device for evaluating the mimicry of breast milk fat globule membrane proteins based on multidimensional similarity weighted fusion, characterized in that, The device includes: The sample collection module is used to construct a breast milk MFGM proteomics database covering different regions and different lactation stages, and to acquire proteomics data of samples to be evaluated. The reference data processing module is used to obtain a reference matrix based on the breast milk MFGM proteomics database through row and column filtering, perform logarithmic transformation on the reference matrix, calculate the mean and standard deviation of each feature in the reference matrix, and construct the Z-score matrix and missing mask matrix corresponding to the reference matrix. The data processing module is used to obtain the sample matrix to be evaluated based on the proteomics data of the sample to be evaluated through row and column filtering, perform logarithmic transformation on the sample matrix to be evaluated, calculate the mean and standard deviation of each feature in the sample matrix to be evaluated, and construct the Z-score matrix and missing mask matrix corresponding to the sample matrix to be evaluated. The similarity calculation module is used to calculate the abundance distribution similarity and missing pattern similarity of the samples to be evaluated. The comprehensive scoring calculation module is used to set the shrinkage coefficient based on the actual number of overlapping features, obtain the fusion similarity through weighted fusion, obtain the comprehensive score of each sample to be scored through Top-K scoring, filter the core protein set according to the set threshold, and compare to find the types of missing MFGM proteins in the sample to be tested. The evaluation result generation module is used to construct Top-K neighbor centers, calculate the feature deviation contribution and abundance difference of the samples to be evaluated, and output key defective proteins and formulation optimization suggestions based on preset feature importance indicators.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Method for multi-dimensionally evaluating similarity between sample and breast milk
CN116646023A
Method for multidimensional evaluation of the similarity of samples to breast milk
US20240345098A1