Evaluation method and device for polypeptide sequence

By constructing a zero-distribution and optimal matching method, the similarity and content of peptide sequences are evaluated, which solves the problem that existing technologies cannot accurately assess the similarity between peptides in infant formula milk powder and breast milk, and realizes comprehensive evaluation of peptide sequences and product development.

CN120895098AActive Publication Date: 2025-11-04MEIWEISHI (BEIJING) HEALTH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511416224.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-04
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing sequence similarity algorithms cannot effectively assess the similarity and content differences of peptide sequences, resulting in an inability to accurately assess the similarity between peptides in infant formula and those in breast milk.

Method used

By constructing zero distribution and optimal matching, combined with a preset sequence similarity algorithm, the similarity and content of peptide sequences are evaluated, the proportion of significantly detected peptide sequences is determined, and a comprehensive evaluation of peptide sequences is achieved.

Benefits of technology

It can assess the similarity between different raw material peptides and breast milk peptides from the perspectives of content and sequence, guide the development of protein raw materials, and develop products with peptide sequences and contents that are closer to those of breast milk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895098A_ABST
    Figure CN120895098A_ABST
Patent Text Reader

Abstract

The invention relates to an evaluation method and device for a polypeptide sequence. Relates to the technical field of polypeptide detection. The method comprises the following steps: detecting a detection polypeptide sequence set of each to-be-detected protein in a to-be-evaluated sample and the content of the detection polypeptide sequence set; determining a first similarity between each detection polypeptide sequence and a target polypeptide sequence based on a preset sequence similarity algorithm, the content of each detection polypeptide sequence, and the content and zero distribution of the target polypeptide sequence in a target polypeptide sequence set corresponding to the reference sample, and further constructing optimal matching; determining a first proportion of a significant detection polypeptide sequence in the to-be-evaluated sample according to the optimal matching and a preset significance level; determining the total content of the significant detection polypeptide sequence in the reference sample based on the content of the target polypeptide sequence corresponding to the significant detection polypeptide sequence; and evaluating the similarity between the to-be-evaluated sample and the reference sample under the preset significance level based on the first proportion and the total content. The similarity between different raw materials and breast milk is evaluated from the angles of polypeptide content and sequence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of polypeptide detection, and particularly relates to an evaluation method and device for polypeptide sequences. BACKGROUND

[0002] In the related art, sequence similarity algorithms are usually used to quantify the similarity between biological sequences (such as DNA, protein sequences) or text sequences, and play a key role in the fields of bioinformatics and natural language processing. Although the sequence similarity algorithm can evaluate the amino acid sequence similarity of the polypeptide sequence, the content of the added or introduced polypeptide in the raw material also needs to be considered, which makes the above algorithm no longer applicable, because the polypeptide with high sequence similarity may have a large content difference. How to provide a method for evaluating the similarity and content of polypeptides in milk powder is a technical problem to be solved. SUMMARY

[0003] Therefore, the present disclosure provides an evaluation method and device for polypeptide sequences.

[0004] According to an aspect of the present disclosure, an evaluation method for polypeptide sequences is provided, which comprises:

[0005] detecting polypeptides in a sample to be evaluated to obtain a set of detected polypeptide sequences corresponding to each protein to be detected and the content of each detected polypeptide sequence in the set of detected polypeptide sequences;

[0006] determining, based on a preset sequence similarity algorithm, the content of each detected polypeptide sequence, the content of a target polypeptide sequence in a set of target polypeptide sequences corresponding to the reference sample, and a zero distribution, a first similarity between each detected polypeptide sequence and the target polypeptide sequence in each target source in the sample to be evaluated;

[0007] determining, according to the optimal matching between the set of detected polypeptide sequences and the set of target polypeptide sequences and a preset significance level, a first proportion of significant detected polypeptide sequences in the sample to be evaluated whose first similarity is greater than or equal to a similarity threshold;

[0008] determining, based on the content of the target polypeptide sequence corresponding to the significant detected polypeptide sequence, the total content of the significant detected polypeptide sequence in the reference sample;

[0009] evaluating, based on the first proportion and the total content, the similarity between the sample to be evaluated and the reference sample at the preset significance level.

[0010] In a possible implementation, each set of detected polypeptide sequences includes a plurality of detected polypeptide sequences derived from the same protein to be detected.

[0011] The zero distribution is constructed based on random polypeptide sequences in each random polypeptide sequence set corresponding to each protein to be detected in the sample to be evaluated and target polypeptide sequences in a target polypeptide sequence set corresponding to the reference sample.

[0012] In a possible implementation, the method further includes:

[0013] Based on the first similarity between the detected polypeptide sequence and the target polypeptide sequence, an optimal match between the detected polypeptide sequence set and the target polypeptide sequence set is constructed.

[0014] In a possible implementation, the method further includes:

[0015] For each protein to be detected in the sample to be evaluated, a random polypeptide sequence set for each protein to be detected is constructed, and each random polypeptide sequence set includes a plurality of random polypeptide sequences derived from the same protein to be detected.

[0016] Based on the preset sequence similarity algorithm, the zero distribution is constructed for a random polypeptide sequence in each random polypeptide sequence set and a target polypeptide sequence in a target polypeptide sequence set corresponding to the reference sample.

[0017] In a possible implementation, the method further includes:

[0018] The reference sample is detected for each target protein to determine a target polypeptide sequence set corresponding to each target protein in the reference sample and a content of each target polypeptide sequence in the target polypeptide sequence set.

[0019] The content of each target polypeptide sequence includes an abundance distribution of the target polypeptide sequence in the reference sample.

[0020] In a possible implementation, for each protein to be detected in the sample to be evaluated, a random polypeptide sequence set for each protein to be detected is constructed, including:

[0021] The amino acid sequence of each protein to be detected in the sample to be evaluated is determined.

[0022] Based on a preset length range and the amino acid sequence of each protein to be detected, a random polypeptide sequence set for each protein to be detected is generated in an enumeration manner.

[0023] In a possible implementation, based on the preset sequence similarity algorithm, the zero distribution is constructed for a random polypeptide sequence in each random polypeptide sequence set and a target polypeptide sequence in a target polypeptide sequence set corresponding to the reference sample, including:

[0024] calculate, based on the preset sequence similarity algorithm, a similarity of a random polypeptide sequence in each of the set of random polypeptide sequences to a target polypeptide sequence corresponding to a target protein matched in the reference sample;

[0025] based on the similarity of the random polypeptide sequence to the corresponding target polypeptide sequence, construct an optimal matching between the set of random polypeptide sequences and the set of target polypeptide sequences;

[0026] based on the optimal matching between the set of random polypeptide sequences and the set of target polypeptide sequences, construct a zero distribution for polypeptide sequences.

[0027] In a possible implementation, based on the preset sequence similarity algorithm, the content of each detection polypeptide sequence, the content of the target polypeptide sequence in the set of target polypeptide sequences corresponding to the reference sample, and the zero distribution, the first similarity between each detection polypeptide sequence and the target polypeptide sequence in each target source in the sample to be evaluated is determined, comprising:

[0028] calculate, based on the preset sequence similarity algorithm, a similarity of a random polypeptide sequence in each of the set of random polypeptide sequences to a target polypeptide sequence corresponding to a target protein matched in the reference sample;

[0029] calculate, based on the preset sequence similarity algorithm, a similarity of a random polypeptide sequence in each of the set of random polypeptide sequences to a target polypeptide sequence corresponding to a target protein matched in the reference sample;

[0030] In the case where the original similarity is greater than or equal to the filtering threshold, based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the set of target polypeptide sequences, and the content of each detection polypeptide sequence, the first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated.

[0031] In the case where the original similarity is less than the filtering threshold, zero is determined as the first similarity between the corresponding detection polypeptide sequence and the target polypeptide sequence.

[0032] In a possible implementation, in the case where the original similarity is greater than or equal to the filtering threshold, based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the set of target polypeptide sequences, and the content of each detection polypeptide sequence, the first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated, comprising:

[0033] In the case where the original similarity is greater than or equal to the filtering threshold, the original similarity is converted into a percentile in the zero distribution to obtain a normalized sequence similarity;

[0034] calculate a content similarity between each of the detection polypeptide sequences and the target polypeptide sequences based on the content of each of the target polypeptide sequences in the set of target polypeptide sequences and the content of each of the detection polypeptide sequences;

[0035] calculate a raw comprehensive score between each of the detection polypeptide sequences and the target polypeptide sequences based on the content similarity and the normalized sequence similarity;

[0036] calculate a first similarity between the detection polypeptide sequence and the target polypeptide sequence based on a target multiplier of each of the target polypeptide sequences and the raw comprehensive score; the target multiplier is calculated based on a preset weight parameter.

[0037] In a possible implementation, according to the optimal matching between the set of detection polypeptide sequences and the set of target polypeptide sequences and a preset significance level, a first proportion of significant detection polypeptide sequences with a first similarity greater than or equal to a similarity threshold in the sample to be evaluated is determined, including:

[0038] According to the optimal matching between the set of detection polypeptide sequences and the set of target polypeptide sequences, a set of raw similarities in the optimal matching is determined, and the set of raw similarities includes the first similarity between the matched detection polypeptide sequence and the target polypeptide sequence in the optimal matching;

[0039] Based on the preset significance level, a similarity threshold is determined from the zero distribution;

[0040] Based on the set of raw similarities, a first number of significant detection polypeptide sequences with a first similarity greater than or equal to the similarity threshold is determined, and a first proportion of the significant detection polypeptide sequences in all detection polypeptide sequences is determined based on the first number.

[0041] According to another aspect of the present disclosure, an evaluation device for polypeptide sequences is provided, and the device includes:

[0042] a polypeptide detection module configured to detect polypeptides in a sample to be evaluated, to obtain a set of detection polypeptide sequences corresponding to each of the polypeptides to be detected and a content of each of the detection polypeptide sequences in the set of detection polypeptide sequences;

[0043] a first similarity determination module configured to determine a first similarity between each of the detection polypeptide sequences and a target polypeptide sequence in each of the target sources in the sample to be evaluated based on a preset sequence similarity algorithm, a content of each of the detection polypeptide sequences, a content of the target polypeptide sequence in a set of target polypeptide sequences corresponding to the reference sample, and a zero distribution;

[0044] The proportion determination module is configured to determine, according to the optimal matching between the set of detected polypeptide sequences and the set of target polypeptide sequences and the preset significance level, a first proportion of significant detected polypeptide sequences in the to-be-evaluated sample that have a first similarity greater than or equal to a similarity threshold.

[0045] The total content determination module is configured to determine, based on a content of the target polypeptide sequence corresponding to the significant detected polypeptide sequence, a total content of the significant detected polypeptide sequence in the reference sample.

[0046] The evaluation module is configured to evaluate, based on the first proportion and the total content, a similarity between the to-be-evaluated sample and the reference sample at the preset significance level.

[0047] In a possible implementation, each of the set of detected polypeptide sequences includes a plurality of detected polypeptide sequences derived from a same to-be-detected protein.

[0048] The zero distribution is constructed in advance based on a random polypeptide sequence in each of a set of random polypeptide sequences corresponding to each of the to-be-detected proteins of the to-be-evaluated sample and a target polypeptide sequence in a set of target polypeptide sequences corresponding to the reference sample.

[0049] In a possible implementation, the device further includes:

[0050] The matching construction module is configured to construct, based on a first similarity between the detected polypeptide sequence and the target polypeptide sequence, an optimal matching between the set of detected polypeptide sequences and the set of target polypeptide sequences.

[0051] In a possible implementation, the device further includes:

[0052] The set construction module is configured to, for a to-be-detected protein in a to-be-evaluated sample, construct a set of random polypeptide sequences for each of the to-be-detected proteins, each of the set of random polypeptide sequences including a plurality of random polypeptide sequences derived from a same to-be-detected protein, the to-be-detected protein being a protein that has a common evolutionary origin and a similar function with a target protein in a reference sample.

[0053] The zero distribution construction module is configured to, based on the preset sequence similarity algorithm, construct, for a random polypeptide sequence in each of the set of random polypeptide sequences and a target polypeptide sequence in a set of target polypeptide sequences corresponding to the reference sample, the zero distribution.

[0054] In a possible implementation, the device further includes:

[0055] The sample detection module is configured to detect each target protein in the reference sample, determine a set of target polypeptide sequences corresponding to each target protein in the reference sample, and determine a content of each target polypeptide sequence in the set of target polypeptide sequences.

[0056] The content of each target polypeptide sequence includes an abundance distribution of the target polypeptide sequence in the reference sample.

[0057] In a possible implementation, for each target protein in the sample to be evaluated, a set of random polypeptide sequences for each target protein is constructed, including:

[0058] The amino acid sequence of each target protein in the sample to be evaluated is determined.

[0059] Based on the preset length range and the amino acid sequence of each target protein, a set of random polypeptide sequences of each target protein is generated by enumeration.

[0060] In a possible implementation, based on the preset sequence similarity algorithm, for each random polypeptide sequence in the set of random polypeptide sequences and each target polypeptide sequence in the set of target polypeptide sequences corresponding to the reference sample, the zero distribution is constructed, including:

[0061] Based on the preset sequence similarity algorithm, the similarity between each random polypeptide sequence in the set of random polypeptide sequences and each target polypeptide sequence in the set of target polypeptide sequences corresponding to the matching target protein in the reference sample is calculated.

[0062] Based on the similarity between the random polypeptide sequence and the corresponding target polypeptide sequence, the optimal matching between the set of random polypeptide sequences and the set of target polypeptide sequences is constructed.

[0063] The zero distribution for polypeptide sequences is constructed based on the optimal matching between the set of random polypeptide sequences and the set of target polypeptide sequences.

[0064] In a possible implementation, based on the preset sequence similarity algorithm, the content of each detection polypeptide sequence, the content of each target polypeptide sequence in the set of target polypeptide sequences corresponding to the reference sample, and the zero distribution, the first similarity between each detection polypeptide sequence and each target polypeptide sequence in each target source in the sample to be evaluated is determined, including:

[0065] Based on the preset sequence similarity algorithm, the original similarity between each detection polypeptide sequence and each target polypeptide sequence is calculated.

[0066] Based on the preset threshold and the zero distribution, a filtering threshold is calculated.

[0067] In a case where the original similarity is greater than or equal to the filtering threshold, a first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated based on the original similarity, the zero distribution, a content of each target polypeptide sequence in the set of target polypeptide sequences, and a content of each detection polypeptide sequence.

[0068] In a case where the original similarity is less than the filtering threshold, zero is determined as the first similarity between the corresponding detection polypeptide sequence and the target polypeptide sequence.

[0069] In a possible implementation, in a case where the original similarity is greater than or equal to the filtering threshold, a first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated based on the original similarity, the zero distribution, a content of each target polypeptide sequence in the set of target polypeptide sequences, and a content of each detection polypeptide sequence, including:

[0070] In a case where the original similarity is greater than or equal to the filtering threshold, the original similarity is converted into a percentile in the zero distribution to obtain a normalized sequence similarity.

[0071] A content similarity between each detection polypeptide sequence and the target polypeptide sequence is calculated based on a content of each target polypeptide sequence in the set of target polypeptide sequences and a content of each detection polypeptide sequence.

[0072] An original comprehensive score between each detection polypeptide sequence and the target polypeptide sequence is calculated based on the content similarity and the normalized sequence similarity.

[0073] A first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated based on a target multiplier of each target polypeptide sequence and the original comprehensive score, the target multiplier being calculated based on a preset weight parameter.

[0074] In a possible implementation, a first proportion of significant detection polypeptide sequences with a first similarity greater than or equal to a similarity threshold in the sample to be evaluated is determined according to an optimal matching between the set of detection polypeptide sequences and the set of target polypeptide sequences and a preset significance level, including:

[0075] An original similarity set in the optimal matching is determined according to the optimal matching between the set of detection polypeptide sequences and the set of target polypeptide sequences, the original similarity set including a first similarity between a matched detection polypeptide sequence and a target polypeptide sequence in the optimal matching.

[0076] A similarity threshold is determined from the zero distribution based on a preset significance level.

[0077] A first number of the significant detected polypeptide sequences whose first similarity is greater than or equal to the similarity threshold is determined based on the original similarity set, and a first proportion of the significant detected polypeptide sequences in all detected polypeptide sequences is determined based on the first number.

[0078] According to another aspect of the present disclosure, there is provided an evaluation device for polypeptide sequences, comprising a memory, a processor and a computer program stored in the memory, the processor executing the computer program to implement the steps of the above method.

[0079] According to another aspect of the present disclosure, there is provided a non-volatile computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the above method.

[0080] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program, or a non-volatile computer readable storage medium carrying a computer program, the computer program being executed by a processor to implement the steps of the above method.

[0081] The evaluation method and device for polypeptide sequences provided by the embodiments of the present disclosure can evaluate the closeness of different raw material polypeptides and breast milk polypeptides from the perspective of content and sequence, guide the development of protein raw materials, and thus develop products with polypeptide sequences and content closer to breast milk.

[0082] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0083] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.

[0084] Figure 1 A flow chart of an evaluation method for polypeptide sequences according to an embodiment of the present disclosure is shown.

[0085] Figure 2 A flow chart of an evaluation method for polypeptide sequences according to an embodiment of the present disclosure is shown.

[0086] Figure 3 A zero distribution histogram constructed in Example 1 is shown.

[0087] Figure 4 A zero distribution histogram constructed in Example 2 is shown.

[0088] Figure 5 A block diagram of a device 1900 for evaluation of polypeptide sequences according to an exemplary embodiment is shown. DETAILED DESCRIPTION

[0089] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in different drawings represent the same or similar elements / function. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.

[0090] As used herein, the terms "include," "comprise," "have," or their variants are open-ended, and include one or more stated features, integers, elements, steps, components or functions but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions or groups thereof.

[0091] When an element is referred to as being "connected", "coupled", "responsive", or "correlated" to another element, it can be directly connected, coupled, responsive, or correlated to the other element, or intervening elements can be present.

[0092] Although the terms first, second, third, and the like can be used herein to describe various elements / operations, such elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Thus, a first element / operation in some embodiments can be termed a second element / operation in other embodiments without departing from the teachings of the present inventive concept.

[0093] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0094] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions, procedures, components, and circuits can be omitted so as not to unnecessarily obscure the teachings of the present disclosure. It should be noted that the use of the term "or" in the context of describing alternative embodiments is not to be construed as a complete list of alternatives.

[0095] Endogenous free polypeptides in breast milk are a class of active substances with important biological functions, which are mainly derived from milk proteins such as casein and whey protein. Among them, various bioactive polypeptides such as casein phosphopeptide released by specific enzymatic hydrolysis of casein, osteonectin and polymeric immunoglobulin receptor in whey protein are important sources of endogenous free polypeptides. In terms of efficacy, endogenous free polypeptides in breast milk show multidimensional biological activity. On the one hand, some polypeptides have antibacterial activity and can inhibit the growth of pathogenic bacteria such as Escherichia coli and Staphylococcus aureus, thereby building an immune barrier for infants to resist the invasion of external pathogens. On the other hand, some polypeptides can regulate the intestinal flora structure of infants, promote the proliferation of beneficial bacteria such as Bifidobacterium, improve the intestinal microecological environment, and thus improve the intestinal digestive and absorptive function. In addition, endogenous free polypeptides also participate in the regulation of the immune response process of infants, enhance the activity of immune cells, promote the synthesis of immunoglobulin, and help the development and improvement of the immune system of infants. At the same time, some polypeptides can bind to the surface receptors of intestinal cells, promote the growth and repair of intestinal epithelial cells, and play a significant role in the establishment of the intestinal mucosal barrier function of infants.

[0096] Breast milk is the gold standard for the design of infant formula, and infant formula contains a certain amount of polypeptides through addition (such as direct addition of casein phosphopeptide) and / or raw material introduction (such as introduction by hydrolyzed whey protein powder, whey protein powder and raw cow milk). However, because infant formula is usually based on cow milk and goat milk as basic raw materials, there are inherent differences in protein sequences between cow milk and goat milk and breast milk, which inevitably leads to differences in polypeptide sequences between infant formula and breast milk, making it difficult to evaluate the similarity of polypeptides between breast milk and infant formula.

[0097] In related technologies, sequence similarity algorithms are usually used to quantify the similarity between biological sequences (such as DNA, protein sequences) or text sequences, and play a key role in the fields of bioinformatics and natural language processing. Among them, the global matching algorithm is represented by the Needleman-Wunsch algorithm, which is based on the dynamic programming principle and performs global optimal alignment on two complete sequences by constructing a two-dimensional score matrix. This algorithm assigns appropriate scores to base or amino acid matching, mismatching, insertion and deletion at each position, and obtains the global best alignment result by traversing the matrix and backtracking the path. It is suitable for analyzing sequences with similar lengths and high overall similarity. The local matching algorithm is typically represented by the Smith-Waterman algorithm, which also uses the dynamic programming strategy, but allows alignment to start and end at any position in the sequence, focusing on finding the highest similarity local region in the sequence. This algorithm sets a threshold to control the start and end conditions of the alignment, effectively identifying conserved domains or functional fragments in the sequence, and is particularly suitable for analyzing sequences with local similarity, such as finding conserved motifs in protein families.

[0098] While the above algorithms can assess the amino acid sequence similarity of peptide sequences, the content of peptides added to or introduced into the raw materials must also be considered. This renders the algorithms unsuitable because peptides with high sequence similarity may exhibit significant content differences. Therefore, providing a method to assess both the similarity and content of peptides in milk powder is a pressing technical problem that needs to be solved.

[0099] To address the aforementioned technical problems, this disclosure provides a method and apparatus for evaluating peptide sequences. It can assess the similarity between different raw material peptides and breast milk peptides from both content and sequence perspectives, guiding the development of protein raw materials and thereby developing products with peptide sequences and contents more closely resembling those of breast milk.

[0100] like Figure 1 , Figure 2 As shown, the method for evaluating polypeptide sequences provided in this disclosure includes steps S100-S108.

[0101] In step S100, the reference sample is tested for each target protein to determine the set of target polypeptide sequences corresponding to each target protein in the reference sample and the content of each target polypeptide sequence in the set of target polypeptide sequences. The number of reference samples can be one or more, and this disclosure does not limit this.

[0102] In the case where the sample to be evaluated is infant formula milk powder, the reference sample can be breast milk. The set of target polypeptide sequences for each target protein can then be represented as:

[0103]

[0104] in, This is a set of target polypeptide sequences derived from the target protein p in a reference sample (such as breast milk). For the first One derived from the target protein The target polypeptide sequence, the target polypeptide sequence It is composed of multiple amino acids. It is derived from the target protein in the reference sample. The number of target polypeptide sequences. Target protein. This could be proteins such as casein and whey protein in the reference sample that can serve as sources of bioactive substances, i.e. P1 and P2 represent various proteins in the reference sample that can provide active substances.

[0105] In some embodiments, the content of each target polypeptide sequence may include the abundance distribution of the target polypeptide sequences in the reference samples. When the number of reference samples is... In the case that each target polypeptide sequence , the abundance of the target polypeptide sequence in the reference samples can be represented as:

[0106]

[0107] Further, the abundance distribution of each target polypeptide sequence can be calculated as:

[0108]

[0109] wherein, represents the lower limit value of the abundance of the target polypeptide sequence , . represents the upper limit value of the abundance of the target polypeptide sequence , . represents the average value or the median of the abundance of the target polypeptide sequence , then or . Wherein, and may be the percentiles specified by the user, such as l = 5 representing 5% and r = 95 representing 95%.

[0110] In step S101, for each protein to be detected in the sample to be evaluated, a set of random polypeptide sequences for each protein to be detected is constructed, and each set of random polypeptide sequences includes a plurality of random polypeptide sequences derived from the same protein to be detected. The protein to be detected can be multiple.

[0111] In some embodiments, the protein to be detected can be part or all of the various proteins present in the sample to be evaluated, so as to ensure comprehensive and detailed evaluation of the sample to be evaluated. In some embodiments, the protein to be detected can also be each homologous protein in the sample to be evaluated. The homologous protein can be a protein that has a common evolutionary origin and similar function as the target protein in the reference sample. In this way, the evaluation can be simplified while improving the efficiency and speed of the evaluation.

[0112] In one possible implementation, if the protein to be detected is each protein present in the sample to be evaluated, step S101 can include: determining the amino acid sequence of each protein to be detected R in the sample to be evaluated; and then generating a set of random polypeptide sequences of each protein to be detected R by enumeration based on a preset length range and the amino acid sequence of each protein to be detected R ​If the protein to be tested is a homologous protein, step S101 may include: based on the target protein in the reference sample. The corresponding target proteins in the sample to be evaluated were identified. Protein to be tested The amino acid sequence; then, based on a preset length range and each of the proteins to be detected; The amino acid sequence is used to generate each of the proteins to be tested by enumeration. A collection of random polypeptide sequences .

[0113] Each protein to be tested A collection of random polypeptide sequences It can be represented as:

[0114]

[0115] in, The protein to be tested in the sample to be evaluated (such as infant formula milk powder) The source of the target polypeptide sequence set. For the first One originating from the protein to be tested A random polypeptide sequence, the random polypeptide sequence Includes multiple amino acids, It is derived from the protein to be tested in the sample to be evaluated. The number of random polypeptide sequences. Protein to be tested. It can be the target protein in the reference sample. Proteins with a common evolutionary origin and similar functions. Proteins to be tested in the sample to be evaluated. It refers to proteins in the sample to be evaluated that can provide bioactive substances, that is... R1, R2... represent various proteins in the sample to be evaluated that can provide active substances.

[0116] In some embodiments, a preset length range can be set in advance based on the amino acid sequence length of the target peptide sequence in the reference sample, so that the length of the amino acid sequence of the random peptide sequence generated by enumeration is within the preset length range. For example, the preset length range can be 5-20 amino acids, that is, the amino acid sequence of each random peptide sequence consists of 5-20 amino acids.

[0117] In step S102, based on a preset sequence similarity algorithm, for each of the random polypeptide sequence sets... random polypeptide sequences in The set of target polypeptide sequences corresponding to the reference sample The target polypeptide sequence , the zero distribution of polypeptide sequences is constructed .

[0118] In the embodiment, the preset sequence similarity algorithm can be set according to actual needs, such as Smith-Waterman algorithm, Needleman Wunsch algorithm, BLAST algorithm, etc., and the present disclosure does not limit this.

[0119] In a possible implementation, step S102 can include: based on the preset sequence similarity algorithm, calculating the similarity between the random polypeptide sequence in each random polypeptide sequence set and the target polypeptide sequence corresponding to the matching target protein in the reference sample; based on the similarity between the random polypeptide sequence and the corresponding target polypeptide sequence, constructing the optimal matching between the random polypeptide sequence set and the target polypeptide sequence set , the zero distribution of polypeptide sequences is constructed .

[0120] In some embodiments, the similarity calculated based on the preset sequence similarity algorithm can be represented by the following function:

[0121]

[0122] wherein, is a parameter for calculating the similarity between and . The parameter may include a substitution matrix and a gap penalty to accurately quantify sequence similarity to obtain the alignment score between and . That is, represents the similarity between the target polypeptide sequence and the random polypeptide sequence calculated based on the parameter .

[0123] The optimal matching may be the matching that maximizes the similarity (i.e., alignment score) between and , and the optimal matching between the random polypeptide sequence set and the target polypeptide sequence set can be determined based on the following formula: :

[0124]

[0125] and then based on the optimal matching Constructing a zero distribution for the polypeptide sequence .

[0126] The zero distribution may be represented by the following formula:

[0127]

[0128] In step S103, the polypeptides in the sample to be evaluated are detected to obtain a set of detection polypeptide sequences corresponding to each of the proteins to be detected and the content of each of the detection polypeptide sequences in the set of detection polypeptide sequences, each of the set of detection polypeptide sequences including a plurality of detection polypeptide sequences derived from the same protein to be detected.

[0129] The set of detection polypeptide sequences can be represented as .

[0130] wherein, is the detection polypeptide sequence derived from the protein to be detected in the target source k of the sample to be evaluated, the detection polypeptide sequence is composed of a plurality of amino acids and actually exists in the sample to be evaluated, is the number of detection polypeptide sequences derived from the protein to be detected in the target source k of the sample to be evaluated, i.e. . The sample to be evaluated can include at least one target source k. For the set of detection polypeptide sequences , the content of each detection polypeptide sequence

[0131] may be determined by detection. In some embodiments, the content of the detection polypeptide sequence may be represented in the following manner: , wherein, represents the content of . represents the content of .

[0132] In step S104, based on the preset sequence similarity algorithm, the content of each of the detection polypeptide sequences, the content of the target polypeptide sequence in the set of target polypeptide sequences corresponding to the reference sample, and the zero distribution, the first similarity between each of the detection polypeptide sequences and the target polypeptide sequence in each of the target sources in the sample to be evaluated is determined.

[0133] In this embodiment, the first similarity can be a weighted comprehensive similarity between each of the detection polypeptide sequences and the target polypeptide sequence . Wherein,​ , The first similarity can be calculated by the following steps 1-4:

[0134] Step 1, based on a preset sequence similarity algorithm, calculate the original similarity between the detected polypeptide sequence and the target polypeptide sequence. , , .

[0135] Step 2, based on a preset threshold, calculate a filtering threshold. . Wherein, is the quantile of , that is, is the quantile of . For example, if , take the 95% quantile of , such as 80, as the filtering threshold 80.

[0136] Step 3, in the case where the original similarity is greater than or equal to the filtering threshold, based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the target polypeptide sequence set, and the content of each detected polypeptide sequence, calculate the first similarity between the detected polypeptide sequence and the target polypeptide sequence.

[0137] In some embodiments, the calculation process of step 3 can be:

[0138] In the case where the original similarity is greater than or equal to , convert the original similarity to the percentile in the zero distribution to obtain a normalized sequence similarity. Wherein, the normalized sequence similarity . Then, based on the content of each target polypeptide sequence in the target polypeptide sequence set and the content of each detected polypeptide sequence, calculate the content similarity between each detected polypeptide sequence and the target polypeptide sequence .

[0139] Wherein, in the case where the content information of the target polypeptide sequence is the content distribution , the content similarity between and is the abundance similarity. Then the process of calculating the abundance similarity may be:

[0140] Let , calculate the abundance similarity​​ :

[0141] If , then .

[0142] If , then .

[0143] If , then .

[0144] If , then .

[0145] Further, based on the abundance similarity (i.e., the content similarity) and the normalized sequence similarity , the original comprehensive score between the detection polypeptide sequence and the target polypeptide sequence is calculated; , the target multiplier for the detection polypeptide sequence is calculated based on the preset weight parameter β; , and the first similarity (i.e., the weighted comprehensive similarity) is calculated based on the target multiplier and the original comprehensive score. .

[0146] The original comprehensive score may be the product between the abundance similarity and the normalized sequence similarity , i.e., .

[0147] The target multiplier may be calculated based on the following formula: .

[0148] Further, the first similarity may be the product between the target multiplier and the original comprehensive score , i.e., .

[0149] Step 4, in the case that the original similarity is less than , zero is determined as the first similarity between the detection polypeptide sequence and the target polypeptide sequence.

[0150] In step S105, based on the first similarity between the detection polypeptide sequence and the target polypeptide sequence, the optimal matching between the detection polypeptide sequence set and the target polypeptide sequence set is constructed.

[0151] Among them, the optimal match can be achieved. It can make and First similarity between The longest match, that is:

[0152]

[0153] In step S106, based on the optimal match between the detection peptide sequence set and the target peptide sequence set... and the preset significance level The first similarity in the samples to be evaluated is determined. Greater than or equal to the similarity threshold The first proportion of significant detection of peptide sequences In this way, the proportion of detectable peptide sequences in the sample to be evaluated that are similar to the target peptide sequences in the reference sample can be reflected by the first proportion. The similarity can be reflected by the degree of sequence similarity, and the higher the value of the first proportion, the stronger the conservation of the sequence.

[0154] In some embodiments, step S106 may include: determining the optimal match. The original similarity set in Based on the preset significance level Distribution from zero Calculate the similarity threshold Based on the original similarity set The first similarity is determined to be greater than or equal to the similarity threshold. The first significant detection of peptide sequences and according to the first quantity The significant detection peptide sequence was determined to have the highest proportion among all detection peptide sequences. .

[0155] Among them, the original similarity set It can be represented as: .

[0156] Among them, significance level This is a preset value, representing the salience level specified by the user. quantiles ( ), that is yes of Quantiles.

[0157] The significant detection of the polypeptide sequence corresponds to the first similarity. and The detection of polypeptide sequences.

[0158] Further,

[0159] wherein, is the number of matches in .

[0160] In step S107, based on the content of the target polypeptide sequence corresponding to the significantly detected polypeptide sequence, the total content of the significantly detected polypeptide sequence in the reference sample is determined . In this way, the total content reflects the total abundance level of the significantly detected polypeptide sequence in the reference sample, and the similarity is reflected from the similarity degree of the content. The higher the value of the total content, the more the main functional polypeptides matched to the reference sample.

[0161] wherein, the total content may be represented as:

[0162]

[0163] In step S108, based on the first proportion and the total content , the similarity between the to-be-evaluated sample and the reference sample at the preset significance level is evaluated.

[0164] To further illustrate the method provided by the embodiments of the present disclosure, the following will be illustratively described through Examples 1 and 2.

[0165] Firstly, both the reference sample and the to-be-evaluated sample need to be processed as follows, and hereinafter, both the reference sample and the to-be-evaluated sample are referred to as samples, and specifically,

[0166] First step, sample defatting treatment and polypeptide extraction and purification:

[0167] Take each sample , centrifuge at 4°C, 16,000 × g for 10 minutes. Absorb the defatted milk layer below the fat layer. Repeat the above centrifugation and absorption steps until no visible fat layer is formed.

[0168] Second step, protein precipitation:

[0169] Add trichloroacetic acid (TCA) solution to 200 μL of defatted milk. After mixing well with a vortex mixer, centrifuge at 4°C, 3,000 × g for 10 minutes. Collect the supernatant.

[0170] Third step, polypeptide adsorption concentration and impurity removal:

[0171] The supernatant (enriched in peptide fragments) was treated with a 200 mg bed volume of C18 solid phase extraction (SPE) cartridge to adsorb and concentrate the polypeptides and to purify from impurities (mainly oligosaccharides and salts). The polypeptides were eluted from the SPE cartridge with a solution containing 80% acetonitrile (ACN) and 0.1% trifluoroacetic acid (TFA). The eluate sample was collected and vacuum freeze-dried. The sample was reconstituted before mass spectrometry detection.

[0172] Fourth step, liquid chromatography-mass spectrometry (LC-MS / MS) analysis:

[0173] Liquid chromatography separation: analysis was performed using an EASY-nLC 1200 liquid chromatography system, etc. The elution gradient was set as follows: the proportion of mobile phase B (containing acetonitrile with 0.1% formic acid) was linearly increased from 5% to 30% in 50 minutes (mobile phase A was aqueous solution containing 0.1% formic acid). Subsequently, the proportion of mobile phase B was rapidly increased to 50% in 3 minutes for column cleaning (both mobile phases A and B contained 50% water and 0.1% formic acid).

[0174] Mass spectrometry detection: data were collected in positive ion mode using a Thermo Scientific Orbitrap Fusion Lumos mass spectrometer, etc. The key parameters of the Thermo Scientific Orbitrap Fusion Lumos mass spectrometer were set as follows: the electrospray voltage was 2400 V, the mass spectrometry scanning range was 400 m / z - 1500 m / z, the first mass spectrometry resolution was 120,000 (defined at m / z 200), the automatic gain control (AGC) target value was 5e5, the maximum injection time was 50 ms, the fragmentation mode was collision-induced dissociation (CID), the collision energy was 35%, the mass spectrometry cycle time was 3 seconds (data-dependent acquisition mode, automatic selection of precursor ions), the precursor ion exclusion time was 60 seconds after fragmentation (mass tolerance ± 10 ppm), and the fragmentation precursor ion selection criterion was the ion with the highest signal intensity; the ion intensity threshold was 2e4; the charge state was 1 to 5; the fragment ion detection was ion trap automatic scanning range detection.

[0175] ​​​​​​​Database search and data analysis: The raw spectra were searched against databases using the software Thermo Proteome Discoverer (v2.4). The human protein database in uniprot was used to analyze breast milk samples, and the bovine milk protein database in uniprot was used to analyze formula milk powder and raw material samples. The search settings were as follows: "Min. Precursor Mass" was 300 Da, "Max Precursor Mass" was 5000 Da, the enzyme type was No-Enzyme (Unspecific), the shortest length of polypeptide was set to 4, the longest length was set to 144, "Precursor Mass Tolerance" was 10ppm, "fragment Mass Tolerance" was 0.8 Da, methionine oxidation and serine and threonine phosphorylation were set as variable modifications, and there was no fixed modification. Only polypeptides with high confidence were included (P < 0.01), and polypeptide sequences with multiple modifications were grouped into a polypeptide for counting. The value measured was the number of unique polypeptide sequences determined in a sample. The abundance measure was the area under the curve of the elution peak (ion intensity). Among them, the polypeptide content was analyzed according to the percentage of abundance in total abundance.

[0176] Example 1: Similarity evaluation analysis of different samples and breast milk β-casein peptides using the above method

[0177] 1. Reference sample, sample to be evaluated

[0178] 200 breast milk samples (sample coverage in Beijing, Guangzhou, Weihai, Jinhua, Lanzhou, Chengdu, Wuhan and Harbin) were randomly selected from the breast milk sample library as reference samples, and the content of free polypeptides in breast milk was detected according to the above method to obtain a target polypeptide sequence set corresponding to the target protein (β-casein) and the content of each target polypeptide sequence in the target polypeptide sequence set. Formula bovine milk powder A (adding ordinary casein phosphopeptide), B (not adding casein phosphopeptide) and casein phosphopeptide raw material C (β-casein phosphopeptide) were selected as samples to be evaluated, and the same method was used to detect the content of free polypeptides to obtain a detection polypeptide sequence set and the content of each detection polypeptide sequence in the detection polypeptide sequence set.

[0179] 2. Construction of bovine milk β-casein random polypeptide sequence and breast milk β-casein peptide similarity zero distribution

[0180] The FASTA file containing the amino acid sequence of bovine (Bos taurus) protein was obtained from uniprot.org. The FASTA file was parsed using the biopython library in Python. A script was written to enumerate bovine β-casein peptide sequences within a preset length range (5-20), resulting in 3400 unique random peptide sequences. Peptides derived from β-casein in 200 breast milk samples were screened, yielding 1464 unique target peptide sequences (breast milk β-casein peptides). A similarity index was constructed using a self-written Python script. Figure 3 The zero distribution D shown null (See step S102). This example uses the Smith-Waterman algorithm as the preset sequence similarity algorithm to calculate sequence similarity (gap opening penalty is set to -10, gap extension penalty is set to -0.5, and amino acid substitution matrix is ​​set to BLOSUM62). Figure 3 As shown, when the preset length range is set to 5~20, the minimum value of the random polypeptide sequence of bovine β-casein and the similarity of breast milk β-casein peptide to the zero distribution is 11, the maximum value is 85, the median is 39, and the 95% confidence interval is [21, 66].

[0181] 3. Calculation of weighted comprehensive similarity (i.e., first similarity)

[0182] Calculate the percentage abundance of each target polypeptide sequence derived from β-casein in the breast milk sample relative to the total β-casein abundance. The 25th and 75th percentiles are used as the content range for that target polypeptide sequence. Calculate the percentage of the detected β-casein polypeptide sequences in formula milk powders A and B and casein phosphopeptide raw material C relative to the total β-casein peptides in that sample. Then, based on the above step S104, construct a weighted comprehensive similarity score. The weighted overall similarity is obtained by using the Hungarian algorithm. Maximum value of the sum and its optimal matching .

[0183] The above method was then used to assess the free polypeptide content in formula milk powders A, B, and C, obtaining the number of β-casein polypeptide sequences in the samples to be evaluated. The number of detected β-casein polypeptide sequences varied considerably among different samples, with sample A having the highest number of 428, sample C having the lowest number of only 116, and sample B having a middle number of 168.

[0184] Furthermore, Set to 10, Set to 50, set different The given calculations were further performed using a self-written Python script. Below and The results are shown in Table 1 below. As can be seen from Table 1, no matter what the value of is, the abundance similarity is always C much higher than A and B, indicating that the abundance similarity of the β-casein peptides from C is much higher than that of the two formula powders, which is consistent with the characteristics of C (C is a casein phosphopeptide raw material derived from β-casein hydrolysis). In addition, the abundance similarity is always , which is also consistent with the formula of the two (A adds ordinary casein phosphopeptide, and B does not add casein phosphopeptide). When , the sequence similarity is , which is consistent with the abundance similarity result.

[0185] Table 1: Calculation results of Example 1

[0186]

[0187] Example 2: Similarity analysis of different samples and breast milk osteopontin (OPN) peptides using the above method

[0188] 1. Reference sample, sample to be evaluated

[0189] Randomly select 200 breast milk samples (samples covering Beijing, Guangzhou, Weihai, Jinhua, Lanzhou, Chengdu, Wuhan and Harbin) from a breast milk sample library as reference samples, and detect the content of free polypeptides in the breast milk according to the above method to obtain a target polypeptide sequence set corresponding to the target protein OPN and the content of each target polypeptide sequence in the target polypeptide sequence set. Select formula milk powder D (adding whey protein containing osteopontin), E (not adding whey protein containing osteopontin), ordinary raw cow milk (F), concentrated whey protein (WPC1) and concentrated whey protein (WPC2) as samples to be evaluated, and detect the content of free polypeptides using the same method to obtain a detection polypeptide sequence set and the content of each detection polypeptide sequence in the detection polypeptide sequence set.

[0190] 2. Construction of breast milk OPN peptide similarity zero distribution of random OPN polypeptide sequences in cow milk

[0191] The fasta file of the amino acid sequence of the bovine (Bos taurus) protein was obtained from uniprot.org, the fasta file was parsed using the biopython library of python, and the script was written to enumerate the bovine milk OPN polypeptide sequence according to the preset length range (5-20). A total of 4264 unique polypeptides (random polypeptide sequences) were obtained. Screening the OPN-derived polypeptides in 200 breast milk samples, a total of 628 unique target polypeptide sequences (breast milk beta-casein peptides) were obtained. The python script independently written was used to construct the similarity zero distribution D as shown in Figure 4 . null The example uses the Smith-Waterman algorithm to calculate the sequence similarity (gap opening penalty is set to -10, gap extension penalty is set to -0.5, and amino acid substitution matrix is set to BLOSUM62). As shown in Figure 4 , when the preset length range is set to 5-20, the minimum value of the similarity zero distribution of the bovine milk OPN random polypeptide sequence and the breast milk OPN target polypeptide sequence is 16, the maximum value is 102, the median is 52, and the 95% confidence interval is [22, 91].

[0192] 3. Weighted comprehensive similarity (i.e. first similarity) calculation

[0193] The abundance of each OPN-derived target polypeptide sequence in the breast milk sample accounted for the percentage of the total beta-casein abundance, and the 5th and 95th percentiles of the percentage were used as the content range of the target polypeptide sequence. The percentage of the detected OPN-derived polypeptide sequence in the total OPN peptide in formula milk powder D (added with bone bridge protein-containing whey protein), E (not added with bone bridge protein-containing whey protein), ordinary raw bovine milk, concentrated whey protein WPC1 and concentrated whey protein WPC2 was calculated. Further based on the above step S104, the weighted comprehensive similarity was constructed. The Hungarian algorithm was used to solve, and the maximum value of the weighted comprehensive similarity sum and its optimal match .

[0194] Further, the above method was used to evaluate the free polypeptide content in formula milk powder D (added with bone bridge protein-containing whey protein), E (not added with bone bridge protein-containing whey protein), ordinary raw bovine milk (F), concentrated whey protein (WPC1) and concentrated whey protein (WPC2). The number of OPN detection polypeptide sequences in different samples is quite different, with 68 in sample D, 46 in sample E, 55 in sample F, and 88 in sample WPC2.

[0195] Further, the is set to 1, is set to 50, and different Further calculation by self-written python script gives the results shown in Table 2 below. As can be seen from Table 2, the abundance similarity is the highest when the sequence similarity is set to 0.25 (i.e. the sequence similarity is higher than the 75th percentile of the zero distribution), and the abundance similarity is the highest for WPC2 and the lowest for F. The abundance similarity of formula milk powder D fortified with osteopontin is higher than that of formula milk powder E without osteopontin at different levels.

[0196] Table 2: Calculation results of Example 2

[0197]

[0198] The embodiments of the present disclosure also provide an evaluation device for polypeptide sequences, which comprises:

[0199] a polypeptide detection module, configured to detect polypeptides in a sample to be evaluated, to obtain a set of detected polypeptide sequences corresponding to each protein to be detected and a content of each detected polypeptide sequence in the set of detected polypeptide sequences;

[0200] a first similarity determination module, configured to determine a first similarity between each detected polypeptide sequence and a target polypeptide sequence in a set of target polypeptide sequences corresponding to a reference sample in each target source in the sample to be evaluated based on a preset sequence similarity algorithm, the content of each detected polypeptide sequence, the content of the target polypeptide sequence in the set of target polypeptide sequences corresponding to the reference sample and a zero distribution;

[0201] a proportion determination module, configured to determine a first proportion of significant detected polypeptide sequences in the sample to be evaluated based on an optimal matching between the set of detected polypeptide sequences and the set of target polypeptide sequences and a preset significance level, wherein the first proportion of significant detected polypeptide sequences is a proportion of detected polypeptide sequences with a first similarity greater than or equal to a similarity threshold value;

[0202] a total content determination module, configured to determine a total content of the significant detected polypeptide sequences in the reference sample based on a content of a target polypeptide sequence corresponding to the significant detected polypeptide sequence;

[0203] an evaluation module, configured to evaluate a similarity between the sample to be evaluated and the reference sample at the preset significance level based on the first proportion and the total content.

[0204] In a possible implementation, each set of detected polypeptide sequences includes a plurality of detected polypeptide sequences derived from the same protein to be detected.

[0205] ​​​The zero distribution is constructed based on random polypeptide sequences in each random polypeptide sequence set corresponding to each protein to be detected in the sample to be evaluated and target polypeptide sequences in a target polypeptide sequence set corresponding to the reference sample.

[0206] In a possible implementation, the apparatus further includes:

[0207] The matching construction module is configured to construct an optimal match between the detection polypeptide sequence set and the target polypeptide sequence set based on the first similarity between the detection polypeptide sequence and the target polypeptide sequence.

[0208] In a possible implementation, the apparatus further includes:

[0209] The set construction module is configured to, for each protein to be detected in the sample to be evaluated, construct a random polypeptide sequence set for each protein to be detected, each random polypeptide sequence set including a plurality of random polypeptide sequences derived from the same protein to be detected, the protein to be detected being a protein having a common evolutionary origin and similar function as a target protein in the reference sample.

[0210] The zero distribution construction module is configured to, for each random polypeptide sequence in each random polypeptide sequence set and each target polypeptide sequence in the target polypeptide sequence set corresponding to the reference sample, construct the zero distribution based on the preset sequence similarity algorithm.

[0211] In a possible implementation, the apparatus further includes:

[0212] The sample detection module is configured to detect each target protein in the reference sample, determine a target polypeptide sequence set corresponding to each target protein in the reference sample and a content of each target polypeptide sequence in the target polypeptide sequence set.

[0213] The content of each target polypeptide sequence includes an abundance distribution of the target polypeptide sequence in the reference sample.

[0214] In a possible implementation, for each protein to be detected in the sample to be evaluated, a random polypeptide sequence set for each protein to be detected is constructed, including:

[0215] The amino acid sequence of each protein to be detected in the sample to be evaluated is determined.

[0216] The random polypeptide sequence set of each protein to be detected is generated in an enumeration manner based on a preset length range and the amino acid sequence of each protein to be detected.

[0217] In a possible implementation, based on the preset sequence similarity algorithm, the zero distribution is constructed for each random polypeptide sequence in the random polypeptide sequence set and each target polypeptide sequence in the target polypeptide sequence set corresponding to the reference sample, including:

[0218] Based on the preset sequence similarity algorithm, the similarity between each random polypeptide sequence in the random polypeptide sequence set and each target polypeptide sequence in the target polypeptide sequence set corresponding to the matching target protein in the reference sample is calculated.

[0219] Based on the similarity between each random polypeptide sequence and each target polypeptide sequence corresponding thereto, the optimal matching between the random polypeptide sequence set and the target polypeptide sequence set is constructed.

[0220] The zero distribution for polypeptide sequences is constructed based on the optimal matching between the random polypeptide sequence set and the target polypeptide sequence set.

[0221] In a possible implementation, based on the preset sequence similarity algorithm, the content of each detection polypeptide sequence, the content of each target polypeptide sequence in the target polypeptide sequence set corresponding to the reference sample, and the zero distribution, the first similarity between each detection polypeptide sequence and each target polypeptide sequence in each target source in the sample to be evaluated is determined, including:

[0222] Based on the preset sequence similarity algorithm, the original similarity between each detection polypeptide sequence and each target polypeptide sequence is calculated.

[0223] Based on the preset threshold and the zero distribution, a filtering threshold is calculated.

[0224] In a case where the original similarity is greater than or equal to the filtering threshold, the first similarity between each detection polypeptide sequence and each target polypeptide sequence is calculated based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the target polypeptide sequence set, and the content of each detection polypeptide sequence.

[0225] In a case where the original similarity is less than the filtering threshold, zero is determined as the first similarity between the corresponding detection polypeptide sequence and the target polypeptide sequence.

[0226] In a possible implementation, in a case where the original similarity is greater than or equal to the filtering threshold, the first similarity between each detection polypeptide sequence and each target polypeptide sequence is calculated based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the target polypeptide sequence set, and the content of each detection polypeptide sequence, including:

[0227] In a case where the original similarity is greater than or equal to the filtering threshold, the original similarity is converted into a percentile in the zero distribution, to obtain a normalized sequence similarity;

[0228] Based on the content of each target polypeptide sequence in the target polypeptide sequence set and the content of each detection polypeptide sequence, a content similarity between each detection polypeptide sequence and the target polypeptide sequence is calculated.

[0229] Based on the content similarity and the normalized sequence similarity, an original comprehensive score between each detection polypeptide sequence and the target polypeptide sequence is calculated.

[0230] Based on a target multiplier of each target polypeptide sequence and the original comprehensive score, a first similarity between the detection polypeptide sequence and the target polypeptide sequence is calculated. The target multiplier is calculated based on a preset weight parameter.

[0231] In a possible implementation, according to the optimal matching between the detection polypeptide sequence set and the target polypeptide sequence set and a preset significance level, a first proportion of significant detection polypeptide sequences with a first similarity greater than or equal to a similarity threshold in the sample to be evaluated is determined, including:

[0232] According to the optimal matching between the detection polypeptide sequence set and the target polypeptide sequence set, an original similarity set in the optimal matching is determined, and the original similarity set includes a first similarity between a matched detection polypeptide sequence and a target polypeptide sequence in the optimal matching.

[0233] Based on a preset significance level, a similarity threshold is determined from the zero distribution.

[0234] Based on the original similarity set, a first number of significant detection polypeptide sequences with a first similarity greater than or equal to the similarity threshold is determined, and based on the first number, a first proportion of the significant detection polypeptide sequences in all detection polypeptide sequences is determined.

[0235] It should be noted that although the above-mentioned embodiments are introduced as examples for the evaluation method and device for polypeptide sequences, those skilled in the art can understand that the present disclosure should not be limited thereto. In fact, the user can flexibly set each step and module according to personal preferences and / or actual application scenarios, as long as it conforms to the technical solutions of the present disclosure. For example, after the polypeptide samples (i.e. samples to be evaluated) obtained under different process conditions are detected by the polypeptide group, the raw material development engineer can use the above-mentioned evaluation method and device for polypeptide sequences to evaluate the similarity of different polypeptide samples and target source polypeptides (i.e. reference samples) such as breast milk, thereby guiding the raw material development; further, the raw material development engineer can establish a prediction model of the similarity of polypeptide samples under different process conditions and breast milk polypeptides, accurately predict the closeness of polypeptide samples under different processes and breast milk, thereby improving the efficiency of raw material development. The product development engineer and the raw material engineer can use the method to evaluate the closeness of polypeptide raw materials from different suppliers and breast milk polypeptides. If the raw material contains polypeptides from multiple protein sources, the similarity can be calculated respectively and then the average value can be calculated for evaluation.

[0236] It should be emphasized that the polypeptide sequence similarity function mentioned in the present method The user can freely choose according to the needs and research targets. In addition to the commonly used Smith-Waterman algorithm, Needleman Wunsch algorithm, BLAST algorithm, other arbitrary types of sequence similarity functions based on pre-trained models (such as ProteinBert) can also be used, and the present disclosure does not limit this. Further, the user can also customize the similarity function, for example, the function can be added to the peptide segment translation modification (phosphorylation and glycosylation, etc.) and other information, and the present disclosure does not limit this.

[0237] In some embodiments, the device provided by the embodiments of the present disclosure has functions or contains modules which can be used to execute the methods described in the above method embodiment, and the specific implementation can refer to the description of the above method embodiment. For the sake of brevity, it will not be repeated here.

[0238] The embodiments of the present disclosure also provide an evaluation device for polypeptide sequences, comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the computer program to realize the steps of the above-mentioned method.

[0239] The embodiments of the present disclosure also provide a non-volatile computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the steps of the above-mentioned method.

[0240] The embodiment of the present disclosure further provides a computer program product, comprising a computer program or a nonvolatile computer readable storage medium carrying the computer program, and the computer program is executed by a processor to implement the steps of the above method.

[0241] Figure 5 is a block diagram of an apparatus 1900 for evaluation for polypeptide sequences according to an exemplary embodiment. For example, the apparatus 1900 can be provided as a server or a terminal device. Referring to Figure 5 , the apparatus 1900 comprises a processing component 1922, which further comprises one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can comprise one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.

[0242] The apparatus 1900 can further comprise a power supply component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input output interface 1958 (I / O interface). The apparatus 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0243] In an exemplary embodiment, a nonvolatile computer readable storage medium, such as the memory 1932 comprising computer program instructions, is also provided, and the above computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method.

[0244] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0245] Computer programs (or computer readable program instructions) described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device from a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0246] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0247] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0248] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0249] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0250] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0251] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements in the art, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for evaluating polypeptide sequences, characterized in that, The method includes: The peptides in the sample to be evaluated are detected to obtain a set of detection peptide sequences corresponding to each protein to be tested and the content of each detection peptide sequence in the set of detection peptide sequences; Based on a preset sequence similarity algorithm, the content of each of the detected peptide sequences, the content and zero distribution of the target peptide sequences in the target peptide sequence set corresponding to the reference sample, the first similarity between each of the detected peptide sequences in each target source of the sample to be evaluated and the target peptide sequence is determined; Based on the optimal matching between the set of detected polypeptide sequences and the set of target polypeptide sequences and the preset significance level, the first proportion of significant detected polypeptide sequences in the sample to be evaluated with a first similarity greater than or equal to a similarity threshold is determined. Based on the content of the target polypeptide sequence corresponding to the significant detection polypeptide sequence, the total content of the significant detection polypeptide sequence in the reference sample is determined; Based on the first proportion and the total content, the similarity between the sample to be evaluated and the reference sample is evaluated at the preset significance level.

2. The method according to claim 1, characterized in that, Each of the aforementioned detection polypeptide sequence sets includes multiple detection polypeptide sequences derived from the same protein to be detected; The zero distribution is constructed in advance based on random polypeptide sequences in the set of random polypeptide sequences of each protein to be tested corresponding to the sample to be evaluated, and target polypeptide sequences in the set of target polypeptide sequences corresponding to the reference sample.

3. The method according to claim 1, characterized in that, The method further includes: Based on the first similarity between the detected peptide sequence and the target peptide sequence, an optimal match is constructed between the set of detected peptide sequences and the set of target peptide sequences.

4. The method according to claim 1, characterized in that, The method further includes: For each protein to be tested in the sample to be evaluated, a set of random polypeptide sequences for each protein to be tested is constructed, wherein each set of random polypeptide sequences includes multiple random polypeptide sequences derived from the same protein to be tested. Based on the preset sequence similarity algorithm, the zero distribution is constructed for the random polypeptide sequences in each set of random polypeptide sequences and the target polypeptide sequences in the set of target polypeptide sequences corresponding to the reference sample.

5. The method according to claim 1, characterized in that, The method further includes: The reference sample is tested for each target protein to determine the set of target polypeptide sequences corresponding to each target protein in the reference sample and the content of each target polypeptide sequence in the set of target polypeptide sequences; The content of each target polypeptide sequence includes the abundance distribution of the target polypeptide sequences in the reference sample.

6. The method according to claim 2, characterized in that, For each protein to be tested in the sample to be evaluated, a set of random polypeptide sequences for each protein to be tested is constructed, including: The amino acid sequences of each of the proteins to be tested in the sample to be evaluated were determined; Based on a preset length range and the amino acid sequence of each of the proteins to be tested, a random polypeptide sequence set of each of the proteins to be tested is generated by enumeration.

7. The method according to claim 4, characterized in that, Based on the preset sequence similarity algorithm, the null distribution is constructed for the random polypeptide sequences in each of the random polypeptide sequence sets and the target polypeptide sequences in the target polypeptide sequence set corresponding to the reference sample, including: Based on a preset sequence similarity algorithm, the similarity between the random polypeptide sequence in each set of random polypeptide sequences and the target polypeptide sequence in the set of target polypeptide sequences corresponding to the target protein in the reference sample is calculated. Based on the similarity between the random polypeptide sequence and the corresponding target polypeptide sequence, an optimal match is constructed between the set of random polypeptide sequences and the set of target polypeptide sequences; A zero distribution for peptide sequences is constructed based on the optimal matching between the set of random peptide sequences and the set of target peptide sequences.

8. The method according to claim 1, characterized in that, Based on a preset sequence similarity algorithm, the content of each of the detected peptide sequences, the content and zero distribution of the target peptide sequences in the target peptide sequence set corresponding to the reference sample, the first similarity between each of the detected peptide sequences in each of the target sources in the sample to be evaluated and the target peptide sequence is determined, including: Based on a preset sequence similarity algorithm, the original similarity between each detected polypeptide sequence and the target polypeptide sequence is calculated. The filtering threshold is calculated based on the preset threshold and the zero distribution; If the original similarity is greater than or equal to the filtering threshold, a first similarity between the detected polypeptide sequence and the target polypeptide sequence is calculated based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the target polypeptide sequence set, and the content of each detected polypeptide sequence. If the original similarity is less than the filtering threshold, zero is determined as the first similarity between the detected polypeptide sequence and the target polypeptide sequence.

9. The method according to claim 8, characterized in that, If the original similarity is greater than or equal to the filtering threshold, a first similarity between the detected polypeptide sequence and the target polypeptide sequence is calculated based on the original similarity, the zero distribution, the content of each target polypeptide sequence in the target polypeptide sequence set, and the content of each detected polypeptide sequence, including: If the original similarity is greater than or equal to the filtering threshold, the original similarity is converted to the percentile in the zero distribution to obtain the normalized sequence similarity. Based on the content of each target polypeptide sequence in the target polypeptide sequence set and the content of each detection polypeptide sequence, the content similarity between each detection polypeptide sequence and the target polypeptide sequence is calculated; The original comprehensive score between each detected polypeptide sequence and the target polypeptide sequence is calculated based on the content similarity and the normalized sequence similarity. Based on the target multipliers of each target polypeptide sequence and the original comprehensive score, the first similarity between the detected polypeptide sequence and the target polypeptide sequence is calculated; the target multipliers are calculated based on preset weight parameters.

10. The method according to claim 1, characterized in that, Based on the optimal match and preset significance level between the set of detected peptide sequences and the set of target peptide sequences, the first proportion of significant detected peptide sequences in the sample to be evaluated with a first similarity greater than or equal to a similarity threshold is determined, including: Based on the optimal match between the set of detected peptide sequences and the set of target peptide sequences, an original similarity set in the optimal match is determined, wherein the original similarity set includes the first similarity between the detected peptide sequences and the target peptide sequences matched in the optimal match; A similarity threshold is determined from the null distribution based on a preset significance level. Based on the original similarity set, a first number of significant detection polypeptide sequences with a first similarity greater than or equal to the similarity threshold is determined, and based on the first number, a first proportion of the significant detection polypeptide sequences in all detection polypeptide sequences is determined.

11. An evaluation device for polypeptide sequences, characterized in that, The device includes: The peptide detection module is used to detect peptides in the sample to be evaluated, and to obtain a set of detection peptide sequences corresponding to each protein to be tested and the content of each detection peptide sequence in the set of detection peptide sequences. The first similarity determination module is used to determine the first similarity between each of the detected peptide sequences in each target source of the sample to be evaluated and the target peptide sequence based on a preset sequence similarity algorithm, the content of each of the detected peptide sequences, the content of the target peptide sequences in the target peptide sequence set corresponding to the reference sample, and the zero distribution. The proportion determination module is used to determine the first proportion of significant detection peptide sequences in the sample to be evaluated whose first similarity is greater than or equal to a similarity threshold, based on the optimal matching between the detection peptide sequence set and the target peptide sequence set and a preset significance level. The total content determination module is used to determine the total content of the significant detection peptide sequence in the reference sample based on the content of the target peptide sequence corresponding to the significant detection peptide sequence; An evaluation module is used to evaluate the similarity between the sample to be evaluated and the reference sample at the preset significance level, based on the first proportion and the total content.

12. An evaluation device for polypeptide sequences, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.

13. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Methods and apparatus for comparing, aligning, and optimizing protein sequences

    CA2415968A1

  • Prediction method for microscene risk causal relationship of multivariate sequence of infant formula milk powder production sequence

    CN115796343A

  • Method for multi-dimensionally evaluating similarity between sample and breast milk

    CN116646023A

  • Protein multi-sequence alignment method and device, storage medium and electronic equipment

    CN117037913A

  • Milk quality detection system and method based on gene sequence

    CN117805326A