Method for processing protein sequencing data, related methods and apparatuses

By acquiring protein sequencing data from target slices and calculating protein signal intensity thresholds, the problem of distinguishing between positive and negative signals in protein sequencing data was solved using methods such as inter-class variance and information entropy, thus improving the accuracy of foreground sites.

CN120164523BActive Publication Date: 2026-04-21SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
Filing Date
2023-12-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The lack of reasonable and accurate methods in the current technology to distinguish between positive and negative signals in protein sequencing data makes it difficult to accurately determine foreground sites.

Method used

By acquiring protein sequencing data from the target slice, the frequency and probability of protein signal intensity values ​​are determined, the protein signal intensity threshold is calculated, and methods such as inter-class variance and information entropy are used to classify protein signal intensity and determine foreground sites.

Benefits of technology

This enables a reasonable and accurate distinction between positive and negative signals, improves the precision of foreground sites, and enhances the accuracy of spatial proteomics research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164523B_ABST
    Figure CN120164523B_ABST
Patent Text Reader

Abstract

This application provides a method, related methods, and apparatus for processing protein sequencing data, relating to the field of protein sequencing technology. The method for processing protein sequencing data involves acquiring protein sequencing data of a target protein from a target slice, determining the frequency of each protein signal intensity value in the sequencing data, and calculating the probability of each protein signal intensity value based on the frequency; calculating a protein signal intensity threshold to classify protein signal intensity based on the protein signal intensity value and probability; and identifying sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold as foreground sites corresponding to the target protein. Since the protein signal intensity threshold is calculated based on the protein signal intensity value and probability, using the protein signal intensity value and probability as defining factors for the protein signal intensity threshold helps to reasonably and accurately determine the threshold used to distinguish positive and negative signals, thereby facilitating more precise identification of foreground sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of protein sequencing technology, specifically to a method for processing protein sequencing data, related methods, and apparatus. Background Technology

[0002] Space proteomics sequencing is a novel technique for studying the spatial expression patterns of proteins in tissues. Based on differences in quantitative methods, space proteomics techniques can be divided into three main categories: label-based imaging, mass spectrometry, and sequencing-based methods. Among these, sequencing-based space proteomics has gained wider application due to its high research efficiency and high resolution.

[0003] It should be noted that the detected antibody protein signal intensity includes two cases: first, a protein-positive signal, i.e., specific binding; and second, a protein-negative signal, i.e., non-specific binding. It should be pointed out that the industry often uses a threshold to distinguish between protein-positive and protein-negative signals in order to identify foreground sites characterizing specific binding and thus analyze related data of these foreground sites. However, in related technologies, there is no reasonable and accurate method for determining the threshold used to distinguish between positive and negative signals. Therefore, the accurate identification of foreground sites has become a major problem that urgently needs to be solved in the industry. Summary of the Invention

[0004] The main objective of this application is to propose a method, related methods and apparatus for processing protein sequencing data, which aims to reasonably and accurately determine the threshold for distinguishing positive and negative signals, so as to more accurately determine the foreground sites.

[0005] To achieve the above objectives, a first aspect of this application provides a method for processing protein sequencing data, the method comprising:

[0006] Obtain protein sequencing data of the target protein in the target slice, wherein the protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites;

[0007] Determine the frequency of each protein signal intensity value in the protein sequencing data, and calculate the probability of each protein signal intensity value occurring based on the frequency;

[0008] Based on the protein signal intensity value and the probability, a protein signal intensity threshold is calculated to classify the protein signal intensity into strong and weak categories.

[0009] Sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold are identified as foreground sites corresponding to the target protein.

[0010] In some embodiments of this application, the protein signal intensity threshold for classifying the strength of the protein signal based on the protein signal intensity value and the probability calculation includes:

[0011] The protein signal intensity values ​​are divided into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​using any target protein signal intensity value as a threshold.

[0012] Based on the probability, calculate the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein;

[0013] Based on the inter-class variance value, a protein signal intensity threshold is determined among multiple target protein signal intensity values ​​to classify the protein signal intensity into strong and weak categories.

[0014] In some embodiments of this application, calculating the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein based on the probability includes:

[0015] A first probability sum is obtained by calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the first type of protein, and a second probability sum is obtained by calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the second type of protein;

[0016] The average signal intensity of the first protein is calculated based on the first probability sum, the signal intensity value of the first type of protein, and the corresponding probability; and the average signal intensity of the second protein is calculated based on the second probability sum, the signal intensity value of the second type of protein, and the corresponding probability.

[0017] The inter-class variance between the signal intensity values ​​of the first type of protein and the second type of protein is calculated based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity.

[0018] In some embodiments of this application, the step of dividing the protein signal intensity value into a first type of protein signal intensity value and a second type of protein signal intensity value based on any target protein signal intensity value as a threshold includes:

[0019] Determine any one of the multiple protein signal intensity values ​​as the target protein signal intensity value;

[0020] Protein signal intensity values ​​with values ​​less than the target protein signal intensity value are classified as first-class protein signal intensity values, and protein signal intensity values ​​with values ​​not less than the target protein signal intensity value are classified as second-class protein signal intensity values.

[0021] In some embodiments of this application, determining the protein signal intensity threshold for classifying protein signal intensity based on the inter-class variance value among multiple target protein signal intensity values ​​includes:

[0022] Determine the inter-class variance value corresponding to the signal intensity value of each target protein;

[0023] The target protein signal intensity value with the largest inter-class variance is determined as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values.

[0024] In some embodiments of this application, the step of calculating the inter-class variance between the signal intensity values ​​of the first class of proteins and the signal intensity values ​​of the second class of proteins based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity includes:

[0025] Calculate the squared difference between the mean signal intensity of the first protein and the mean signal intensity of the second protein;

[0026] The inter-class variance between the signal intensity values ​​of the first type of protein and the second type of protein is obtained by multiplying the sum of the first probability, the sum of the second probability, and the square of the difference.

[0027] In some embodiments of this application, the protein signal intensity threshold for classifying the strength of the protein signal based on the protein signal intensity value and the probability calculation includes:

[0028] The protein signal intensity values ​​are divided into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​using any target protein signal intensity value as a threshold.

[0029] Based on the probability, calculate the first information entropy corresponding to the signal intensity value of the first type of protein and the second information entropy corresponding to the signal intensity value of the second type of protein;

[0030] Based on the sum of the first information entropy and the second information entropy, a protein signal intensity threshold is determined from multiple target protein signal intensity values ​​to classify the strength of the protein signal.

[0031] To achieve the above objectives, a second aspect of this application provides a method for detecting antigen protein distribution regions, the method comprising:

[0032] Obtain biological slices to be tested;

[0033] The biological slices were immersed in a solution of a pre-set antibody protein, and then washed to obtain the target slices.

[0034] Protein sequencing was performed on the target slice to obtain protein sequencing data;

[0035] The protein sequencing data is processed using the protein sequencing data processing method described in any one of the first aspect embodiments to obtain the foreground sites of the preset antibody protein;

[0036] The distribution area of ​​the antigen protein in the biological slice is determined based on the pre-defined foreground site of the antibody protein.

[0037] To achieve the above objectives, a third aspect of this application provides a protein sequencing data processing apparatus, comprising:

[0038] The first acquisition unit is used to acquire protein sequencing data of the target protein in the target slice, wherein the protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites;

[0039] The first determining unit is used to determine the frequency of each protein signal intensity value in the protein sequencing data, and to calculate the probability of each protein signal intensity value appearing based on the frequency.

[0040] A threshold division unit is used to calculate a protein signal intensity threshold for classifying the strength of a protein signal based on the protein signal intensity value and the probability.

[0041] The second determining unit is used to determine the sequencing sites whose protein signal intensity values ​​are greater than the protein signal intensity threshold as the foreground sites corresponding to the target protein.

[0042] To achieve the above objectives, a fourth aspect of this application provides an antigen protein distribution region detection device, comprising:

[0043] The second acquisition unit is used to acquire the biological slice to be tested;

[0044] The slide cleaning unit is used to wet the biological slide with a preset antibody protein solution and to clean the wetted biological slide to obtain the target slide.

[0045] A protein sequencing unit is used to perform protein sequencing on the target slice to obtain protein sequencing data.

[0046] A sequencing data processing unit is used to process the protein sequencing data according to the protein sequencing data processing method described in any one of the first aspect embodiments to obtain the foreground site of the preset antibody protein;

[0047] The third determining unit is used to determine the distribution area of ​​the antigen protein in the biological slice based on the foreground site of the preset antibody protein.

[0048] To achieve the above objectives, a fifth aspect of this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the protein sequencing data processing method of the first aspect or the antigen protein distribution region detection method of the second aspect.

[0049] To achieve the above objectives, a sixth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the protein sequencing data processing method of the first aspect or the antigen protein distribution region detection method of the second aspect.

[0050] To achieve the above objectives, a seventh aspect of this application provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the protein sequencing data processing method of the first aspect or the antigen protein distribution region detection method of the second aspect.

[0051] The protein sequencing data processing method, related methods, and apparatus proposed in this application obtain protein sequencing data of the target protein in a target slice, determine the frequency of each protein signal intensity value in the protein sequencing data, and calculate the probability of each protein signal intensity value based on the frequency. The protein sequencing data includes protein signal intensity values ​​of the target protein measured at multiple sequencing sites. Further, a protein signal intensity threshold for classifying protein signal intensity based on the protein signal intensity value and probability is calculated. Further still, sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold are identified as foreground sites corresponding to the target protein. Since the protein signal intensity threshold for classifying protein signal intensity is calculated based on the protein signal intensity value and probability, using the protein signal intensity value and probability as defining factors for the protein signal intensity threshold helps to reasonably and accurately determine the threshold used to distinguish positive and negative signals, thereby facilitating more precise identification of foreground sites. Attached Figure Description

[0052] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0053] Figure 1 This is a flowchart of the protein sequencing data processing method provided in the embodiments of this application;

[0054] Figure 2 yes Figure 1 The flowchart of step S103 in the process;

[0055] Figure 3 yes Figure 2 The flowchart of step S201 in the text;

[0056] Figure 4 yes Figure 2 The flowchart of step S202 in the document;

[0057] Figure 5 yes Figure 4 The flowchart of step S403 in the process;

[0058] Figure 6 yes Figure 2 The flowchart of step S203 in the process;

[0059] Figure 7 This is a schematic diagram showing the correspondence between protein signal intensity thresholds and inter-class variance values;

[0060] Figure 8 yes Figure 1 The flowchart of step S103 in the process;

[0061] Figure 9 This is a flowchart of the antigen protein distribution region detection method provided in the embodiments of this application;

[0062] Figure 10 This is a schematic diagram of the spatial distribution of the target slice after removing the negative protein signal based on the protein signal intensity threshold;

[0063] Figure 11 This is a flowchart of the protein sequencing data processing apparatus provided in the embodiments of this application;

[0064] Figure 12 This is a flowchart of the antigen protein distribution region detection device provided in the embodiments of this application;

[0065] Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0067] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0068] Space proteomics sequencing is a novel technique for studying the spatial expression patterns of proteins in tissues. Based on differences in quantitative methods, space proteomics techniques can be divided into three main categories: label-based imaging, mass spectrometry, and sequencing-based methods. Among these, sequencing-based space proteomics has gained wider application due to its high research efficiency and high resolution.

[0069] Spatial proteomics is the scientific field that studies the structure, assembly, and function of proteins in three-dimensional space. Proteins are essential components of living organisms, and their structure and function are crucial for maintaining life activities. The aim of spatial proteomics research is to reveal the expression patterns of proteins in physiological and pathological processes through comprehensive analysis of their expression on the cell surface in tissues. It should be understood that spatial proteomics is a novel omics approach to studying the spatial expression patterns of proteins in tissues.

[0070] It's important to note that the principle of spatial proteomics lies in using high-precision laser capture microdissection (LCM) to extract regions of interest (ROIs) from tissues or cells. The ROI is defined in machine vision and image processing, using shapes such as rectangles, circles, ellipses, and irregular polygons to delineate the area to be processed from the image. Furthermore, optimized ultra-small samples are non-destructively extracted and enzyme-digested into peptides, which are then analyzed using high-sensitivity mass spectrometry to determine the expression characteristics of these proteins at different spatial locations.

[0071] It should be noted that in spatial proteomics, to analyze the expression characteristics of cell surface antigen proteins at different spatial locations, it is necessary to use labeled antibody proteins to bind to epitopes of cell surface antigen proteins to mark the location of the antigen proteins. During this process, specific and non-specific binding of antibody proteins to antigen proteins may occur. Therefore, the detected antibody protein signal intensity includes two cases: first, a protein-positive signal, i.e., specific binding; and second, a protein-negative signal, i.e., non-specific binding, i.e., a negative signal. It should be noted that the industry often uses a threshold to distinguish between protein-positive and protein-negative signals in order to identify foreground sites characterizing specific binding and thus analyze relevant data of foreground sites. However, in related technologies, there is no reasonable and accurate method for determining the threshold used to distinguish between positive and negative signals. Therefore, the accurate identification of foreground sites has become a major problem that urgently needs to be solved in the industry.

[0072] The main objective of this application is to propose a method, related methods and apparatus for processing protein sequencing data, which aims to reasonably and accurately determine the threshold for distinguishing positive and negative signals, so as to more accurately determine the foreground sites.

[0073] The following explanation, in conjunction with the accompanying drawings, will provide further details.

[0074] To achieve the above objectives, this application provides a method for processing protein sequencing data. The method for processing protein sequencing data may include, but is not limited to, steps S101 to S104 described below.

[0075] Step S101: Obtain protein sequencing data of the target protein in the target slice. The protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites.

[0076] Step S102: Determine the frequency of each protein signal intensity value in the protein sequencing data, and calculate the probability of each protein signal intensity value based on the frequency.

[0077] Step S103: Calculate the protein signal intensity threshold for classifying protein signal intensity based on protein signal intensity value and probability.

[0078] Step S104: Identify sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold as foreground sites corresponding to the target protein.

[0079] The protein sequencing data processing method shown in steps S101 to S104 involves acquiring protein sequencing data of the target protein in the target slice, determining the frequency of each protein signal intensity value in the protein sequencing data, and calculating the probability of each protein signal intensity value based on the frequency. The protein sequencing data includes protein signal intensity values ​​of the target protein measured at multiple sequencing sites. Further, a protein signal intensity threshold for classifying protein signal intensity based on the protein signal intensity value and probability is calculated. Further still, sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold are identified as foreground sites corresponding to the target protein. Since the protein signal intensity threshold for classifying protein signal intensity is calculated based on the protein signal intensity value and probability, using the protein signal intensity value and probability as defining factors for the protein signal intensity threshold helps to reasonably and accurately determine the threshold used to distinguish between positive and negative signals, thus facilitating more precise identification of foreground sites.

[0080] In step S101 of some embodiments, protein sequencing data of the target protein in the target slice is obtained. The protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites. It is important to emphasize that in spatial proteomics, to analyze the expression characteristics of cell surface antigen proteins at different spatial locations, it is necessary to use antibody proteins with labeled signals to bind to cell surface antigen protein epitopes in the target slice to mark the location of the antigen protein. It should be noted that the target slice is a tissue slice used for sequencing, and the target protein is an antibody protein with a labeled signal. The labeled signal can be a fluorescent labeling signal or other types of labeling signals. Furthermore, it should be clarified that the protein sequencing data is obtained after the antigen protein and antibody protein have fully bound. The target slice contains multiple sequencing sites; by detecting the intensity of the labeled signal in the target protein at these sequencing sites, the protein signal intensity values ​​of the target protein at multiple sequencing sites can be obtained. In this way, the protein sequencing data of the target protein in the target slice can be obtained.

[0081] In step S102 of some embodiments, the frequency of each protein signal intensity value in the protein sequencing data is determined, and the probability of each protein signal intensity value occurring is calculated based on the frequency. It should be noted that the protein signal intensity values ​​measured at different sequencing sites in the target slice may be the same or different. Therefore, based on the protein signal intensity values ​​at each detection site, the frequency of each protein signal intensity value in the protein sequencing data can be determined. Having determined the frequency of each protein signal intensity value, the probability of each protein signal intensity value occurring can then be calculated based on the frequency. It should be pointed out that determining each protein signal intensity value and the probability of each protein signal intensity value occurring in the protein sequencing data is intended to facilitate the determination of the protein signal intensity threshold in subsequent steps.

[0082] In some specific embodiments, the protein signal intensity values ​​of some sequencing sites are relatively close, belonging to the same intensity range; while the protein signal intensity values ​​of other sequencing sites differ significantly, belonging to different intensity ranges. Based on this, the embodiments of this application can determine the frequency of each intensity range in the protein sequencing data and calculate the probability of each intensity range based on the frequency. In this way, the data processing steps can be simplified while maintaining high accuracy, further improving efficiency.

[0083] In some more specific embodiments, if the number of protein signal intensity values ​​is m, then the multiple protein signal intensity values ​​can be represented as (n1, n2, n3, ..., n m If (n1, n2, n3, ..., n) is a given number, then (n1, n2, n3, ..., n) is a given number. m The frequencies of occurrence corresponding to these m protein signal intensity values ​​can be represented as (f1, f2, f3, ..., f...). m). The frequencies of occurrence corresponding to m protein signal intensity values ​​are determined to be (f1, f2, f3, ..., f...). m After that, the probability of occurrence of the m protein signal intensity values ​​can be determined based on the following analytical formulas (1) and (2): (p1, p2, p3, ..., p...). m ):

[0084]

[0085]

[0086] In step S103 of some embodiments, a protein signal intensity threshold is calculated to classify the strength of protein signals based on protein signal intensity values ​​and probabilities. It should be noted that the protein signal intensity threshold is used to classify the strength of protein signals. Through the protein signal intensity threshold, positive and negative protein signals can be distinguished. A positive protein signal represents a labeling signal of a target protein that has specifically bound to an antigen protein in the target slice, while a negative protein signal represents a labeling signal of a target protein that has not specifically bound. It should be understood that the purpose of calculating the protein signal intensity threshold based on protein signal intensity values ​​and probabilities is to determine foreground sites characterizing specific binding, thereby analyzing relevant data of the foreground sites and clarifying the expression characteristics of cell surface antigen proteins at different spatial locations.

[0087] In step S104 of some embodiments, sequencing sites with protein signal intensity values ​​greater than a protein signal intensity threshold are identified as foreground sites corresponding to the target protein. It should be noted that because the protein signal intensity threshold is used to classify protein signal intensity, sequencing sites with protein signal intensity values ​​greater than the threshold can be identified as foreground sites corresponding to the target protein. This allows for the analysis of relevant data from these foreground sites, clarifying the expression characteristics of cell surface antigen proteins at different spatial locations.

[0088] The following is a detailed description of step S103.

[0089] Reference Figure 2 In some embodiments of this application, step S103, which calculates the protein signal intensity threshold for classifying the strength of the protein signal based on the protein signal intensity value and probability, may include, but is not limited to, the following steps S201 to S203.

[0090] Step S201: Using any target protein signal intensity value as a threshold, the protein signal intensity values ​​are divided into first-class protein signal intensity values ​​and second-class protein signal intensity values;

[0091] Step S202: Calculate the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein based on probability;

[0092] Step S203: Determine the protein signal intensity threshold for classifying the strength of protein signals among multiple target protein signal intensity values ​​based on the inter-class variance value.

[0093] In step S201 of some embodiments, the protein signal intensity values ​​are divided into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​using any target protein signal intensity value as a threshold. It should be noted that among multiple target protein signal intensity values, using any target protein signal intensity value as a threshold can classify these target protein signal intensity values ​​into a category with stronger target protein signals and a category with weaker target protein signals; these two categories of target protein signal intensity values ​​are also known as the first category of protein signal intensity values ​​and the second category of protein signal intensity values.

[0094] In some more specific embodiments, the stronger target protein signal among the first type of protein signal intensity value and the second type of protein signal intensity value can be identified as the positive protein signal intensity value, thereby representing the target protein signal intensity value for specific binding between the antigen protein and the antibody protein; the weaker target protein signal among the first type of protein signal intensity value and the second type of protein signal intensity value can be identified as the negative protein signal intensity value, thereby representing the target protein signal intensity value for non-specific binding between the antigen protein and the antibody protein.

[0095] Reference Figure 3 According to some specific embodiments of this application, step S201, which divides the protein signal intensity value into a first type of protein signal intensity value and a second type of protein signal intensity value based on any target protein signal intensity value as a threshold, may include, but is not limited to, steps S301 to S302 below.

[0096] Step S301: Determine any one of the multiple protein signal intensity values ​​as the target protein signal intensity value;

[0097] Step S302: Protein signal intensity values ​​with values ​​less than the target protein signal intensity value are classified as first-class protein signal intensity values, and protein signal intensity values ​​with values ​​not less than the target protein signal intensity value are classified as second-class protein signal intensity values.

[0098] In some embodiments, step S301 involves determining any one of the multiple protein signal intensity values ​​as the target protein signal intensity value. It should be noted that if the number of protein signal intensity values ​​is m, then the multiple protein signal intensity values ​​can be represented as (n1, n2, n3, ..., n...). m From (n1,n2,n3,…….,n) mFrom these m protein signal intensity values, any one protein signal intensity value is selected as the target protein signal intensity value t, that is, t is selected from (n1, n2, n3, ..., nn...). m It is determined in ) .

[0099] In some embodiments, step S302 involves classifying protein signal intensity values ​​that are less than the target protein signal intensity value into a first category of protein signal intensity values, and classifying protein signal intensity values ​​that are not less than the target protein signal intensity value into a second category of protein signal intensity values. It should be noted that in the range (n1, n2, n3, ..., n... m After determining any protein signal intensity value as the target protein signal intensity value t from these m protein signal intensity values, further, in (n1, n2, n3, ..., n m Among these m protein signal intensity values, those with values ​​less than the target protein signal intensity value t are classified as the first class of protein signal intensity values ​​C1, and those in (n1, n2, n3, ..., n...) are further classified as... m Among these m protein signal intensity values, the protein signal intensity values ​​that are not less than the target protein signal intensity value t are classified as the second type of protein signal intensity values ​​C2.

[0100] Through the embodiments of this application shown in steps S301 to S302, any one of the multiple protein signal intensity values ​​is determined to be the target protein signal intensity value. Then, protein signal intensity values ​​with values ​​less than the target protein signal intensity value are classified into a first category of protein signal intensity values, and protein signal intensity values ​​with values ​​not less than the target protein signal intensity value are classified into a second category of protein signal intensity values. In this way, multiple protein signal intensity values ​​can be efficiently classified into a weaker first category and a stronger second category based on the target protein signal intensity value t.

[0101] In some embodiments, step S202 involves calculating the inter-class variance between the signal intensity values ​​of the first and second classes of proteins based on probability. It should be noted that data within the same class are more similar, resulting in a smaller intra-class variance; data from different classes are less similar, resulting in a larger inter-class variance. Therefore, it is necessary to calculate the inter-class variance between the signal intensity values ​​of the first and second classes of proteins based on probability. Because data from different classes are less similar, the inter-class variance is larger. Thus, the inter-class variance can be used in subsequent steps to determine the protein signal intensity threshold for classifying protein signal intensity from multiple target protein signal intensity values. It should be understood that by summing the analytical expressions for the inter-class variance and the intra-class variance, it can be proven that the sum of the inter-class variance and the intra-class variance is a constant value. Therefore, it is determined that the larger the inter-class variance, the smaller the intra-class variance. Based on the inter-class variance, the protein signal intensity threshold for classifying the strength of protein signals can be determined among multiple target protein signal intensity values. The determination of the protein signal intensity threshold does not require the intervention of the intra-class variance.

[0102] Reference Figure 4 According to some specific embodiments of this application, step S202, which calculates the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein based on probability, may include, but is not limited to, steps S401 to S403 below.

[0103] Step S401: Calculate the sum of multiple probabilities corresponding to the signal intensity values ​​of the first type of protein to obtain a first probability sum, and calculate the sum of multiple probabilities corresponding to the signal intensity values ​​of the second type of protein to obtain a second probability sum;

[0104] Step S402: Calculate the mean value of the first protein signal intensity based on the first probability sum, the signal intensity value of the first type of protein and the corresponding probability; and calculate the mean value of the second protein signal intensity based on the second probability sum, the signal intensity value of the second type of protein and the corresponding probability.

[0105] Step S403: Calculate the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity.

[0106] In some embodiments, step S401 involves calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the first type of protein to obtain a first probability sum, and calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the second type of protein to obtain a second probability sum. It should be noted that (n1, n2, n3, ..., n m The probability of occurrence of these m protein signal intensity values ​​can be represented as (p1, p2, p3, ..., p...). mTherefore, the sum of multiple probabilities corresponding to the signal intensity value C1 of the first type of protein to obtain the first probability P1(t) can be expressed as analytical equation (3):

[0107]

[0108] Furthermore, since all protein signal intensity values ​​are divided into two categories, the first type protein signal intensity value C1 and the second type protein signal intensity value C2, the sum of multiple probabilities corresponding to the second type protein signal intensity value C2 is used to obtain the second probability P2(t), which can be expressed as analytical formula (4):

[0109]

[0110] In some embodiments, step S402 involves calculating the average value of a first protein signal intensity based on a first probability sum, the signal intensity value of a first type of protein, and the corresponding probability; and calculating the average value of a second protein signal intensity based on a second probability sum, the signal intensity value of a second type of protein, and the corresponding probability. It should be noted that the average value of the first protein signal intensity refers to the average value of the protein signal intensity values ​​assigned to the first type of protein signal intensity value C1, expressed as... The mean signal intensity of the second protein refers to the average value of the protein signal intensity assigned to the second protein signal intensity value C2, denoted as:

[0111] It should be noted that after calculating the sum of multiple probabilities corresponding to the signal intensity value C1 of the first type of protein to obtain the first probability P1(t), the mean value of the first protein signal intensity can be calculated using analytical formula (5) based on the first probability P1(t), the signal intensity value C1 of the first type of protein, and the corresponding probabilities (p1, p2, p3, ..., t).

[0112]

[0113] Similarly, after calculating the sum of multiple probabilities corresponding to the signal intensity value C2 of the second type of protein to obtain the second probability P2(t), the analytical formula (6) can be used to determine the second probability P2(t), the signal intensity value C2 of the second type of protein, and the corresponding probabilities (t, ..., p). m Calculate the mean signal intensity of the second protein.

[0114]

[0115] In some embodiments, step S403 involves calculating the inter-class variance between the signal intensity values ​​of the first class of proteins and the signal intensity values ​​of the second class of proteins based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity. It should be noted that after obtaining the mean of the first protein signal intensity... Mean signal intensity of the second protein Subsequently, further analysis can be performed based on the first probability and P1(t), the second probability and P2(t), and the mean intensity of the first protein signal. Mean signal intensity of the second protein Calculate the inter-class variance between the signal intensity values ​​C1 of the first type of protein and C2 of the second type of protein.

[0116] According to the embodiments of this application shown in steps S401 to S403, a first probability sum is first calculated by summing multiple probabilities corresponding to the signal intensity values ​​of the first type of protein, and a second probability sum is calculated by summing multiple probabilities corresponding to the signal intensity values ​​of the second type of protein. Then, the mean value of the first protein signal intensity is calculated based on the first probability sum, the signal intensity values ​​of the first type of protein, and their corresponding probabilities; and the mean value of the second protein signal intensity is calculated based on the second probability sum, the signal intensity values ​​of the second type of protein, and their corresponding probabilities. Further, the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein is calculated based on the first probability sum, the second probability sum, the mean value of the first protein signal intensity, and the mean value of the second protein signal intensity. In this way, the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein can be calculated efficiently based on each protein signal intensity value and its frequency, which helps to accurately determine the threshold used to distinguish positive and negative signals, thus facilitating more precise determination of the foreground site.

[0117] Reference Figure 5 According to some more specific embodiments of this application, step S403 calculates the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity, which may include, but is not limited to, steps S501 to S502 below.

[0118] Step S501: Calculate the square of the difference between the mean signal intensity of the first protein and the mean signal intensity of the second protein;

[0119] Step S502: Calculate the product of the first probability sum, the second probability sum, and the squared difference to obtain the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein.

[0120] In some embodiments, step S501 involves calculating the squared difference between the mean signal intensity of the first protein and the mean signal intensity of the second protein. It should be noted that this is to calculate the inter-class variance σ between the signal intensity values ​​of the first and second types of proteins. 2 The mean signal intensity of the first protein needs to be calculated using analytical formula (7). Mean signal intensity of the second protein Squared difference between them:

[0121]

[0122] In some embodiments, step S502 involves calculating the product of the first probability sum, the second probability sum, and the squared difference to obtain the inter-class variance between the signal intensity values ​​of the first and second types of proteins. It should be noted that after calculating the mean signal intensity of the first protein... Mean signal intensity of the second protein Squared difference between Then, the first probability and P1(t), the second probability and P2(t), and the squared difference can be calculated further using analytical formula (8). The product of these values ​​yields the inter-class variance between the signal intensity values ​​of the first type of protein, C1, and the signal intensity values ​​of the second type of protein, C2.

[0123]

[0124] According to the embodiments of this application shown in steps S501 to S502, the squared difference between the mean signal intensity of the first protein and the mean signal intensity of the second protein is first calculated. Then, the product of the first probability sum, the second probability sum, and the squared difference is calculated to obtain the inter-class variance between the signal intensity values ​​of the first and second types of proteins. In this way, the inter-class variance between the signal intensity values ​​of the first and second types of proteins can be calculated efficiently, which helps to accurately determine the threshold used to distinguish positive and negative signals, so as to more accurately determine the foreground site.

[0125] In step S203 of some embodiments, a protein signal intensity threshold for classifying protein signal intensity into strong and weak classes is determined among multiple target protein signal intensity values ​​based on the inter-class variance value. It should be noted that after calculating the inter-class variance value between the first and second class protein signal intensity values ​​based on the frequency of each protein signal intensity value in the protein sequencing data, since the data from different classes are less similar and have larger inter-class variance values, the protein signal intensity threshold for classifying protein signal intensity into strong and weak classes can be determined based on the inter-class variance value.

[0126] Reference Figure 6 According to some specific embodiments of this application, step S203, which determines the protein signal intensity threshold for classifying the strength of protein signals among multiple target protein signal intensity values ​​based on the inter-class variance value, may include, but is not limited to, steps S601 to S602 below.

[0127] Step S601: Determine the inter-class variance value corresponding to the signal intensity value of each target protein;

[0128] Step S602: Determine the target protein signal intensity value with the largest inter-class variance as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values.

[0129] In some embodiments, step S601 involves determining the inter-class variance value corresponding to each target protein signal intensity value. It should be noted that the target protein signal intensity value t ranges from (n1, n2, n3, ..., n...). m For each value in the corresponding range, determine the inter-class variance (σ1) corresponding to the signal intensity value t of each target protein. 2 ,σ2 2 ,σ3 2 ,…….,σ m 2 ).

[0130] In some embodiments, step S602 involves determining the target protein signal intensity value with the largest inter-class variance as a protein signal intensity threshold for classifying protein signal intensities into strong and weak values. It should be noted that the target protein signal intensity value t ranges from (n1, n2, n3, ..., n...). m The corresponding m inter-class variance values ​​(σ1) in ) 2 ,σ2 2 ,σ3 2 ,…….,σ m 2 In the process, the largest inter-class variance value σ is determined. max 2 The largest inter-class variance value σ max 2 The corresponding target protein signal intensity value is determined as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values.

[0131] In the embodiments of this application shown in steps S601 to S602, the inter-class variance value corresponding to each target protein signal intensity value is first determined, and then the target protein signal intensity value with the largest inter-class variance value is determined as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values. In this way, the target protein signal intensity value with the largest inter-class variance value can be determined from multiple target protein signal intensity values ​​and used as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values. This helps to accurately determine the threshold used to distinguish between positive and negative signals, thereby facilitating more precise identification of foreground sites.

[0132] In the embodiments of this application shown in steps S201 to S203, protein signal intensity values ​​are first divided into a first category and a second category based on any target protein signal intensity value as a threshold. Then, the inter-class variance between the first and second category protein signal intensity values ​​is calculated based on probability. Further, a protein signal intensity threshold for classifying protein signal intensity based on the inter-class variance value is determined among multiple target protein signal intensity values. Since data from different categories are less similar, the inter-class variance value is larger. After calculating the inter-class variance value between the first and second category protein signal intensity values, this value can be used to determine the protein signal intensity threshold. In this way, a threshold for distinguishing positive and negative signals can be reasonably and accurately determined, facilitating more precise identification of foreground sites.

[0133] Reference Figure 7 This provides a more specific embodiment of one of the above steps S201 to S203. Figure 6 The graph reflects the correspondence between protein signal intensity thresholds and inter-class variance values, including three schematic lines: the first is the inter-class variance curve for the mouse_CD4 sample; the second is the schematic line for the maximum inter-class variance value; and the third is the schematic line for the optimal threshold. It should be noted that a large inter-class variance value indicates greater dissimilarity between data from different classes. Therefore, the protein signal intensity corresponding to the maximum inter-class variance value of 27.02111 is selected as the optimal threshold, i.e., a protein signal intensity threshold of 12.0. Based on this, the protein signal intensity threshold is used to classify protein signal intensity into strong and weak categories, identify foreground sites representing specific binding, and thus analyze the relevant data of foreground sites to clarify the expression characteristics of the mouse_CD4 antigen protein on the cell surface at different spatial locations.

[0134] Reference Figure 8In some embodiments of this application, the protein signal intensity threshold for classifying the strength of protein signals based on the protein signal intensity value and probability calculation in step S103 is not limited to the embodiments shown in steps S201 to S203 above, and may also include steps S801 to S803 below.

[0135] Step S801: Using any target protein signal intensity value as a threshold, the protein signal intensity values ​​are divided into first-class protein signal intensity values ​​and second-class protein signal intensity values;

[0136] Step S802: Calculate the first information entropy corresponding to the signal intensity value of the first type of protein and the second information entropy corresponding to the signal intensity value of the second type of protein based on probability.

[0137] Step S803: Determine the protein signal intensity threshold for classifying the strength of protein signals from multiple target protein signal intensity values ​​based on the sum of the first information entropy and the second information entropy.

[0138] It should be noted that step S103, which calculates the protein signal intensity threshold based on the protein signal intensity value and probability to classify the protein signal intensity into strong and weak categories, is not limited to the embodiments shown in steps S201 to S203 above. Alternatively, it can first divide the protein signal intensity value into a first category and a second category using any target protein signal intensity value as a threshold, calculate the average of the first category and the average of the second category, and then, when the sum of the two averages remains stable, the corresponding target protein signal intensity value is the protein signal intensity threshold. Furthermore, it may also include the embodiments shown in steps S801 to S803.

[0139] In step S801 of some embodiments, the protein signal intensity value is divided into a first type of protein signal intensity value and a second type of protein signal intensity value using any target protein signal intensity value as a threshold. It should be noted that any protein signal intensity value is determined from multiple protein signal intensity values ​​as the target protein signal intensity value t, and the protein signal intensity value is divided into a first type of protein signal intensity value and a second type of protein signal intensity value using any target protein signal intensity value t as a threshold.

[0140] In some embodiments, step S802 involves calculating the first information entropy corresponding to the signal intensity value of the first type of protein and the second information entropy corresponding to the signal intensity value of the second type of protein based on probability. It should be noted that if the signal intensity value of the i-th protein is represented as p... i The probability of being assigned to the signal intensity value of the first type of protein is expressed as p. A ,in Then, the first information entropy H(A) corresponding to the signal intensity value of the first type of protein can be calculated according to analytical formula (9):

[0141]

[0142] Similarly, the probability of being assigned to the signal intensity value of the second type of protein is expressed as p. B ,in The second information entropy H(B) corresponding to the signal intensity value of the second type of protein can then be calculated according to analytical formula (10):

[0143]

[0144] In step S803 of some embodiments, a protein signal intensity threshold for classifying protein signal intensity into strong and weak values ​​is determined from multiple target protein signal intensity values ​​based on the sum of the first information entropy and the second information entropy. It should be noted that after calculating the first information entropy H(A) and the second information entropy H(B), it is necessary to further determine the protein signal intensity threshold for classifying protein signal intensity into strong and weak values ​​from multiple target protein signal intensity values ​​based on the sum of the first information entropy H(A) and the second information entropy H(B).

[0145] In some more specific embodiments, the sum of the first information entropy H(A) and the second information entropy H(B) It can be represented as:

[0146]

[0147]

[0148] Furthermore, the sum of the first information entropy H(A) and the second information entropy H(B) It can be expressed by analytical expression (11):

[0149]

[0150] Based on this, by defining each protein signal intensity value as the target protein signal intensity value, and by iterating through all protein signal intensity values ​​sequentially, it is possible to find the sum of the first information entropy H(A) and the second information entropy H(B). The target protein signal intensity value that yields the maximum value can be determined as the protein signal intensity threshold.

[0151] In the embodiments of this application shown in steps S801 to S803, protein signal intensity values ​​are first divided into a first type of protein signal intensity value and a second type of protein signal intensity value using any target protein signal intensity value as a threshold. Further, a first information entropy corresponding to the first type of protein signal intensity value and a second information entropy corresponding to the second type of protein signal intensity value are calculated based on probability. Further still, a protein signal intensity threshold for classifying protein signal intensity based on the sum of the first and second information entropies is determined among multiple target protein signal intensity values. In this way, a threshold for distinguishing positive and negative signals can be reasonably and accurately determined, facilitating more precise identification of foreground sites.

[0152] The method for detecting the distribution area of ​​antigen proteins in this application is described below.

[0153] Reference Figure 9 This application provides a method for detecting antigen protein distribution regions, which may include, but is not limited to, the following steps S901 to S905.

[0154] Step S901: Obtain the biological slice to be tested;

[0155] Step S902: The biological slice is soaked in a pre-set antibody protein solution and then washed to obtain the target slice.

[0156] Step S903: Perform protein sequencing on the target slice to obtain protein sequencing data;

[0157] Step S904: Process the protein sequencing data based on the protein sequencing data processing method to obtain the pre-defined foreground sites of the antibody protein;

[0158] Step S905: Determine the distribution area of ​​the antigen protein in the biological slice based on the preset foreground sites of the antibody protein.

[0159] It is important to emphasize that in spatial proteomics, to analyze the expression characteristics of cell surface antigen proteins at different spatial locations, it is necessary to use antibody proteins with labeled signals to bind to cell surface antigen protein epitopes in the target section to mark the location of the antigen proteins. It should be noted that the target section is a tissue section used for sequencing.

[0160] In some embodiments, steps S901 to S905 involve acquiring a biological slide to be tested, then wetting the biological slide with a solution of a pre-defined antibody protein, and washing the wetting slide to obtain a target slide. Protein sequencing is then performed on the target slide to obtain protein sequencing data. Further, the protein sequencing data is processed using a method based on the protein sequencing data to obtain the foreground sites of the pre-defined antibody protein. Further still, the distribution area of ​​the antigen protein in the biological slide is determined based on the foreground sites of the pre-defined antibody protein. It should be noted that, to obtain the protein sequencing data of the target slide, the biological slide to be tested can be acquired first, then the biological slide can be wetting with a solution of a pre-defined antibody protein, and the wetting slide can be washed to obtain the target slide. It should be emphasized that the target protein is an antibody protein with a labeled signal, wherein the labeled signal can be a fluorescent labeling signal or other types of labeling signals. It should be clarified that protein sequencing data is obtained after the antigen protein and antibody protein have fully bound. The target slice contains multiple sequencing sites. By detecting the intensity of the labeled signal in the target protein at these sequencing sites, the protein signal intensity values ​​of the target protein at multiple sequencing sites can be obtained. In this way, the protein sequencing data of the target protein in the target slice can be obtained.

[0161] After processing the protein sequencing data using the protein sequencing data processing method based on the embodiments of this application to obtain the pre-set foreground sites for specific binding of antibody proteins, the distribution area of ​​antigen proteins in biological slices can be determined based on the pre-set foreground sites of antibody proteins. This helps to analyze the data of the foreground sites for specific binding in a targeted manner.

[0162] Reference Figure 10 , Figure 10 This diagram illustrates the spatial distribution of target slices corresponding to mouse_CD4 samples after removing negative protein signals based on a protein signal intensity threshold in some embodiments of this application. It is important to emphasize that the protein signal intensity threshold is used to classify the strength of protein signals. Through this threshold, positive and negative protein signals can be distinguished. A positive protein signal represents a labeling signal of a target protein that has specifically bound to the antigen protein of the cells in the target slice, while a negative protein signal represents a labeling signal of a target protein that has not specifically bound. The foreground site is... Figure 10 Bright spot regions in the image. Identifying foreground sites that characterize specific binding helps in analyzing relevant data from these foreground sites and clarifying the expression characteristics of the mouse_CD4-corresponding antigen protein on the cell surface at different spatial locations.

[0163] The following describes the apparatus for processing protein sequencing data in this application.

[0164] Reference Figure 11 According to some embodiments, this application provides a protein sequencing data processing apparatus 1100, comprising:

[0165] The first acquisition unit 1101 is used to acquire protein sequencing data of the target protein in the target slice. The protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites.

[0166] The first determining unit 1102 is used to determine the frequency of each protein signal intensity value in the protein sequencing data, and calculate the probability of each protein signal intensity value appearing based on the frequency.

[0167] Threshold division unit 1103 is used to calculate the protein signal intensity threshold for dividing the protein signal intensity into strong and weak based on the protein signal intensity value and probability.

[0168] The second determining unit 1104 is used to determine sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold as foreground sites corresponding to the target protein.

[0169] It is evident that the content of the above-described protein sequencing data processing method embodiments is applicable to the embodiments of this protein sequencing data processing device. The specific functions implemented by this protein sequencing data processing device embodiment are the same as those of the above-described protein sequencing data processing method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described protein sequencing data processing method embodiments.

[0170] The following describes a device for detecting the distribution area of ​​antigen proteins according to this application.

[0171] Reference Figure 12 According to some embodiments, this application provides an antigen protein distribution region detection device 1200, comprising:

[0172] The second acquisition unit 1201 is used to acquire the biological slice to be detected;

[0173] The slide cleaning unit 1202 is used to wet biological slides with a preset antibody protein solution and to clean the wetted biological slides to obtain the target slides.

[0174] Protein sequencing unit 1203 is used to perform protein sequencing on the target slice to obtain protein sequencing data;

[0175] The sequencing data processing unit is used to process protein sequencing data based on protein sequencing data processing methods to obtain the pre-defined foreground sites of antibody proteins.

[0176] The third determining unit 1204 is used to determine the distribution area of ​​the antigen protein in the biological slice based on the preset foreground site of the antibody protein.

[0177] It is evident that the content of the above-described antigen protein distribution region detection method embodiments is applicable to the embodiments of this antigen protein distribution region detection device. The specific functions implemented by this antigen protein distribution region detection device embodiment are the same as those of the above-described antigen protein distribution region detection method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described antigen protein distribution region detection method embodiments.

[0178] Reference Figure 13 , Figure 13 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0179] The processor 1301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0180] The memory 1302 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1302 and is called and executed by the processor 1301 to execute the protein sequencing data processing method or antigen protein distribution region detection method of the embodiments of this application.

[0181] The input / output interface 1303 is used to implement information input and output;

[0182] The communication interface 1304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0183] Bus 1305 transmits information between various components of the device (e.g., processor 1301, memory 1302, input / output interface 1303, and communication interface 1304);

[0184] The processor 1301, memory 1302, input / output interface 1303 and communication interface 1304 are connected to each other within the device via bus 1305.

[0185] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the protein sequencing data processing method or the antigen protein distribution region detection method described above.

[0186] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0187] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0188] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0189] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0191] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0193] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0194] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A method for processing protein sequencing data, characterized in that, The method includes: Obtain protein sequencing data of the target protein in the target slice, wherein the protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites; Determine the frequency of each protein signal intensity value in the protein sequencing data, and calculate the probability of each protein signal intensity value occurring based on the frequency; A protein signal intensity threshold is calculated to classify the strength of protein signals based on the protein signal intensity values ​​and the probability of each protein signal intensity value occurring. Sequencing sites with protein signal intensity values ​​greater than the protein signal intensity threshold are identified as foreground sites corresponding to the target protein; The step of calculating the protein signal intensity threshold for classifying protein signal intensity based on the protein signal intensity value and the probability of occurrence of each protein signal intensity value includes: The protein signal intensity values ​​are divided into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​using any target protein signal intensity value as a threshold. Based on the probability of occurrence of each protein signal intensity value, the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein is calculated. The protein signal intensity threshold is determined from multiple target protein signal intensity values ​​based on the inter-class variance value; or, Based on the probability of occurrence of each protein signal intensity value, calculate the first information entropy corresponding to the first type of protein signal intensity value and the second information entropy corresponding to the second type of protein signal intensity value. The protein signal intensity threshold is determined from multiple target protein signal intensity values ​​based on the sum of the first information entropy and the second information entropy.

2. The method according to claim 1, characterized in that, The calculation of the inter-class variance between the signal intensity values ​​of the first class of proteins and the signal intensity values ​​of the second class of proteins based on the probability of occurrence of each of the protein signal intensity values ​​includes: A first probability sum is obtained by calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the first type of protein, and a second probability sum is obtained by calculating the sum of multiple probabilities corresponding to the signal intensity values ​​of the second type of protein; The average signal intensity of the first protein is calculated based on the first probability sum, the signal intensity value of the first type of protein, and the corresponding probability; and the average signal intensity of the second protein is calculated based on the second probability sum, the signal intensity value of the second type of protein, and the corresponding probability. The inter-class variance between the signal intensity values ​​of the first type of protein and the second type of protein is calculated based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity.

3. The method according to claim 1, characterized in that, The step of dividing the protein signal intensity values ​​into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​based on any target protein signal intensity value as a threshold includes: Determine any one of the multiple protein signal intensity values ​​as the target protein signal intensity value; Protein signal intensity values ​​with values ​​less than the target protein signal intensity value are classified as first-class protein signal intensity values, and protein signal intensity values ​​with values ​​not less than the target protein signal intensity value are classified as second-class protein signal intensity values.

4. The method according to claim 1, characterized in that, The step of determining the protein signal intensity threshold for classifying protein signal intensity based on the inter-class variance value among multiple target protein signal intensity values ​​includes: Determine the inter-class variance value corresponding to the signal intensity value of each target protein; The target protein signal intensity value with the largest inter-class variance is determined as the protein signal intensity threshold for classifying protein signal intensity into strong and weak values.

5. The method according to claim 2, characterized in that, The calculation of the inter-class variance between the signal intensity values ​​of the first class of proteins and the signal intensity values ​​of the second class of proteins based on the first probability sum, the second probability sum, the mean of the first protein signal intensity, and the mean of the second protein signal intensity includes: Calculate the squared difference between the mean signal intensity of the first protein and the mean signal intensity of the second protein; The inter-class variance between the signal intensity values ​​of the first type of protein and the second type of protein is obtained by multiplying the sum of the first probability, the sum of the second probability, and the square of the difference.

6. A method for detecting the distribution region of an antigen protein, characterized in that, The method includes: Obtain biological slices to be tested; The biological slices were immersed in a solution of a pre-set antibody protein, and then washed to obtain the target slices. Protein sequencing was performed on the target slice to obtain protein sequencing data; The protein sequencing data is processed according to the protein sequencing data processing method according to any one of claims 1 to 5 to obtain the foreground site of the preset antibody protein; The distribution area of ​​the antigen protein in the biological slice is determined based on the pre-defined foreground site of the antibody protein.

7. A protein sequencing data processing device, characterized in that, include: The first acquisition unit is used to acquire protein sequencing data of the target protein in the target slice, wherein the protein sequencing data includes the protein signal intensity values ​​of the target protein measured at multiple sequencing sites; The first determining unit is used to determine the frequency of each protein signal intensity value in the protein sequencing data, and to calculate the probability of each protein signal intensity value appearing based on the frequency. A threshold division unit is used to calculate a protein signal intensity threshold for classifying protein signal intensity based on the protein signal intensity value and the probability of occurrence of each protein signal intensity value; wherein, the calculation of the protein signal intensity threshold for classifying protein signal intensity based on the protein signal intensity value and the probability of occurrence of each protein signal intensity value includes: The protein signal intensity values ​​are divided into a first category of protein signal intensity values ​​and a second category of protein signal intensity values ​​using any target protein signal intensity value as a threshold. Based on the probability of occurrence of each protein signal intensity value, the inter-class variance between the signal intensity values ​​of the first type of protein and the signal intensity values ​​of the second type of protein is calculated. The protein signal intensity threshold is determined from multiple target protein signal intensity values ​​based on the inter-class variance value; or, Based on the probability of occurrence of each protein signal intensity value, calculate the first information entropy corresponding to the first type of protein signal intensity value and the second information entropy corresponding to the second type of protein signal intensity value. The protein signal intensity threshold is determined from multiple target protein signal intensity values ​​based on the sum of the first information entropy and the second information entropy. The second determining unit is used to determine the sequencing sites whose protein signal intensity values ​​are greater than the protein signal intensity threshold as the foreground sites corresponding to the target protein.

8. A device for detecting the distribution region of an antigen protein, characterized in that, include: The second acquisition unit is used to acquire the biological slice to be tested; The slide cleaning unit is used to wet the biological slide with a preset antibody protein solution and to clean the wetted biological slide to obtain the target slide. A protein sequencing unit is used to perform protein sequencing on the target slice to obtain protein sequencing data. A sequencing data processing unit is used to process the protein sequencing data according to the protein sequencing data processing method according to any one of claims 1 to 5 to obtain the foreground site of the preset antibody protein; The third determining unit is used to determine the distribution area of ​​the antigen protein in the biological slice based on the foreground site of the preset antibody protein.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the protein sequencing data processing method according to any one of claims 1 to 5 or the antigen protein distribution region detection method according to claim 6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the protein sequencing data processing method according to any one of claims 1 to 5 or the antigen protein distribution region detection method according to claim 6.

11. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the protein sequencing data processing method of any one of claims 1 to 5 or the antigen protein distribution region detection method of claim 6.

Citation Information

Patent Citations

  • Protein sequence hydrolysis site prediction method and device, equipment and storage medium

    CN115798595A