Fluid sample classification
By using a two-dimensional analysis method based on the similarity of relative quantity and component profiles in the biopharmaceutical manufacturing process, similarity scores are generated, and the problem of difficult analysis of complex sample separation profiles is solved, and rapid and automated sample comparison and classification is achieved, reducing costs and time.
Patent Information
- Application Number
- CN202080070768.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-09
- Filing Date
- 2020-10-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-07
AI Technical Summary
In the biopharmaceutical manufacturing process, complex sample separation profiles are difficult to accurately analyze, resulting in increased cost and processing time, and are prone to introduce bias from individual operators, requiring a fast and automated sample comparison method.
Different samples in the biopharmaceutical process were used to compare different samples by introducing similarity scores based on the similarity to the relative amount and composition profile similarity to the reference sample. This method generates similarity scores for the classification and analysis of samples by calculating the similarity between the components and separation profiles of the samples and the reference samples.
A fast, non-operator-dependent sample classification scheme is implemented, allowing automated analysis, reducing the time and cost of manufacturing biopharmaceuticals, and able to process incomplete sample separation data.
Smart Images

Figure CN114450588B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to methods and computer programs for the classification of samples that have had their components separated, for example, by chromatography or electrophoresis, and more particularly, to methods for classifying such samples based on similarity of relative amounts and compositional profiles to reference samples. Background Art
[0002] In the manufacture of biopharmaceuticals (such as vaccines, antibodies, recombinant proteins, gene therapy vectors, etc.), several chromatographic separation steps are typically required to remove various contaminants and impurities from the product. During each step of the manufacturing process, it is necessary to check both the amount and purity compared to a reference sample.
[0003] However, separation profiles typically show multiple molecular peak bands that may overlap. Analysis of such complex separation profiles significantly increases cost and processing time. In addition, the separation profiles of complex samples may be difficult to accurately analyze, which may introduce individual operator bias. Therefore, there is significant interest in rapid automated analysis to eliminate personal bias and reduce the time to manufacture biopharmaceuticals. Thus, there is a need for methods and computer programs for sample comparison that can be both rapid and automated if desired. Summary of the Invention
[0004] One aspect of the present invention is to provide a method, which can be implemented by a computer program, for comparing different samples in a biopharmaceutical process. This is achieved by introducing a similarity score based on two-dimensional analysis of similarity of relative amounts and compositional profiles to a reference sample. The relative amount is a measure of the amount values of different chemical components of a sample compared to the amount values of the components of the reference sample. The compositional profile similarity calculation is a measure of the similarity of comparing the spatial or temporal profile of the separated components produced by a separation process with the profiles of one or more reference samples that have undergone the same separation process. By providing a similarity score for each sample, the resulting two-dimensional data set forms the basis for sample classification, and the similarity score allows an estimate of the similarity of the sample of interest to the reference sample. After classification criteria have been set, computer analysis software and suitable hardware can be used to automate and implement the analysis method.
[0005] One advantage is that such an analytical method allows for a fast, operator-independent classification scheme where limits for grouping samples can be easily set for automated analysis. The method allows for decisions such as whether a manufacturing process is running satisfactorily or whether separation parameters need to be changed. Additionally, the separation of sample components is often incomplete; in other words, the measured component bands overlap, resulting in data that is difficult to analyze. By comparing the measured profiles to a reference in the manner just described above rather than just looking at certain peaks that are measured, the proposed method allows for such incomplete separations.
[0006] Other suitable embodiments of the invention are described in the dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Images showing Coomassie-stained gels that have been analyzed. Different samples (1 - 10) were separated on a polyacrylamide gel and subsequently stained with Coomassie stain. The gels were analyzed using the analytical software Image Quant TL (available from Cytiva LifeSciences).
[0008] Figure 2 Showing Figure 1 the electrophoretic lane profile of sample 5 in
[0009] Figure 3 Showing Figure 1 and Figure 2 a two-dimensional similarity score plot of the samples in . The relative amounts of the samples (y-axis) and the lane profile similarity scores (x-axis) are shown in one plot. In this example, sample 8 is the reference sample.
[0010] Figure 4 Showing one way to group Figure 3 the different samples in into three groups (A - C).
[0011] Figure 5 Showing chromatograms of two protein samples.
[0012] Figure 6 Showing a 2D similarity scatter plot of 185 chromatograms for different cycles in a protein purification process generated using an embodiment of the invention.
[0013] Figure 7 Showing a graphical user interface (GUI) for presenting the results generated using an embodiment of the invention.
[0014] Figure 8 Showing the method according to various embodiments of the invention.
[0015] DEFINITIONS
[0016] As used herein, the terms "comprises," "comprising," "contains," "having," etc. can have the meanings ascribed to them and can mean "includes," "including," etc.; "consisting essentially of" or "consists essentially" likewise have the meanings ascribed in patent law, and the terms are open-ended, allowing for more than those listed, provided that the basic or novel characteristics of those listed are not changed by the presence of more than those listed, but excluding prior art embodiments. Detailed Description
[0017] In one aspect, the present invention discloses a computer-aided method for automated analysis of a sample that has undergone separation, the term including forming at least some components of the sample into higher concentration aliquots, fractions, or streams. The term separation includes partial separation.
[0018] In certain embodiments, the sample can be an intermediate or final product in a biopharmaceutical manufacturing process. Additionally, a reference sample can undergo chemical separation and subsequent analysis during the creation process, or be saved for chemical separation analysis at a later stage. Those skilled in the art of biopharmaceutical manufacturing know well how to store reference samples for future analysis. The saved reference sample data can also be used for comparison.
[0019] In some embodiments, separation is performed by electrophoresis, and the separated molecules are detected using a color stain or a fluorescent dye. Subsequently, the lane profile of the electrophoretic separation is compared in terms of both how many sample components are present in the lane and the degree of similarity of the lane space profile compared to a reference sample.
[0020] In some embodiments, separation is performed by chromatography, and the separated molecules are detected using absorbance measurements. Subsequently, the separation profile is compared in terms of both how much of the sample elutes from the chromatographic column and the degree of similarity of the sample profile (also known as a chromatogram). In this way, the reproducibility of different batches in a bioprocess manufacturing process can be quickly evaluated, and a decision to continue the process can be easily made. Examples
[0021] Example 1
[0022] Different protein samples were analyzed using SDS-PAGE electrophoresis. The resulting Coomassie-stained gels are shown in Figure 1Among them. Some samples contain multiple proteins, and the lane profiles obtained after electrophoretic separation are complex, that is, they exhibit many overlapping peaks. Such lane profiles are difficult to compare manually by eye. The opacity of each or multiple target lanes can be measured relative to the spatial position along the lane, also known as optical density. In Figure 2 the figure shown, such a measurement is given only for Figure 1 lane 5, but the same measurement can be performed for other lanes shown in Figure 1 . Using the data generated from the Figure 2 figure (which does not need to be graphically represented and can exist purely as data), using the Pearson correlation coefficient, analyze the lane profiles in terms of both the amount of protein (in this case, opacity = optical density) and how the different lane profiles are spatially correlated. This calculation compares two data arrays, in this case the sample lane profile data of the type illustrated in Figure 2 with the reference sample lane profile (in this case lane 8). The calculated value can vary from -1 (negative correlation) to 0 (no correlation) to 1 (positive correlation). Pearson pairwise correlation calculation is a way to correlate different lane profiles using the reference lane profile or the reference average of several lane profiles. However, there are many other ways available to correlate lane profiles.
[0023] Figure 1 The results of the correlation calculations for each of the ten lanes shown in Figure 3 are shown in the two-dimensional plot in Figure 1 . It is obvious that in terms of the amount of sample components, some lanes are well correlated with the reference sample, in which the separated proteins in the lane produce Figure 1 the opaque bands shown in Figure 1 , and give a score close to 1 (y-axis), and the entire lane profile of some lanes is also well correlated with the reference lane and thus gives a score close to 1 (x-axis). These data provide two similarity comparisons: similarity in the amount of chemical components (such as proteins) (y-axis); and similarity in the actual chemical component profiles (x-axis). It is also obvious that the samples in this embodiment fall into different groups based on the separation profile similarity scores.
[0024] As shown in Figure 4 , the data points allow for easy classification by grouping of samples. Such classification allows for rapid analysis and does not rely on the user's estimation of the degree of similarity of different profiles. The limitations for grouping samples can be based on the following settings:
[0025] A) Only similarity in the amount of the sample;
[0026] B) Only similarity in chemical separation profiles;
[0027] C) Similarity in the amount of sample and chemical separation profile.
[0028] Thus, depending on which samples have been analyzed, different rules for grouping samples can be applied. For example, it is considered that Figure 4 only the samples in group C are classified as falling within the acceptable similarity to the reference sample.
[0029] Example 2
[0030] Techniques similar to those described above can be applied to chromatographic separations as shown in Figure 5 , where an absorbance meter has been used to measure the bands of chemical components (such as proteins) emerging from a column containing a chromatographic medium (typical UV absorbance for measuring protein concentration). The resulting data (referred to as a chromatogram) is shown in Figure 5 , where two chromatograms are superimposed, and the data set contains photodetector output readings (equivalent to Figure 2 opacity / component amount measurements) and time (time-related) data (equivalent to the spatial / distance data also shown in Figure 2 ). This data set is processed in the same manner as described above, again providing a comparison score of the component amount and chemical similarity to the reference sample or average sample data. Subsequently, this processed data classifies the samples, for example, as pass or fail results in a biopharmaceutical manufacturing process. In Figure 5 , two chromatograms of two protein samples with the same starting concentration were compared using an Äkta pure25 chromatographic system (available from Cytiva Life Sciences). The samples differed in the type of Strep-Tactin tag used for separation. As a result, the chromatograms (i.e., the time separation profiles) showed different peak shapes and minor variations in the separation profiles. For this part of the chromatogram, similarity score analysis gave a similarity score for sample 1 of (0.90; 0.87), where 0.90 is the separation profile similarity score and 0.87 is the relative amount, which shows that sample 1 differs from the reference sample in terms of both the elution amount and, to some extent, the chemical binding to the separation matrix.
[0031] Those skilled in the art will realize that various analyses can be performed according to the present invention. For example, when considering chromatographic examples, additional dimensions in chromatographic analysis can be, for example, one or more of the following: pH, conductivity, and / or pressure. Thus, chromatographic run data (such as time-related data corresponding to pressure, pH, and / or conductivity data) can be used for sample classification. Such parameters can replace the relative amount or be used additionally as the y-axis data (ordinate data) of the similarity plot.
[0032] Figure 6Displays a 2D similarity scatter plot of 185 chromatograms from different cycles during protein purification using embodiments of the present invention. The plot shows how the relative peak areas of the chromatograms correlate with the separation profile similarity scores. Such plot data demonstrate that many runs can be analyzed at high resolution using embodiments of the present invention, and data points that may not be easily determined by a skilled operator can be further identified. Thus, various embodiments of the present invention can be used to alert an operator of outliers and / or can be used in an automated processing system to optimize control parameters for a chromatography-based bioprocessing system.
[0033] In Figure 6 , in an Äkta Pure 25™ system with pH step elution, using a Fibro HiTrapPrismA™ chromatography unit, data was derived for each of the 185 chromatograms from the purification cycles. A reference sample was selected as the first chromatogram in the series. All the chromatograms were clearly very similar, but Pearson correlation coefficient calculations were used to detect two runs that differed in peak shape. These two values appeared as outliers on the Figure 6 left side. After inspection, all the cyclic chromatograms were judged to be acceptable. (Note: Some of the values in these chromatograms were saturated and thus the peak areas were not proportional to the relative amounts for this particular data set.)
[0034] Figure 7 Displays a graphical user interface for presenting results generated using embodiments of the present invention. The GUI includes a window that provides a 2D visualization of the data. In this example, a plot of similarity scores versus relative amounts is displayed. Such a 2D scatter plot provides easy user visualization. The GUI may also or alternatively enable the user to select a reference profile in order to display profiles, such as electrophoresis lanes or chromatograms, and display the two-dimensional sample data in the scatter plot.
[0035] Furthermore, in various embodiments, the 2D scatter plot can enable the user to select regions of interest for analysis, remove all data points that have reached the maximum limit of the detector from the analysis, and / or use the scatter plot to set limits for grouping samples into different groups. Additionally, various embodiments may also allow the scatter plot to track the protein purification process, for example, by using trend lines or color gradients.
[0036] Figure 8A method showing various embodiments of the present invention. Method 100 includes the following steps: a. Separating at least partially 110 one or more chemical components of a fluid sample; b. Measuring and recording 120 the amount of the separated chemical components of the sample during or after chemical separation; c. Measuring and recording 130 the spatial or temporal separation profile of the sample components during or after separation and providing a data set thereof; d. Comparing 140 the amount of the separated components with one or more reference samples; e. Comparing 150 the spatial or temporal separation profile with the corresponding profile of the reference sample or each reference sample; f. Designating 160 a similarity score for the sample based on the similarity of the comparison of the amount or profile of the separated components as performed in step d. 140 and / or e. 150 above; g. Providing 170 a classification of the sample based on the similarity score.
[0037] In various embodiments involving electrophoresis, method 100 may include one or more of the following steps:
[0038] 1a. Selecting a region of interest in an image. For electrophoresis, the user can, for example, use a GUI of the above type to create a lane box.
[0039] 2a. Optionally, saturated regions of the lane profile can subsequently be excluded from the analysis, i.e., regions where the detector has reached its maximum value. This can be automatic or user-driven.
[0040] 3a. Correcting for non-uniform migration. The user can adjust the lane box and the lane to correct for non-uniform migration through the electrophoresis gel. Optionally, the lane profile scale can be corrected by comparison with a labeled sample (i.e., a sample with a known molecular weight) in other lanes, either by the user or automatically.
[0041] In various embodiments involving chromatography, method 100 may include one or more of the following steps:
[0042] 1b. Manually or automatically selecting a region of interest in a chromatogram by analysis software. This step is optional, or the entire chromatogram can be analyzed.
[0043] 2b. Optionally, saturated regions of the lane profile are excluded from the analysis, i.e., regions where the detector has reached its maximum value.
[0044] 3b. Performing alignment of the chromatograms for comparison. Chromatogram alignment can be automated by software algorithms or performed by the user. Alignment is typically based on the chromatographic operation performed, such as the elution time of a known reference sample, or the start of a phase. This step is optional and in some cases alignment is not required.
[0045] Subsequently, the following steps can be applied to both electrophoretic analysis and chromatogram analysis:
[0046] 4. If the data sets have different numbers of data points, the individual electrophoretic lane profiles or chromatograms can be sampled or interpolated to obtain the same number of data points for each sample.
[0047] 5. Analysis of all data points or a user-defined analysis range can then be performed. For example, in some cases, where there is only one target peak, the analysis range can be automated or adjusted by the user accordingly.
[0048] 6. Calculate the integrated signal of the lane or chromatogram sample. Alternatively, for each lane, sum the total amounts of all detected bands or peaks.
[0049] 7. In a preferred embodiment, all possible pairwise comparisons of N samples are made. This results in an N×N array. For example, one array for relative amounts and one array for profile similarity scores. This can also be used to compare pressure, pH, and / or conductivity data sets for chromatography.
[0050] 8. The GUI allows the user to select a reference lane or chromatogram, and subsequently a 2D scatter plot can be generated, showing relative amounts on the y-axis and profile similarity scores on the x-axis.
[0051] Thus, various embodiments and features of the invention have been described. This written description further uses examples to disclose the invention, including the preferred mode, and is provided to enable a person skilled in the art to practice the invention, including making and using any device or system and performing any incorporated method. The patentable scope of the invention is defined by the claims and may include other examples that occur to a person skilled in the art. If such other examples have structural elements that are not different from the literal language of the claims, or if they include equivalent structural elements that are not materially different from the literal language of the claims, then such other examples are intended to be within the scope of the claims. Any patent or patent application or commercially available product (such as a system or software) mentioned in the text herein is incorporated herein by reference in its entirety, as if it were incorporated separately, to the extent permitted.
Claims
1. A method for classifying a fluid sample, the method comprising the following steps: a. Separating at least partially one or more chemical components of the fluid sample; b. Measuring and recording the amounts of the separated chemical components of the fluid sample during or after chemical separation; c. Measuring and recording the spatial or temporal separation profile of the sample components during or after separation, and providing a data set thereof; d. Comparing the amounts of the separated chemical components with one or more reference samples; e. Comparing the spatial or temporal separation profile with the corresponding profile of the reference sample or each reference sample; f. Based on the similarity of the comparison of the amounts and profiles of the separated chemical components as carried out in steps d and e above, assigning a similarity score to the fluid sample, respectively using the amounts and profiles of the reference sample or each reference sample; and g. Providing a classification of the fluid sample based on the similarity score.
2. The method according to claim 1, wherein a plurality of samples are classified, wherein step b. comprises providing a data set of the separated components or the amounts of each separated component of each of the samples, wherein step c. comprises providing a data set of the spatial separation profile or the temporal separation profile of each sample, and wherein the two data sets are processed by an algorithm to provide a two-dimensional sample data set for each sample used in steps d. to g.
3. The method according to claim 2, wherein the samples of different classifications represent pass or fail results in quality control of a biopharmaceutical manufacturing process.
4. The method according to claim 2 or 3, wherein the presence or absence of samples in different classifications causes changes in the biopharmaceutical manufacturing process.
5. The method according to any one of claims 1-3, wherein the chemical separation is electrophoresis.
6. The method according to any one of claims 1-3, wherein the chemical separation is chromatography.
7. The method according to claim 6, wherein a time series of chromatography run pressure, pH and / or conductivity data, or a combination thereof, is used for classification of the sample.
8. The method according to any one of claims 1-3, wherein the spatial or temporal separation profile similarity score is calculated using a Pearson correlation function.
9. A computer program comprising program code for performing the method according to any one of claims 1-8 when the program is run on a computer.
10. The computer program according to claim 9, which is further operable to provide a graphical user interface for user selection and / or presentation of reference profiles and / or regions and / or data for analysis purposes, and / or electrophoresis lanes and / or chromatograms and / or two-dimensional scatter plots.
11. The computer program according to claim 10, wherein the graphical user interface is further configured to enable the user to remove data points that have reached the maximum limit of the detector from the analysis.
12. The computer program according to claim 10 or 11, wherein the graphical user interface is configured to present a two-dimensional scatter plot that can be used to set limits for grouping samples into different groups and / or to track a protein purification process.
13. The computer program according to claim 12, wherein the graphical user interface is configured to present a trend line and / or a color gradient to assist in tracking a protein purification process.
14. A biopharmaceutical manufacturing device configured to implement the method as defined in any one of claims 1-8 to inspect and / or control a biopharmaceutical manufacturing process.
Citation Information
Patent Citations
Method and device for characterising an analyte
US20180031530A1
Automated analysis of analytical GELS and blots
WO2019126693A1