A method for improving the throughput of compound-protein interaction experiments

CN116819086BActive Publication Date: 2026-05-29SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2022-06-07
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for measuring compound-target protein interactions are time-consuming and costly, with low detection throughput.

Method used

Multiple test compounds are mixed into several mixture systems according to certain mixing rules. The corresponding relationship between the interaction between each test compound and the target protein is constructed by an optimization algorithm. The interaction between the compound and the target protein is then analyzed in high throughput using existing measurement methods.

Benefits of technology

It significantly increases the throughput of compound-target protein assays by more than 10 times, while saving more than 90% of experimental costs and time, which has significant economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116819086B_ABST
    Figure CN116819086B_ABST
Patent Text Reader

Abstract

The application discloses a method for improving the flux of compound-protein interaction experiment. The method of the application adopts the method of mixing a plurality of test compounds according to a certain mixing rule to form a plurality of mixture systems, and establishing the corresponding relationship between the interaction ability of each test compound and the target protein and the mixture system, and then high-throughput analyzing the corresponding target protein of the test compound. The analysis method of the application can improve the existing test compound-target protein experiment detection flux by more than 10 times, save more than 90% of the experiment cost and time, greatly reduce the labor, time and experiment cost of consumables, and has significant economic significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of experimental design technology for drug target discovery, specifically relating to a method for increasing the throughput of experiments involving the interaction of compounds and proteins. Background Technology

[0002] Compound molecules typically modulate cellular processes through physical interactions with proteins in the body, thereby producing toxic and therapeutic effects. Identifying the binding targets of compound molecules is usually addressed using affinity-based or activity-based proteomics methods (e.g., ABPP, PAL, KinoBeads), which require prior chemical derivatization of the compound molecules, making the procedures relatively complex. While derivatization-free mass spectrometry (MS) methods such as SPROX, TPP, DARTS, LiP-MS, and CPP can be extended to a wider range of compounds, they still require longer sample preparation and instrument measurement times. Summary of the Invention

[0003] This application provides a method to increase the throughput of compound-protein interaction experiments, in order to solve the technical problems of high time and economic costs and low detection throughput in existing compound-target protein interaction measurement experiments.

[0004] To achieve the aforementioned objectives, this application provides a method for increasing the throughput of experiments involving the interaction of compounds and proteins. The method comprises the following steps:

[0005] The n test compounds are arranged into m mixture systems. Each mixture system includes at least two of the n test compounds. The types of test compounds contained in different mixture systems are different, and the difference in the types of test compounds contained in different mixture systems is within a first preset range. The same test compound exists in at least two different mixture systems, and the difference in the number of mixture systems in which each test compound exists is within a second preset range. The difference in the number of test compounds contained in each mixture system is within a third preset range.

[0006] m portions of target solution are prepared for each of the mixture systems, the target solution comprising all of the test compounds and target proteins included in the mixture system;

[0007] Measure the response value of the interaction ability between each of the mixture systems in each of the target solutions and the target protein; determine the interaction between any of the test compounds included in each of the mixture systems and the target protein based on the response value of the interaction ability between each of the mixture systems and the target protein.

[0008] Further, determining the interaction between any one of the test compounds included in each of the mixture systems and the target protein based on the response value of the interaction ability of each of the mixture systems with the target protein includes:

[0009] Based on the response value of the interaction ability between each of the mixture systems and the target protein, determine the contribution of each of the test compounds in each of the mixture systems to the response value of the mixture system;

[0010] If the contribution of the first test compound in any of the mixture systems to the response value of any of the mixture systems is greater than or equal to a preset threshold, then it is determined that the first test compound interacts with the target protein, wherein the first test compound is one of all the test compounds included in any of the mixture systems.

[0011] Furthermore, the step of assembling n test compounds into m mixture systems includes:

[0012] The n test compounds are arranged into m mixture systems according to an m×n permutation matrix S. Each row of the permutation matrix S represents one of the mixture systems, and each column represents one of the test compounds. The permutation matrix S includes m×n indicators, which are used to indicate whether the mixture system contains the test compound corresponding to the column in which the indicator is located. The same test compound exists in at least two of the mixture systems.

[0013] Furthermore, the permutation matrix S is obtained through the following method:

[0014] The n test compounds are mixed into m mixture systems, and each test compound exists in a different mixture system, where m≥3, n≥4, a≥2, and m, n, and a are all integers; the m mixture systems are denoted as an m×n initial arrangement matrix A of the test compounds, and the value of each element in the initial arrangement matrix A is a random number between 0 and 1;

[0015] Perform a binary conversion on the initial permutation matrix A of the test compounds: find the a values ​​in each column: X1, X2, X... i ... X a Any one of the a values ​​X i All are greater than the other values ​​in the column, where 1≤i≤a, and a is an integer; convert the a values ​​in each column to binary 1, and the other values ​​to binary 0, to obtain the transformation matrix S;

[0016] The first preset range, the second preset range, and the third preset range are controlled by optimizing the initial permutation matrix A of the test compounds. The optimization steps include:

[0017] Through the following objective function:

[0018] L=Sum(S·S T -I)+Sum(RS-Mean(RS)) 2

[0019] Find the objective function value L; where RS is the sum of each row of the transformation matrix S, I is the identity matrix, and S T It is the transpose of S; the initial permutation matrix A of the test compound is optimized by an optimization algorithm to minimize the objective function value L; the initial permutation matrix A of the test compound with the minimum objective function value L is subjected to the binary conversion to obtain the permutation matrix S.

[0020] Further, the measurement of the response value of the interaction ability between each of the mixture systems in each sample of the target solution and the target protein includes:

[0021] The interaction between the test compound and the target protein is measured using a method based on binding energy or activity, to measure the response value of the interaction ability between each mixture system and the target protein.

[0022] Furthermore, the target protein is derived from purified protein or cell lysate containing the target protein, and the response value of the interaction ability between each mixture system and the target protein is quantitatively measured by any one of ABPP, PAL, TPP, or LiP-MS.

[0023] Further, determining the contribution of each analyte compound in each mixture system to the response value of the mixture system based on the response value of the interaction ability between each mixture system and the target protein includes:

[0024] Based on the response value of the interaction ability between each of the mixture systems and the target protein, determine the response vector corresponding to each of the mixture systems;

[0025] Each of the response vectors is normalized so that its value is between 0 and 1; the response vector Y i The following relationship exists between Y and the permutation matrix S: i =S×β j +R,

[0026] Among them, Y iβ represents the response value of the interaction ability between the mixture system identified as i in m mixture systems and the target protein. j R represents the contribution of the analyte compound identified j in the mixture system identified i to the response value of the mixture system identified i, and R is a residual vector of length n.

[0027] Use traditional statistical methods or machine learning methods to build a regression model and optimize the solution of β. j The value of the residual R is minimized.

[0028] Furthermore, the optimization algorithm includes genetic algorithm and ant colony algorithm.

[0029] Furthermore, the traditional statistical methods include least squares method, LASSO regression method; and / or

[0030] The machine learning methods mentioned include support vector machines and random forests.

[0031] Compared with the prior art, this application has the following technical effects:

[0032] This application discloses a method for increasing the throughput of compound-protein interaction experiments. This method involves assembling multiple test compounds into several mixture systems according to a specific mixing rule, establishing a correspondence between the interaction ability of each test compound and its target protein and the mixture system, thereby enabling high-throughput resolution of the target protein corresponding to the test compound. This method can increase the throughput of existing compound-target protein detection experiments by more than 10 times, while saving more than 90% of experimental costs and time, significantly reducing labor, time, and experimental consumable costs, thus demonstrating significant economic benefits. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0034] Figure 1 A flowchart illustrating the optimization of the initial permutation matrix A provided in this embodiment of the application;

[0035] Figure 2 A schematic diagram of the permutation matrix S of a certain intermediate state in the optimization process provided in Embodiment 1 of this application;

[0036] Figure 3 A schematic diagram of the optimized permutation matrix S of the final state during the optimization process provided in Embodiment 1 of this application;

[0037] Figure 4 This is a schematic diagram showing how the objective function value L decreases to a constant as the number of iterations increases, as provided in Embodiment 1 of this application.

[0038] Figure 5 A schematic diagram of the optimized permutation matrix S of the final state provided in Embodiment 2 of this application;

[0039] Figure 6 This is a schematic diagram showing how the objective function value L decreases to a constant as the number of iterations increases, as provided in Embodiment 2 of this application. Detailed Implementation

[0040] To make the technical problems, technical solutions, and beneficial effects of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0041] This application provides a method for increasing the throughput of experiments involving compound-protein interactions. The method includes the following steps:

[0042] (1) n kinds of test compounds are combined into m mixture systems. Each mixture system includes at least 2 kinds of n test compounds. The types of test compounds contained in different mixture systems are different. The difference in the types of test compounds contained in different mixture systems is within a first preset range. The same test compound exists in at least 2 different mixture systems. The difference in the number of mixture systems in which each test compound exists is within a second preset range. The difference in the number of test compounds contained in each mixture system is within a third preset range.

[0043] (2) Prepare m portions of target solution for each mixture system. The target solution includes all the test compounds and target proteins included in the mixture system.

[0044] (3) Measure the response value of the interaction ability between each mixture system in each target solution and the target protein; based on the response value of the interaction ability between each mixture system and the target protein, determine the interaction between any one of the test compounds in each mixture system and the target protein.

[0045] This application discloses a method for increasing the throughput of compound-protein interaction experiments. This method involves assembling multiple test compounds into several mixture systems according to a specific mixing rule, establishing a correspondence between the interaction ability of each test compound and its target protein and the mixture system, thereby enabling high-throughput resolution of the target protein corresponding to the test compound. This method can increase the throughput of existing compound-target protein detection experiments by more than 10 times, while saving more than 90% of experimental costs and time, significantly reducing labor, time, and experimental consumable costs, thus demonstrating significant economic benefits.

[0046] In step (1) above, the "first preset range" is controlled by making the types of compounds contained in each mixture system as different as possible; the "second preset range" is controlled by making the number of test compounds contained in each mixture system as consistent as possible; and the "third preset range" is controlled by making the number of mixture systems in which each test compound exists as consistent as possible.

[0047] Further, step (1) above, "combining n test compounds into m mixture systems," specifically includes: combining n test compounds into m mixture systems according to an m×n optimized permutation matrix S. Each row in the permutation matrix S represents a mixture system, and each column represents a test compound. The permutation matrix S includes m×n indicators, which are used to indicate whether the mixture system contains the test compound corresponding to the column in which the indicator is located. The same test compound exists in at least two mixture systems. For example, in Embodiment 1 of this application, in a 9×15 permutation matrix S, the permutation matrix S represents combining 15 test compounds into 9 mixture systems. Figure 2 , Figure 3 The vertical axis (1-9) represents the nine mixture systems formed by mixing, and the horizontal axis (1-15) represents the 15 test compounds. Furthermore, Figure 2 , Figure 3 The black squares represent the presence of the corresponding test compound in the mixture system represented by that row, while the white squares represent the absence of the corresponding test compound in the mixture system represented by that row.

[0048] For example, in Figure 3 In the text, mixture system 1 represents the mixture system containing test compound 4, test compound 5, test compound 7, test compound 12 and test compound 13.

[0049] Furthermore, the permutation matrix S described above can be obtained using the following method:

[0050] The n test compounds are mixed into m mixture systems, and each test compound exists in a different mixture system, where m≥3, n≥4, a≥2, and m, n, and a are all integers; the m mixture systems are denoted as an m×n initial arrangement matrix A of the test compounds, and the value of each element in the initial arrangement matrix A is a random number between 0 and 1;

[0051] Perform a binary conversion on the initial permutation matrix A of the test compounds: find the a values ​​in each column: X1, X2, X... i ... X a Any one of the a values ​​X i All values ​​are greater than the other values ​​in the column, where 1 ≤ i ≤ a, and a is an integer. The a values ​​in each column are converted to binary 1, and the other values ​​are converted to binary 0, resulting in a transformation matrix S. The numbers 0 and 1 in the transformed transformation matrix S are the indicators mentioned above. In the transformation matrix S, a value of 1 represents that the corresponding mixture system contains the test compound, and a value of 0 represents that the corresponding mixture system does not contain the test compound.

[0052] The aforementioned "first preset range", "second preset range" and "third preset range" can be controlled by optimizing the initial arrangement matrix A of the test compounds. Specific optimization can be performed through the following steps:

[0053] Through the following objective function:

[0054] L=Sum(S·S T -I)+Sum(RS-Mean(RS)) 2

[0055] Find the objective function value L; where RS is the sum of each row of the transformation matrix S, I is the identity matrix, and S T This is the transpose of S. The correlation between columns in the transformation matrix S can be expressed by (S·S) T -I) is obtained, denoted as the correlation matrix. The first term Sum(S·S) in the objective function L T -I) is used to ensure that the correlation between columns in the transformation matrix S is minimized. The second term is Sum(RS-Mean(RS)). 2 This is used to ensure that the number of test compounds contained in each mixture system is approximately the same.

[0056] Then, the initial permutation matrix A of the test compounds can be optimized using an optimization algorithm to minimize the objective function value L. The optimized matrix A, where the objective function value L is minimized, is then subjected to the aforementioned binary transformation to obtain the optimized permutation matrix S. The optimization algorithm includes, but is not limited to, genetic algorithms and ant colony algorithms.

[0057] The optimization process for the initial permutation matrix A in this embodiment is as follows: Figure 1 As shown.

[0058] In this way, this application maps the n test compounds into m mixture systems to an m×n permutation matrix S, and then establishes a correspondence between the interaction ability of the test compounds and the target protein and the permutation matrix S, thereby resolving the target protein corresponding to the test compound located at a specific position in the permutation matrix S in a high-throughput manner.

[0059] Furthermore, the step (3) above, "determining the interaction between any one of the test compounds in each mixture system and the target protein based on the response value of the interaction ability of each mixture system with the target protein," can be specifically determined by the following method:

[0060] Based on the response value of the interaction ability between each mixture system and the target protein, determine the contribution of each test compound in each mixture system to the response value of the mixture system;

[0061] If the contribution of the first test compound in any mixture system to the response value of any mixture system is greater than or equal to a preset threshold, then the first test compound is determined to interact with the target protein, wherein the first test compound is one of all test compounds included in any mixture system.

[0062] Further, "determining the contribution of each of the test compounds in each of the mixture systems to the response value of the mixture system" includes:

[0063] Based on the response value of the interaction ability between each mixture system and the target protein, the response vector corresponding to each mixture system is determined. Specifically, any existing measurement method based on binding energy or activity of the analyte and the target protein can be used to measure the response value of the interaction ability between each mixture system and the target protein. For example, the response value of the interaction ability between each mixture system and the target protein can be quantitatively measured by any of the following methods: ABPP (activity-directed proteomic analysis), PAL (photocrosslinking affinity), TPP (thermal proteomics analysis), or LiP-MS (limited protein hydrolysis mass spectrometry). In specific embodiments 1 and 2 of this application, the interaction between each mixture system and the target protein is measured using the single-temperature TPP method.

[0064] Each response vector is normalized so that its value is between 0 and 1; response vector Y i The following relationship exists between Y and the permutation matrix S: i =S×β j +R,

[0065] Among them, Y i β represents the response value of the mixture system identified as i in m mixture systems, indicating its interaction ability with the target protein. j R represents the contribution of the analyte compound labeled j in the mixture system labeled i to the response value of the mixture system labeled i, and R is a residual vector of length n.

[0066] Use traditional statistical methods or machine learning methods to build a regression model and optimize the solution of β. j The goal is to find the numerical value of the residual R and minimize it. Traditional statistical methods include, but are not limited to, least squares and LASSO regression, while machine learning methods include, but are not limited to, support vector machines and random forest regression.

[0067] In a specific embodiment of this application, the optimal solution for β can be obtained using the LASSO regression method according to the following formula:

[0068]

[0069] Where λ is a penalty term used to adjust the degree of compression of β. In specific embodiments 1 and 2 of this application, λ is set to 0.1, and the threshold is set to 0.1. If the calculated β corresponding to a certain test compound is... i If the value is higher than 0.1, it can be considered that the corresponding test compound interacts with the target protein.

[0070] The following specific embodiments illustrate a method for analyzing the interaction between a compound and a protein according to an example of this application.

[0071] Example 1

[0072] This embodiment 1 provides a method for increasing the throughput of experiments involving the interaction of compounds and proteins, comprising the following steps:

[0073] S01: Given 15 drugs to be tested, i.e., test compounds: Palbociclib, Panobinostat, Raltitrexed, Methotrexate, Vemurafenib, Fimepinostat, SCIO-469, SL-327, 5-Fluorouracil, Olaparib, Belumosudil, OTS964, Parthenolide, CCT137690, Belumosudil, numbered sequentially as test compound 1, test compound 2, test compound 3, ..., test compound 15. Assume that the above 15 test compounds are mixed into 9 mixture systems: mixture system 1, mixture system 2, mixture system 3, ..., mixture system 9, each test compound exists in 3 different mixture systems, i.e., m = 9, n = 15, a = 3. Denote the 9 mixture systems as a 9×15 initial permutation matrix A of test compounds, and randomly initialize it so that all its values ​​are random floating-point numbers between 0 and 1;

[0074] S02: Perform binary conversion on the initial permutation matrix A of the test compound: find three values ​​in each column: X1, X2, X3, any one of these three values ​​is greater than the other six values ​​in the column; convert these three values ​​in each column to binary 1, and convert the other values ​​to binary 0, to obtain the conversion matrix S;

[0075] S03: Through the following objective function:

[0076] L=Sum((S·S T -I)+Sum(RS-Mean(RS)) 2

[0077] Find the objective function value L; where RS is the sum of each row of the transformation matrix S, I is the identity matrix, and S T It is the transpose of S.

[0078] The initial permutation matrix A of the test compounds is iteratively optimized using a genetic algorithm to minimize the objective function value L. The permutation matrix A, where L is minimized, undergoes a binary conversion in step S02 to obtain the optimized permutation matrix S (e.g., ...). Figure 3 As shown in the figure, the optimized permutation matrix S is the final permutation matrix. Figure 2 A schematic diagram of the permutation matrix S of a certain intermediate state during the optimization process is shown. Figure 2 , Figure 3 In the diagram, black squares represent the value 1, and white squares represent the value 0. During the iteration process, the objective function value L decreases until it becomes constant as the number of iterations increases, as shown below. Figure 4As shown. The 15 test compounds were arranged according to the final obtained... Figure 3 The 9×15 optimized permutation matrix S shown is mixed to obtain 9 mixture systems.

[0079] S04: Add DMSO as a solvent to each mixture system in step S03 to achieve a concentration of 40 μM for each drug. Then, use the single-temperature point TPP method to measure the interaction between each mixture system and the protein. The specific experimental steps are as follows: Mix equal volumes of K562 cell lysate and drug mixture systems, incubate at room temperature for 10 minutes, then heat at 52°C for 3 minutes, and then rapidly cool to 4°C on a PCR machine. Centrifuge the samples at 21000 rcf for 20 minutes at 4°C and collect the supernatant. According to the mass spectrometry-based whole proteomics quantitative method, after enzymatic digestion and TMT labeling of each sample, the content of each protein is measured using LC-MS mass spectrometry. According to the principle of the TPP method, binding to the compound can improve the thermal stability of the protein. Therefore, if a protein can interact with the mixture system, the protein content measured by mass spectrometry will be higher, and vice versa.

[0080] S05: Nine measurements Y obtained from the interaction of nine mixture systems with the target protein. i Let Y be the response vector; normalize the response vector Y so that its value is between 0 and 1; the response vector Y and the permutation matrix S have the following relationship:

[0081] Y i =S×β j +R,

[0082] Among them, Y i This represents the response value of the mixture system labeled i out of the nine mixture systems, indicating its ability to interact with the target protein, β. j β represents the contribution of the analyte compound labeled j in the mixture system labeled i to the response value of the mixture system labeled i. j Let R be a coefficient vector of length n, and let R be a residual vector of length n.

[0083] The optimal solution for β is obtained using the LASSO regression method as follows:

[0084]

[0085] Here, λ is a penalty term used to adjust the degree of compression of β. In this embodiment, the value of λ is 0.1, and the threshold is 0.1. If the calculated β corresponding to a certain test drug... i A value higher than 0.1 indicates that the corresponding drug compound being tested interacts with the target protein.

[0086] Ultimately, by preparing nine mixture systems, the targets of 15 test compounds were identified, of which 11 were successfully identified (see Table 1 below), with a success rate of 73.3%. On average, only 0.6 samples were needed to prepare each test compound.

[0087] In contrast, the conventional single-temperature TPP method (Ball et al., Commun. Biol. 2020(3).75) requires four dosing group samples and four control group samples for target identification of each drug compound, and eight samples are needed to identify the target of one drug. Therefore, the method of this application increases the detection throughput by 8 / 0.6 = 13.3 times compared with the conventional method, and reduces the experimental cost by (8-0.6) / 8 = 92.5%.

[0088] Table 1. Target identification results of 15 drug compounds to be tested

[0089]

[0090] The method in Example 1 of this application is an effective method to improve the throughput of drug compound target identification. First, an optimized permutation matrix S is constructed through an optimization algorithm, and multiple test compounds are arranged into a mixture system according to the optimized permutation matrix S. Then, combined with existing compound-target interaction measurement methods, the correspondence of compound-target interactions is analyzed using statistical methods. This method can increase the number of compounds that can be analyzed for target identification by more than 10 times under the same experimental cost, significantly reducing manpower, time, and experimental consumable costs, and has significant economic benefits.

[0091] Example 2

[0092] This embodiment 2 provides a method for increasing the throughput of experiments involving the interaction of compounds and proteins, comprising the following steps:

[0093] Based on the 15 test compounds given in Example 1, Tioxolone, Parthenolide, Abemaciclib, Caffeic acid phenethyl ester, RG2833, Encorafenib, TAK-285, CNX-774, Dienogest, and ZM241385 were added, for a total of 25 test compounds, numbered sequentially as test compound 1, test compound 2, test compound 3, ..., test compound 25. It is assumed that the above 25 test compounds are mixed into 14 mixture systems: mixture system 1, mixture system 2, mixture system 3, ..., mixture system 14, with each test compound existing in 4 different mixture systems, i.e., m = 14, n = 25, a = 4. Following the steps in Example 1, the optimized permutation matrix S was obtained using a genetic algorithm, as shown below. Figure 5 As shown. Figure 5 In the diagram, black squares represent the value 1, and white squares represent the value 0. During the iteration process, the objective function value L decreases until it becomes constant as the number of iterations increases, as shown below. Figure 6 As shown. The 25 test compounds were arranged according to the final obtained... Figure 5 The optimized permutation matrix S of 14×25 shown is mixed to obtain 14 mixture systems. The remaining operation steps are the same as in Example 1.

[0094] Ultimately, by preparing 14 mixture systems, the targets of 25 test compounds were identified, of which 14 were successfully identified (see Table 2 below), with a success rate of 56%. On average, only 0.56 samples were prepared for each test compound.

[0095] In contrast, the conventional single-temperature TPP method (Ball et al, Commun. Biol. 2020(3).75) requires four dosing group samples and four control group samples for target identification of each drug, and eight samples are needed to identify the target of one drug. Therefore, the method of this application increases the detection throughput by 8 / 0.56 = 14.3 times compared with the conventional method, and reduces the experimental cost by (8-0.56) / 8 = 93.0%.

[0096] Table 2. Target identification results of 25 drug compounds to be tested

[0097]

[0098] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for increasing the throughput of experiments involving the interaction of compounds and proteins, characterized in that, Includes the following steps: The n test compounds are arranged into m mixture systems. Each mixture system includes at least two of the n test compounds. The types of test compounds contained in different mixture systems are different, and the difference in the types of test compounds contained in different mixture systems is within a first preset range. The same test compound exists in at least two different mixture systems, and the difference in the number of mixture systems in which each test compound exists is within a second preset range. The difference in the number of the test compounds contained in each of the mixture systems is within a third preset range; m portions of target solution are prepared for each of the mixture systems, the target solution comprising all of the test compounds and target proteins included in the mixture system; Measuring the response value of the interaction ability between each of the mixture systems in each of the target solutions and the target protein includes: quantitatively measuring the response value of the interaction ability between each of the mixture systems and the target protein using a measurement method based on binding energy or activity of the test compound and the target protein; Based on the response value of the interaction ability of each of the mixture systems with the target protein, the interaction between any of the test compounds included in each of the mixture systems and the target protein is determined; The step of assembling n test compounds into m mixture systems includes: According to the m×n permutation matrix S The n test compounds are arranged into m mixture systems, and the permutation matrix is... S Each row in the matrix represents one of the mixture systems, and each column represents one of the test compounds. S It includes m×n indicators, which are used to indicate whether the mixture system contains the test compound corresponding to the column of the indicator, and the same test compound exists in at least two mixture systems; The permutation matrix S It can be obtained through the following methods: The n test compounds are mixed into m mixture systems, and each test compound exists in a different mixture system, where m≥3, n≥4, a≥2, and m, n, and a are all integers; the m mixture systems are denoted as an m×n initial arrangement matrix A of the test compounds, and the value of each element in the initial arrangement matrix A is a random number between 0 and 1; Perform a binary conversion on the initial permutation matrix A of the test compounds: find the a values ​​in each column: X1, X2, X... i ... X a Any one of the a values ​​X i All values ​​are greater than the other values ​​in the column, where 1 ≤ i ≤ a, and a is an integer; convert the a values ​​in each column to binary 1 and the other values ​​to binary 0 to obtain the transformation matrix. S ; The first preset range, the second preset range, and the third preset range are controlled by optimizing the initial permutation matrix A of the test compounds. The optimization steps include: Through the following objective function: Find the objective function value L ;in, RS It is the transformation matrix S The sum of each row, I It is the identity matrix. ST yes S The transpose matrix; the initial permutation matrix A of the test compounds is optimized using an optimization algorithm to improve the objective function value. L Minimize; reduce the objective function value L The initial permutation matrix A of the test compounds at its smallest value is obtained by binary conversion to obtain the permutation matrix. S The objective function L The first item This is used to ensure that the correlation between columns in the transformation matrix S is minimized; the second term... This is used to ensure that the number of test compounds contained in each mixture system is approximately the same.

2. The method according to claim 1, characterized in that, The step of determining the interaction between any one of the test compounds in each of the mixture systems and the target protein based on the response value of the interaction ability of each of the mixture systems with the target protein includes: Based on the response value of the interaction ability between each of the mixture systems and the target protein, determine the contribution of each of the test compounds in each of the mixture systems to the response value of the mixture system; If the contribution of the first test compound in any of the mixture systems to the response value of any of the mixture systems is greater than or equal to a preset threshold, then it is determined that the first test compound interacts with the target protein, wherein the first test compound is one of all the test compounds included in any of the mixture systems.

3. The method according to claim 1 or 2, characterized in that, The target protein is derived from purified protein or cell lysate containing the target protein, and the response value of the interaction ability between each mixture system and the target protein is quantitatively measured by any one of ABPP, PAL, TPP or LiP-MS.

4. The method according to claim 1, characterized in that, The step of determining the contribution of each analyte compound in each mixture system to the response value of the mixture system based on the response value of the interaction ability between each mixture system and the target protein includes: Based on the response value of the interaction ability between each of the mixture systems and the target protein, determine the response vector corresponding to each of the mixture systems; Each of the response vectors is normalized so that its value is between 0 and 1; the response vector Y i With the permutation matrix S The following relationship exists between them: Y i = S ×β j +R, Among them, Y i β represents the response value of the interaction ability between the mixture system identified as i in m mixture systems and the target protein. j R represents the contribution of the analyte compound identified j in the mixture system identified i to the response value of the mixture system identified i, and R is a residual vector of length n. Use traditional statistical methods or machine learning methods to build a regression model and optimize the solution of β. j The value of the residual R is minimized.

5. The method according to claim 1, characterized in that, The optimization algorithms include genetic algorithms and ant colony algorithms.

6. The method according to claim 4, characterized in that, The traditional statistical methods include least squares regression and LASSO regression.

7. The method according to claim 4, characterized in that, The machine learning methods mentioned include support vector machines and random forests.