Non-equilibrium drug reaction data resampling method and device based on swarm intelligence optimization, storage medium and equipment

By introducing the resampling framework R-SIOIC based on population intelligence optimization into drug reaction data classification, the problem of unbalanced data processing in drug reaction multiomics data sets is solved, and the accuracy and generalization ability of drug reaction data classification is improved.

CN120220845APending Publication Date: 2025-06-27DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381056.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of unbalanced data in multiomics data sets of drug responses, resulting in a decrease in the accuracy of drug redirection.

Method used

A resampling framework based on population intelligence optimization is proposed. By dynamically adjusting the sample selection strategy, the continuous population intelligence optimization algorithm (such as PSO, GWO, WOA) is used to iteratively update the position of the drug reaction data candidate set based on the Heming distance measurement mechanism to generate the global optimal drug reaction data resampling results.

Benefits of technology

It significantly improves the accuracy and generalization ability of drug response data classification, especially in the identification of a few samples, improves the accuracy of drug redirection, and shows stability and robustness on different data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220845A_ABST
    Figure CN120220845A_ABST
Patent Text Reader

Abstract

The invention provides an unbalanced drug response data resampling method and device based on swarm intelligence optimization, a storage medium and equipment, and belongs to the field of unbalanced data classification in data mining. The method comprises the following steps: firstly, generating an initial sampled drug reaction data candidate set; training the classification performance of a classifier according to each drug reaction data candidate set; dynamically adjusting a sample selection strategy, and iteratively updating samples contained in each candidate set through a position updating formula; and finally, outputting to obtain a globally optimal drug reaction resampling data set. The method can be compatible with continuous swarm intelligent optimization algorithms (such as a particle swarm optimization algorithm and a grey wolf optimization algorithm), and is suitable for drug reaction unbalanced data classification tasks of different scales and distribution characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of imbalanced data classification in data mining, and the implementation result is the optimized resampling of an imbalanced drug response data set. In particular, the present invention relates to a method, device, storage medium and equipment for resampling imbalanced drug response data based on swarm intelligence optimization. Background Art

[0002] In real-world applications, there is a widespread phenomenon of severely imbalanced data class distributions in fields such as medical diagnosis, equipment fault detection, and credit risk assessment, that is, the "imbalanced data problem". In such scenarios, the identification of minority class samples (such as diseased individuals, equipment fault records) is significantly more important than that of the majority class. However, traditional classification methods tend to favor the majority class in imbalanced data, resulting in a significant reduction in the classification accuracy of the minority class, which seriously affects the practical application value. For example, in the drug repositioning scenario, it is necessary to determine which cell lines a drug can react with. However, in experimental data, the number of cell lines that react with the drug is often much lower than that of non-reacting cell lines, resulting in a decrease in the accuracy of the model to identify reacting cell lines, and thus seriously affecting the accuracy of drug repositioning.

[0003] As a common imbalanced data set, the drug response multi-omics data set poses a great challenge to existing imbalanced data processing methods. Due to its complex feature relationships and high feature dimensions, existing imbalanced classification methods usually have difficulty capturing its effective features. Excessive features will reduce the classification accuracy of the model, which is particularly serious in imbalanced data sets, especially when the number of data points in the smallest class is very small and it is difficult to find a stable decision boundary in the high-dimensional space. Currently, drug response multi-omics data sets are widely used in the field of drug applications, including the discovery of new drug targets, the repositioning of existing drugs, etc. The imbalanced problem of drug response data affects the accurate classification of drug response data.

[0004] The related technologies involved in the present invention include swarm intelligence optimization methods and resampling technologies, etc. Swarm intelligence optimization algorithms (such as Particle Swarm Optimization PSO, Ant Colony Optimization ACO, Artificial Bee Colony Algorithm ABC) show high-efficiency global search capabilities in solving complex optimization problems by simulating the intelligent behaviors of biological groups. In recent years, such algorithms have begun to be introduced into the imbalanced classification problem of drug response data, but the existing applications have significant limitations: in the resampling optimization scenario, related research mainly focuses on combinatorial optimization algorithms (such as ant colony algorithms), while more continuous space optimization algorithms (such as PSO, GWO) are difficult to be directly applied to sample selection optimization due to the lack of an effective discrete decision mapping mechanism. In addition, existing methods usually bind swarm intelligence individuals to a single sampling strategy and fail to construct a unified framework implementation scheme, resulting in high algorithm transplantation costs and limited optimization dimensions.

[0005] Resampling techniques (including undersampling, oversampling, and hybrid sampling) have been proposed to balance data distribution. Classical resampling methods usually optimize the class distribution of a dataset through a single sampling of the dataset. For example, the SMOTE (Synthetic Minority Over-sampling Technique) oversampling algorithm generates new samples through linear interpolation, and the RUS (Random Under Sampling) undersampling randomly deletes majority-class samples. Such methods have inherent defects. The data distribution generated by preset rules will cause obvious distribution shift problems for classifiers when facing imbalanced datasets with high-dimensional characteristics. Moreover, due to the lack of a model performance feedback mechanism, data distribution optimization cannot be achieved through dynamic adjustment, ultimately resulting in insufficient generalization ability of the classifier.

[0006] The purpose of the present invention is to innovatively construct a resampling framework R-SIOIC (Resampling framework based on Swarm Intelligence Optimization for Imbalanced data Classification) for the task of classifying imbalanced drug response data. This framework dynamically adjusts the sample selection strategy through a swarm intelligence optimization mechanism, including an initialization module, a verification module, a swarm intelligence optimization module, and an output module. It supports the seamless embedding of most continuous swarm intelligence optimization algorithms and finally outputs the globally optimal resampling result of drug response data. Summary of the Invention

[0007] The R-SIOIC framework sequentially includes an initialization module, a verification module, a swarm intelligence optimization module, and an output module: The initialization module generates an initial candidate set of sampled drug response data; the verification module evaluates the classification performance of the classifier trained by each candidate set of drug response data; the swarm intelligence optimization module dynamically adjusts the sample selection strategy and iteratively updates the position of each candidate set through a position update formula; the output module outputs the globally optimal resampled drug response dataset. This framework realizes compatibility with continuous swarm intelligence optimization algorithms (such as particle swarm optimization, grey wolf optimization algorithm, etc.) through modular design and is applicable to imbalanced drug response data classification tasks with different scales and distribution characteristics.

[0008] The technical solution of the present invention:

[0009] A method for resampling imbalanced drug response data based on swarm intelligence optimization specifically includes the following steps:

[0010] Step 1, resample to generate a candidate set of drug response data and train a base classifier

[0011] The original drug response unbalanced dataset includes two types of cell lines, one type is the cell lines that can respond to the drug, and the other type is the cell lines that cannot respond to the drug;

[0012] During resampling, n initial drug response data candidate sets are generated from the original drug response unbalanced dataset. Each drug response data candidate set is a balanced subset composed of an equal number of cell lines selected from each class. The specific number of cell lines selected from each class is determined by the minimum number of cell lines in the class, min_class_amount. The formula is as follows:

[0013] num = sample_size * min_class_amount

[0014] Among them, sample_size represents the proportion of cell lines selected from the smallest class, which is set to 0.6 - 0.8. The specific value needs to ensure the quantity of training data while retaining a certain degree of randomness.

[0015] After resampling, the original drug response unbalanced dataset is encoded as a binary vector X = [x1, x2,..., x i ,..., x n , where x i ∈ {0, 1} represents the cell line selection status during resampling. "1" means the cell line is selected, and "0" means it is ignored. The encoded binary vector is used as the position vector of each drug response data candidate set.

[0016] Each drug response data candidate set corresponds to a base classifier (such as an SVM classifier), and the base classifier is trained using the drug response data candidate set;

[0017] Step 2: Evaluate the classification performance of the drug response data candidate sets

[0018] The classification performance of all drug response data candidate sets is evaluated using the AUC metric. The top m candidate sets with the best performance in the current iteration are selected to guide the optimization process. If there are candidate sets with better performance than the current top m candidate sets in subsequent iteration processes, the optimal record is updated. The unselected candidate sets proceed to Step 3 for position update;

[0019] Step 3: Iteratively update the positions of the drug response data candidate sets

[0020] This step includes two-phase operations:

[0021] In the first phase, the Hamming distance between the drug response data candidate sets is calculated to measure the difference between different resampling results.

[0022] In the second phase, the positions of the drug response data candidate sets are updated according to the selected swarm intelligence algorithm. The position update formula is defined as:

[0023] X new = Update(X current , X top-m , D Hamming )

[0024] where X top-m is the current optimal candidate set of m drug response data, D Hamming is the Hamming distance, X current represents the position of the current candidate set, and Update(·) represents the position update function.

[0025] Furthermore, the process of updating the position of the drug response data candidate set by different swarm intelligence algorithms in the second stage is as follows:

[0026] In this study, three continuous SIO methods - Grey Wolf Optimization (GWO), Particle Swarm Optimization (PSO), and Whale Optimization Algorithm (WOA) - are used as test algorithms, and their specific update methods are as follows:

[0027] (1) The Grey Wolf Optimization (GWO) method simulates the social hierarchy and cooperative hunting behavior of a grey wolf pack, and guides the drug response data candidate set population to update its position through three current optimal candidate sets, namely α, β, and δ. The specific implementation involves three stages of operations:

[0028] In the first stage, the Hamming distance is used to calculate the distances between the current candidate set and the optimal candidate sets respectively:

[0029]

[0030] where: represents the position of the current candidate set (i.e., X in the second stage of step 3 current ), represents the positions of the three current optimal candidate sets (i.e., X in the second stage of step 3 top-m ), Hamming(·) represents the Hamming distance, represents the distance between the current candidate set and the three optimal candidate sets.

[0031] In the second stage, candidate positions are generated based on the distances between the current candidate set and each optimal candidate set:

[0032]

[0033] where: is the candidate position updated according to the positions of the three current optimal candidate sets;

[0034] The parameter is a parameter that controls the search pattern of the optimal drug response data candidate set. When Perform a global search when and perform a local search when

[0035] In the third stage, the candidate positions are weighted and averaged to obtain the position of the next iteration round of the current drug response data candidate set:

[0036]

[0037] (2) The whale optimization algorithm (WOA) is based on the spiral bubble-net hunting strategy of humpback whales and includes a special behavioral pattern of spirally approaching the optimal drug response resampling data set:

[0038]

[0039] where b is the spiral shape coefficient; l ∈ [-1, 1] controls the spiral step size, represents the distance between the current candidate set and the optimal candidate set, represents the distance between the current candidate set and the random candidate set, represents the optimal candidate set, represents the random candidate set. The update formula switches different behavioral patterns according to the probability p and the parameter (the same as the parameter defined in GWO):

[0040] 1) When p > 0.5, the drug response data candidate set approaches the possible optimal drug response resampling data set along a spiral path.

[0041] 2) When p ≤ 0.5 and , the drug response data candidate set optimizes in the direction of the current optimal candidate set, aiming to conduct a detailed search near the current optimal candidate set to find the optimal drug response resampling data set.

[0042] 3) When p ≤ 0.5 and , the drug response data candidate set randomly selects a candidate set (non-optimal candidate set) in the population as a reference and conducts a random search by updating its own position.

[0043] (3) The particle swarm optimization (PSO) method simulates the foraging behavior of bird flocks. In this algorithm, the position of each drug response data candidate set represents a candidate solution in the continuous solution space. Each drug response data candidate set updates its own position through velocity, and the update of velocity is affected by both the historical optimal position of each drug response data candidate set and the historical best position of the drug response data candidate set population.

[0044] The velocity update formula of the drug response data candidate set is as follows:

[0045]

[0046] Among them, r1 and r2 are random numbers within the interval [0, 1], respectively. is the speed of the previous round of drug response data candidate set iteration (i.e., the inertial speed). ω, c1, and c2 are all preset hyperparameters. D7 and D8 represent the distances between the current drug response data candidate set and the self-historical optimal position and the population historical best position of the current candidate set, respectively:

[0047]

[0048] Among them, x pbest represents the historical optimal position of the current candidate set. represents the position of the current candidate set in the current iteration round. x gbest represents the overall historical optimal position among all candidate sets.

[0049] Step 4: Generate the optimal resampled dataset of drug response data

[0050] When the preset maximum number of iterations is reached, terminate the optimization process, and select the historical optimal drug response data candidate set as the output of the optimal resampled dataset of drug response data.

[0051] Furthermore, the effectiveness of the obtained optimal resampled dataset of drug response data is verified in the following way: Use the test set to evaluate the classification performance of the classifier trained by the optimal resampled dataset of drug response data, and compare it with the drug response data classification method MOLI.

[0052] A non-equilibrium drug response data resampling device based on swarm intelligence optimization, the non-equilibrium drug response data resampling device includes:

[0053] Initialization module: used to resample and generate a drug response data candidate set;

[0054] Base classifier: Use the resampled drug response data candidate set to train the base classifier;

[0055] Verification module: Evaluate the classification performance of the classifier trained by each drug response data candidate set;

[0056] Swarm intelligence optimization module: Dynamically adjust the sample selection strategy, and iteratively update the position of each candidate set through the position update formula;

[0057] Output module: Output the globally optimal drug response resampled dataset.

[0058] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the non-equilibrium drug response data resampling method based on swarm intelligence optimization.

[0059] A server includes a processor and a memory. The memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the non-equilibrium drug response data resampling method based on swarm intelligence optimization.

[0060] Advantageous effects of the present invention:

[0061] (1) Innovation of the method

[0062] The present invention first proposes a framework for solving the non-equilibrium classification problem of drug response data by combining a continuous swarm intelligence optimization algorithm with a resampling technique. By defining the position vector of the drug response data candidate set as a binary vector and introducing a Hamming distance metric mechanism, the limitations of traditional combinatorial optimization problems are broken through, and the direct application of continuous optimization algorithms such as PSO and GWO in the selection problem of discrete drug response data candidate sets is realized.

[0063] (2) Improvement of the classification performance of the non-equilibrium drug response data set

[0064] The present invention first uses a continuous swarm intelligence optimization algorithm and a resampling technique simultaneously in the non-equilibrium classification problem of drug response data, and significantly alleviates the problem of the decrease in the classification accuracy of the minority class (response data) caused by the non-equilibrium problem in the field of drug response data classification. In the test, R-SIOIC ranks first in the AUC index among seven different drug response data sets, verifying the excellent classification effect and the robustness of its cross-data set scenario.

[0065] (3) Breakthrough in generalization ability

[0066] Aiming at the problem of insufficient generalization ability of traditional undersampling methods in high-dimensional data scenarios, R-SIOIC significantly improves the recognition accuracy of the classifier for minority class samples by dynamically optimizing the sample selection strategy, and maintains stable performance on different drug response data sets, providing a solution with both flexibility and reliability for the non-equilibrium classification task of drug response data. Description of the drawings

[0067] Figure 1 It is a schematic modular structure diagram of the R-SIOIC framework of the present invention. Detailed implementation manners

[0068] The specific implementation manners and experimental results of the present invention will be further described below in conjunction with the accompanying drawings and technical solutions.

[0069] The operation process of the R-SIOIC method consists of four major modules: initialization, verification, swarm intelligence optimization (SIO), and output. First, n initial candidate sets of drug response data are generated from the original unbalanced drug response dataset. Each candidate set of drug response data constructs a balanced subset by equally selecting various cell lines, and the cell line selection status is marked with binary encoding. Subsequently, an independent classifier is trained for each candidate set. The verification module evaluates the performance of the classifier and screens the optimal candidate set as the optimization guiding population. The SIO module iteratively updates the position of each candidate set of drug response data by calculating the Hamming distance and applying the continuous SIO algorithm to form a new candidate set population for the next round of optimization. When the maximum number of iterations is reached, the output module selects the optimal resampled dataset of drug response data obtained by the model to train the final classifier, and uses the test set to test the final classifier as the output result of the framework. The whole process realizes the collaborative optimization of data resampling and classifier performance.

[0070] As Figure 1 shown, the R-SIOIC framework includes an initialization module, a verification module, an SIO swarm intelligence optimization module, and an output module. This framework takes the unbalanced drug response dataset as the input. The initialization module generates a population of balanced candidate sets with binary encoding. The verification module evaluates and guides the optimization direction through the AUC metric. The SIO module relies on the continuous optimization algorithm to update the population of candidate sets of drug response data. Finally, the output module extracts the optimal resampled dataset of drug response data to train the classifier. Each module forms a closed-loop link of "generation → evaluation → optimization → output". Among them, the SIO module supports flexible replacement of different continuous optimization algorithms (such as particle swarm, grey wolf, etc.), and realizes efficient search of the solution space through the Hamming distance metric and position update formula.

[0071] In terms of experimental settings, seven drug response multi-omics datasets from the GDSC database were selected as the training set. The multi-omics data includes copy number variation data, gene expression data, and somatic mutation data. The number of cell lines in the dataset ranges from 362 to 856, and the imbalance rate ranges from 4.64 to 14.61, all showing obvious imbalance. Seven drug response multi-omics datasets corresponding to the drugs selected in the training set from the PDX and TCGA databases were selected as the test set. The classifier uses a linear kernel support vector machine (SVM), and the regularization parameter C is fixed at 5. The population size of the candidate sets of drug response data is uniformly set to 20 to balance the calculation speed and model accuracy.

[0072] Regarding the hyperparameter settings of the three algorithms (GWO, WOA, PSO) in the SIO module, a dynamic adjustment strategy is adopted to achieve a balance between exploration and exploitation, and all hyperparameters change linearly with the number of iterations. Taking the parameters in GWO and WOA as an example, their value ranges gradually shrink from [0, 2] to [0, 0] during the iteration process. Specifically, in the initial stage of iteration (the first 50% stage), the upper bound of is linearly reduced from 2 to 1. At this time, the method maintains a balance between global exploration (the candidate set of drug response data deviates from the optimal candidate set with a 50% probability) and local search (approaches the optimal candidate set with a 50% probability); when the iteration exceeds 50%, the upper bound of is further reduced from 1 to 0, forcing the method to focus on local search around the optimal candidate set. For the PSO algorithm, as the number of iteration rounds increases, three key parameters (inertia weight ω, individual cognitive coefficient c1, social cognitive coefficient c2) change collaboratively: ω gradually decreases from the initial 0.9 to 0.4, reducing the influence of the inertial velocity on the current update; c1 decreases from 2.5 to 0.5, weakening the dependence of each candidate set of drug response data on its own historical optimal position; c2 increases from 0.5 to 2.5, enhancing the guiding role of the optimal candidate set of drug response data. This parameter configuration makes each candidate set of drug response data tend to global exploration in the initial stage of iteration and then turn to local search around the optimal candidate set of drug response data in the later stage.

[0073] Since the position vectors of the candidate sets of drug response data in R - SIOIC are represented by binary vectors, the Hamming distance is used to measure the differences between candidate sets. In the position update of GWO and WOA, the distance (Hamming distance value) between the current candidate set of drug response data and the optimal candidate set of drug response data determines the number of binary digits to be flipped. For example, if the distance is 3, then 3 bits are randomly selected from the position vector of the optimal candidate set for value flipping (0 → 1 or 1 → 0). For the PSO algorithm, the position update requires special processing: first, the historical optimal position of each candidate set of drug response data and the historical optimal position of the population need to be recorded; second, during the update process, the binary difference bits between the current candidate set position and the optimal position need to be compared, and the number of flipped bits is determined according to the velocity v of the current candidate set.

[0074] In terms of evaluation metrics, the AUC metric is used to evaluate the discrimination ability of the model for different classes. This metric can effectively reflect the comprehensive performance of the model in unbalanced drug response data and avoid misleading results caused by traditional metrics such as accuracy due to skewed class distributions.

[0075] Table 1 shows the output results of R - SIOIC under the AUC metric

[0076]

[0077]

[0078] Method / Data

[0079] P.Paclitaxel P.Gemcitabine P.Cetuximab P.Erlotinib T.DocetaxelT.Cisplatin T.Gemcitabine Set

[0080]

[0081] Table 1 and Table 2 show the results of the embodiments of the present invention and the comparative experiment results. Table 1 shows the R-SIOIC output results under the AUC index; Table 2 shows the comparative test results under the AUC index. They are all for verifying the effectiveness of the R-SIOIC framework. We re-trained the drug response data classification method MOLI in the field and used the publicly available code to train on the same seven GDSC database drug response multi-omics datasets and test on the same seven PDX and TCGA database drug response multi-omics datasets.

Claims

1. A non-equilibrium drug response data resampling method based on swarm intelligence optimization, characterized in that: The specific steps include: Step 1: Resample to generate candidate sets of drug response data and train the base classifier The original drug response non-equilibrium dataset includes two types of cell lines, one is the cell lines that can respond to drugs, and the other is the cell lines that cannot respond to drugs; During resampling, n initial candidate drug response data sets are generated from the original unbalanced drug response data set. Each candidate drug response data set is a balanced subset consisting of an equal number of cell lines extracted from each class. The specific number of extractions from each class is determined by the minimum number of cell lines in each class, min_class_amount, and the formula is as follows: num=sample_size*min_class_amount Among them, sample_size indicates the proportion of cell lines selected from the smallest class. The specific value needs to ensure the number of training data while retaining a certain degree of randomness; After resampling, the original drug response imbalanced dataset is encoded into a binary vector X = [x1, x2, ..., x i ,...,x n ], where x i ∈{0,1} represents the cell line selection status during resampling, "1" means the cell line is selected, and "0" means ignored; the encoded binary vector is used as the position vector of each candidate set of drug response data; Each drug response data candidate set corresponds to a base classifier, and the base classifier is trained using the drug response data candidate set; Step 2: Evaluate the classification performance of the candidate set of drug response data The classification performance of all drug reaction data candidate sets is evaluated by the AUC indicator; the top m candidate sets with the best performance in the current iteration are selected to guide the optimization process. If the performance of a candidate set exceeds the current top m candidate sets in the subsequent iterations, the optimal record is updated; the position of the unselected candidate sets is updated in step 3; Step 3: Iteratively update the candidate set position of drug response data This step consists of two phases: In the first stage, the Hamming distance between candidate drug response data sets is calculated to measure the differences between different resampling results; In the second stage, the position of the candidate set of drug reaction data is updated according to the selected swarm intelligence algorithm, and its position update formula is defined as: X new =Update(X current ,X top-m ,D Hamming ) Where X top-m is the current optimal m drug response data candidate set, D Hamming is the Hamming distance, X current represents the position of the current candidate set, and Update(·) represents the position update function; Step 4: Generate the optimal resampled data set for drug response data When the preset maximum number of iterations is reached, the optimization process is terminated, and the candidate set of historical optimal drug response data is selected as the optimal resampling data set output for drug response data.

2. The method for resampling non-equilibrium drug response data based on swarm intelligence optimization according to claim 1, characterized in that: The process of updating the positions of the candidate drug response data sets by different swarm intelligence algorithms in the second stage is as follows: One of three continuous SIO methods is used as the test algorithm: Grey Wolf Optimization GWO, Particle Swarm Optimization PSO, and Whale Optimization WOA. Their specific update methods are as follows: (1) The gray wolf optimization GWO method simulates the social hierarchy and collaborative hunting behavior of gray wolf groups, and guides the drug response data candidate cluster to update its position through the three current optimal candidate sets α, β, and δ. The specific implementation includes three stages of operation: In the first stage, the distance between the current candidate set and the optimal candidate set is calculated using the Hamming distance: in: Indicates the position of the current candidate set, represents the positions of the three current optimal candidate sets, Hamming(·) represents the Hamming distance, Represents the distance between the current candidate set and the three best candidate sets; In the second stage, candidate positions are generated based on the distance between the current candidate set and each optimal candidate set: in: It is the candidate position updated according to the positions of the three current optimal candidate sets; parameter is the parameter of the model controlling the search mode of the optimal drug response data candidate set. Perform a global search when the parameter Perform local search when ; weighted average the candidate positions to obtain the next iteration position of the current drug reaction data candidate set: (2) The whale-optimized WOA method is based on the spiral bubble net feeding strategy of the humpback whale, which includes a special spiral approximation behavior pattern of the optimal drug response resampling dataset: Where b is the spiral shape coefficient; l∈[-1,1] controls the spiral step size, represents the distance between the current candidate set and the optimal candidate set, represents the distance between the current candidate set and the random candidate set, represents the optimal candidate set, represents a random candidate set; the update formula is based on the probability p and parameter (same as the parameters in GWO Same definition) switch different behavior modes: 1) When p>0.5, the candidate set of drug response data approaches the possible optimal drug response resampled data set in a spiral path; 2) When p≤0.5 and When , the candidate set of drug response data will be optimized towards the direction of the current optimal candidate set, and the goal is to conduct a detailed search near the current optimal candidate set to find the optimal drug response resampling data set; 3) When p≤0.5 and When , the drug reaction data candidate set will randomly select a candidate set in the population as a reference and perform a random search by updating its own position; (3) The particle swarm optimization (PSO) method simulates the foraging behavior of bird flocks. In this algorithm, the position of each candidate set of drug response data represents a candidate solution in the continuous solution space. Each candidate set of drug response data updates its own position through speed, and the speed update is affected by both the historical optimal position of each candidate set of drug response data and the historical optimal position of the candidate set of drug response data. The formula for updating the drug reaction data candidate set speed is as follows: Among them, r1 and r2 are random numbers in the interval [0,1], is the iteration speed of the previous round of drug reaction data candidate set, ω, c1, c2 are all preset hyperparameters, D7, D8 represent the distance between the current drug reaction data candidate set and its own historical optimal position and the group historical optimal position respectively: Among them, x pbest Indicates the best historical position of the current candidate set, Indicates the current iteration position of the current candidate set, x gbest Represents the total historical optimal position in all candidate sets.

3. The method for resampling non-equilibrium drug response data based on swarm intelligence optimization according to claim 1, characterized in that: The optimal resampling dataset of drug response data obtained in step 4 is verified to be effective in the following way: the classification performance of the classifier trained by the optimal resampling dataset of drug response data is evaluated using a test set, and compared with the drug response data classification method MOLI.

4. A non-equilibrium drug reaction data resampling device based on swarm intelligence optimization, characterized in that: Used to implement the method for resampling unbalanced drug reaction data based on swarm intelligence optimization as described in any one of claims 1 to 3, the unbalanced drug reaction data resampling device comprises: Initialization module: used to resample and generate candidate sets of drug response data; Base classifier: Use resampling to generate a candidate set of drug response data to train the base classifier; Validation module: evaluates the classification performance of the classifier trained by each drug response data candidate set; Swarm intelligence optimization module: dynamically adjusts the sample selection strategy and iteratively updates the position of each candidate set through the position update formula; Output module: Output the globally optimal drug response resampling data set.

5. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the non-equilibrium drug response data resampling method based on swarm intelligence optimization as described in any one of claims 1-3.

6. A server, characterized in that: The server includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement a non-equilibrium drug response data resampling method based on swarm intelligence optimization as described in any one of claims 1-3.