Artificial intelligence screening method for multi-mode carrier-free nano-drug combination

By using multimodal data fusion and machine learning algorithms to screen carrier-free nanodrug combinations, the systematic and precision issues in the discovery process of carrier-free nanodrug combinations have been resolved, enabling efficient drug combination screening and clear research and development directions.

CN120877945APending Publication Date: 2025-10-31KUNMING MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511051119.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The discovery process of carrier-free nanomedicines lacks systematicness and precision, resulting in unclear research directions and low efficiency.

Method used

An artificial intelligence screening method based on multimodal data fusion is adopted. By constructing an efficient artificial intelligence model, machine learning algorithms are used to extract and combine features from multimodal carrier-free nanomedicine data to screen out drug combinations with synergistic or combined therapeutic effects.

Benefits of technology

It significantly improves the screening efficiency of carrier-free nanomedicine combinations, reduces randomness, clarifies the research and development direction, and promotes the development of this field towards high efficiency and precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877945A_ABST
    Figure CN120877945A_ABST
Patent Text Reader

Abstract

The invention relates to the field of drug screening, and discloses an artificial intelligence screening method of a multi-mode carrier-free nano-drug combination, and the method comprises the following steps: obtaining data information of multi-mode carrier-free nano-drugs; performing feature extraction on different modal data by adopting a feature extraction algorithm to obtain extracted features; combining the extracted features to obtain a comprehensive feature vector; adopting a machine learning algorithm to judge the probability of the drug combination in the comprehensive feature vector, and outputting a screening result; through fusion and analysis of multi-modal data, the method not only significantly improves the screening efficiency of the carrier-free nano-drug combination and reduces the contingency in the drug combination discovery process, but also opens up a new thought and method for research and application of carrier-free nano-drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug screening, and more particularly to an artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies. Background Technology

[0002] In recent years, the rapid development of nanotechnology has driven the rise of carrier-free nanomedicines, an emerging field. Carrier-free nanomedicines eliminate the dependence on complex chemical processes and functional carrier molecules found in traditional nanomedicines, forming stable nanostructures solely through intermolecular forces between drug molecules. This unique drug form not only exhibits synergistic or combined therapeutic effects but also demonstrates strong potential for translational and clinical applications. However, the current drug combination discovery process for carrier-free nanomedicines suffers from significant limitations, often relying on chance and lacking systematicity and efficiency, which greatly restricts its widespread research and application. Therefore, developing an efficient and systematic method for screening carrier-free nanomedicine combinations is particularly urgent. Currently, the development of carrier-free nanomedicines lacks precise drug compatibility screening methods, resulting in unclear research directions and low efficiency. Summary of the Invention

[0003] The purpose of this invention is to propose an artificial intelligence screening method for multimodal carrier-free nanomedicine combinations, which solves the technical problem that the development of carrier-free nanomedicines lacks precise drug compatibility screening methods, resulting in unclear research directions and low efficiency.

[0004] This invention proposes an artificial intelligence screening method based on multimodal data fusion. By constructing an efficient artificial intelligence model, it can accurately screen carrier-free nanomedicine combinations with synergistic or combined therapeutic effects, thereby clarifying the research and development direction, significantly improving the development efficiency of carrier-free nanomedicines, and promoting the development of this field towards efficiency and precision.

[0005] Specifically, this invention provides an artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies, comprising the following steps: S1. Obtain data information on multimodal carrier-free nanomedicines; S2. Feature extraction algorithms are used to extract features from different modal data to obtain the extracted features; S3. Combine the extracted features to obtain a comprehensive feature vector; S4. Use machine learning algorithms to determine the probability of drug combinations in the comprehensive feature vector and output the screening results.

[0006] A storage device that stores instructions and data for implementing an artificial intelligence screening method for multimodal carrier-free nanodrug assemblies.

[0007] An artificial intelligence screening device for multimodal carrier-free nanomedicine combinations includes: a processor and a storage device; the processor loads and executes instructions and data in the storage device to implement an artificial intelligence screening method for multimodal carrier-free nanomedicine combinations.

[0008] The beneficial effects provided by this invention are as follows: The artificial intelligence screening method for multimodal carrier-free nanomedicine combinations proposed in this invention fully utilizes machine learning large-scale model technology, combines massive amounts of drug molecule data, and employs efficient algorithms and model training methods to rapidly screen drug combinations with potential synergistic or combined therapeutic effects, and constructs corresponding carrier-free nanomedicines. Through the fusion and analysis of multimodal data, this method not only significantly improves the screening efficiency of carrier-free nanomedicine combinations and reduces the randomness in the drug combination discovery process, but also opens up new ideas and methods for the research and application of carrier-free nanomedicines. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies of the present invention; Figure 2 This is the feature extraction flowchart for Mode 1; Figure 3 This is a flowchart of feature extraction for mode 2; Figure 4 This is a flowchart of feature extraction for mode 3; Figure 5 This is a flowchart of feature extraction for modes 4 and 5; Figure 6 This is a schematic diagram illustrating the performance evaluation of the machine learning model obtained from the example; Figure 7 This is a schematic diagram of the hardware device used in this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0011] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.

[0012] Please refer to Figure 1 The present invention provides an artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies, comprising the following steps: S1. Obtain data information on multimodal carrier-free nanomedicines; It should be noted that this step involves five different input modes, each targeting different properties of the nanomedicine: Mode 1: Physicochemical parameters, including the physical and chemical properties of the drug.

[0013] Modality 2: SMILES sequences, an ASCII string used to represent molecular structures.

[0014] Modality 3: Molecular fingerprint, a digital code used to describe molecular structure.

[0015] Modal 4: 2D structure diagram, a two-dimensional structural representation of drug molecules.

[0016] Modal 5: 3D structure diagram, a three-dimensional structural representation of drug molecules.

[0017] S2. Feature extraction algorithms are used to extract features from different modal data to obtain the extracted features; It should be noted that different artificial intelligence algorithms are used for feature extraction for each input modality: In this invention, step S2 is specifically as follows: S21. Construct the Algorithm-Modal Matching Network (AMMNet) for the adaptation and evaluation of modal-feature extraction algorithms. It should be noted that this model is a virtual model and does not have a corresponding structure. For the sake of convenience, this invention refers to it as the Modality-Feature Extraction Algorithm Adaptation Evaluation Model AMMNet. Its core processing procedure is in step S22. This invention refers to the detailed processing procedure of step S22 as "Modality-Feature Extraction Algorithm Adaptation Evaluation Model AMMNet".

[0018] S22. Input the modal data into the modality-feature extraction algorithm adaptation evaluation model AMMnet to obtain the feature extraction algorithm for the corresponding modality data; It should be noted that step S22 is as follows: S221. The modality-feature extraction algorithm adaptation evaluation model AMMNet classifies modality data into parametric data and structural data. Parametric data includes: physicochemical parameters, SMILES sequences, and molecular fingerprints. Structural data includes: 2D structural diagrams and 3D structural diagrams. Specifically, for the input modal data, there is first a data type identification process. For numerical data, which is collectively referred to as parameter data in this invention, it may include the aforementioned physicochemical parameters, SMILES sequences, and molecular fingerprints; while for structural data, it includes corresponding structural diagrams, which may include the aforementioned 2D structural diagrams and 3D structural diagrams; its classification can be detected and judged by data type, suffix, or format.

[0019] S222. For parameter data, it matches the feature extraction algorithm in the first preset algorithm library; for structure data, it matches the feature extraction algorithm in the second preset algorithm library. It should be noted that the first preset algorithm library includes: Principal Component Analysis (PCA), Natural Language Processing (NLP), and Convolutional Neural Network (CNN); the second preset algorithm library includes: Graph Convolutional Network (GCN) and Equivariant Graph Neural Network (EGNN).

[0020] S223. Different categories of modal data are used to calculate similarity scores in the corresponding preset algorithm library, and the feature extraction algorithm with the highest score is selected as the final algorithm.

[0021] It should be noted that the similarity score calculation formula in step S223 is as follows: Score A i, D j )= α ⋅CompScore+ β ⋅TimeScore+ γ ⋅AccScore in A i Represents modal data; D j Indicates the feature extraction algorithm, α, β, γ Here, is the weight parameter; CompScore represents the degree of matching between data complexity and algorithm processing capability; TimeScore represents the execution efficiency of the algorithm on historically similar data; and AccScore represents the feature discrimination of the algorithm on the validation set.

[0022] For example, if a physicochemical parameter mode in the existing parameter data needs to find the most suitable feature extraction algorithm from the first preset algorithm library, then the mode needs to be matched with all the algorithms in the first preset algorithm library, their similarity is calculated, and finally the feature extraction algorithm with the highest similarity score is selected as the final algorithm for the mode.

[0023] The similarity score calculation primarily considers three dimensions. The first dimension is CompScore, which represents the degree of matching between data complexity and algorithm processing capability. It characterizes the matching degree between the complexity of the physicochemical parameters and the processing capability of the current matching algorithm. For example, if the physicochemical parameter data complexity is high, but the processing capability of natural language processing and convolutional neural network algorithms is relatively poor, then their matching degree is not high, and the corresponding score for this item will be low. This data is manually set through prior knowledge or testing. The second dimension is TimeScore, which represents the algorithm's execution efficiency on historical similar data. This data is obtained by examining historical data. The third dimension is AccScore, which represents the algorithm's feature discrimination on the validation set. This data is also obtained by examining historical data.

[0024] The above examples are based solely on physicochemical data. The calculation process for other modal data is similar and will not be elaborated upon here.

[0025] S23. Use the feature extraction algorithm of the corresponding modal data to extract features from the modal data to obtain the extracted features.

[0026] As one embodiment, the final feature processing in this invention is as follows: Principal component analysis (PCA): used to extract key features from physicochemical parameters and generate feature combinations F1.

[0027] Natural Language Processing (NLP): Used to extract features from SMILES sequences and generate feature combinations F2.

[0028] Convolutional Neural Network (CNN): Used to extract features from molecular fingerprints and generate feature combinations F3.

[0029] Graph Convolutional Networks (GCNs): Used to extract features from 2D structure graphs and generate feature combinations F4.

[0030] Equivariant Graph Neural Network (EGNN): Used to extract features from 3D structure graphs and generate feature combinations F5.

[0031] Of course, the above are only embodiments provided by the present invention. In other embodiments, or for other modal data, the method of the present invention can also be used. The present invention is not intended to limit the above.

[0032] To better illustrate the results of the present invention, please refer to the following embodiment: Figure 2 , Figure 2 This is a flowchart of Principal Component Analysis (PCA) extracting key features from physicochemical parameters and generating feature combinations F1, as shown below: (1) Data acquisition Eleven physicochemical parameters of carrier-free nanomedicines were automatically extracted from the DrugBank database, including: melting point, water solubility, logP, logS, pKa, charge, number of hydrogen bond donor groups, number of hydrogen bond acceptor groups, polarizable area, polarizability, and number of ring structures.

[0033] (2) Data preprocessing In the data preprocessing stage, we first meticulously handled missing and outlier values ​​in the physicochemical parameters. This step ensured the integrity and accuracy of the data. Subsequently, we performed min-max normalization on each column of data to ensure that all data were compared on the same scale, thereby improving the consistency and reliability of the analysis.

[0034] (3) Principal component analysis (PCA) In the principal component analysis (PCA) phase, our goal was to screen for physicochemical parameters that contributed more than 90% to the drug combination. Through the above analysis, a key feature vector F1 was extracted from 11 physicochemical parameters. This vector concentrates the most significant variation information in the data, providing important reference for subsequent drug research.

[0035] As one example, please refer to Figure 3 , Figure 3 This paper demonstrates an N-Grams algorithm based on Natural Language Processing (NLP) technology for extracting feature vectors F2 from SMILES sequences. The specific steps are as follows: (1) Data acquisition The SMILES sequences of carrier-free nanomedicines were retrieved from the PubChem database. For example, the SMILES sequence of aspirin is: CC(=O)OC1=CC=CC=C1C(=O)O.

[0036] (2) Data preprocessing The acquired SMILES sequences are preprocessed to remove spaces, newlines, and other invalid characters. The preprocessed strings are then uniformly converted to uppercase to ensure data consistency and standardization. For example, the preprocessed aspirin SMILES sequence is: CC(=O)OC1=CC=CC=C1C(=O)O.

[0037] (3) N-Grams feature extraction With N set to 1, a sliding window technique is used to process the preprocessed SMILES sequence, extracting subsequences of length 1 in sequence. The occurrence frequency of each subsequence of length 1 in the SMILES sequence is counted, and a feature vector F2 is constructed based on this statistical result.

[0038] For example, the eigenvector F2 of the aspirin SMILES sequence is: [9, 2, 4, 4, 2, 2] corresponding to [C, (, =,O, ), 1].

[0039] The above methods can effectively extract the characteristic information of carrier-free nanomedicine SMILES sequences, providing reliable data support for subsequent drug analysis, classification and prediction studies.

[0040] As one example, please refer to Figure 4 , Figure 4 The process of extracting the molecular fingerprint feature vector F3 based on a convolutional neural network (CNN) is demonstrated, and the specific steps are as follows: (1) Data acquisition SMILES sequences of carrier-free nanomedicines were retrieved from the PubChem database.

[0041] (2) Data preprocessing: The acquired SMILES sequences were parsed using the RDKit chemical toolkit and converted into molecular objects to achieve an accurate representation of the molecular structure. Based on these molecular objects, their PubChem fingerprints were calculated. The calculated PubChem fingerprints were further converted into binary vector form to facilitate adaptation to convolutional neural network models and provide standardized input data for the feature extraction process.

[0042] (3) Convolutional Neural Network (CNN) processing The binary vector is flattened to convert it into a format suitable for convolutional neural network (CNN) processing. The flattened data is then input into the CNN model, passing through convolutional layers, dropout layers, and fully connected layers. The convolutional layers extract local features, the dropout layers prevent overfitting, and the fully connected layers integrate the extracted features. Finally, the feature vector F3 is extracted from the CNN model.

[0043] As one example, please refer to Figure 5 , Figure 5 Image (a) illustrates the process of extracting 2D structure graph feature vectors F4 based on a graph convolutional network (GCN). The specific steps are as follows: (1) Data acquisition Two-dimensional (2D) structural diagrams of carrier-free nanomedicines were retrieved from the PubChem database.

[0044] (2) Graph Convolutional Network (GCN) Processing Graph convolution operation (GCNConv): Graph convolution layers (GCNConv) are used to process 2D molecular graphs and extract feature information from molecular nodes. Graph convolution updates the features of the current node by aggregating information from neighboring nodes, effectively capturing local features in the molecular structure. The specific formula is as follows:

[0045] in This represents the feature matrix of the nodes in the l-th layer. The weight matrix is ​​a learnable matrix.

[0046] Add the identity matrix I to the adjacency matrix A (to introduce self-connection). for The degree matrix, This represents the activation function.

[0047] Dropout prevents overfitting: After graph convolution operations, a Dropout layer is introduced to randomly discard feature information from some nodes to prevent the model from overfitting and enhance the model's generalization ability.

[0048] ReLU Nonlinear Transformation: The extracted features are nonlinearly transformed using the ReLU (Rectified Linear Unit) activation function, introducing nonlinear factors that enable the model to learn more complex feature representations.

[0049] Deep feature extraction: The graph convolutional layer (GCNConv) is used again to further extract the features after Dropout and ReLU processing to obtain the deep features of the 2D molecule.

[0050] Further feature optimization: The extracted features are further optimized by performing Dropout and ReLU operations again, enhancing their robustness and discriminative power.

[0051] (3) Output layer processing Linear Transformation: Features processed by the graph convolutional network are input into a linear layer, where a linear transformation is used to integrate the feature information. The specific formula is as follows: +

[0052] in, and These are the weight matrix and bias vector of the linear layer, respectively. This is the feature matrix output by the graph convolutional network.

[0053] ReLU integration: The features after linear transformation are further processed by the ReLU activation function and finally integrated into the feature vector F4.

[0054] Figure 5 (b) Demonstrates the process of extracting 3D structure graph feature vectors F54 based on Equivariant Graph Neural Network (EGNN). The specific steps are as follows: (1) Data acquisition 3D structural images of carrier-free nanomedicines were retrieved from the PubChem database.

[0055] (2) Processing of Equivariant Graph Neural Networks (EGNN) Equivariant Graph Convolution Layer (EGCL): This layer processes 3D molecular graphs to extract the initial geometric and topological features of molecular nodes. EGCL operations simultaneously consider the spatial geometry and topological structure of molecules, effectively capturing their initial features. Its core formula is as follows:

[0056]

[0057] in and Let N(i) represent the features and coordinates of node i at level l, respectively, and let N(i) represent the set of neighboring nodes of node i. and These are functions for feature update and coordinate update, respectively.

[0058] Dropout prevents overfitting: After the convolution operation on the isomorphic graph, a Dropout layer is introduced to randomly discard feature information of some nodes to prevent the model from overfitting and enhance the model's generalization ability.

[0059] Deep feature extraction: The features processed by Dropout are further extracted using an equivariant graph convolutional layer (EGCL) to obtain deeper features inside the 3D molecule.

[0060] Dropout again to prevent overfitting: After deep feature extraction, a Dropout layer is introduced again to further prevent the model from overfitting and ensure the robustness of the features.

[0061] Pooling Dimensionality Reduction and Aggregation: The Pooling layer performs dimensionality reduction and aggregation operations on the extracted features, integrating high-dimensional node feature information into a low-dimensional global feature representation, providing a foundation for subsequent feature integration.

[0062] (3) Output layer processing Linear Transformation: The global features processed by Pooling are input into a linear layer, and the feature information is further integrated through linear transformation.

[0063] ReLU integration: The features after linear transformation are processed nonlinearly using the ReLU activation function and finally integrated into a feature vector F5.

[0064] S3. Combine the extracted features to obtain a comprehensive feature vector; It should be noted that the features extracted from different modalities, F1 to F5, will be integrated to form a comprehensive feature vector.

[0065] S4. Use machine learning algorithms to determine the probability of drug combinations in the comprehensive feature vector and output the screening results.

[0066] It should be noted that in step S4, based on the comprehensive feature vector, support vector machine, K-nearest neighbor and logistic regression algorithms are used respectively to judge the probability of drug combination, thereby determining the feasibility of drug combination. According to the judgment results of the above algorithms, the conclusion of whether the two drugs can be combined is output.

[0067] As one example, three advanced machine learning techniques are used to conduct in-depth analysis of the comprehensive feature set extracted from multimodal data in order to identify and screen drug combinations with potential therapeutic effects.

[0068] Support Vector Machine (SVM): SVM achieves efficient classification by finding the optimal decision boundary to maximize the margin between different classes. In this process, SVM is used to evaluate the classification potential of different drug combinations, helping us identify the combinations most likely to succeed.

[0069] K-Nearest Neighbors (KNN) algorithm: Predicts the class of a new data point by calculating the distance between the new data point and each point in the training dataset. The simplicity and intuitiveness of KNN make it a powerful tool for exploring the properties of drug combinations, especially when dealing with non-linearly separable data.

[0070] Logistic Regression (LR): LR can assess the probability of success of different drug combinations, thereby providing support for decision-making.

[0071] By comprehensively utilizing these three algorithms, we can evaluate the potential effects of drug combinations from multiple perspectives, ensuring that the selected drug combinations are not only theoretically feasible but also highly efficient and reliable in practical applications. This multi-algorithm fusion strategy significantly improves the accuracy and efficiency of drug combination screening.

[0072] Furthermore, the specific judgment process for step S4 is as follows: A voting mechanism is used for judgment. If at least two of the outputs of the support vector machine, K-nearest neighbor, and logistic regression algorithms can be combined, then the combination screening result is considered to be a combination.

[0073] In another embodiment, the specific judgment process of step S4 is as follows: The average mechanism is used to determine the combination probability. The average of the combination probabilities output by the support vector machine, K nearest neighbor, and logistic regression algorithms is taken and compared with the preset value. If the result exceeds the preset value, the combination screening result is considered acceptable.

[0074] The present invention provides a comprehensive embodiment as follows: (1) Data acquisition and feature extraction We collected publicly available physicochemical information on existing carrier-free nanomedicines and NSAID / chemotherapeutic drug combinations from the DrugBank public database, including: melting point, water solubility, logP, logS, pKa, charge, number of hydrogen bond donor groups, number of hydrogen bond acceptor groups, polarizable area, polarizability, and number of ring structures. We also collected publicly available SMILES sequences, 2D structure diagrams, and 3D structure diagrams of existing carrier-free nanomedicines and NSAID / chemotherapeutic drug combinations from the PubChem public database, where molecular fingerprints were obtained by processing SMILES sequence data.

[0075] (2) Screening of drug combinations ① Randomly divide the drug combinations into training and test sets in a ratio of 35:52. Input the drug combination features (F1+F2+F3+F4+F5) from the training set into the Support Vector Machine (SVM) module of the Python software for training. ② Randomly divide the drug combinations into training and test sets in a ratio of 35:52. Input the drug combination features (F1+F2+F3+F4+F5) from the training set into the K-Nearest Neighbors (KNN) module of the Python software for training. ③ Randomly divide the drug combinations into training and test sets in a ratio of 35:52. Input the drug combination features (F1+F2+F3+F4+F5) from the training set into the logistic regression module of the Python software for training.

[0076] (3) Performance evaluation of results The physicochemical properties of the drugs in the test set are input into three vector machines after training to predict whether nanomedicines will form. The prediction results are compared with the actual preparation results to form a confusion matrix, and the accuracy, F-score, and recall are calculated to comprehensively evaluate the model.

[0077] Please refer to Figure 6 , Figure 6 This is a schematic diagram illustrating the performance evaluation of the machine learning model obtained from the example.

[0078] Figure 6In the diagram, (a) is the confusion matrix of the three models; (b) is the evaluation of the accuracy, recall, and F1-score results of the three machine learning models; (c) is the learning curve of the three machine learning models; and (d) is the receiver operating characteristic curve of the three models. Figure 6 As can be seen, all three machine learning models can be used to predict drug combinations.

[0079] Therefore, based on the results of the above analysis method, the final combination screening results in this application are divided into two categories: Combinable: This means that these drug combinations are theoretically feasible.

[0080] Cannot be combined: This means that these drug combinations are theoretically not feasible.

[0081] Please see Figure 7 , Figure 7 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: an artificial intelligence screening device 401 for multimodal carrier-free nanodrug combinations, a processor 402, and a storage device 403.

[0082] An artificial intelligence screening device 401 for multimodal carrier-free nanomedicine assemblies: The artificial intelligence screening device 401 for multimodal carrier-free nanomedicine assemblies implements the artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies.

[0083] Processor 402: The processor 402 loads and executes the instructions and data in the storage device 403 to implement the artificial intelligence screening method for multimodal carrier-free nanomedicine combinations.

[0084] Storage device 403: The storage device 403 stores instructions and data; the storage device 403 is used to implement the artificial intelligence screening method for multimodal carrier-free nanomedicine combinations.

[0085] The key point of this invention is: 1. Multimodal data fusion and targeted feature extraction This invention integrates five different modalities of carrier-free nanomedicine data (physicochemical parameters, SMILES sequences, molecular fingerprints, 2D structure maps, and 3D structure maps), and employs a specialized feature extraction algorithm for each modality to achieve comprehensive feature extraction of drug molecules. This multimodal data fusion strategy fully leverages the advantages of different modalities and avoids the information limitations of single-modal data. For example, SMILES sequences provide linear information about molecular structure, while 3D structure maps reflect the spatial geometric features of the molecule. In this way, this invention can more comprehensively capture the characteristics of drug molecules, providing a richer information foundation for subsequent drug combination screening.

[0086] 2. A combination of efficient feature extraction algorithms This invention employs advanced algorithms for feature extraction on different modalities of data, including Principal Component Analysis (PCA), Natural Language Processing (NLP), Convolutional Neural Networks (CNN), Graph Convolutional Networks (GCN), and Equivariant Graph Neural Networks (EGNN). The selection of these algorithms is based on a deep understanding of the characteristics of different modalities of data. For example, for structured data such as physicochemical parameters, PCA can effectively extract key features and reduce data dimensionality; while for textual data such as SMILES sequences, NLP technology can uncover the potential information within the sequence. This combination of algorithms not only improves the efficiency of feature extraction but also enhances the quality and representativeness of the features, providing more accurate feature input for subsequent drug combination screening.

[0087] 3. Construction of comprehensive feature vectors and screening of drug combinations This invention combines features F1, F2, F3, F4, and F5 extracted from different modalities to form a comprehensive feature vector. Then, it uses Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Logistic Regression (LR) algorithms to determine the probability of the drug combination, thereby assessing its feasibility. This multi-algorithm fusion strategy can evaluate the potential effects of drug combinations from multiple perspectives, avoiding the biases and limitations that may arise from a single algorithm. Compared to traditional drug combination screening methods that rely on chance and experience, this invention can systematically and efficiently screen carrier-free nanomedicine combinations with synergistic or combined therapeutic effects, significantly improving the accuracy and efficiency of screening.

[0088] 4. Systematic drug development process support This invention provides a systematic method for screening carrier-free nanomedicine combinations, forming a complete process from multimodal data acquisition, feature extraction, feature combination to drug combination screening. This method provides scientific basis and technical support for the research and development of carrier-free nanomedicines, clarifies the research direction, and significantly improves the development efficiency of carrier-free nanomedicines. Compared with traditional drug development methods, this invention can more quickly screen drug combinations with potential synergistic effects, reduce research and development time and costs, and promote the development of this field towards higher efficiency and precision.

[0089] 5. Optimization of real-time performance and data processing efficiency In the feature extraction and drug combination screening processes, this invention employs various optimization strategies. For example, a Dropout layer is introduced in the CNN to prevent overfitting; in the EGNN, a Pooling layer is used to reduce the dimensionality and aggregate features, further optimizing the feature representation. These optimization strategies not only improve the model's performance but also enhance its robustness and generalization ability, enabling it to better adapt to different types of carrier-free nanomedicine data. Furthermore, by rationally designing the data processing flow, this invention can improve data processing efficiency while ensuring data integrity and accuracy, meeting the real-time requirements of carrier-free nanomedicine development.

[0090] In summary, the beneficial effects of this invention are as follows: The artificial intelligence screening method for multimodal carrier-free nanomedicine combinations proposed in this invention fully utilizes machine learning large-scale model technology, combines massive amounts of drug molecule data, and employs efficient algorithms and model training methods to rapidly screen drug combinations with potential synergistic or combined therapeutic effects, and constructs corresponding carrier-free nanomedicines. Through the fusion and analysis of multimodal data, this method not only significantly improves the screening efficiency of carrier-free nanomedicine combinations and reduces the randomness in the drug combination discovery process, but also opens up new ideas and methods for the research and application of carrier-free nanomedicines.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies, characterized in that: The method includes the following steps: S1. Obtain data information on multimodal carrier-free nanomedicines; S2. Feature extraction algorithms are used to extract features from different modal data to obtain the extracted features; S3. Combine the extracted features to obtain a comprehensive feature vector; S4. Use machine learning algorithms to determine the probability of drug combinations in the comprehensive feature vector and output the screening results.

2. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 1, characterized in that: The data information of the multimodal carrier-free nanomedicine in step S1 includes: physicochemical parameters, SMILES sequence, molecular fingerprint, 2D structure diagram and 3D structure diagram.

3. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 2, characterized in that: Step S2 is as follows: S21. Construct the AMMNet model for adapting and evaluating modality-feature extraction algorithms; S22. Input the modal data into the modality-feature extraction algorithm adaptation evaluation model AMMnet to obtain the feature extraction algorithm for the corresponding modal data; S23. Use the feature extraction algorithm of the corresponding modal data to extract features from the modal data to obtain the extracted features.

4. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 3, characterized in that: Step S22 is as follows: S221. The modality-feature extraction algorithm adaptation evaluation model AMMNet classifies modality data into parametric data and structural data. Parametric data includes: physicochemical parameters, SMILES sequences, and molecular fingerprints. Structural data includes: 2D structural diagrams and 3D structural diagrams. S222. For parameter data, it matches the feature extraction algorithm in the first preset algorithm library; for structure data, it matches the feature extraction algorithm in the second preset algorithm library. S223. Different categories of modal data are used to calculate similarity scores in the corresponding preset algorithm library, and the feature extraction algorithm with the highest score is selected as the final algorithm.

5. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 4, characterized in that: The similarity score calculation formula in step S223 is as follows: Score( A i, D j )= α ⋅CompScore ij + β ⋅TimeScore ij + γ ⋅AccScore ij in A i Represents modal data; D j Indicates the feature extraction algorithm, i Index representing modal data, j An index representing the feature extraction algorithm; α, β, γ Here, is the weight parameter; CompScore represents the degree of matching between data complexity and algorithm processing capability; TimeScore represents the execution efficiency of the algorithm on historically similar data; and AccScore represents the feature discrimination of the algorithm on the validation set.

6. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 1, characterized in that: In step S4, based on the comprehensive feature vector, support vector machine, K-nearest neighbor and logistic regression algorithms are used respectively to judge the probability of drug combination, thereby determining the feasibility of drug combination. According to the judgment results of the above algorithms, the conclusion of whether the two drugs can be combined is output.

7. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 6, characterized in that: The specific judgment process for step S4 is as follows: A voting mechanism is used for judgment. If at least two of the outputs of the support vector machine, K-nearest neighbor, and logistic regression algorithms can be combined, then the combination screening result is considered to be a combination.

8. The artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in claim 6, characterized in that: The specific judgment process for step S4 is as follows: The average mechanism is used to determine the combination probability. The average of the combination probabilities output by the support vector machine, K nearest neighbor, and logistic regression algorithms is taken and compared with the preset value. If the result exceeds the preset value, the combination screening result is considered acceptable.

9. A storage device, characterized in that: The storage device stores instructions and data for implementing an artificial intelligence screening method for multimodal carrier-free nanomedicine combinations as described in any one of claims 1 to 8.

10. An artificial intelligence screening device for multimodal carrier-free nanomedicine assemblies, characterized in that: include: A processor and a storage device; the processor loads and executes instructions and data in the storage device to implement the artificial intelligence screening method for multimodal carrier-free nanomedicine assemblies as described in any one of claims 1 to 8.