Antiviral drug screening method and system based on high-dimensional features, and storage medium

By constructing an antiviral drug screening method with high-dimensional features, using adjacency matrix and graph convolutional network to extract high-order features and constructing a hypergraph structure, the problems of high time consumption and incomplete relationship capture of existing methods are solved, and efficient and accurate drug screening is achieved.

CN120510956BActive Publication Date: 2025-10-17GENERAL HOSPITAL OF PLA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511000958.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-17
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing antiviral drug screening methods rely on time-consuming and costly laboratory experiments, and existing computer drug repositioning methods are unable to fully capture the complex high-order relationships between drugs and viruses.

Method used

An antiviral drug screening method based on high-dimensional features is adopted. By constructing an adjacency matrix, calculating the similarity matrix, combining the autoencoder and graph convolutional network to extract high-order features, building a hypergraph structure and defining the objective function, solving the projection matrix, and calculating the prediction score matrix, the candidate drugs with the strongest potential correlation are screened out.

Benefits of technology

It improves the predictive accuracy and reliability of antiviral drug screening, reduces computational complexity, enhances the generalization ability of the model, and can more comprehensively capture the drug-virus relationship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510956B_ABST
    Figure CN120510956B_ABST
Patent Text Reader

Abstract

The application provides an antiviral drug screening method and system based on high-dimensional features and a storage medium, and belongs to the technical field of bioinformatics, computational biology and artificial intelligence. The method comprises the following steps: constructing an adjacency matrix; calculating a drug integration similarity matrix and a virus integration similarity matrix; based on an automatic encoder and a graph convolution network, combining the adjacency matrix, the virus integration similarity matrix and the drug integration similarity matrix, constructing a high-dimensional feature set; constructing a hypergraph with the high-dimensional feature set as a vertex set, and defining an objective function based on the hypergraph to obtain a projection matrix; calculating a prediction score matrix based on the projection matrix and the high-dimensional feature set; and screening the score of the row where the target virus is located based on the prediction score matrix to obtain a final prediction result after sorting. The application can more comprehensively depict the internal relationship between viruses and drugs by fusing multiple similarities, and can improve the prediction performance. The automatic encoder can learn the potential nonlinear relationship between samples, and can enhance the generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of antiviral drug screening, in particular to an antiviral drug screening method and system based on high-dimensional features and a storage medium. BACKGROUND

[0002] In order to cope with the outbreak of infectious diseases caused by various viruses, it is urgent to find effective antiviral drugs. At present, although some drugs have been approved, with the continuous emergence of new variants, the clinical effectiveness of existing treatment regimens is still controversial, and there is an urgent need to develop new and efficient antiviral drug screening methods.

[0003] Traditional antiviral drug development mainly relies on laboratory experiments, which are usually time-consuming and costly. For example, laboratory screening techniques require complex cell culture systems, biosafety facilities and a large amount of manpower and material resources. At the same time, computer-aided drug repositioning methods are gradually becoming an important complementary means for antiviral drug discovery due to their high efficiency and low cost. Studies have shown that computational methods can effectively reduce trial and error costs and accelerate the identification process of potential drug candidates.

[0004] Existing computer drug repositioning methods for viruses can be divided into two categories: one is network-driven methods that construct heterogeneous networks by analyzing the similarity and relevance between drug, virus, protein and other nodes for prediction; the other is molecular structure-based methods that analyze the structural matching between drug molecules and virus protein targets for screening. However, the above methods often only use single or limited similarity measures, making it difficult to fully capture the complex high-order relationships between drugs and viruses.

[0005] With the development of bioinformatics and artificial intelligence technologies, especially the application of advanced network analysis methods such as hypergraph learning, we are provided with new drug discovery ideas. Hypergraph structure can represent the relationship between three or more nodes, which is more suitable for capturing high-order associations in complex systems compared to traditional graph models. However, existing research (CN202310910294.0 Antiviral drug screening method, system and storage medium based on hypergraph learning) has not fully explored the application potential of hypergraph learning in antiviral drug screening.

[0006] In summary, developing an antiviral drug screening method that can integrate multi-source similarity information, capture high-order structural relationships and has high computational efficiency is of great significance for accelerating the antiviral drug discovery process. Based on this background, the present application proposes a hypergraph prediction model based on high-dimensional features, aiming to provide a new technical path for efficient antiviral drug screening. SUMMARY

[0007] The application provides an antiviral drug screening method and system based on high-dimensional features and a storage medium to solve one or more of the above problems.

[0008] To achieve the above object, the application adopts the following technical solutions:

[0009] The antiviral drug screening method based on high-dimensional features comprises the following steps:

[0010] An adjacency matrix is constructed based on known virus-drug association information;

[0011] Based on the adjacency matrix, Gaussian distance similarity and cosine similarity of the drugs are calculated to obtain a drug integrated similarity matrix, and Gaussian distance similarity and cosine similarity of the viruses are calculated to obtain a virus integrated similarity matrix;

[0012] Based on an autoencoder and a graph convolution network, the adjacency matrix, the virus integrated similarity matrix and the drug integrated similarity matrix are combined to construct a high-dimensional feature set;

[0013] The high-dimensional feature set is used as a vertex set to construct a hypergraph, and a target function is defined based on the hypergraph to obtain a projection matrix;

[0014] A prediction score matrix is calculated based on the projection matrix and the high-dimensional feature set;

[0015] Based on the prediction score matrix, the scores of the rows where the target viruses are located are screened, and the final prediction result is obtained after sorting.

[0016] In the specification, the construction step of the high-dimensional feature set comprises: setting the initial feature of the virus as the concatenation of the row vector of the adjacency matrix and the virus integrated similarity matrix, extracting high-order feature information through a two-layer graph convolution network, obtaining the latent variable distribution parameter through an encoder, and then obtaining the latent variable through reparameterization, wherein the latent variable constitutes the high-dimensional feature set.

[0017] In the specification, the high-order feature information of the virus node is extracted through the graph convolution network, and the specific process is as follows: in the graph convolution network, the propagation of the signal at the t-1 layer is to operate the initial feature matrix and the virus integrated similarity adjacency matrix and its degree matrix after adding a self-loop, and then the model parameter processing and linear rectifier function activation are performed to obtain the high-order feature information.

[0018] In the specification, the feature is gradually reduced in dimension through the encoder to obtain the distribution parameter of the latent variable, and the specific process is as follows: taking the initial feature matrix and the virus integrated similarity matrix as input, the parameters set in the encoder are operated to obtain the mean and standard deviation of the latent variable, which are two distribution parameters, and through the reparameterization method, the random variable subject to the standard normal distribution is multiplied by the standard deviation and then added to the mean to obtain the final latent variable.

[0019] In the specification, the decoder reconstructs the virus-drug network structure by inner product operation of latent variables, outputs the correlation probability of both, and adopts a double loss function during training: binary cross-entropy loss is used to measure the deviation of the predicted correlation probability from the actual label, and KL divergence loss is used to constrain the latent variable distribution to be close to the standard normal distribution.

[0020] In the specification, each sample in the high-dimensional feature set is taken as a vertex, and a hyperedge is generated according to the k-neighbor rule, each hyperedge connecting multiple similar vertices, and the weight of the hyperedge being calculated according to the vertex distance; a normalized hypergraph Laplacian matrix is constructed to depict the high-order correlation structure between samples.

[0021] In the specification, a target function is defined to balance the hypergraph structure reservation, data reconstruction error and model regularization constraint, in which the projection matrix is used to map the high-dimensional features to a low-dimensional space, and the regularization coefficient controls the weight of each part of the loss, and the optimal projection matrix is obtained by solving the closed-form solution.

[0022] In the specification, the high-dimensional feature set is mapped to a low-dimensional space by using the optimal projection matrix, and a predicted score matrix is calculated, each element of which representing the potential correlation strength between the corresponding virus and drug; for a target virus, its row vector in the score matrix is extracted and arranged in descending order of score, and the candidate drug with the strongest potential correlation is screened out, providing a ranking result for antiviral drug screening.

[0023] The antiviral drug screening system based on high-dimensional features applies the antiviral drug screening method based on high-dimensional features in any one of the above embodiments, and the antiviral drug screening system based on high-dimensional features comprises:

[0024] An adjacency matrix construction module is configured to construct an adjacency matrix based on known virus-drug correlation information;

[0025] An integrated similarity matrix calculation module is configured to calculate the Gaussian distance similarity and the cosine similarity of drugs based on the adjacency matrix to obtain an integrated similarity matrix of drugs, and calculate the Gaussian distance similarity and the cosine similarity of viruses to obtain an integrated similarity matrix of viruses;

[0026] A high-dimensional feature set construction module is configured to construct a high-dimensional feature set based on an autoencoder and a graph convolutional network in combination with the adjacency matrix, the integrated similarity matrix of viruses and the integrated similarity matrix of drugs;

[0027] A projection matrix solving module is configured to construct a hypergraph with the high-dimensional feature set as a vertex set, and solve a projection matrix based on a target function defined by the hypergraph;

[0028] A predicted score matrix calculation module is configured to calculate a predicted score matrix based on the projection matrix and the high-dimensional feature set;

[0029] A prediction module is configured to filter out scores of a row where the target virus is located based on the prediction score matrix, and sort the scores to obtain a final prediction result.

[0030] A computer-readable storage medium stores computer instructions, when a computer reads the computer instructions, the computer executes the anti-virus drug screening method based on high-dimensional features as described in any one of the above.

[0031] In summary, the present application has at least the following beneficial effects:

[0032] The present application constructs a drug and virus integrated similarity matrix by integrating the adjacency matrix, Gaussian distance similarity and cosine similarity, effectively fuses multiple biological information, captures the complex pattern of drug-virus interaction, and improves the prediction accuracy; adopts two-layer GCN combined with ReLU activation function to extract high-order structure features, and simultaneously uses an autoencoder to generate latent variables, which can automatically learn the nonlinear feature representation of drugs and viruses, and enhance the representation ability of the model for high-dimensional sparse data; constructs a hypergraph structure based on a high-dimensional feature set, connects multiple vertices through a hyperedge, more comprehensively captures the high-order association between data points, and more accurately models the drug-virus relationship compared with a traditional graph structure; defines a target function to solve a projection matrix, maps high-dimensional features to a low-dimensional space, effectively preserves the internal structure information of the data, reduces the computational complexity, and improves the generalization ability of the model; calculates the final score by combining the bidirectional prediction matrix from the virus perspective and the drug perspective, comprehensively considers the information of the two modalities, reduces the prediction bias, and improves the reliability of the anti-virus drug screening. BRIEF DESCRIPTION OF DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0034] Figure 1 A schematic diagram of the anti-virus drug screening method based on high-dimensional features involved in the present application.

[0035] Figure 2 A flowchart of the anti-virus drug screening method based on high-dimensional features involved in the present application.

[0036] Figure 3 A schematic diagram of the performance comparison results with the baseline method involved in the present application. DETAILED DESCRIPTION

[0037] In the following, certain example embodiments are simply described. As those skilled in the art will realize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the inventive embodiments. The drawings and description are therefore to be considered an exemplification of the application, and not a limitation thereof.

[0038] The following disclosure provides many different embodiments, or examples, for implementing different structures of the inventive embodiments. For the purpose of simplicity, the elements and settings of particular examples in the following are described. Of course, they are merely examples and are not intended to limit the inventive embodiments. Furthermore, the inventive embodiments can repeat reference numerals and / or letters in various examples, and this repetition is for the purpose of simplicity and clarity, and does not in itself dictate a relationship between the various embodiments and / or settings discussed.

[0039] The embodiments of the present application will be described in detail below with reference to the drawings.

[0040] As Figure 1 shown, the present embodiment provides an antiviral drug screening method based on high-dimensional features, comprising:

[0041] Constructing an adjacency matrix based on known virus-drug association information;

[0042] Based on the adjacency matrix, calculating the Gaussian distance similarity and cosine similarity of the drugs to obtain a drug integrated similarity matrix, and calculating the Gaussian distance similarity and cosine similarity of the viruses to obtain a virus integrated similarity matrix;

[0043] Based on the auto-encoder and the graph convolution network, combining the adjacency matrix, the virus integrated similarity matrix and the drug integrated similarity matrix, constructing a high-dimensional feature set;

[0044] Constructing a hypergraph with the high-dimensional feature set as the vertex set, and defining an objective function based on the hypergraph to obtain a projection matrix;

[0045] Calculating a predicted score matrix based on the projection matrix and the high-dimensional feature set;

[0046] Based on the predicted score matrix, screening the scores of the rows where the target viruses are located, and obtaining the final prediction result after sorting.

[0047] In some embodiments, the step of constructing the high-dimensional feature set comprises: setting the initial features of the viruses as the concatenation of the row vectors of the adjacency matrix and the virus integrated similarity matrix, extracting high-order features through a two-layer graph convolution network, obtaining latent variable distribution parameters through an encoder, and then obtaining latent variables through reparameterization, wherein the latent variables constitute the high-dimensional feature set.

[0048] In some embodiments, high-order feature information of the virus nodes is extracted by a graph convolution network, and the specific process is as follows: in the graph convolution network, the propagation of signals at the t-1 layer is to operate the initial feature matrix and the virus integrated similarity adjacency matrix after adding a self-loop and its degree matrix, and then the high-order feature information is obtained through model parameter processing and linear rectifier function activation.

[0049] In some embodiments, the features are processed by the encoder to gradually reduce the dimension, so as to obtain the distribution parameters of the latent variables, and the specific process is as follows: taking the initial feature matrix and the virus integrated similarity matrix as inputs, the parameters set in the encoder are operated to obtain the mean and standard deviation of the latent variables, which are two distribution parameters, and through the reparameterization method, the final latent variable is obtained by multiplying the random variable subject to the standard normal distribution by the standard deviation and adding the mean.

[0050] In some embodiments, the decoder reconstructs the virus-drug network structure by the inner product operation of the latent variables, and outputs the association probability of the two, and a double loss function is used for training: the binary cross-entropy loss is used to measure the deviation of the predicted association probability and the actual label, and the KL divergence loss is used to constrain the distribution of the latent variables to be close to the standard normal distribution.

[0051] In some embodiments, each sample in the high-dimensional feature set is taken as a vertex, and a hyperedge is generated according to the k-nearest neighbor rule, each hyperedge connects multiple similar vertices, and the weight of the hyperedge is calculated according to the distance between the vertices; a normalized hypergraph Laplacian matrix is constructed to depict the high-order association structure between samples.

[0052] In some embodiments, a target function is defined to balance the hypergraph structure reservation, data reconstruction error and model regularization constraint, in the target function, the projection matrix is used to map the high-dimensional features to the low-dimensional space, the regularization coefficient controls the weight of each part of the loss, and the optimal projection matrix is obtained by solving the closed-form solution.

[0053] In some embodiments, the high-dimensional feature set is mapped to a low-dimensional space by using the optimal projection matrix, and a predicted score matrix is calculated, each element in the matrix representing the potential association strength between the corresponding virus and drug; for a target virus, the row vector of the virus in the score matrix is extracted and arranged in descending order of the score, and the candidate drug with the strongest potential association is selected, providing a ranking result for the anti-virus drug screening.

[0054] The anti-virus drug screening system based on high-dimensional features applies the anti-virus drug screening method based on high-dimensional features in any one of the above embodiments, and the anti-virus drug screening system based on high-dimensional features comprises:

[0055] An adjacency matrix construction module is configured to construct an adjacency matrix based on known virus-drug association information.

[0056] An integrated similarity matrix calculation module is used to calculate the Gaussian distance similarity and cosine similarity of drugs based on the adjacency matrix to obtain a drug integrated similarity matrix; and to calculate the Gaussian distance similarity and cosine similarity of viruses to obtain a virus integrated similarity matrix;

[0057] A high-dimensional feature set construction module, which is used to construct a high-dimensional feature set based on an autoencoder and a graph convolutional network, combining the adjacency matrix, the virus integration similarity matrix, and the drug integration similarity matrix;

[0058] A projection matrix solving module is used to construct a hypergraph using the high-dimensional feature set as a vertex set, define an objective function based on the hypergraph, and solve to obtain a projection matrix;

[0059] A prediction score matrix calculation module, configured to calculate a prediction score matrix based on the projection matrix and the high-dimensional feature set;

[0060] The prediction module is used to filter out the scores of the rows where the target virus is located based on the prediction score matrix, and obtain the final prediction results after sorting.

[0061] A computer-readable storage medium stores computer instructions. When a computer reads the computer instructions, the computer executes any one of the above-described antiviral drug screening methods based on high-dimensional features.

[0062] The technical ideas of the present invention are as follows:

[0063] 1. Construct an adjacency matrix using known association information

[0064] Record known viruses as a set , drugs are recorded as a set Then, the virus-drug adjacency matrix was constructed using known virus-drug associations ,in Indicates the number of viruses, Indicates the amount of medicine, The meanings of the elements are as follows:

[0065] ;

[0066] 2. Use the adjacency matrix to calculate Gaussian similarity and cosine distance similarity

[0067] Take the medicine Represented as its associated fingerprint in the virus-drug adjacency matrix A , i.e. the length of column i in A A vector of 0 or 1, For drugs d ( j ), i.e., the jth column of A, and then calculate the drug Gaussian distance similarity:

[0068] ;

[0069] The normalized kernel width parameter is:

[0070] ;

[0071] Calculate the cosine similarity of drugs:

[0072] ;

[0073] Calculate drug integration similarity matrix Diagonal elements are forced to 1.

[0074] 3. Calculate virus similarity

[0075] Put the virus Fingerprint Let the i-th row of A (length ), calculate the virus Gaussian distance similarity:

[0076] ;

[0077] The normalized kernel width parameter is:

[0078] ;

[0079] Calculate virus cosine similarity:

[0080] ;

[0081] Calculation of viral integration similarity matrix Diagonal elements are forced to 1.

[0082] 4. Construct a high-dimensional feature set: Based on the autoencoder, input the virus-drug adjacency matrix A and the virus integration similarity matrix S v , drug integration similarity matrix S d Get the latent variables;

[0083] Feature initialization and construction input, let the initial feature matrix X, the feature vector of each virus is taken from a row of the virus-drug adjacency matrix A, and the virus integration similarity matrix S v Next, splice S v and A serve as input data for the main model.

[0084] The high-order feature information of virus nodes is extracted from the virus-drug network through the graph convolutional network (GCN) layer. The propagation formula of the graph convolutional network at the t-1 layer is:

[0085] ;

[0086] where X(t-1) is the input node feature matrix of the t-1 layer graph convolutional network, taking the row or column of the virus-drug adjacency matrix A, is the virus integration similarity adjacency matrix after adding self-loops, that is, the virus integration similarity matrix with all diagonal elements valued as 1, is the degree matrix of is the model weight parameter, which can be optimized during the training process, and ReLU is the activation function.

[0087] The formula for obtaining the latent variable distribution parameter after step-by-step dimension reduction by the encoder is:

[0088] ;

[0089] wherein, is the input node feature matrix, is the parameter matrix of the graph convolutional layer, corresponding to the weight of each layer respectively, used for convolution operation and feature transformation.

[0090] Finally, the latent variable is obtained, wherein represents the mean vector output by the GCN, which is the latent variable mean parameter in the variational autoencoder (VAE), represents the standard deviation vector output by the GCN, which is the variance parameter of the VAE latent variable, is noise following the standard normal distribution N(0, 1), used for VAE reparameterization sampling, and * represents multiplication.

[0091] Through the decoder, the latent variable z is used to reconstruct the virus network structure, and the virus-drug association probability is output:

[0092] ;

[0093] wherein sigmoid() represents the sigmoid function, z represents the latent variable (hidden variable), and represents the virus or drug node embedding obtained through the VAE; T represents matrix transposition.

[0094] The GCN loss function includes binary cross-entropy loss and latent variable KL divergence loss.

[0095] The row samples of the high-dimensional feature set are taken as vertices, and the k-nearest neighbors of each vertex generate a hyperedge set , and the hyperedge weight adopts ; the normalized hypergraph Laplacian matrix is calculated as: wherein L represents the normalized hypergraph Laplacian matrix, used to reflect the high-order structural relationship between samples; I represents the unit matrix, with the same dimension as the number of vertices; D v ​denotes the vertex degree diagonal matrix, the i-th element on the diagonal line denotes the sum of the weights of all hyperedges that vertex i belongs to; D e denotes the hyperedge degree diagonal matrix, the i-th element on the diagonal line denotes the number (or weight sum) of vertices contained in the hyperedge; H denotes the incidence matrix of the hypergraph, and a value of 1 indicates that a vertex belongs to a hyperedge, otherwise 0; W denotes the hyperedge weight diagonal matrix, and the diagonal element denotes the weight of the hyperedge, which is taken as a unit diagonal matrix, that is, all hyperedge weights are equal.

[0096] The objective function is defined as: , wherein P is a to-be-solved projection matrix, is the objective function of the to-be-solved projection matrix P, is the transpose of the matrix P, X is the splicing matrix of the above-mentioned latent variable z, L is the normalized hypergraph Laplacian matrix, and tr() is the trace of the matrix, is a regularization coefficient, which controls the weight of the reconstruction error term in the objective function; is a weight parameter, which controls the weight of the regularization penalty of P in the objective function; is the square of the F norm, which prevents overfitting.

[0097] Then the closed-form solution is obtained as:

[0098] ;

[0099] wherein, is a unit matrix;

[0100] Finally, the predicted score matrix S is calculated as:

[0101] ;

[0102] wherein S denotes the predicted score matrix, and the size of the predicted score matrix is the same as that of the virus-drug adjacency matrix.

[0103] The effectiveness of the application is verified as follows:

[0104] Based on the process of the high-dimensional feature-based antiviral drug screening method shown in FIG. Figure 2 The specific implementation is as follows: first, all known virus-drug associations are randomly and evenly divided into 5 groups, and then each group in the 5 groups is set as a test sample, and the other groups are set as training samples. The training samples are used as the input of the method to obtain the prediction result, and finally the prediction score of each test sample in the group is compared with the candidate score. In order to reduce the influence of random division in the process of obtaining the test sample on the result, 50 times of five-fold cross-validation are performed.

[0105] After calculation using Matlab, the following data are obtained, as shown in FIG. Figure 3AUROC (Area Under the ROC curve) values between the present method and several virus-drug screening models reported so far are shown. AUROC value of 0.8798 ± 0.0014 was achieved in 5-fold cross-validation, showing better prediction performance than several classic models.

[0106] In another aspect, the present method is used to predict and screen drug scoring matrix for a specific virus, such as SARS-CoV-2. S The row corresponding to SARS-CoV-2 in the above table is the prediction score of the relevant drug, and the top 20 drugs in descending order have 18 drugs supported by reported literature.

[0107] The second and third columns in the above table represent the names of the top 20 predicted drugs and the supporting literature PMID numbers, respectively.

[0108]

[0109] The above-described embodiments are used to illustrate the present application and are not intended to limit the present application, so changes in example values or replacement of equivalent elements should still belong to the scope of the present application.

[0110] From the above detailed description, it is clear to those skilled in the art that the present application can achieve the above-mentioned purposes, and has met the requirements of the Patent Law.

[0111] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to these embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including all the preferred embodiments and all changes and modifications falling within the scope of the present application. The above description is only the preferred embodiments of the present application and is not intended to limit the present application. It should be noted that any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

[0112] It should be noted that the above description of the process is only for example and illustration, and does not limit the scope of the present application. Those skilled in the art can make various modifications and changes to the process under the guidance of the present application. However, these modifications and changes are still within the scope of the present application.

[0113] Having now described the basic concept, it will be apparent to those skilled in the art after reading this patent disclosure that numerous modifications, substitutions and changes can be made to the application as described without departing from the scope of the application as recited in the claims. Although specific embodiments of the application have been described herein, it will be understood by those skilled in the art that various modifications, improvements and / or reconstructions can be made to the application as described herein. Such modifications, improvements and / or reconstructions are suggested by this application and are still within the spirit and scope of the exemplary embodiments of the application.

[0114] Also, certain terms have been used herein for the purpose of reference only and are not intended to be limiting. For example, "one embodiment", "an embodiment", and / or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" or "in some embodiments" in various places throughout this specification are not necessarily referring to the same embodiment of the application. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner on an embodiment or embodiments of the application.

[0115] Moreover, those skilled in the art will appreciate that the various aspects of the application can be illustrated and described by means of a number of illustrative examples or scenarios involving any new and useful processes, machines, products or compositions of matter, or any new and useful improvements thereof, including any new and useful processes, machines, products or compositions of matter, or any new and useful improvements thereof. Thus, the various aspects of the application can be embodied in whole or in part in hardware, in software (including firmware, resident software, micro-code, etc.), or in a combination of hardware and software. The hardware or software can be referred to as a "unit", "module" or "system". Furthermore, the various aspects of the application can take the form of a computer program product on one or more computer-readable media having computer readable program code embodied in the medium.

[0116] Computer program code for carrying out operations of various aspects of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, conventional procedural programming languages, such as the C programming language, Visual Basic, Fortran 2103, Perl, COBOL 2102, PHP, ABAP, dynamic programming languages, such as Python, Ruby and Groovy, or another programming language. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any form of network, such as a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet) or within a cloud computing environment or as a service, such as software as a service (SaaS).

[0117] Furthermore, the order of presentation of the processing elements and sequences, unless specifically stated otherwise, is not intended to be construed as a limitation, but is presented for purposes of illustrative clarity. Although the present application has been described in terms of certain implementations, it is to be understood that the detailed disclosure provided herein, including the descriptions of the various embodiments of the application, is to be taken in an illustrative and conceptual sense, as being exemplary, rather than as being limiting of the present application. For example, although the implementation of the various components described above can be realized in hardware devices, it can also be realized as a pure software solution, for example, as an installation on existing servers or mobile devices.

[0118] Similarly, it is to be noticed that the term "comprising", used in the description, is not to be interpreted as being restricted to the means listed thereafter. It is to be understood that the features named after the comma are complementary, and that the unit "comprising" can be interpreted as specifying the possession of at least one of the stated features. It is to be further understood that the features recited after the term "comprising" can be some or all of the features of one of the various embodiments of the present application, and that the unit "comprising" can be interpreted to mean that the described features are included in the claimed subject matter, but that not all of the features need to be present in the claimed subject matter.

Claims

1. An antiviral drug screening method based on high-dimensional features, characterized in that: include: Construct an adjacency matrix based on known virus-drug association information; Based on the adjacency matrix, the Gaussian distance similarity and cosine similarity of the drugs are calculated to obtain the drug integrated similarity matrix; And calculate the Gaussian distance similarity and cosine similarity of the virus to obtain the virus integrated similarity matrix; Based on autoencoders and graph convolutional networks, a high-dimensional feature set is constructed by combining adjacency matrix, virus integration similarity matrix, and drug integration similarity matrix; Constructing a hypergraph using the high-dimensional feature set as a vertex set, defining an objective function based on the hypergraph, and solving to obtain a projection matrix; Calculating a prediction score matrix based on the projection matrix and the high-dimensional feature set; Based on the prediction score matrix, the scores of the target virus row are filtered out and sorted to obtain the final prediction results; The steps of constructing the high-dimensional feature set include: setting the initial features of the virus as the concatenation of the row vectors of the adjacency matrix and the virus integrated similarity matrix, extracting high-order features through a graph convolutional network, obtaining latent variable distribution parameters through an encoder, and then obtaining latent variables through reparameterization, wherein the latent variables constitute the high-dimensional feature set; The high-order feature information of virus nodes is extracted through the graph convolutional network. The specific process is as follows: In the graph convolutional network, the signal propagation in the t-1 layer is to calculate the initial feature matrix with the virus integrated similarity adjacency matrix and its degree matrix after adding self-loops. Then, after model parameter processing and linear rectification function activation, high-order feature information is obtained. The propagation formula of the graph convolutional network at layer t-1 is: ; where X (t-1) is the input node feature matrix of the t-1 layer graph convolutional network, and takes the rows or columns of the virus-drug adjacency matrix A. is the virus integration similarity adjacency matrix after adding the self-loop, that is, the virus integration similarity matrix with all diagonal elements assigned to 1, for The degree matrix of is the model weight parameter and ReLU is the activation function.

2. The antiviral drug screening method based on high-dimensional features according to claim 1, characterized in that: The encoder gradually reduces the dimensionality of the features to obtain the distribution parameters of the latent variables. Specifically, the initial feature matrix and the virus integration similarity matrix are used as input, and the parameters set in the encoder are used for calculation to obtain the two distribution parameters of the latent variable, the mean and standard deviation. Through the reparameterization method, the random variable that obeys the standard normal distribution is multiplied by the standard deviation and then added to the mean to obtain the final latent variable.

3. The antiviral drug screening method based on high-dimensional features according to claim 2, characterized in that: The decoder reconstructs the virus-drug network structure through the inner product operation of the latent variables and outputs the association probability between the two. A dual loss function is used during training: binary cross entropy loss is used to measure the deviation between the predicted association probability and the actual label, and KL divergence loss constrains the latent variable distribution to be close to the standard normal distribution.

4. The antiviral drug screening method based on high-dimensional features according to claim 3, characterized in that: Taking each sample in the high-dimensional feature set as a vertex, hyperedges are generated according to the k-nearest neighbor rule. Each hyperedge connects multiple similar vertices, and the hyperedge weight is calculated based on the vertex distance. A normalized hypergraph Laplacian matrix is ​​constructed to characterize the high-order correlation structure between samples.

5. The antiviral drug screening method based on high-dimensional features according to claim 4, characterized in that: An objective function is defined to balance the preservation of the hypergraph structure, data reconstruction error, and model regularization constraints. In the objective function, the projection matrix is ​​used to map high-dimensional features to low-dimensional space, and the regularization coefficient controls the weight of the loss of each part. The optimal projection matrix is ​​obtained by solving the closed-form solution.

6. The antiviral drug screening method based on high-dimensional features according to claim 5, characterized in that: The optimal projection matrix is ​​used to perform dimensionality reduction mapping on the high-dimensional feature set, and a prediction score matrix is ​​calculated. Each element in the matrix represents the potential correlation strength between the corresponding virus and the drug. For the target virus, its row vector in the score matrix is ​​extracted and arranged in descending order by score to screen out the candidate drugs with the strongest potential correlation, providing ranking results for antiviral drug screening. 7.Antiviral drug screening system based on high-dimensional features, characterized by: The antiviral drug screening method based on high-dimensional features according to any one of claims 1 to 6 is applied, wherein the antiviral drug screening system based on high-dimensional features comprises: An adjacency matrix construction module, used to construct an adjacency matrix based on known virus-drug association information; An integrated similarity matrix calculation module is used to calculate the Gaussian distance similarity and cosine similarity of drugs based on the adjacency matrix to obtain a drug integrated similarity matrix; and to calculate the Gaussian distance similarity and cosine similarity of viruses to obtain a virus integrated similarity matrix; A high-dimensional feature set construction module, which is used to construct a high-dimensional feature set based on an autoencoder and a graph convolutional network, combining the adjacency matrix, the virus integration similarity matrix, and the drug integration similarity matrix; A projection matrix solving module is used to construct a hypergraph using the high-dimensional feature set as a vertex set, define an objective function based on the hypergraph, and solve to obtain a projection matrix; A prediction score matrix calculation module, configured to calculate a prediction score matrix based on the projection matrix and the high-dimensional feature set; The prediction module is used to filter out the scores of the rows where the target virus is located based on the prediction score matrix, and obtain the final prediction results after sorting.

8. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions, the computer executes the antiviral drug screening method based on high-dimensional features according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Microorganism-disease relevance prediction method based on graph attention network and system thereof

    CN113345523A

  • Antiviral drug screening method and system based on hypergraph learning and storage medium

    CN116631502A

  • Chinese herbal medicine indication discovery method and system based on hypergraph convolution

    CN119601257A