Deep learning sleep-aiding molecular drug design method and system based on disturbance phenotype inversion

By inverting gene expression changes using deep learning models, small molecular structures that can induce target sleep phenotypes are generated, solving the problems of low efficiency and high cost in traditional sleep aid drug development and achieving efficient and controllable drug design.

CN121838931APending Publication Date: 2026-04-10SHANDONG PETROCHEMICAL INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In the current development of sleep aid drugs, traditional screening strategies cannot achieve reverse design from "target phenotype to small molecule drug", resulting in low development efficiency, high cost, and lack of effective biological feedback mechanisms, making it difficult to generate candidate molecules with real biological effects.

Method used

A deep learning method based on perturbation phenotype inversion is adopted to infer gene expression changes through a deep learning model, generate small molecule structures that can induce target sleep phenotypes, establish a unified mapping relationship between drug structure, gene expression, and phenotype, and realize end-to-end reverse design.

Benefits of technology

This significantly improves the efficiency and success rate of sleep aid drug discovery, reduces the research and development cycle and cost, and directly generates highly reliable candidate molecular structures to meet clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838931A_ABST
    Figure CN121838931A_ABST
Patent Text Reader

Abstract

The invention relates to the cross technical field of artificial intelligence, bioinformatics and drug design, and provides a deep learning sleep-aiding molecular drug design method and system based on perturbation phenotype inversion. The method comprises the steps of obtaining a transcriptome expression profile before disturbance and target sleep-aiding phenotype data, and performing preprocessing; based on the preprocessed pre-disturbance transcriptome expression profile and the target sleep-aiding phenotype data, obtaining a target transcriptome expression vector by adopting an encoder; calculating a target expression change vector based on a difference value between the target transcriptome expression vector and the preprocessed pre-perturbation transcriptome expression profile; and based on the target expression change vector, a decoder is adopted to obtain a candidate micromolecule structure sequence. Gene expression change which should be generated is deduced through a deep learning model, a candidate small molecule structure capable of inducing the change is further generated, and end-to-end reverse design from a sleep aiding target phenotype to the candidate small molecule structure is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, bioinformatics and drug design, in particular to a deep learning assisted sleep molecule drug design method and system based on perturbed phenotype inversion. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] In the research and development process of sleep-aiding drugs, candidate molecule screening is a core and key link directly determining the research and development efficiency and the final success rate, and its technical level directly affects the research and development process and industrialization landing speed of new sleep-aiding drugs. At present, the mainstream candidate molecule screening method in the industry still relies on the mode of combining traditional laboratory synthesis and in vitro activity test, which needs to invest a large number of professional research and development personnel, precise experimental equipment and various experimental consumables, and consumes a very long experimental period. There are such outstanding problems as long research and development period (the average period of single drug research and development is 5-10 years), high research and development cost (the cost of single drug research and development has exceeded 1 billion US dollars) and low candidate molecule hit rate, which have made it difficult to meet the urgent research and development needs and market application expectations of the pharmaceutical industry for safe and efficient new sleep-aiding drugs.

[0004] In the prior art, the traditional sleep-aiding drug discovery process mainly relies on two types of screening strategies: one is the target-based strategy, which needs to first identify the biological target related to sleep regulation, and then based on the molecular structure and functional characteristics of the target, small molecule compounds that can specifically act on the target are screened from a large number of compound libraries; the other is the phenotypic screening strategy, the core of which is to construct a cell model or animal model that can activate or inhibit the sleep regulation pathway, and by directly observing the sleep-related phenotype changes produced after the compound acts on the model, the screening of candidate molecules is realized. However, the above two types of screening strategies and related technical methods derived therefrom have key technical bottlenecks that are difficult to break through in actual application.

[0005] Firstly, the traditional screening technology cannot realize the reverse design process of "target phenotype → small molecule drug", that is, there is no mature technology that can directly design a candidate small molecule compound structure with corresponding function according to the specific sleep biological effect expected by the research and development personnel, resulting in a blind screening state of drug research and development, which seriously restricts the research and development efficiency. Secondly, sleep phenotype itself involves extremely complex multi-gene regulation network, and the start, maintenance and termination of sleep process need to be realized through the synergistic action of multiple signal pathways such as circadian rhythm pathway, calcium signaling pathway and GABAergic signaling pathway, while in the traditional quantitative structure-activity relationship (QSAR) model or conventional molecular generation model, there is generally a lack of clear biological constraint relationship between phenotype characteristics and gene expression, resulting in a disconnection between the model output result and the actual biological activity requirement. Finally, the current mainstream molecular generation model, such as diffusion model (Diffusion), variational autoencoder (VAE) and generative adversarial network (GAN), lacks an effective biological feedback mechanism, and cannot ensure that the compound molecules generated by it can induce the specific post-sleep phenotype required by research and development in the biological body, further reducing the effective hit rate of candidate molecules and increasing the difficulty and cost pressure of sleep aid drug research and development. SUMMARY

[0006] In order to solve the technical problems existing in the background art, the present application provides a deep learning sleep aid molecule drug design method and system based on perturbed phenotype inversion, which infers the gene expression changes that should be produced by a deep learning model, and further generates candidate small molecule structures that can induce the changes, realizing end-to-end reverse design from sleep aid target phenotype to candidate small molecule structure.

[0007] In order to achieve the above purpose, the present application adopts the following technical solutions: The first aspect of the present application provides a deep learning sleep aid molecule drug design method based on perturbed phenotype inversion.

[0008] A deep learning sleep aid molecule drug design method based on perturbed phenotype inversion, comprising: Obtaining the transcriptome expression profile before perturbation, the phenotype data after perturbation, the transcriptome expression profile after perturbation and the small molecule structure inducing the perturbation, and preprocessing; inputting the preprocessed phenotype data after perturbation into a deep learning model to obtain the gene expression profile after perturbation The input encoder obtains a target transcriptome expression vector, a first loss function is designed according to the target transcriptome expression vector and the transcriptome expression profile after the perturbation, and the hyperparameters of the encoder are optimized; the preprocessed small molecule structure of the induced perturbation is input into the decoder, and then mapped to the gene space to obtain an induced gene expression change vector; a second loss function is designed according to the induced gene expression change vector, and the hyperparameters of the decoder are optimized; the target transcriptome expression vector and the induced gene expression change vector are input into the expression phenotype alignment module, so that the difference between the target transcriptome expression vector and the induced gene expression change vector is equal to the transcriptome expression profile before the perturbation, and the alignment of the phenotype pathway and the small molecule pathway is completed. The transcriptome expression profile before the perturbation and the target sleep aid phenotype data are obtained and preprocessed; based on the preprocessed transcriptome expression profile before the perturbation and the target sleep aid phenotype data, an encoder is used to obtain a target transcriptome expression vector; based on the difference between the target transcriptome expression vector and the preprocessed transcriptome expression profile before the perturbation, a target expression change vector is calculated; based on the target expression change vector, a decoder is used to obtain a candidate small molecule structure sequence.

[0009] Further, the preprocessed post-perturbation phenotype data is input into the encoder to obtain a target transcriptome expression vector; the method comprises: The preprocessed post-perturbation phenotype data is mapped to the same feature space to obtain a first post-perturbation phenotype feature by splicing; The first post-perturbation phenotype feature is input into a multi-head self-attention to obtain a second post-perturbation phenotype feature; The second post-perturbation phenotype feature is input into a feedforward network to obtain a third post-perturbation phenotype feature; The third post-perturbation phenotype feature is back-projected to the gene dimension to obtain a target transcriptome expression vector.

[0010] Further, in the training process of the expression phenotype alignment module, a third loss function is designed to optimize the hyperparameters of the expression phenotype alignment module; the third loss function is represented by the following formula:

[0011] Wherein, L3 represents the third loss function; X represents the target transcriptome expression vector; Y represents the transcriptome expression profile before the perturbation; D represents the induced gene expression change vector.

[0012] Further, the first loss function is represented by the following formula:

[0013] The second loss function is represented by the following formula:

[0014]

[0015] wherein, represents a first loss function; represents a perturbed transcriptomic expression profile; represents a target transcriptomic expression vector; represents a second loss function; represents an induced gene expression change vector; represents a perturbed transcriptomic expression profile and a pre-perturbation transcriptomic expression profile between them.

[0016] Further, the pre-processing process comprises: normalizing and log1p transforming the pre-perturbation transcriptomic expression profile, the target sleep-aiding phenotype data, the post-perturbation phenotype data, and the post-perturbation transcriptomic expression profile; converting the structure of the small molecule inducing the perturbation into an atomic sequence.

[0017] Further, based on the target expression change vector, a decoder is used to obtain a candidate small molecule structure sequence; the following formula is used to represent:

[0018] wherein, represents a candidate small molecule structure sequence; represents a target expression change vector; represents noise.

[0019] A second aspect of the present application provides a deep learning sleep-aiding molecule drug design system based on perturbation phenotype inversion.

[0020] A deep learning sleep-aiding molecule drug design system based on perturbation phenotype inversion, comprising: a training module configured to: obtain a pre-perturbation transcriptomic expression profile, post-perturbation phenotype data, post-perturbation transcriptomic expression profile, and structure of a small molecule inducing the perturbation, and perform pre-processing; and input the pre-processed post-perturbation phenotype data The input encoder obtains a target transcriptome expression vector, a first loss function is designed according to the target transcriptome expression vector and the transcriptome expression profile after perturbation, and the hyperparameters of the encoder are optimized; the preprocessed small molecule structure of the induced perturbation is input into the decoder, and then mapped to the gene space to obtain an induced gene expression change vector; a second loss function is designed according to the induced gene expression change vector, and the hyperparameters of the decoder are optimized; the target transcriptome expression vector and the induced gene expression change vector are input into the expression-phenotype alignment module, so that the difference between the target transcriptome expression vector and the induced gene expression change vector is equal to the transcriptome expression profile before perturbation, and the alignment of the phenotype path and the small molecule path is completed. The reasoning module is configured to: obtain the transcriptome expression profile before perturbation and target sleep-aiding phenotype data, and perform preprocessing; based on the preprocessed transcriptome expression profile before perturbation and the target sleep-aiding phenotype data, an encoder is adopted to obtain a target transcriptome expression vector; based on the difference between the target transcriptome expression vector and the preprocessed transcriptome expression profile before perturbation, a target expression change vector is calculated; based on the target expression change vector, a decoder is adopted to obtain a candidate small molecule structure sequence.

[0021] The third aspect of the present application provides a computer device, which comprises: a processor adapted to execute a computer program; a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being executed by the processor to implement the steps in the deep learning sleep-aiding molecule drug design method based on perturbation phenotype inversion according to the first aspect.

[0022] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, the computer program being adapted to be loaded and executed by a processor to implement the steps in the deep learning sleep-aiding molecule drug design method based on perturbation phenotype inversion according to the first aspect.

[0023] The fifth aspect of the present application provides a computer program product or a computer program.

[0024] The present application provides a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the steps in the deep learning sleep-aiding molecule drug design method based on perturbation phenotype inversion according to the first aspect.

[0025] Compared with the prior art, the present application has the following beneficial effects: The present application realizes end-to-end reverse design from sleep aid target phenotype to candidate small molecule structure, that is, input the sleep improvement effect expected to be achieved, and the present application can output the small molecule drug structure capable of inducing the effect, realizing the change of drug design direction from 'target-based' to 'phenotype demand-based'.

[0026] The present application reverses the gene regulation path of the target phenotype by using the transcriptome expression information before and after disturbance, and makes the model have the ability to infer the gene expression changes that should occur according to the phenotype by constructing a disturbance inversion Transformer, so as to realize the deep modeling of the complex sleep regulation mechanism.

[0027] The present application establishes a unified mapping relationship between drug structure, gene expression and phenotype, and makes the expression changes induced by small molecules consistent with the phenotype inference path through the expression phenotype alignment module, so as to ensure that the small molecules generated have real biological effects on sleep-related pathways (such as GABAergic signaling, circadian rhythm, etc.).

[0028] The present application significantly improves the efficiency, hit rate and controllability of sleep aid drug discovery. Traditional sleep aid drug screening relies on large-scale experiments, and the present application directly generates high-credibility candidate structures through a deep learning model for reference by R&D personnel, which can greatly reduce the drug R&D cycle and cost. BRIEF DESCRIPTION OF DRAWINGS

[0029] The drawings accompanying the specification of the present application form a part thereof and serve to further understand the present application, the illustrative embodiments of the present application and the description thereof serve to explain the present application and do not constitute an improper limitation of the present application.

[0030] Figure 1 is a flowchart of the deep learning sleep aid molecule drug design method based on disturbance phenotype inversion according to the embodiment of the present application; Figure 2 is a structural diagram of the deep learning sleep aid molecule drug design system based on disturbance phenotype inversion according to the embodiment of the present application; Figure 3 is a structural diagram of the computer device according to the embodiment of the present application. DETAILED DESCRIPTION

[0031] The present application will be further described below in conjunction with the drawings and embodiments.

[0032] It should be pointed out that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the present application belongs.

[0033] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0034] As introduced in the background, the prior art cannot directly infer from the target sleep phenotype to generate small molecule structures with sleep-aiding effect, nor can it guarantee that the generated molecules can biologically truly induce the target phenotype. In order to break through the limitations of traditional target-driven and phenotype screening strategies, the present application provides a deep learning sleep-aiding molecule drug design method and system based on perturbed phenotype inversion, which can automatically invert the gene expression perturbation that should be generated according to the gene expression profile before perturbation and the expected sleep-aiding phenotype change, and further generate small molecule compounds that can induce the perturbation. The scheme of the present application is described in detail through several embodiments below.

[0035] Figure 1 is a flowchart of the deep learning sleep-aiding molecule drug design method based on perturbed phenotype inversion according to an embodiment of the present application; referring to Figure 1 , the method comprises: obtaining the transcriptome expression profile before perturbation, the phenotype data after perturbation, the transcriptome expression profile after perturbation, and the small molecule structure inducing the perturbation, and preprocessing; inputting the preprocessed transcriptome expression profile before perturbation and the phenotype data after perturbation into an encoder to obtain a target transcriptome expression vector, designing a first loss function according to the target transcriptome expression vector and the transcriptome expression profile after perturbation, and optimizing the hyperparameters of the encoder; inputting the preprocessed small molecule structure inducing the perturbation into a decoder, and then mapping to the gene space to obtain an induced gene expression change vector; designing a second loss function according to the induced gene expression change vector, and optimizing the hyperparameters of the decoder; inputting the target transcriptome expression vector and the induced gene expression change vector into an expression phenotype alignment module, so that the difference between the target transcriptome expression vector and the induced gene expression change vector is equal to the transcriptome expression profile before perturbation, and the alignment of the phenotype path and the small molecule path is completed; obtaining the transcriptome expression profile before perturbation and the target sleep-aiding phenotype data, and preprocessing; based on the preprocessed transcriptome expression profile before perturbation and the target sleep-aiding phenotype data, using an encoder to obtain a target transcriptome expression vector; based on the difference between the target transcriptome expression vector and the preprocessed transcriptome expression profile before perturbation, calculating a target expression change vector; based on the target expression change vector, using a decoder to obtain a candidate small molecule structure sequence.

[0036] The application constructs a new sleep-aiding drug design paradigm, takes perturbation phenotype inversion as the core driving, takes gene expression changes as the biological basis, realizes molecular design by a deep generation model, so as to realize the automatic sleep-aiding drug generation process of “phenotype→expression→small molecule” three-layer linkage, and provide a new, safe and effective sleep-aiding drug research and development scheme for the clinic.

[0037] In some embodiments, the pre-perturbation transcriptome expression profile, post-perturbation phenotype data, post-perturbation transcriptome expression profile and structure of the small molecule inducing perturbation are obtained and preprocessed; the implementation process comprises: obtaining the pre-perturbation transcriptome expression profile , post-perturbation phenotype data (sleep improvement signal characteristics) , post-perturbation transcriptome expression profile and structure of the small molecule inducing perturbation , normalizing the original pre-perturbation and post-perturbation experimental data (transcriptome FPKM / UMI, sleep-aiding score), and small molecule SMILES (based on DrugBank, LINCS query) and unifying the data format: In the preprocessing stage, the post-perturbation phenotype data (here, corresponding to “sleep-aiding score”) and other information are standardized and log1p converted, and then based on the processed “sleep-aiding score” data, the corresponding phenotype vector is constructed; the small molecule SMILES is converted into an atomic sequence (Tokenizer), and finally the phenotype vector and other data (such as transcriptome expression profile, small molecule structure) are combined to form four-tuple data for training .

[0038] In some embodiments, the pre-perturbation transcriptome expression profile and post-perturbation phenotype data after preprocessing are input into an encoder to obtain a target transcriptome expression vector ; the implementation process comprises: adopting a Transformer Encoder structure to infer “given phenotype and initial expression→expression changes that should be produced”, that is, to predict the post-perturbation expression profile that should be produced: .

[0039] Specifically, the pre-processed post-perturbation phenotype data is mapped to the same feature space , and the first post-perturbation phenotype feature is obtained by splicing: ; the following formula is used to represent:

[0040] , wherein represents the pre-perturbation transcriptome expression profile split into multiple subparts, in order to facilitate model processing, it will be split into k smaller dimension sub-vectors; the first perturbed phenotype feature inputting the multi-head self-attention, to obtain a second perturbed phenotype feature ; the following formula is used to represent:

[0041] the second perturbed phenotype feature inputting a feedforward network FFN for transformation, to obtain a third perturbed phenotype feature ; the following formula is used to represent:

[0042] the third perturbed phenotype feature back-projecting to the gene dimension, to obtain a target transcriptome expression vector ; the following formula is used to represent:

[0043] wherein, denotes back-projection.

[0044] In some embodiments, the first loss function of the encoder is represented by the following formula:

[0045] wherein, denotes the first loss function; denotes the perturbed transcriptome expression profile; denotes the target transcriptome expression vector.

[0046] In this embodiment, the perturbed phenotype is injected into the Transformer as conditional information, so that the model has the ability to infer the perturbation causal path in reverse. Using a multi-modal attention mechanism of gene patches and phenotype tokens, fine mapping of phenotype to transcriptome in the phenotype path is achieved.

[0047] The present application uses the transcriptome expression information before and after perturbation to reverse the gene regulation path of the target phenotype, and by constructing a perturbation inversion Transformer, the model has the ability to infer the gene expression changes that should occur according to the phenotype, thereby realizing deep modeling of the complex sleep regulation mechanism.

[0048] In some embodiments, the small molecule structure of the induced perturbation after preprocessing inputting the decoder, and then mapping to the gene space, to obtain an induced gene expression change vector The implementation process includes predicting the gene expression changes induced by the input small molecule structure as a second path of expression-phenotype alignment. The input is a small molecule structure M (token sequence or graph embedding), and the output is the predicted gene expression changes; the SMILES token embedding is obtained using nn.embedding, the chemical structure sequence is modeled using a Transformer Decoder, and the preprocessed small molecule structure induced by the perturbation is converted into an embedding to obtain a small molecule vector induced by the perturbation and mapped to the gene space:

[0049] wherein MLP represents a multi-layer perceptron.

[0050] In some embodiments, the second loss function of the decoder is represented by the following formula:

[0051]

[0052] wherein represents the second loss function; represents the induced gene expression change vector.

[0053] In this embodiment, the small molecule path of "structure → expression change" is proposed to form the biological mechanism constraint of the model. The drug-induced expression is taken as strong supervision to guide the generation of more biologically credible small molecules.

[0054] In some embodiments, the target transcriptome expression vector and the induced gene expression change vector are input into the expression-phenotype alignment module, so that the difference between the target transcriptome expression vector and the induced gene expression change vector is equal to the transcriptome expression profile before the perturbation , and the alignment of the phenotype path and the small molecule path is completed; the implementation process includes forcing the two paths: the phenotype path and the small molecule path to be consistent, i.e.,

[0055] wherein the input is (from the encoder) and (from the decoder), and the output is a shared expression space subject to consistency constraints.

[0056] In some embodiments, the third loss function of the expression-phenotype alignment module is represented by the following formula:​

[0057] wherein, represents a third loss function; represents a target transcriptomic expression vector; represents a pre-perturbation transcriptomic expression profile; represents an induced gene expression change vector.

[0058] In the present embodiment, the expression phenotype alignment module first ensures that the phenotype inference and drug mechanism prediction two paths are biologically consistent, realizing the deep integration of the counteracting mechanism. It solves the core limitation of the existing generative model that the small molecule structure is inconsistent with the phenotype.

[0059] In some embodiments, in the model inference stage, the pre-perturbation transcriptomic expression profile and the target sleep-aiding phenotype data are obtained and pre-processed; the pre-processing process is consistent with the above, and will not be repeated here.

[0060] According to the pre-perturbation expression profile , the target sleep-aiding phenotype data , a small molecule structure satisfying the target phenotype is automatically generated ; that is, the input is , and the output is a small molecule structure sequence (SMILES).

[0061] Specifically, the pre-processed pre-perturbation transcriptomic expression profile and the target sleep-aiding phenotype data are input into the encoder to obtain the target transcriptomic expression vector ; that is, .

[0062] Taking the target expression change as a condition vector, based on the difference between the target transcriptomic expression vector and the pre-processed pre-perturbation transcriptomic expression profile , the target expression change vector is calculated; that is: .

[0063] A Transformer-based conditional molecule generator is proposed, which takes as a conditional constraint to generate SMILES tokens step by step:

[0064] wherein, represents a candidate small molecule structure sequence; represents a target expression change vector; represents noise; This indicates a transformer-based decoder.

[0065] This invention establishes a unified mapping relationship between drug structure, gene expression, and phenotype. Through an expression phenotype alignment module, it ensures that the expression changes induced by small molecules are consistent with the phenotype inference path, thereby ensuring that the generated small molecules have real biological effects on sleep-related pathways (such as GABAergic signaling and circadian rhythm).

[0066] This embodiment proposes for the first time a three-level reverse generation chain of "phenotype-gene expression-small molecule structure". New molecules directly target phenotypic needs, eliminating the need for target screening, and realizing an end-to-end generation process from "sleep-aid target phenotype → candidate small molecule structure". This significantly improves the efficiency, hit rate, and controllability of sleep-aid drug discovery. Traditional sleep-aid drug screening relies on large-scale experiments; by using a deep learning model to directly generate high-confidence candidate structures for researchers' reference, it can significantly reduce drug development cycles and costs.

[0067] The above combination Figure 1 The method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion using deep learning, as provided in the embodiments of the present invention, has been described in detail. Next, the system for designing sleep-aid molecular drugs based on perturbation phenotypic inversion using deep learning, as provided in the embodiments of the present invention, will be introduced with reference to the accompanying drawings.

[0068] Figure 2 This is a schematic diagram of the structure of a deep learning-based sleep aid molecular drug design system based on perturbation phenotypic inversion, as shown in an embodiment of the present invention. (Refer to...) Figure 2 The system described in this invention includes: The training module is configured to: acquire pre-perturbation transcriptome expression profiles, post-perturbation phenotypic data, post-perturbation transcriptome expression profiles, and induced perturbation small molecule structures, and perform preprocessing; and process the preprocessed post-perturbation phenotypic data. The encoder is input to obtain the target transcriptome expression vector. Based on the target transcriptome expression vector and the perturbed transcriptome expression profile, a first loss function is designed to optimize the encoder's hyperparameters. The preprocessed induced perturbation small molecule structure is input into the decoder and mapped to the gene space to obtain the induced gene expression change vector. Based on the induced gene expression change vector, a second loss function is designed to optimize the decoder's hyperparameters. The target transcriptome expression vector and the induced gene expression change vector are input into the expression phenotype alignment module so that the difference between the target transcriptome expression vector and the induced gene expression change vector equals the unperturbed transcriptome expression profile, thus completing the alignment of the phenotypic path with the small molecule path. The reasoning module is configured to: obtain pre-disturbance transcriptome expression profiles and target sleep aid phenotype data, and pre-process the same; based on the pre-processed pre-disturbance transcriptome expression profiles and target sleep aid phenotype data, obtain a target transcriptome expression vector by using an encoder; based on a difference between the target transcriptome expression vector and the pre-processed pre-disturbance transcriptome expression profiles, calculate a target expression change vector; and based on the target expression change vector, obtain a candidate small molecule structure sequence by using a decoder.

[0069] In some embodiments, inputting the pre-processed post-disturbance phenotype data into the encoder to obtain the target transcriptome expression vector comprises: mapping the pre-processed post-disturbance phenotype data to the same feature space to splice a first post-disturbance phenotype feature; inputting the first post-disturbance phenotype feature into a multi-head self-attention to obtain a second post-disturbance phenotype feature; inputting the second post-disturbance phenotype feature into a feedforward network to obtain a third post-disturbance phenotype feature; and projecting the third post-disturbance phenotype feature to the gene dimension to obtain the target transcriptome expression vector.

[0070] In some embodiments, during the training process of the expression phenotype alignment module, a third loss function is designed to optimize the hyperparameters of the expression phenotype alignment module; the third loss function is represented by the following formula:

[0071] wherein, represents the third loss function; represents the target transcriptome expression vector; represents the pre-disturbance transcriptome expression profiles; represents the induced gene expression change vector.

[0072] In some embodiments, the first loss function is represented by the following formula:

[0073] The second loss function is represented by the following formula:

[0074]

[0075] wherein, represents the first loss function; represents the post-disturbance transcriptome expression profiles; represents the target transcriptome expression vector; represents the second loss function; represents the induced gene expression change vector; represents a difference between the post-disturbance transcriptome expression profiles and the pre-disturbance transcriptome expression profiles .

[0076] In some embodiments, the process of preprocessing comprises: normalizing and log1p transforming the pre-disturbance transcriptomic expression profile, the target sleep-aiding phenotype data, the post-disturbance phenotype data, and the post-disturbance transcriptomic expression profile; and converting the structure of the small molecule inducing the disturbance into an atomic sequence.

[0077] In some embodiments, based on the target expression change vector, a decoder is adopted to obtain a candidate small molecule structure sequence; and the following formula is adopted to represent:

[0078] wherein, represents the candidate small molecule structure sequence; represents the target expression change vector; represents noise.

[0079] The deep learning sleep-aiding molecule drug design system based on disturbance phenotype inversion according to the embodiments of the present application can correspond to execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module of the deep learning sleep-aiding molecule drug design system based on disturbance phenotype inversion are respectively for realizing Figure 1 the corresponding processes of each method in the above formula, for the sake of brevity, will not be described here.

[0080] Referring to Figure 3 the structural diagram of the computer device, the computer device comprises a processor, a communication interface, and a computer readable storage medium. The processor, the communication interface, and the computer readable storage medium can be connected through a bus or other means. The communication interface is used to receive and send data. The computer readable storage medium can be stored in the memory of the computer device, and the computer readable storage medium is used to store a computer program, the computer program comprises program instructions, and the processor is used to execute the program instructions stored in the computer readable storage medium. The processor (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is particularly suitable for loading and executing one or more instructions to realize the corresponding steps in the deep learning sleep-aiding molecule drug design method based on disturbance phenotype inversion embodiment.

[0081] The embodiment provides a computer readable storage medium (Memory), which is a memory device in a computer device, and is used to store programs and data. It can be understood that the computer readable storage medium herein can include an internal storage medium in the computer device, and of course can include an extended storage medium supported by the computer device. The computer readable storage medium provides a storage space, and the storage space stores a processing system of the computer device. Also, stored in the storage space are one or more instructions adapted to be loaded and executed by the processor, which can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.

[0082] In an embodiment, the computer-readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned deep learning-based sleep-aiding molecule drug design method embodiment based on perturbation phenotype inversion.

[0083] The embodiment provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to make the computer device execute the corresponding steps in the above-mentioned deep learning-based sleep-aiding molecule drug design method embodiment based on perturbation phenotype inversion.

[0084] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0085] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The function of the device specified in one flow or multiple flows and / or blocks Figure 1 The function of the device specified in one flow or multiple flows and / or blocks

[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0088] Those of ordinary skill in the art can understand that all or part of the flow of the above-mentioned embodiment method can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the flow of the above-mentioned embodiment method. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0089] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion, characterized in that, include: We obtained the pre-perturbation transcriptome expression profile, post-perturbation phenotypic data, post-perturbation transcriptome expression profile, and induced perturbation small molecule structures, and performed preprocessing. The preprocessed perturbed phenotypic data The encoder is input to obtain the target transcriptome expression vector. Based on the target transcriptome expression vector and the perturbed transcriptome expression profile, a first loss function is designed to optimize the encoder's hyperparameters. The preprocessed induced perturbation small molecule structure is input into the decoder and mapped to the gene space to obtain the induced gene expression change vector. Based on the induced gene expression change vector, a second loss function is designed to optimize the decoder's hyperparameters. The target transcriptome expression vector and the induced gene expression change vector are input into the expression phenotype alignment module so that the difference between the target transcriptome expression vector and the induced gene expression change vector equals the unperturbed transcriptome expression profile, thus completing the alignment of the phenotypic path with the small molecule path. The pre-perturbation transcriptome expression profile and target sleep-aid phenotype data were acquired and preprocessed. Based on the preprocessed pre-perturbation transcriptome expression profile and target sleep-aid phenotype data, an encoder was used to obtain the target transcriptome expression vector. Based on the difference between the target transcriptome expression vector and the preprocessed pre-perturbation transcriptome expression profile, the target expression change vector was calculated. Based on the target expression change vector, a decoder was used to obtain candidate small molecule structural sequences.

2. The deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion according to claim 1, characterized in that, The preprocessed, perturbed phenotypic data is input into the encoder to obtain the target transcriptome expression vector; the method includes: The preprocessed perturbed phenotypic data are mapped to the same feature space and concatenated to obtain the first perturbed phenotypic feature. The first perturbation phenotypic feature is input into the multi-head self-attention to obtain the second perturbation phenotypic feature; The second perturbation phenotypic feature is input into the feedforward network to obtain the third perturbation phenotypic feature; The phenotypic features after the third perturbation are back-projected onto the gene dimension to obtain the target transcriptome expression vector.

3. The deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion according to claim 1, characterized in that, During the training of the phenotypic alignment module, a third loss function is designed to optimize the hyperparameters of the phenotypic alignment module; the third loss function is expressed by the following formula: in, Represents the third loss function; Represents the target transcriptome expression vector; This represents the pre-perturbation transcriptome expression profile; This represents a vector representing changes in induced gene expression.

4. The deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion according to claim 1, characterized in that, The first loss function is expressed by the following formula: The second loss function is expressed by the following formula: in, Represents the first loss function; This indicates the transcriptome expression profile after perturbation; Represents the target transcriptome expression vector; This represents the second loss function; Represents a vector of induced gene expression changes; Indicates the transcriptome expression profile after perturbation Compared with the pre-perturbation transcriptome expression profile The difference between them.

5. The deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion according to claim 1, characterized in that, The preprocessing process includes: Normalization and log1p transformation were performed on the pre-perturbation transcriptome expression profile, the target sleep-aid phenotype data, the post-perturbation phenotype data, and the post-perturbation transcriptome expression profile. The induced perturbation of small molecular structures is converted into atomic sequences.

6. The deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion according to claim 1, characterized in that, Based on the target expression change vector, a decoder is used to obtain candidate small molecule structure sequences; Expressed using the following formula: in, Indicates the candidate small molecule structural sequence; This represents the target expression change vector; Indicates noise.

7. A deep learning-based molecular drug design system for sleep aids based on perturbation phenotypic inversion, characterized in that, include: The training module is configured to: acquire pre-perturbation transcriptome expression profiles, post-perturbation phenotypic data, post-perturbation transcriptome expression profiles, and induced perturbation small molecule structures, and perform preprocessing; and process the preprocessed post-perturbation phenotypic data. The encoder is input to obtain the target transcriptome expression vector. Based on the target transcriptome expression vector and the perturbed transcriptome expression profile, a first loss function is designed to optimize the encoder's hyperparameters. The preprocessed induced perturbation small molecule structure is input into the decoder and mapped to the gene space to obtain the induced gene expression change vector. Based on the induced gene expression change vector, a second loss function is designed to optimize the decoder's hyperparameters. The target transcriptome expression vector and the induced gene expression change vector are input into the expression phenotype alignment module so that the difference between the target transcriptome expression vector and the induced gene expression change vector equals the unperturbed transcriptome expression profile, thus completing the alignment of the phenotypic path with the small molecule path. The inference module is configured to: acquire the pre-perturbation transcriptome expression profile and target sleep-aid phenotype data, and preprocess them; based on the preprocessed pre-perturbation transcriptome expression profile and target sleep-aid phenotype data, use an encoder to obtain the target transcriptome expression vector; calculate the target expression change vector based on the difference between the target transcriptome expression vector and the preprocessed pre-perturbation transcriptome expression profile; and use a decoder based on the target expression change vector to obtain candidate small molecule structure sequences.

8. A computer device, characterized in that, A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the steps of the deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and to execute the steps of the deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps in the deep learning-based method for designing sleep-aid molecular drugs based on perturbation phenotypic inversion as described in any one of claims 1-6.