A protein structure alignment method and device based on a coherent ising machine

By constructing a two-dimensional mesh graph and using a coherent Ising machine system to search for the ground state of the Hamiltonian, the NP-hard challenge in the protein structure alignment problem is solved, achieving a fast and accurate global optimal solution, which is applicable to the alignment of large-scale protein datasets.

CN120388602BActive Publication Date: 2025-12-26PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510458909.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-12-26
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing protein structure alignment methods have limitations in terms of high computational resource consumption, alignment quality, and processing speed when processing large protein datasets. In particular, traditional methods struggle to effectively solve NP-hard problems when dealing with complex protein structures that require high-precision alignment.

Method used

A coherent Ising machine-based approach is adopted to generate a protein structure alignment map by constructing a two-dimensional mesh graph and using the coherent Ising machine system to search for the ground state of the Hamiltonian. The parallelism and speed of the coherent Ising machine system are used to solve the protein structure alignment problem.

Benefits of technology

It enables fast and accurate finding of the global optimum on large-scale protein datasets, improving computational efficiency and alignment quality, and is suitable for handling large-scale parallel pairing problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388602B_ABST
    Figure CN120388602B_ABST
Patent Text Reader

Abstract

The application provides a protein structure alignment method and device, the method comprises the following steps: firstly, obtaining two protein files to be aligned; then, for each protein file, determining a corresponding contact pair set; then, according to the two contact pair sets, generating a two-dimensional grid graph, and establishing a connection edge between the grid vertices which can form an effective alignment; then, using a coherent Ising machine system, searching the solution space represented by the two-dimensional grid graph, and solving the ground state of Hamiltonian H; finally, generating a protein structure alignment graph according to the ground state of Hamiltonian H. The protein structure alignment method provided by the application can systematically search all possible alignment schemes, and ensure that the found solution is globally optimal. In addition, the application is based on a coherent Ising machine system, and the calculation process has inherent parallelism, which can more quickly explore the solution space, and is particularly suitable for processing large-scale parallel pairing problems in protein alignment problems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computational biology, and particularly relates to a protein structure alignment method and device based on a coherent Ising machine. BACKGROUND

[0002] In bioinformatics, protein structure alignment is a core task. Protein structure alignment problem can be formulated as: finding the optimal correspondence between the three-dimensional structures of two proteins, so that their spatial structures match as much as possible. The purpose is to identify and compare the spatial structure similarity between different proteins, which is crucial for understanding protein function, guiding drug design, and revealing biological evolutionary relationships. Protein structure alignment problem is a typical NP-hard problem, which means that as the size of the protein increases, the computing resources and time required to find the best or near-optimal structure alignment solution will increase dramatically.

[0003] For the protein structure alignment problem, existing solutions include:

[0004] 1. Distance matrix-based method: protein structure alignment is achieved by calculating and comparing the similarity of distance matrices between protein pairs. Although this method can provide accurate alignment in theory, it is computationally expensive and difficult to scale when dealing with large protein datasets.

[0005] 2. Dynamic programming-based method: Dynamic programming algorithms, such as Smith-Waterman and Needleman-Wunsch algorithms, can be used to find local or global alignment between protein structures, but their time and space complexity grow rapidly with the length of the protein sequence, limiting their application to large-scale problems.

[0006] 3. Machine learning-based method: With the development of machine learning technology, neural networks and other machine learning models are used to predict protein structure alignment. Although machine learning shows potential in predicting protein structure alignment, they usually require a large amount of labeled data for training, and may lack generalization ability in dealing with the complexity and diversity of protein structures.

[0007] 4. Branch and bound and branch and cut algorithms: Branch and bound and branch and cut algorithms are a method for solving integer programming problems, also applied to protein structure alignment problems, by systematically exploring the solution space and pruning infeasible branches to find the optimal solution. Although branch and bound and branch and cut algorithms can theoretically find the optimal solution, they may become very time-consuming in practice due to the huge solution space, especially for the NP-hard problem of protein structure alignment.

[0008] These methods have improved the efficiency and accuracy of protein structure alignment to varying degrees, but they still face the challenges posed by the NP-hard problem, especially when dealing with large protein datasets. These traditional methods have limitations in terms of computational resource consumption, alignment quality, and processing speed, particularly for complex protein structures requiring high-precision alignment. Therefore, a more efficient and accurate method applicable to large protein datasets is needed to solve the protein structure alignment problem. Summary of the Invention

[0009] To address the aforementioned problems, this invention provides a protein structure alignment method and apparatus based on a coherent Ising machine.

[0010] According to the first aspect, a protein structure alignment method based on the coherent Ising machine is provided, comprising the following steps:

[0011] S1. Obtain two protein files that need to be aligned. Each protein file includes the spatial coordinates of multiple atoms in the protein. The multiple atoms form multiple amino acid residues, which are arranged in the order of the original sequence.

[0012] S2. For each protein file, identify the corresponding multiple contact pairs, where each contact pair includes two amino acid residues.

[0013] S3. Generate a two-dimensional mesh diagram based on the multiple contact pairs corresponding to the two protein files, where rows represent contact pairs corresponding to one protein file and columns represent contact pairs corresponding to the other protein file.

[0014] In the two-dimensional mesh diagram, connecting edges are established between mesh vertices that can form effective alignment. The effective alignment must satisfy the following condition: after alignment, the amino acid residues are arranged in the order of the original sequence.

[0015] S4. Using a coherent Ising machine system, search the solution space represented by the two-dimensional mesh diagram to find the ground state of the Hamiltonian H, where the Hamiltonian H is:

[0016]

[0017] in, Representing the Line number Whether the grid vertices of the column are selected. for The weight, Representing the Line number The grid vertices of the column and the first Line number Are there connecting edges between the grid vertices of the column? K1 is The coefficient.

[0018] S5、generating a protein structure alignment graph according to the ground state of the Hamiltonian H.

[0019] In some embodiments, the distance between the two amino acid residues is less than a certain threshold.

[0020] In some embodiments, the searching of the solution space represented by the two-dimensional grid graph by the coherent Ising machine system to solve the ground state of the Hamiltonian H specifically comprises:

[0021] According to the Hamiltonian H, initialize the coherent Ising machine, and start the coherent Ising machine system evolution process;

[0022] During the coherent Ising machine system evolution process, the phase and intensity of the optical field in the coherent Ising machine system are measured multiple times, and the parameters of the coherent Ising machine system are adjusted in real time according to the measurement results, and finally the collective oscillation mode is obtained, and the quantum spin state under this mode corresponds to the ground state of the Hamiltonian H.

[0023] In some embodiments, the coherent Ising machine system comprises:

[0024] An optical parametric oscillator network is used to simulate the spin interaction in the Ising model.

[0025] A phase-sensitive amplifier is used to amplify signals of a certain phase.

[0026] A field programmable gate array is used to control the optical parametric oscillator network.

[0027] A phase / intensity measurer is used to measure the phase and intensity of the optical field in the cavity.

[0028] An optical modulator is used to change the phase difference between the light beams, thereby simulating the interaction strength between different spin states.

[0029] A beam splitter is used to produce interference effects to simulate the connection relationship between spins in the Ising model.

[0030] An optical fiber is used as a transmission medium for optical signals.

[0031] According to a second aspect, a protein structure alignment device is provided, characterized in that it comprises:

[0032] An acquisition module is configured to acquire two protein files that need to be aligned, each protein file including the spatial coordinates of a plurality of atoms in a protein, the plurality of atoms forming a plurality of amino acid residues, and the amino acid residues being arranged in the order of the original sequence.

[0033] a contact pair generation module configured to determine, for each protein file, a respective plurality of contact pairs, wherein each contact pair comprises two amino acid residues.

[0034] a graph generation module configured to generate a two-dimensional grid graph according to the two sets of contact pairs of the two protein files, wherein rows represent contact pairs of one protein file and columns represent contact pairs of the other protein file; and connecting edges are established between grid vertices in the two-dimensional grid graph that can form an effective alignment, which requires that, after the alignment is completed, the amino acid residues are arranged in the order of the original sequence.

[0035] a solution module configured to search a solution space represented by the two-dimensional grid graph using a coherent Ising machine system to solve a ground state of a Hamiltonian H, the Hamiltonian H being:

[0036]

[0037] wherein, represents whether a grid vertex in the i-th row and the j-th column is selected, is a weight of the grid vertex, represents whether there is a connecting edge between the grid vertex in the i-th row and the j-th column and the grid vertex in the i’-th row and the j’-th column, and K1 is a coefficient of the grid vertex.

[0038] an output module configured to generate a protein structure alignment graph according to the ground state of the Hamiltonian H.

[0039] In some embodiments, the contact pairs comprise two amino acid residues that are non-adjacent and have a distance less than a first threshold value.

[0040] In some embodiments, the searching of the solution space represented by the two-dimensional grid graph using the coherent Ising machine system to solve the ground state of the Hamiltonian H specifically comprises:

[0041] initializing the coherent Ising machine according to the Hamiltonian H and starting an evolution process of the coherent Ising machine system.

[0042] In the evolution process of the coherent Ising machine system, the phase and intensity of the light field in the coherent Ising machine system are measured multiple times, and the parameters of the coherent Ising machine system are adjusted in real time according to the measurement results, and finally a collective oscillation mode is obtained, and the quantum spin state in this mode corresponds to the ground state of the Hamiltonian H.

[0043] In some embodiments, the coherent Ising machine system comprises: ​​​​​​​​

[0044] Optical parametric oscillator network for simulating spin interactions in Ising model.

[0045] Phase sensitive amplifier for amplifying signals of specific phase.

[0046] Field programmable gate array for controlling optical parametric oscillator network.

[0047] Phase / intensity measurer for measuring phase and intensity of intra-cavity optical field.

[0048] Optical modulator for changing phase difference between optical beams to simulate interaction strength between different spin states.

[0049] Beam splitter for generating interference effect to simulate connection relationship between spins in Ising model.

[0050] Optical fiber for serving as transmission medium of optical signals.

[0051] The protein structure alignment method provided by the present application can systematically search all possible alignment schemes, and ensure that the solution found is globally optimal. In addition, the present application is based on a coherent Ising machine system, and the calculation process has inherent parallelism, can handle the interaction of multiple spin states at one time, can more quickly explore the solution space, and improve the solving speed, and is particularly suitable for processing large-scale parallel pairing problems in protein alignment problems. Therefore, the method provided by the present application provides a more rapid and accurate solution for the protein structure alignment problem. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description are briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0053] Figure 1 A protein structure alignment method flow provided by the present application is shown;

[0054] Figure 2 A coherent Ising machine system structure provided by the embodiment of the present application is shown;

[0055] Figure 3 A protein structure alignment device structure provided by the present application is shown. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme in the embodiments of the present application will be described below with reference to the drawings.

[0057] In order to be able to quickly and accurately provide results when facing the structure alignment problem of large-scale or complex structure proteins, the present application provides a protein structure alignment method based on a coherent Ising machine (CIM).

[0058] The coherent Ising machine is a computing platform based on quantum optics simulation of the Ising model, which has the advantages of fast computing speed and strong scalability, and is suitable for processing various large-scale combinatorial optimization problems, especially in processing large-scale quantum bits and NP-hard problems. As a computing model, the unique parallelism of the coherent Ising machine provides potential advantages for solving NP-hard problems. CIM can simultaneously explore multiple possible solutions in the coherent evolution of quantum states, thereby showing potential beyond traditional algorithms in the protein structure alignment problem. This method is expected to significantly improve the computational efficiency of alignment, and at the same time, handle larger-scale protein data sets while maintaining high accuracy. Therefore, the technical solution of the present application aims to utilize the computing advantages of CIM to provide an innovative and efficient solution for the protein structure alignment problem, in order to achieve breakthrough progress in the field of bioinformatics.

[0059] Based on this, the present application first designs the Hamiltonian of the coherent Ising machine to convert the protein structure alignment problem into an Ising model optimization problem, so that it can be efficiently calculated; then uses the coherent Ising machine to solve the protein structure alignment problem, ensuring that the solution found is globally optimal.

[0060] The protein structure alignment method provided by the present application will be described below with reference to the accompanying drawings, the flowchart of the method is as shown in Figure 1 The method includes the following steps:

[0061] S1, obtaining two protein files to be aligned, each protein file including the spatial coordinates of multiple atoms in the protein, the multiple atoms forming multiple amino acid residues, and the amino acid residues being arranged in the order of the original sequence.

[0062] It can be understood that a protein is a molecular chain formed by multiple amino acid residues connected by peptide bonds in a certain order, and this order is the order of the original sequence of the protein, which is determined by genetic information. Each amino acid residue is formed by atom arrangement and bonding, and the atom is the most basic unit of protein.

[0063] S2, determining a plurality of contact pairs corresponding to each protein file, wherein each contact pair includes two amino acid residues.

[0064] It is understood that in the three-dimensional structure of a protein, if the distance between any atoms of two amino acid residues is less than a certain threshold value (usually 5 to 10 angstroms), the two residues are considered to form a contact pair. Contact pairs can be used to assess the spatial proximity between different amino acid residues in a protein molecule. Based on this, the similarity between the three-dimensional structures of two proteins can be determined, thereby solving the protein structure alignment problem.

[0065] In one embodiment of this step, for each protein file, all the amino acid residues contained therein can be exhaustively paired, so that for any pair of amino acid residues, if it is determined that the spatial distance between the two amino acid residues is less than a certain threshold value, the pair of amino acid residues is classified as a contact pair.

[0066] In another embodiment of this step, a corresponding structure graph can be constructed according to each protein file, where each vertex represents an amino acid residue, and each edge connects the vertices corresponding to a pair of amino acid residues with a distance less than a certain threshold value, i.e. each edge represents a contact pair.

[0067] In some embodiments, the two amino acid residues are non-adjacent and have a distance less than a certain threshold value.

[0068] Where adjacent means that the two amino acid residues are connected by a peptide bond, hydrogen bond, etc. In some cases, researchers may only be interested in the contact between non-adjacent amino acid residues, and only consider a pair of non-adjacent amino acid residues with a distance less than a certain threshold value as a contact pair.

[0069] S3, generating a two-dimensional grid graph according to the plurality of contact pairs corresponding to the two protein files respectively, wherein the rows represent the contact pairs corresponding to one protein file, and the columns represent the contact pairs corresponding to the other protein file.

[0070] This step constructs a two-dimensional grid graph G, in which the rows and columns of the grid represent the contact pairs in the two proteins, and the vertices of the grid formed by the intersection of each row and each column represent a potential alignment of the contact pairs of the two proteins.

[0071] For the sake of distinction, the contact pairs corresponding to one of the protein files are referred to as first contact pairs, and the contact pairs corresponding to the other protein file are referred to as second contact pairs. Thus, the grid vertex located at the ith row and jth column in the two-dimensional grid graph corresponds to the potential alignment of the ith first contact pair and the jth second contact pair. It is understood that "first" in "first contact pair" and "second" in "second contact pair" are for the sake of distinction, and do not have other limiting functions such as ordering.

[0072] For example:

[0073] The ith row in the graph G corresponds to a contact pair (a1, a2) in protein A; the jth column corresponds to a contact pair (b1, b2) in protein B, where a1, a2, b1, b2 represent amino acid residues respectively. Then, the intersection of the ith row and the jth column in the graph G forms a grid vertex, which represents the alignment of the two contact pairs (a1, a2) and (b1, b2).

[0074] The above example explains the meaning of the grid vertex in the graph G. The following example further explains the establishment and meaning of the connecting edge in the graph G. The connecting edge is established between the grid vertices that can form an effective alignment in the two-dimensional grid graph. The effective alignment needs to satisfy that, after the alignment, the amino acid residues are arranged in the original sequence.

[0075] As described above, each grid vertex in the graph G represents a potential alignment of the two contact pairs. By adding connecting edges between these grid vertices, the effective alignment that still satisfies the amino acid alignment feasibility rule after the two potential alignments are simultaneously effective can be screened out. The condition of the effective alignment corresponds to the order in the amino acid alignment feasibility rule. Under the premise that all the amino acid alignments comply with the order, the alignment naturally complies with the exclusivity, i.e., an amino acid cannot be aligned with multiple amino acids in another protein. For example:

[0076] The ith row and the kth row in the graph G correspond to the contact pairs (a1, a2) and (a3, a4) in protein A respectively; the jth column and the lth column correspond to the contact pairs (b1, b2) and (b3, b4) in protein B respectively, where a1, a2, a3, a4, b1, b2, b3, b4 represent amino acid residues respectively. Then, the intersection of the ith row and the jth column in the graph G forms a grid vertex 1, which represents the alignment of the two contact pairs (a1, a2) and (b1, b2); the intersection of the kth row and the lth column in the graph G forms a grid vertex 2, which represents the alignment of the two contact pairs (a3, a4) and (b3, b4).

[0077] If the amino acid residues are still arranged in the original sequence after the alignment of the grid vertex 1 and the grid vertex 2, it indicates that the two grid vertices can form an effective alignment, and the connecting edge needs to be established between them. Otherwise, it indicates that the two grid vertices cannot form an effective alignment, and the connecting edge cannot be established between them.

[0078] The above example explains the establishment and meaning of the connecting edge in the graph G. As can be seen from the above, if there is a connecting edge between any two grid vertices in a set of grid vertices in the graph G, it means that the alignment represented by any pair of vertices in the set is feasible, and the alignment represented by the entire set of vertices is also feasible. This set corresponds to the solution of the protein structure alignment problem. In this way, the protein structure alignment problem is transformed into the maximum clique problem of the graph G.

[0079] Next, by constructing the Hamiltonian H of the graph G, the ground state of the Hamiltonian H is solved by using the CIM to solve this maximum clique problem, that is:

[0080] S4, using a coherent Ising machine system, searching the solution space represented by the two-dimensional grid graph, solving the ground state of the Hamiltonian H, the Hamiltonian H is:

[0081] (1)

[0082] Wherein the meaning and possible values of each parameter are as follows:

[0083] represent the grid vertex of the first row and the first column is selected or not, whether the grid vertex is selected corresponds to the node spin state in the coherent Ising machine. For example, it can be set to spin +1, indicating that this grid vertex is selected.

[0084] is the weight of , which can be 1.

[0085] represent whether there is a connection edge between the grid vertex of the first row and the first column and the grid vertex of the first row and the first column. For example, take 1 when there is a connection edge, and take 0 when there is no connection edge.

[0086] K1 is the coefficient of , which can be set to any positive number, such as 1.

[0087] In some embodiments, the Hamiltonian H is solved by using a coherent Ising machine system to search the solution space represented by the two-dimensional grid graph, and specifically includes:

[0088] According to the Hamiltonian H, initialize the coherent Ising machine, and start the coherent Ising machine system evolution process;

[0089] During the evolution process of the coherent Ising machine system, the phase and intensity of the light field in the coherent Ising machine system are measured multiple times, and the parameters of the coherent Ising machine system are adjusted in real time according to the measurement results, and finally the collective oscillation mode is obtained. The quantum spin state under this mode corresponds to the ground state of the Hamiltonian H.

[0090] From the above, the ground state of the Hamiltonian H can be obtained by using the coherent Ising machine system.

[0091] S5, generating a protein structure alignment graph according to the ground state of the Hamiltonian H.

[0092] The ground state of the Hamiltonian H corresponds to the maximum clique of the graph G, and the alignment graph of the grid vertices in the maximum clique inversely coded into amino acid residues is the result of the alignment of the two protein structures.

[0093] The above is the protein structure alignment method based on the coherent Ising machine provided by the application, and the coherent Ising machine system used in the method will be briefly introduced below.

[0094] The coherent Ising machine system uses a doubly resonant optical parametric oscillator (DOPO) to realize artificial spins, and realizes the enhancement of light signals with specific phases by placing a phase sensitive amplifier (PSA) in the optical cavity. The PSA is an optical amplifier based on optical parametric amplification, which can effectively amplify the 0 and π phase components relative to the pump phase. Therefore, the DOPO only uses 0 or π phase above the oscillation threshold; therefore, the discrete phase state can be used to represent the Ising spin state. The interaction between DOPO pulses is realized using a measurement feedback technique, which repeatedly measures the feedback modulation process in the cavity, while increasing the pump amplitude from 0, and finally obtains a "strongest" collective oscillation mode much higher than the threshold, which corresponds to the best solution of the given Ising problem.

[0095] The schematic diagram of the coherent Ising machine system is shown in Figure 2 , which includes:

[0096] (1) Optical parametric oscillator network, used to simulate the spin interaction in the Ising model.

[0097] (2) Phase sensitive amplifier, used to amplify signals with specific phases.

[0098] (3) Field-programmable gate array (FPGA), used to control the optical parametric oscillator network.

[0099] The FPGA can be programmed to configure its internal circuit, which can be used to adjust the parameters of the coherent Ising machine system in real time according to the measurement results during the evolution of the coherent Ising machine system.

[0100] (4) Phase / intensity measurer, used to measure the phase and intensity of the light field in the cavity.

[0101] (5) Optical modulator, used to change the phase difference between light beams, thereby simulating the interaction strength between different spin states.

[0102] (6) Beam splitter, used to produce interference effects to simulate the connection relationship between spins in the Ising model.

[0103] (7) Optical fiber, used as a transmission medium for optical signals.

[0104] From the above, compared with the prior art, the protein structure alignment method provided by the present application has the following beneficial effects:

[0105] 1. Guarantee of global optimal solution: The present method can systematically search all possible alignment schemes, ensuring that the solution found is globally optimal. This is particularly important in the NP-hard protein alignment problem, because many heuristic algorithms can quickly find approximate solutions, but cannot guarantee the optimality of the solution.

[0106] 2. Acceleration potential: The coherent Ising machine can explore the solution space more quickly due to the use of coherence effects, improving the speed of solution, especially for large graph clique problems. Traditional protein alignment algorithms, such as dynamic programming and heuristic algorithms, usually take a long time to solve large-scale graphs. In addition, the calculation process of the coherent Ising machine has inherent parallelism, and can handle the interaction of multiple spin states at once, so it is suitable for handling large-scale parallel pairing in protein alignment problems, further improving the computational efficiency.

[0107] In summary, compared with traditional protein alignment solving methods, the protein structure alignment method based on the coherent Ising machine of the present application solves the problem that heuristic algorithms may consume a long time but cannot derive a standard solution, providing a faster and more accurate solution for protein structure alignment problems, especially for large protein structure alignment.

[0108] The present application also provides a protein structure alignment device 300, a schematic diagram of which is shown in Figure 3 , comprising:

[0109] An acquisition module 301 is configured to acquire two protein files that need to be aligned, each protein file including the spatial coordinates of a plurality of atoms in a protein, the plurality of atoms forming a plurality of amino acid residues, and the amino acid residues being arranged in the order of the original sequence.

[0110] A contact pair generation module 302 is configured to determine, for each protein file, a corresponding set of contact pairs, wherein each contact pair includes two amino acid residues.

[0111] A graph generation module 303 is configured to generate a two-dimensional grid graph according to the two sets of contact pairs of the two protein files, wherein the rows represent the contact pairs in the set of contact pairs corresponding to one protein file, and the columns represent the contact pairs in the set of contact pairs corresponding to the other protein file; the connecting edges are established between the grid vertices that can form an effective alignment in the two-dimensional grid graph, and the effective alignment needs to satisfy that after the alignment is completed, the amino acid residues are arranged in the order of the original sequence.

[0112] a solving module 304, configured to search a solution space of the two-dimensional grid graph representation by using a coherent Ising machine system, and solve a ground state of a Hamiltonian H, the Hamiltonian H being:

[0113]

[0114] wherein, represents whether a grid vertex in the i-th row and the j-th column is selected, is a weight of the i-th row and the j-th column, represents whether there is a connecting edge between a grid vertex in the i-th row and the j-th column and a grid vertex in the i'-th row and the j'-th column, and K1 is a coefficient of the i-th row and the j-th column; an output module 305, configured to generate a protein structure alignment graph according to the ground state of the Hamiltonian H. It should be noted that the apparatus described above can perform the aforementioned protein structure alignment method, and the functions of each module can be referred to the aforementioned description of the method, and will not be described herein. In the description of the embodiments of the present application, the words "exemplary", "for example", or "for instance" are used to mean serving as an example, instance, or illustration. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. In fact, the words "exemplary", "for example", or "for instance" are used to represent the relevant concept in a specific manner. In the description of the embodiments of the present application, the term "and / or" is merely used to represent the association relationship of the associated objects, and means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, B alone, and A and B simultaneously. In addition, unless otherwise specified, the term "multiple" means two or more.

[0115] In addition, the terms "include", "contain", "have", and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0116]

[0117] In the description of the embodiments of the present application, the words "exemplary", "for example", or "for instance" are used to mean serving as an example, instance, or illustration. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. In fact, the words "exemplary", "for example", or "for instance" are used to represent the relevant concept in a specific manner.

[0118] In the description of the embodiments of the present application, the term "and / or" is merely used to represent the association relationship of the associated objects, and means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, B alone, and A and B simultaneously. In addition, unless otherwise specified, the term "multiple" means two or more.

[0119] In addition, the terms "include", "contain", "have", and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0120] ​​​​​The above detailed description of the specific embodiments of the present application has been given to illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.

Claims

1. A method for aligning protein structures, characterized in that, The method comprises the following steps: obtaining two protein files to be aligned, each protein file comprising spatial coordinates of a plurality of atoms in a protein, the plurality of atoms forming a plurality of amino acid residues, the amino acid residues being arranged in the order of original sequences; for each protein file, determining a plurality of contact pairs corresponding thereto, wherein each contact pair comprises two amino acid residues; generating a two-dimensional grid graph according to the plurality of contact pairs corresponding to the two protein files respectively, wherein rows represent the contact pairs corresponding to one protein file, and columns represent the contact pairs corresponding to the other protein file; establishing a connection edge between grid vertices capable of forming an effective alignment in the two-dimensional grid graph, the effective alignment satisfying that, after the alignment is completed, the amino acid residues are arranged in the order of original sequences; searching, by using a coherent Ising machine system, a solution space represented by the two-dimensional grid graph to solve a ground state of a Hamiltonian H, the Hamiltonian H being: wherein, represents the first row and the first column of the grid, is the weight of the first row and the first column of the grid, represents whether there is a connecting edge between the grid vertex of the first row and the first column and the grid vertex of the first row and the second column of the grid, and K1 is the coefficient of the first generating a protein structure alignment graph according to the ground state of the Hamiltonian H.

2. The method of claim 1, wherein, The two amino acid residues are non-adjacent and have a distance less than a specific threshold.

3. The method of claim 1, wherein, The searching, by using the coherent Ising machine system, the solution space represented by the two-dimensional grid graph to solve the ground state of the Hamiltonian H specifically comprises: initializing the coherent Ising machine according to the Hamiltonian H, and starting an evolution process of the coherent Ising machine system; during the evolution process of the coherent Ising machine system, measuring the phase and intensity of the optical field in the coherent Ising machine system multiple times, and adjusting the parameters of the coherent Ising machine system in real time according to the measurement results, so as to finally obtain a collective oscillation mode, and the quantum spin state in the mode corresponds to the ground state of the Hamiltonian H.

4. The method of claim 1, wherein, The coherent Ising machine system comprises: an optical parametric oscillator network for simulating spin interactions in the Ising model; a phase-sensitive amplifier for amplifying signals of specific phases; a field programmable gate array for controlling the optical parametric oscillator network; a phase / intensity measurer for measuring the phase and intensity of the optical field in the cavity; an optical modulator for changing the phase difference between light beams, thereby simulating the interaction strength between different spin states; a beam splitter for generating interference effects to simulate the connection relationship between spins in the Ising model; optical fibers serving as transmission media for optical signals.

5. A protein structure alignment apparatus, characterized by, The method comprises: an obtaining module for obtaining two protein files to be aligned, each protein file comprising spatial coordinates of a plurality of atoms in a protein, the plurality of atoms forming a plurality of amino acid residues, the amino acid residues being arranged in the order of original sequences; a contact pair generation module for, for each protein file, determining a plurality of contact pairs corresponding thereto, wherein each contact pair comprises two amino acid residues; a graph generation module for generating a two-dimensional grid graph according to the plurality of contact pairs corresponding to the two protein files respectively, wherein rows represent the contact pairs corresponding to one protein file, and columns represent the contact pairs corresponding to the other protein file; a connection edge is established between grid vertices capable of forming an effective alignment in the two-dimensional grid graph, the effective alignment satisfying that, after the alignment is completed, the amino acid residues are arranged in the order of original sequences; A solving module is configured to search the solution space of the two-dimensional grid graph representation by using a coherent Ising machine system, and solve the ground state of a Hamiltonian H, where the Hamiltonian H is: wherein, representing the first row and the first column of the grid vertices, the weight of , the weight of representing whether there is a connecting edge between the first row and the first column of the grid vertices and the first row and the first column of the grid vertices, K1 is the coefficient of ; an output module, configured to generate a protein structure alignment graph according to the ground state of the Hamiltonian H.

6. The apparatus of claim 5, wherein, The two amino acid residues are non-adjacent and the distance between them is less than a certain threshold.

7. The apparatus of claim 5, wherein, The searching the solution space of the two-dimensional grid graph representation by using a coherent Ising machine system, and solving the ground state of a Hamiltonian H specifically includes: According to the Hamiltonian H, initializing the coherent Ising machine, and starting the coherent Ising machine system evolution process; During the coherent Ising machine system evolution process, the phase and intensity of the light field in the coherent Ising machine system are measured multiple times, and the coherent Ising machine system parameters are adjusted in real time according to the measurement results, and finally the collective oscillation mode is obtained, and the quantum spin state in this mode corresponds to the ground state of the Hamiltonian H.

8. The apparatus of claim 5, wherein, The coherent Ising machine system includes: An optical parametric oscillator network is configured to simulate the spin interaction in the Ising model; A phase-sensitive amplifier is configured to amplify signals of a specific phase; A field programmable gate array is configured to control the optical parametric oscillator network; A phase / intensity measurer is configured to measure the phase and intensity of the light field in the cavity; An optical modulator is configured to change the phase difference between the light beams, thereby simulating the interaction strength between different spin states; A beam splitter is configured to generate an interference effect to simulate the connection relationship between spins in the Ising model; An optical fiber is used as a transmission medium for optical signals.

Citation Information

Patent Citations

  • Method and apparatus for evolutionary data driven design of protein and other sequence defined biomolecules using machine learning

    CN114651064A

  • Protein structure alignment determination method and related device

    CN118841064A