Methods for antibody complementarity determining sequence alignment and apparatuses and electronic devices therefor
By employing a fully automated antibody complementarity determinant alignment method, the problems of time-consuming and labor-intensive processes and lack of target information in existing technologies are solved. This method achieves efficient antibody sequence alignment and target information extraction, reducing experimental costs and infringement risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2021-10-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies cannot efficiently and automatically compare the complementary determinant sequences of antibodies, and the comparison results lack target information, which is time-consuming and labor-intensive, making it difficult to meet the needs of antibody drug treatment for new diseases.
By acquiring the original antibody sequence, processing and numbering it, comparing it with the antibody sequence library, extracting information on the complementary determinant cluster region, determining the differences, and outputting the target information, the comparison is fully automated using a programming language.
It reduces experimental costs and human operation time, provides more comprehensive information on the differences in complementary determinant cluster regions, reduces the risk of infringement, and improves comparison efficiency.
Smart Images

Figure CN116072211B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological information, and in particular to a method, apparatus and electronic device for aligning antibody complementary determinant cluster sequences. Background Technology
[0002] The complementarity determinants of an antibody are important functional regions of the antibody's variable region, determining whether the antibody can effectively bind to the antigen, thereby producing a targeting effect.
[0003] Current techniques can only compare the full sequence similarity of antibodies. Alignment of complementarity determinants requires manual extraction and comparison, which is time-consuming and labor-intensive, and the results lack target information. However, complementarity determinant protein sequences are often key areas of patent protection, and different antibody targets offer the possibility of antibody drugs treating new diseases. Summary of the Invention
[0004] In view of this, the present invention aims to at least partially solve one of the problems in the related art. Therefore, the object of this application is to provide a method, apparatus, and electronic device for aligning antibody complementarity determinant cluster sequences.
[0005] This application provides a method for aligning antibody complementarity determinant cluster (CDR) sequences. The method includes: acquiring at least one original antibody sequence and processing it to obtain antibody sequence information to be aligned; comparing the antibody sequence information to be aligned with an antibody sequence library to obtain aligned sequence information; numbering the original antibody sequence and the aligned sequence and extracting CDR region information; determining the CDR differences between the original antibody sequence and the aligned sequence; extracting target information from the target file title based on the aligned sequence information; and outputting a difference result based on the antibody sequence information to be aligned, the aligned sequence information, the CDR regions, the CDR differences, and the target information.
[0006] In some embodiments, the method further includes: downloading antibody sequences via big data to establish the antibody sequence library.
[0007] In some embodiments, obtaining at least one original antibody sequence to process and obtain antibody sequence information to be compared includes: obtaining at least one original antibody sequence input by a user; determining the original sequence tag of each original antibody sequence to obtain input sequence information; and processing the input sequence information to obtain the antibody sequence information to be compared.
[0008] In some embodiments, processing the input sequence information to obtain the antibody sequence information to be compared includes: segmenting the original antibody sequence according to a preset symbol; identifying the segmented original antibody tags and corresponding original antibody sequences; and constructing a list of one-to-one correspondences between the original sequence tags, original antibody sequences, and interface input text to obtain the antibody sequence information to be compared.
[0009] In some implementations, the step of comparing the antibody sequence information to be compared with an antibody sequence library to obtain the alignment sequence information includes: calling an analysis tool interface based on similarity comparison, inputting the interface input text to construct a request form and submitting it; obtaining the task number after the request form is submitted; constructing a pull form based on the task number to pull the alignment results; determining the alignment sequence information based on the alignment results; and replacing the non-sequence characters of the alignment sequence in the alignment sequence information.
[0010] In some implementations, the step of constructing a pull form to retrieve comparison results based on the task number includes: after a preset time has elapsed, constructing a pull form to retrieve comparison results based on the task number, wherein the preset time is related to the number of the original antibody sequences.
[0011] In some implementations, the replacement process for the alignment sequence includes replacing non-sequence characters in the alignment sequence with empty strings.
[0012] In some embodiments, numbering the original antibody sequence and the aligned sequence and extracting complementarity determinant cluster (CDR) region information includes: numbering the original antibody sequence and the aligned sequence after replacement according to a preset antibody numbering method, wherein the preset antibody numbering method includes at least one of Kabat, Chothia, IMGT, Gelfand, Aho, and Martin; and extracting CDR regions to construct a list of numbered residues in the CDRs of the original antibody sequence and the aligned sequence in a "number-residue" manner.
[0013] In some embodiments, determining the difference in complementary determinant clusters between the original antibody sequence and the aligned sequence includes: comparing a list of numbered residues in the complementary determinant clusters of the original antibody sequence and the aligned sequence; obtaining the sequence of each complementary determinant cluster region of the aligned sequence, and the number of numbered residues in each complementary determinant cluster region that differ between the aligned sequence and the original antibody sequence.
[0014] In some implementations, the alignment sequence information includes an alignment sequence identity code, and the step of extracting target information from the target file title based on the alignment sequence information includes: extracting the corresponding target file title based on the alignment sequence identity code; and extracting the target information from the target file title using Gnormplus.
[0015] This application also discloses an apparatus for aligning antibody complementarity determinant cluster (CDR) sequences. The apparatus includes: an acquisition module, a comparison module, a first extraction module, a difference determination module, a second extraction module, and an output module. The acquisition module acquires at least one original antibody sequence for processing to obtain antibody sequence information to be aligned; the comparison module compares the antibody sequence information to be aligned with an antibody sequence library to obtain aligned sequence information; the first extraction module numbers the original antibody sequence and the aligned sequence and extracts CDR region information; the difference determination module determines the differences in CDRs between the original antibody sequence and the aligned sequence; the second extraction module extracts target information from the target file title based on the aligned sequence information; and the output module outputs a difference result based on the antibody sequence information to be aligned, the aligned sequence information, the CDR regions, the CDR differences, and the target information.
[0016] This application also provides an electronic device. The electronic device includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the method described in any of the above embodiments.
[0017] This application also provides a non-volatile computer-readable storage medium containing a computer program. When the computer program is executed by one or more processors, it implements the method described in any of the above embodiments.
[0018] This application constructs a fully automated method for aligning antibody complementarity determinants using a programming language, which can reduce experimental costs and the time cost of manually extracting and verifying the similarity of complementarity determinant sequences, and provide more comprehensive information on the differences in complementarity determinant region alignment.
[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0021] Figure 1 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0022] Figure 2 This is a schematic diagram of the structure of an antibody complementarity determinant cluster sequence alignment device according to certain embodiments of this application;
[0023] Figure 3 This is a schematic diagram of the tagging method for antibody complementarity determinant cluster sequence alignment in certain embodiments of this application;
[0024] Figure 4 This is a schematic diagram illustrating the differential results of the antibody complementarity determinant cluster sequence alignment method according to certain embodiments of this application;
[0025] Figure 5 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0026] Figure 6 This is a schematic diagram of the structure of an antibody complementarity determinant cluster sequence alignment device according to certain embodiments of this application;
[0027] Figure 7 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0028] Figure 8 This is a schematic diagram of the acquisition module in the apparatus for antibody complementarity determinant cluster sequence alignment according to certain embodiments of this application;
[0029] Figure 9 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0030] Figure 10 This is a schematic diagram of the processing unit in the acquisition module of the apparatus for antibody complementarity determinant cluster sequence alignment according to certain embodiments of this application;
[0031] Figure 11 This is a schematic diagram of the segmented antibody sequence information to be compared in the antibody complementarity determinant cluster sequence alignment method of certain embodiments of this application;
[0032] Figure 12 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0033] Figure 13 This is a schematic diagram of the structure of the comparison module in the antibody complementarity determinant cluster sequence alignment device according to certain embodiments of this application;
[0034] Figure 14 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0035] Figure 15This is a schematic diagram of the pull unit in the comparison module of the apparatus for antibody complementarity determinant cluster sequence alignment according to certain embodiments of this application;
[0036] Figure 16 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0037] Figure 17 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0038] Figure 18 This is a schematic diagram of the structure of the first extraction module of the apparatus for antibody complementarity determinant cluster sequence alignment according to certain embodiments of this application;
[0039] Figure 19 This is a schematic diagram of a list of CDR1 / CDR2 / CDR3 of a certain alignment sequence in some embodiments of this application, rearranged according to "number-residue" information;
[0040] Figure 20 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0041] Figure 21 This is a schematic diagram of the difference determination module of the apparatus for antibody complementarity determinant cluster sequence alignment in certain embodiments of this application;
[0042] Figure 22 This is a flowchart illustrating the method for aligning antibody complementarity determinant cluster sequences according to certain embodiments of this application.
[0043] Figure 23 This is a schematic diagram of the structure of an electronic device according to certain embodiments of this application;
[0044] Figure 24 This is a schematic diagram of the structure of a computer-readable storage medium according to certain embodiments of this application. Detailed Implementation
[0045] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0046] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0047] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0048] The following disclosure provides many different implementations or examples for carrying out different structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or reference letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various implementations and / or arrangements discussed.
[0049] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0050] Please see Figure 1 This application provides a method for aligning antibody complementarity determinant cluster sequences. The method includes:
[0051] 01: Obtain at least one raw antibody sequence to process and obtain the antibody sequence information to be compared;
[0052] 02: Compare the antibody sequence information to be aligned with the antibody sequence library to obtain the alignment sequence information;
[0053] 03: Number the original antibody sequence and the aligned sequence and extract information on the complementarity determinant cluster region;
[0054] 04: Determine the differences in complementary determinants between the original antibody sequence and the aligned sequence;
[0055] 05: Extract target information from the target file title based on the aligned sequence information;
[0056] 06: Output the difference results based on the antibody sequence information to be compared, the aligned sequence information, the complementarity determinant region, the complementarity determinant difference, and the target information.
[0057] Please see Figure 2 This application also provides an apparatus 10 for aligning antibody complementarity determinant cluster sequences. The apparatus 10 includes: an acquisition module 11, a comparison module 12, a first extraction module 13, a difference determination module 14, a second extraction module 15, and an output module 16.
[0058] Step 01 can be implemented by the acquisition module 11, step 02 by the comparison module 12, step 03 by the first extraction module 13, step 04 by the difference determination module 14, step 05 by the second extraction module 15, and step 06 by the output module 16. That is, the acquisition module 11 is used to acquire at least one original antibody sequence to process and obtain the antibody sequence information to be aligned; the comparison module 12 is used to compare the antibody sequence information to be aligned with the antibody sequence library to obtain aligned sequence information; the first extraction module 13 is used to number the original antibody sequence and the aligned sequence and extract the complementarity determinant cluster (CDR) region information; the difference determination module 14 is used to determine the CDR differences between the original antibody sequence and the aligned sequence; the second extraction module 15 is used to extract the target information in the target file title based on the aligned sequence information; and the output module 16 is used to output the difference results based on the antibody sequence information to be aligned, the aligned sequence information, the CDR region, the CDR difference, and the target information.
[0059] Specifically, the input at least one original antibody sequence may include a variable region (VL or VH) sequence of the light or heavy chain. At least one original antibody sequence can be one or more original antibody sequences. Understandably, the antibody variable region, also known as the FV region, is the most critical region for antigen binding. The antibody variable region contains complementarity determinants and other protein sequence regions, among which the complementarity determinants are the most critical protein sequence regions determining antibody-antigen binding.
[0060] The antibody sequence information to be compared is the information contained in the tagged antibody sequence. That is, the antibody complementarity determinant cluster sequence alignment method of this application can process the obtained original antibody sequences after obtaining at least one original antibody sequence, that is, tag each antibody sequence.
[0061] The tags can be any combination of strings and numbers. There are no specific rules for naming the tags; the tags should be easy to understand and facilitate the labeling of subsequent results.
[0062] For example, such as Figure 3 As shown, this application uses adalimumab and bevacizumab as examples for labeling. Adalimumab is labeled as Sequence 1, and bevacizumab as Sequence 2. The labels for the heavy and light chains of the original antibody sequence 1 are the string followed by the number "seq_1", and the labels for the heavy and light chains of the original antibody sequence 2 are the string followed by the number "seq_2". The heavy chain is labeled with the letter "H", and the light chain is labeled with the letter "L". That is, Sequence 1 is labeled as seq_1H and seq_1L, and the heavy and light chains of Sequence 2 are labeled as seq_2H and seq_2L.
[0063] The antibody sequence information to be compared includes the tag, sequence, and downstream interface input text.
[0064] The antibody sequence library can be a patented antibody sequence library or other antibody libraries, and there are no restrictions here.
[0065] The alignment information includes the alignment sequence label, overall similarity, protein sequence, and sequence ID.
[0066] Then, the difference determination module 14 can determine the complementary determinant cluster differences between the original antibody sequence and the aligned sequence. Furthermore, the second extraction module 15 can extract target information from the target file title based on the aligned sequence information, providing more accurate and comprehensive alignment difference information. It can also determine whether the target of the original sequence is the same as that of a highly similar aligned sequence, exploring the possibility of treating new diseases and further reducing the risk of infringement on existing antibody sequences. Understandably, a target refers to a structure located within an organism that can be recognized or bound by other substances.
[0067] The difference results output by output module 16 can be as follows: Figure 4 The list shown indicates that the differential results specifically include: original sequence label, aligned sequence label, variable region similarity, aligned sequence, aligned sequence ID, visualization of aligned sequence CDR1 / CDR2 / CDR3, number of differences between the original sequence and aligned sequence CDR1 / CDR2 / CDR3, total number of differences in complementarity determination cluster regions, and the target points of the aligned sequence.
[0068] This application constructs a fully automated method for aligning antibody complementarity determinants using a programming language, which can reduce experimental costs and the time cost of manually extracting and verifying the similarity of complementarity determinant sequences, and provide more comprehensive information on the differences in complementarity determinant region alignment.
[0069] Please see Figure 5 In some implementations, the method further includes:
[0070] 001: Antibody sequences are downloaded using big data to establish an antibody sequence library.
[0071] Please combine Figure 6 The apparatus 10 for antibody complementarity determinant cluster sequence alignment includes a library establishment module 101.
[0072] Step 001 can be implemented by the library creation module 101. That is, the library creation module 101 is used to download antibody sequences from large datasets to create an antibody sequence library.
[0073] Specifically, the antibody complementarity determinant cluster sequence alignment method of this application requires downloading antibody sequences via big data before step 02 to establish an antibody sequence library. That is, step 001 can be performed before step 01, or after step 01 and before step 02.
[0074] Preferably, the antibody complementarity determinant cluster sequence alignment method of this application can also localize the alignment procedure, download the patented antibody sequence library in advance and construct the library, and then perform alignment on the constructed patented antibody sequence library, which can utilize localized resources to speed up the alignment process.
[0075] Please see Figure 7 In some implementations, step 01 includes:
[0076] 011: Obtain at least one original antibody sequence input by the user;
[0077] 012: Determine the original sequence tag for each original antibody sequence to obtain the input sequence information;
[0078] 013: Process the input sequence information to obtain the antibody sequence information to be compared.
[0079] Please see Figure 8 The acquisition module 11 includes an acquisition unit 111, a tag determination unit 112, and a processing unit 113.
[0080] Step 011 can be implemented by the acquisition unit 111, step 012 can be implemented by the tag determination unit 112, and step 013 can be implemented by the processing unit 113.
[0081] Specifically, for example, a user inputs two raw antibody sequences, sequence 1 and sequence 2. The labels for these two raw antibody sequences are determined as follows: the raw sequence label for sequence 1 is a string followed by the number "sep_1", and the raw sequence label for sequence 2 is a string followed by the number "sep_2", i.e., ... Figure 3 As shown, the heavy and light chains of sequence 1 are labeled as seq_1H and seq_1L, respectively, and the heavy and light chains of sequence 2 are labeled as seq_2H and seq_2L, respectively, thus obtaining the input sequence information.
[0082] Then, the input sequence information is fed into the processing unit 113 for information processing to obtain the antibody sequence information to be compared. Understandably, the processing unit 113 can perform sequence reading to read the input sequence information and obtain the antibody sequence information to be compared.
[0083] The antibody complementarity determinant cluster sequence alignment method of this application can automatically obtain at least one original antibody sequence input by the user, determine the original sequence tag of each original antibody sequence to obtain input sequence information, and then read the input sequence information to obtain the antibody sequence information to be aligned, thereby automating the acquisition of antibody sequence information to be aligned and facilitating subsequent comparison.
[0084] Please see Figure 9 In some implementations, step 013 includes:
[0085] 0131: Separate the original antibody sequence according to preset symbols;
[0086] 0132: Identify the segmented original antibody tag and the corresponding original antibody sequence;
[0087] 0133: Construct a list of one-to-one corresponding original sequence tags, original antibody sequences, and interface input text to obtain the antibody sequence information to be compared.
[0088] Please see Figure 10 The processing unit 113 includes a segmentation unit 1131, an identification unit 1132, and a list construction unit 1133.
[0089] Step 0131 can be implemented by the segmentation unit 1131, step 0132 can be implemented by the identification unit 1132, and step 0133 can be implemented by the list construction unit 1133. That is, the segmentation unit 1131 is used to segment the original antibody sequence according to preset symbols; the identification unit 1132 is used to identify the segmented original antibody tags and the corresponding original antibody sequences; and the list construction unit 1133 is used to construct a list of one-to-one corresponding original sequence tags, original antibody sequences, and interface input text to obtain the antibody sequence information to be compared.
[0090] Specifically, the preset symbol can be, for example, ">", such as Figure 3 As shown, each original antibody sequence can be segmented using the symbol ">". Other symbols can also be used, and there are no restrictions here. The preset symbol ">" can be placed before the sequence tag to facilitate better segmentation of the original antibody sequence by the segmentation unit 1131.
[0091] Then, the recognition unit 1132 can use R programming software to write a program to recognize the tag and corresponding antibody sequence of each segmented original antibody. For example, seq_1H corresponds to the heavy chain of sequence 1, and seq_2L corresponds to the light chain of sequence 2.
[0092] Finally, the list building unit 1133 can build a list of labels, sequences, and downstream interface input text that correspond one-to-one.
[0093] The above steps can segment the antibody sequence information to be compared into tags (e.g., Figure 11 (as shown in the name), sequence (such as...) Figure 11 (as shown in the sequence), downstream interface input text (such as...) Figure 11 The content shown in [3] is categorized as a list, and the specific list pattern is as follows: Figure 11 As shown. Understandably, displaying the antibody sequence information to be compared in a list format facilitates subsequent sequence alignment and reduces the likelihood of errors.
[0094] Please see Figure 12 In some implementations, step 02 includes:
[0095] 021: Call the interface of the analysis tool based on similarity comparison, input text into the interface to construct a request form and submit it;
[0096] 022: Retrieve the task number after the request form is submitted;
[0097] 023: Build a pull form based on the task number to retrieve the comparison results;
[0098] 024: Determine the alignment sequence information based on the alignment results;
[0099] 025: Replace non-sequence characters in the compared sequence information.
[0100] Please combine Figure 13 The comparison module 12 includes a form construction unit 121, a task number acquisition unit 122, a retrieval unit 123, an information determination unit 124, and a replacement unit 125.
[0101] Step 021 can be implemented by the form construction unit 121, step 022 can be implemented by the task number acquisition unit 122, and step 023 can be implemented by the retrieval unit 123. That is, the form construction unit 121 is used to call the similarity comparison-based analysis tool interface, input text into the interface to construct a request form, and submit it; the task number acquisition unit 122 is used to obtain the task number after the request form is submitted; the retrieval unit 123 is used to construct a retrieval form based on the task number to retrieve the comparison results; the information determination unit 124 is used to determine the comparison sequence information based on the comparison results; and the replacement unit 125 is used to replace non-sequence characters in the comparison sequence information.
[0102] Specifically, the form construction unit 121 first inputs the interface input text of the list in the list construction unit 1133. The form construction unit 121 can call the interface of an analysis tool based on similarity comparison, inputting the interface input text to construct and submit a request form. For example, it can construct and submit a PUT request form for patent sequence alignment based on BLASTP. Specifically, it can call the BLAST interface, use the PUT request method to access the patent sequence database, and submit the sequence as the interface input text. After submission, a task number (RID) will be obtained.
[0103] Next, the task number acquisition unit 122 acquires the task number after the request form is submitted. The task number obtained in this step can provide interface information for downstream steps to obtain results.
[0104] Then, retrieval unit 123 constructs a GET-based retrieval form to retrieve the comparison results in XML format. This step also requires calling the blast interface, using the GET request method to create a request form, providing the previously obtained task number RID, and thus retrieving the comparison results in XML format.
[0105] Next, the information determination unit 124 can use R programming to obtain alignment sequence information from XML. The alignment sequence information includes tags, overall similarity, protein sequence, and sequence ID.
[0106] Finally, the non-sequence characters in the alignment sequence can be replaced. The non-sequence characters include filler sequence identifiers such as "X" and "-". That is, after obtaining the alignment sequence information, the replacement unit 125 can replace filler sequence identifiers such as "X" and "-" in the alignment sequence with the same identifier that is different from the sequence characters, so as to subsequently number and extract the complementary decision clusters.
[0107] Preferably, the blastp program can be localized, allowing the patented antibody sequence library or other antibody libraries to be downloaded in advance and used to construct the alignment library. The constructed patented antibody sequence library can then be aligned, thus utilizing localized resources to speed up the alignment process.
[0108] Please see Figure 14 In some implementations, step 023 includes:
[0109] 0231: After the preset time is reached, a pull form is constructed based on the task number to pull the comparison results. The preset time is related to the number of original antibody sequences.
[0110] Please combine Figure 15 The pull unit 123 includes a timing unit 1231.
[0111] Step 0231 can be implemented by timing unit 1231. That is, timing unit 1231 is used to construct a pull form to pull the comparison results according to the task number after the preset time is reached. The preset time is related to the number of original antibody sequences.
[0112] Specifically, the preset time is related to the number of original antibody sequences. For example, the preset time can be set to the number of original antibody sequences * 20 seconds, where 20 seconds is the processing time for one original antibody sequence retrieval task.
[0113] The antibody complementarity determinant cluster sequence alignment method of this application retrieves the alignment results by constructing a retrieval form based on the task number after a preset time has elapsed. This allows for a certain waiting time for the retrieval task, preventing subsequent retrieval results from reporting errors when the task status is not yet complete.
[0114] Please see Figure 16 In some implementations, step 025 further includes:
[0115] 0251: Replace non-sequence characters in the alignment sequence with empty strings;
[0116] Please combine Figure 13 Step 0251 can be implemented by the replacement unit 125. That is, the replacement unit 125 is used to replace the non-sequence characters of the alignment sequence with empty strings.
[0117] Specifically, non-sequence characters in the alignment sequence are replaced with empty strings. Non-sequence characters include padding sequence identifiers such as "X" and "-". That is, after obtaining the alignment sequence information, the replacement unit 125 can replace padding sequence identifiers such as "X" and "-" in the alignment sequence with empty strings so that the complementary decision clusters can be numbered and extracted subsequently.
[0118] Please see Figure 17 In some implementations, step 03 includes:
[0119] 031: Number the original antibody sequence and the alignment sequence according to a preset antibody numbering method, which includes at least one of Kabat, Chothia, IMGT, Gelfand, Aho and Martin;
[0120] 032: Extract the complementarity determinant region and construct an information list of the original antibody sequence and the aligned sequence in the "number-residue" manner to obtain the complementarity determinant region information.
[0121] Please combine Figure 18 The first extraction module 13 includes a numbering unit 131 and an extraction unit 132.
[0122] Steps 031 and 032 can be implemented by numbering unit 131 and extraction unit 132. That is, numbering unit 131 is used to number the original antibody sequence and the alignment sequence according to a preset antibody numbering method, which includes at least one of Kabat, Chothia, IMGT, Gelfand, Aho, and Martin; extraction unit 132 is used to extract the complementarity determinant region to construct an information list of the original antibody sequence and the alignment sequence in a "number-residue" manner to obtain the complementarity determinant region information.
[0123] The preset antibody numbering method includes at least one of Kabat, Chothia, IMGT, Gelfand, Aho, and Martin. This means that the preset antibody numbering method can be one of Kabat, Chothia, IMGT, Gelfand, Aho, and Martin. This application can adjust in real time according to different numbering methods to meet various numbering requirements.
[0124] Extraction unit 132 can extract complementarity-determining cluster regions and construct information lists of the original sequence and aligned sequence according to a "number-residue" pattern. More specifically, such as Figure 19 The image shows a list of CDR1 / CDR2 / CDR3 of a certain alignment sequence rearranged according to the "number-residue" information. This list contains three complementarity determining regions (CDRs), namely CDR1, CDR2, and CDR3. The residue order of each CDR region may be discontinuous (indicated by "-").
[0125] The extraction unit 132 of this application adjusts each region according to the "number-residue" pattern, that is, it rearranges the information to obtain complementary determinant cluster region information, which can assign each residue to a specific position, which is beneficial to further improve the accuracy of the subsequent alignment.
[0126] Please see Figure 20In some implementations, step 04 includes:
[0127] 041: List of numbered residues in the complementarity determinant clusters of the original antibody sequence and the aligned sequence;
[0128] 042: Obtain the sequences of each complementary determinant cluster region of the aligned sequence, and the number of differences between the aligned sequence and the original antibody sequence number residues in each complementary determinant cluster region.
[0129] Please see Figure 21 The difference determination module 14 includes a comparison unit 141 and a quantity determination unit 142.
[0130] Step 041 can be implemented by alignment unit 141, and step 042 can be implemented by quantity determination unit 142. That is, alignment unit 141 is used to align the original antibody sequence with the list of numbered residues in the complementarity determinant clusters of the alignment sequence; quantity determination unit 142 is used to obtain the sequence of each complementarity determinant cluster region of the alignment sequence, and the number of numbered residues in each complementarity determinant cluster region that differ between the alignment sequence and the original antibody sequence.
[0131] Specifically, firstly, the sequence is compared according to the "number-residue" pattern to obtain the difference information. The comparison unit 141 can obtain the number of different "number-residue" sequences between the compared sequence and the original sequence according to the comparison algorithm, which is to find the different elements in vector x and vector y (only the different elements in x are taken).
[0132] Then, the number of "number-residue" differences between the aligned sequence and the original sequence can be obtained using the alignment algorithm, thus obtaining the sequence of each complementarity-determining cluster region (CDR region), and the sequence of each CDR region. 比 Complementary determinants (CDRs) of the original antibody sequence 原 The number of residue differences in the number of residues.
[0133] Steps 041 and 042 above can identify the differences in complementary determinant clusters of the aligned sequences.
[0134] Please see Figure 22 In some implementations, the alignment sequence information includes alignment sequence identity encoding, and step 05 includes:
[0135] 051: Extract the corresponding target file title based on the alignment sequence identity code;
[0136] 052: Use Gnormplus to extract target information from the title of the target file.
[0137] Please combine Figure 2Steps 051 and 052 can be implemented by the second extraction module 15. That is, the second extraction module 15 is used to extract the corresponding target file title according to the alignment sequence identity code; and to extract the target point information in the target file title using Gnormplus.
[0138] Specifically, the target file title can be, for example, a patent title, which can be obtained based on the sequence ID (identification code) of the comparison sequence in the comparison results. The target file title can also be other types of titles; there are no restrictions on this.
[0139] Then, Gnormplus is used to extract the target points in the patent title. Specifically, the input file can be adjusted according to the format required by Gnormplus, and the target points not matched in the patent title can be marked as NA, thereby extracting the target points in the patent title.
[0140] Please see Figure 23 This application also provides an electronic device 100. The electronic device 100 includes a processor 110 and a memory 120. The memory 120 stores a computer program 121, and the processor 110 implements the method of any of the above embodiments when executing the computer program 121. The electronic device 100 can be a mobile phone, computer, iPad, or other similar device.
[0141] The electronic device 100 of this application constructs a fully automated antibody complementarity determinant cluster alignment method using a programming language, which can reduce experimental costs and the time cost of manually extracting and verifying the similarity of complementarity determinant cluster sequences, and provide more comprehensive information on the differences in complementarity determinant cluster regions.
[0142] Please see Figure 24 This application also provides a non-volatile computer-readable storage medium 200 containing a computer program. When the computer program 210 is executed by one or more processors 220, it implements the method of any of the above embodiments.
[0143] The computer-readable storage medium 200 of this application constructs a fully automated antibody complementarity determinant cluster alignment method using a programming language, which can reduce experimental costs and the time cost of manually extracting and verifying the similarity of complementarity determinant cluster sequences, and provide more comprehensive information on the differences in complementarity determinant cluster regions.
[0144] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for aligning antibody complementarity determinant cluster sequences, characterized in that, include: Obtain at least one raw antibody sequence to process and obtain the antibody sequence information to be compared; The alignment sequence information is obtained by comparing the antibody sequence information to be aligned with the antibody sequence library. The original antibody sequence and the aligned sequence are numbered, and information on the complementarity determination cluster region is extracted. Determine the differences in complementary determinant clusters between the original antibody sequence and the aligned sequence; Target information in the target file title is extracted based on the alignment sequence information; The difference results are output based on the antibody sequence information to be compared, the alignment sequence information, the complementarity determinant region, the complementarity determinant difference, and the target information; Determining the differences in complementary determinant clusters between the original antibody sequence and the aligned sequence includes: The original antibody sequence is compared with the list of numbered residues in the complementarity determinant clusters of the compared sequence; Obtain the sequences of each complementarity determinant cluster region of the aligned sequence, and the number of differences between the aligned sequence and the original antibody sequence number residues in each complementarity determinant cluster region; The alignment sequence information includes the alignment sequence tag, overall similarity, protein sequence, and sequence ID. The difference results include the original antibody sequence marker, alignment sequence tag, variable region similarity, the alignment sequence, the alignment sequence ID, visualization of multiple complementarity determinant cluster regions of the alignment sequence, the number of differences between the original sequence and the alignment sequence in the complementarity determinant cluster regions, the total number of differences in the complementarity determinant cluster regions, and the target of the alignment sequence.
2. The method according to claim 1, characterized in that, The method further includes: Antibody sequences are downloaded using big data to establish the antibody sequence library.
3. The method according to claim 1, characterized in that, The process of obtaining at least one original antibody sequence to process and obtain the antibody sequence information to be compared includes: Obtain at least one original antibody sequence input by the user; Determine the original sequence tag for each of the original antibody sequences to obtain the input sequence information; The input sequence information is processed to obtain the antibody sequence information to be compared.
4. The method according to claim 3, characterized in that, The process of processing the input sequence information to obtain the antibody sequence information to be compared includes: The original antibody sequence is segmented according to preset symbols; Identify the original antibody tags and corresponding original antibody sequences after segmentation; Construct a list of one-to-one correspondences between the original sequence tags, the original antibody sequences, and the interface input text to obtain the antibody sequence information to be compared.
5. The method according to claim 4, characterized in that, The step of comparing the antibody sequence information to be aligned with the antibody sequence library to obtain the alignment sequence information includes: Call the interface of the similarity comparison-based analysis tool, input the interface input text to construct a request form and submit it; Obtain the task number after the request form is submitted; Based on the task number, construct a pull form to retrieve comparison results; The alignment sequence information is determined based on the alignment results; The non-sequence characters in the alignment sequence information are replaced.
6. The method according to claim 5, characterized in that, The step of constructing the pull form pull comparison result based on the task number includes: After the preset time is reached, a pull form is constructed based on the task number to pull the comparison results. The preset time is related to the number of the original antibody sequences.
7. The method according to claim 5, characterized in that, The process of replacing non-sequence characters in the alignment sequence information includes: Replace the non-sequence characters of the alignment sequence with empty strings.
8. The method according to claim 5, characterized in that, The step of numbering the original antibody sequence and the aligned sequence and extracting complementarity determination cluster region information includes: The original antibody sequence and the alignment sequence after replacement are numbered according to a preset antibody numbering method, wherein the preset antibody numbering method includes at least one of Kabat, Chothia, IMGT, Gelfand, Aho, and Martin. The complementarity determinant regions are extracted to construct a list of numbered residues in the complementarity determinants of the original antibody sequence and the aligned sequence in a "number-residue" manner.
9. The method according to claim 1, characterized in that, The alignment sequence information includes an alignment sequence identity code, and the step of extracting target point information from the target file title based on the alignment sequence information includes: Extract the corresponding target file title based on the identity code of the comparison sequence; The target information in the title of the target file was extracted using Gnormplus.
10. An apparatus for aligning antibody complementarity determinant cluster sequences, characterized in that, include: The acquisition module is used to acquire at least one original antibody sequence to process and obtain antibody sequence information to be compared; The comparison module is used to compare the antibody sequence information to be compared with the antibody sequence library to obtain the alignment sequence information; A first extraction module is used to number the original antibody sequence and the aligned sequence and extract information on the complementarity determinant cluster region. A difference determination module, which is used to determine the differences in complementary determinant clusters between the original antibody sequence and the aligned sequence; The second extraction module is used to extract target information from the title of the target file based on the alignment sequence information. The output module is used to output difference results based on the antibody sequence information to be compared, the alignment sequence information, the complementarity determinant region, the complementarity determinant difference, and the target information; The difference determination module is also used to compare the original antibody sequence with the list of numbered residues in the complementary determinant clusters of the aligned sequence; to obtain the sequence of each complementary determinant cluster region of the aligned sequence, and the number of numbered residues in each complementary determinant cluster region that differ between the aligned sequence and the original antibody sequence. The alignment sequence information includes the alignment sequence tag, overall similarity, protein sequence, and sequence ID. The difference results include the original antibody sequence marker, alignment sequence tag, variable region similarity, the alignment sequence, the alignment sequence ID, visualization of multiple complementarity determinant cluster regions of the alignment sequence, the number of differences between the original sequence and the alignment sequence in the complementarity determinant cluster regions, the total number of differences in the complementarity determinant cluster regions, and the target of the alignment sequence.
11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the method of any one of claims 1 to 9.
12. A non-volatile computer-readable storage medium containing a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the method of any one of claims 1 to 9.