A method, system, device, and storage medium for evaluating a double sequence alignment tool

By generating sequencing sequences based on the reference genome and log-normal distribution, the real simulated two-sequence data is constructed, and the dual-sequence alignment tool is combined with the runtime and accuracy evaluation, the limitations of the existing evaluation methods are solved and more accurate tool performance comparison is achieved.

CN119068989BActive Publication Date: 2025-07-18YANTAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411545910.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-07-18
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

The existing two-sequence alignment tool evaluation methods lack systematicity and cannot comprehensively and objectively compare the performance of different algorithms. In simulation experiments, the use of randomly generated sequences cannot truly simulate the wrong distribution characteristics of actual sequencing data, resulting in deviations from the evaluation results and actual applications.

Method used

Sequencing sequences are generated based on reference genome and log-normal distribution parameters, combined with the alignment results of real query sequences, diversified dual-sequence data are constructed, and the performance of the dual-sequence alignment tool is evaluated through runtime, peak memory usage and accuracy.

Benefits of technology

It provides a more accurate tool evaluation method, which can reflect the actual performance of the two-sequence alignment tool in different scenarios and helps to select the optimal tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068989B_ABST
    Figure CN119068989B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of sequence comparison, and specifically provides a method, system, device, and storage medium for evaluating a dual-sequence alignment tool. To solve the technical problem of poor evaluation results of existing tools, when the true query sequence is unknown, the present invention generates sequencing sequences that not only follow the read length characteristics of a lognormal distribution but also can truly simulate the error distribution characteristics of actual sequencing data based on the reference genome, for forming dual-sequence data. When the true query sequence is known, the reference sequence is determined based on the reference genome and the true query sequence, for forming dual-sequence data. Finally, the dual-sequence data is input into the dual-sequence alignment tool for processing, and the alignment effect of the dual-sequence alignment tool is comprehensively and accurately evaluated by combining the running time, peak memory occupancy, and accuracy, fairly comparing the performances of different dual-sequence alignment tools, which helps to select the dual-sequence alignment tool that performs optimally in their respective fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sequence comparison, and specifically provides a method, system, device, and storage medium for evaluating a double-sequence alignment tool. Background Art

[0002] Double-sequence alignment is one of the basic problems in bioinformatics, aiming to find the similarity between two biological sequences (such as DNA, RNA, or protein) to infer their connections in function, structure, or evolutionary relationships. Precise double-sequence alignment can help researchers reveal gene functions, identify conserved regions, infer protein structures, and trace the process of species evolution. Since the accuracy of double-sequence alignment results is crucial for subsequent biological analysis, it is particularly necessary to evaluate double-sequence alignment methods.

[0003] Through systematic evaluation, researchers can compare the performance of different alignment algorithms, including alignment accuracy, calculation speed, and applicability to various sequencing datasets. However, the current double-sequence alignment evaluation methods have certain limitations. First, there is a lack of systematic evaluation tools, making it impossible to comprehensively and objectively compare the performance of different algorithms. Second, existing evaluation methods, especially in simulation experiments, usually use randomly generated sequences as test data. However, these random sequences often cannot truly simulate the error distribution characteristics of actual sequencing data, resulting in a deviation between the tool evaluation results and the performance in actual applications, and reducing the accuracy of double-sequence alignment results. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, system, device, and storage medium for evaluating a double-sequence alignment tool

[0005] The technical solution of the present invention is as follows:

[0006] A method for evaluating a double-sequence alignment tool includes the following operations:

[0007] S1. Determine whether there is a real query sequence; if not, execute S2; if so, execute S3;

[0008] S2. Based on the reference genome, log-normal distribution parameters, and sequence information error rate, obtain a number of sequencing sequences and simulation comparison results; respectively extract the sequence information of the number of sequencing sequences to obtain a number of simulation query sequences; based on the positions and offsets of the number of simulation query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a number of reference sequences; the number of simulation query sequences and the corresponding reference sequences form double-sequence data, and execute S4;

[0009] S3. Align the real query sequence with the reference genome in terms of position, length, and sequence information to obtain an alignment result; based on the alignment result, extract the corresponding sequence from the reference genome as the reference sequence; the real query sequence and the reference sequence form a dual-sequence data, and then execute S4;

[0010] S4. Input the dual-sequence data into a dual-sequence alignment tool for processing to obtain a target alignment result; based on the running time, or / and peak memory occupancy, or / and accuracy during the processing of the dual-sequence alignment tool, obtain an evaluation result.

[0011] In S4, the accuracy is the sequence score similarity between the target alignment result and the simulation comparison result, which is ratio A, or ratio B, or ratio C, or is obtained based on ratio A, ratio B, and ratio C; ratio A is the ratio of the number of simulation query sequences with the same sequence alignment score in the target alignment result and the corresponding sequence in the simulation comparison result to the total number of simulation query sequences; ratio B is the ratio of the number of simulation query sequences with the total sequence alignment score in the target alignment result greater than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulation query sequences; ratio C is the ratio of the number of simulation query sequences with the total sequence alignment score in the target alignment result less than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulation query sequences.

[0012] The sequence alignment score includes: sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-stage affine gap penalty score.

[0013] In S2, the operation of obtaining a number of sequencing sequences is as follows: Input the reference genome into the PBSIM3 tool, and generate a sequence read length that conforms to the lognormal distribution according to the preset lognormal distribution parameters; based on the sequence read length, control the reference genome to generate a number of initial sequencing sequences with read length characteristics that follow the lognormal distribution; then, for each base position in each initial sequencing sequence, assign insertions, or deletions, or substitutions of errors according to the corresponding base error type preference and base error expected probability to obtain a number of sequencing sequences.

[0014] In S2, the simulation comparison result is the alignment result of the sequencing sequence and the reference genome in terms of position, length, and sequence information.

[0015] A dual-sequence alignment tool evaluation system for implementing the above dual-sequence alignment tool evaluation method, including:

[0016] A real query sequence existence judgment module for judging whether there is a real query sequence; if not, execute the first dual-sequence data generation module; if so, execute the second dual-sequence data generation module;

[0017] The first dual-sequence data generation module is used to obtain a plurality of sequencing sequences and simulation comparison results based on a reference genome, lognormal distribution parameters, and sequence information error rate; respectively extract the sequence information of the plurality of sequencing sequences to obtain a plurality of simulated query sequences; based on the positions and offsets of the plurality of simulated query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a plurality of reference sequences; the plurality of simulated query sequences and the corresponding reference sequences form dual-sequence data, and the evaluation result generation module is executed;

[0018] The second dual-sequence data generation module is used to perform alignment processing on the positions, lengths, and sequence information of a real query sequence and a reference genome to obtain an alignment result; based on the alignment result, extract the corresponding sequence from the reference genome as a reference sequence; the real query sequence and the reference sequence form dual-sequence data, and the evaluation result generation module is executed;

[0019] The evaluation result generation module is used to input the dual-sequence data into a dual-sequence alignment tool for processing to obtain a target alignment result; based on the running time, or / and peak memory occupancy, or / and accuracy during the processing of the dual-sequence alignment tool, an evaluation result is obtained.

[0020] A dual-sequence alignment tool evaluation device includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, the above-mentioned dual-sequence alignment tool evaluation method is implemented.

[0021] A computer-readable storage medium is used to store a computer program. Among them, when the computer program is executed by a processor, the above-mentioned dual-sequence alignment tool evaluation method is implemented.

[0022] The beneficial effects of the present invention are as follows:

[0023] The present invention provides a method for evaluating a dual-sequence alignment tool. First, different dual-sequence data generation methods are respectively formulated for the cases where the true query sequence is known and unknown, and diverse dual-sequence data with different complexities are constructed. When the true query sequence is unknown, based on the reference genome, sequencing sequences that not only follow the read length characteristics of the lognormal distribution but also can truly simulate the error distribution characteristics of actual sequencing data are generated. The dual-sequence data formed thereby has better quality score non-uniformity and homopolymer error rate, and can better reflect the comparison effect of the dual-sequence alignment tool. When the true query sequence is known, based on the alignment results of the reference genome and the true query sequence in terms of position, length, and sequence information, the reference sequence is determined, and dual-sequence data is formed, which can more quickly reflect the comparison effect of the dual-sequence alignment tool. Finally, by combining the running time, peak memory occupancy, and accuracy, the comparison effect of the dual-sequence alignment tool is comprehensively and accurately evaluated, and the performance of different dual-sequence alignment tools is fairly compared, which helps to select the dual-sequence alignment tool that performs best in its respective field. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By reading the detailed description of the preferred embodiments below, the solutions and advantages of the present application will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0025] In the drawings:

[0026] Figure 1 is a schematic flowchart of the evaluation method in this embodiment;

[0027] Figure 2 is a summary graph of the running time of different dual-sequence alignment tools when processing long-read sequencing data in this embodiment;

[0028] Figure 3 is a summary graph of the peak memory occupancy of different dual-sequence alignment tools when processing long-read sequencing data in this embodiment;

[0029] Figure 4 is a summary graph of the accuracy of different dual-sequence alignment tools when processing long-read sequencing data in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings.

[0031] This embodiment provides a method for evaluating a dual-sequence alignment tool. Refer to Figure 1 , including the following operations:

[0032] S1. Determine whether there is a true query sequence; if not, execute S2; if so, execute S3;

[0033] S2. Based on the reference genome, log-normal distribution parameters, and sequence information error rate, obtain a number of sequencing sequences and simulation comparison results; respectively extract the sequence information of the number of sequencing sequences to obtain a number of simulated query sequences; based on the positions and offsets of the number of simulated query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a number of reference sequences; the number of simulated query sequences and the corresponding reference sequences form double-sequence data, and execute S4;

[0034] S3. Compare the real query sequence with the reference genome in terms of position, length, and sequence information to obtain a comparison result; based on the comparison result, extract the corresponding sequence from the reference genome as the reference sequence; the real query sequence and the reference sequence form double-sequence data, and execute S4;

[0035] S4. Input the double-sequence data into a double-sequence alignment tool for processing to obtain a target alignment result; based on the running time, and / or peak memory occupancy, and / or accuracy during the processing of the double-sequence alignment tool, obtain an evaluation result.

[0036] S1. Determine whether there is a real query sequence; if not, execute S2; if so, execute S3.

[0037] For the cases where the real query sequence is known and unknown, different double-sequence data generation methods are respectively formulated, which is beneficial to constructing diverse datasets with different complexities, providing comprehensive and rich data support for the subsequent evaluation of tools, helping to comprehensively detect the alignment ability of double-sequence alignment tools, and obtaining more accurate tool evaluation results.

[0038] S2. Based on the reference genome, log-normal distribution parameters, and sequence information error rate, obtain a number of sequencing sequences and simulation comparison results; respectively extract the sequence information of the number of sequencing sequences to obtain a number of simulated query sequences; based on the positions and offsets of the number of simulated query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a number of reference sequences; the number of simulated query sequences and the corresponding reference sequences form double-sequence data, and execute S4;

[0039] When the real query sequence is unknown, based on the reference genome, generate sequencing sequences that not only follow the read length characteristics of the log-normal distribution but also have a similar error pattern to the standard sequencing sequences. The sequencing sequences can truly simulate the error distribution characteristics of actual sequencing data. The double-sequence data formed thereby has better quality score non-uniformity and homopolymer error rate, and can better reflect the comparison effect of the double-sequence alignment tool when used for subsequent double-sequence alignment tool evaluation processing.

[0040] First, based on the reference genome, log-normal distribution parameters, and sequence information error rate, a number of sequencing sequences are obtained. Specifically, first, the Ref file of the reference genome is input into the PBSIM3 tool. According to the preset log-normal distribution parameters (including the location parameter and the dispersion degree of data distribution), a sequence read length that conforms to the log-normal distribution is generated. Then, based on the sequence read length, the reference genome is controlled to generate a number of initial sequencing sequences with read length characteristics that follow the log-normal distribution. Next, for each base position in each initial sequencing sequence, according to the corresponding base error type preference and the expected probability of base error, an error of insertion, deletion, or substitution is assigned to obtain a number of sequencing sequences that not only follow the log-normal distribution read length characteristics but also have a similar error pattern to the standard sequencing sequence, which can better reflect the evaluation performance of the tool for subsequent tool evaluation.

[0041] Then, a number of sequencing sequences are compared with the sequences at the corresponding positions in the reference genome to obtain the simulation comparison result. The simulation comparison result is the comparison result of the sequencing sequence and the reference genome in terms of position, length, and sequence information.

[0042] Next, the sequence information (base information) of a number of sequencing sequences is extracted respectively to obtain a number of simulated query sequences; and based on the positions and offsets of the a number of simulated query sequences on the reference genome (which can be obtained from the simulation comparison result), the corresponding sequence information on the reference genome is extracted respectively to obtain a number of reference sequences.

[0043] Finally, a number of simulated query sequences and the corresponding reference sequences form the dual-sequence data, which is used to perform the evaluation operation of the dual-sequence alignment tool in S4.

[0044] S3. The real query sequence is compared with the reference genome in terms of position, length, and sequence information to obtain the comparison result; based on the comparison result, the corresponding sequence is extracted from the reference genome as the reference sequence; the real query sequence and the reference sequence form the dual-sequence data, and S4 is executed.

[0045] When the real query sequence is known, based on the comparison result of the reference genome and the real query sequence in terms of position, length, and sequence information, the reference sequence is determined, and the dual-sequence data is formed, which improves the calculation efficiency and is used for the subsequent evaluation process of the dual-sequence alignment tool, and can more quickly reflect the comparison effect of the dual-sequence alignment tool.

[0046] S4. The dual-sequence data is input into the dual-sequence alignment tool for processing to obtain the target comparison result; based on the running time, and / or peak memory occupancy, and / or accuracy during the processing of the dual-sequence alignment tool, the evaluation result is obtained.

[0047] Evaluate pairwise sequence alignment tools by combining running time, peak memory usage, and accuracy. This can not only accurately show the alignment effect of pairwise sequence alignment tools, but also facilitate a comprehensive understanding of the performance of pairwise sequence alignment tools in various aspects, fairly compare the performance of different pairwise sequence alignment tools, which helps to select the pairwise sequence alignment tool that performs best in its respective field.

[0048] First, input pairwise sequence data into a pairwise sequence alignment tool for processing. The pairwise sequence alignment tool is an existing tool. The pairwise sequence alignment tool analyzes and compares the positions, lengths, and sequence information of the reference sequence and the query sequence in the pairwise sequence data to obtain the target alignment result.

[0049] Then, based on the running time, and / or peak memory usage, and / or accuracy during the processing of the pairwise sequence alignment tool, obtain the evaluation result.

[0050] Specifically, if the true query sequence is unknown, obtain the evaluation result based on the running time, peak memory usage, and accuracy during the processing of the pairwise sequence alignment tool. If the true query sequence is known, obtain the evaluation result based on the running time and peak memory usage during the processing of the pairwise sequence alignment tool.

[0051] The above accuracy is the sequence score similarity between the target alignment result and the simulation comparison result, which is ratio A, or ratio B, or ratio C, or is obtained based on ratio A, ratio B, and ratio C (is obtained based on weighted processing of ratio A, ratio B, and ratio C).

[0052] Among them, ratio A is the ratio of the number of simulation query sequences with the same sequence alignment score in the target alignment result and the corresponding sequence in the simulation comparison result to the total number of simulation query sequences.

[0053] Ratio B is the ratio of the number of simulation query sequences with the total sequence alignment score in the target alignment result greater than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulation query sequences.

[0054] Ratio C is the ratio of the number of simulation query sequences with the total sequence alignment score in the target alignment result less than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulation query sequences.

[0055] The above sequence alignment scores include: sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, sequence two-segment affine gap penalty score.

[0056] For easy understanding, the following is an example.

[0057] Cases included in the calculation of ratio A: If the sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-segment affine gap penalty score between the simulated query sequence q and the reference sequence t in the target alignment result are the same as those between the sequencing sequence corresponding to the simulated query sequence q and the reference genome in the simulated comparison result, respectively, then the simulated query sequence q is included as a sequence number in the calculation of ratio A.

[0058] Cases included in the calculation of ratio B: If the sum of the sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-segment affine gap penalty score between the simulated query sequence q and the reference sequence t in the target alignment result is greater than the sum of the sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-segment affine gap penalty score between the sequencing sequence corresponding to the simulated query sequence q and the reference genome in the simulated comparison result, then the simulated query sequence q is included as a sequence number in the calculation of ratio B.

[0059] Cases included in the calculation of ratio C: If the sum of the sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-segment affine gap penalty score between the simulated query sequence q and the reference sequence t in the target alignment result is less than the sum of the sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, and sequence two-segment affine gap penalty score between the sequencing sequence corresponding to the simulated query sequence q and the reference genome in the simulated comparison result, then the simulated query sequence q is included as a sequence number in the calculation of ratio C.

[0060] Finally, the evaluation results are output in JSON format for subsequent visual analysis. Specifically, the running time, peak memory occupancy, and accuracy are output as the evaluation results.

[0061] To verify the effectiveness of the evaluation method in this embodiment, the following experiments were conducted.

[0062] Experiment 1: The purpose was to evaluate the running time of different tools when processing long-read sequencing data. In the experiment, dual-sequence data with sequence lengths of 100bp and error rates (sequence information error rates) of 5%, 10%, 15%, and 20% were generated, and 4 existing advanced methods - the edlib tool, bitpal tool, ksw2 tool, and wfa2 tool were selected as dual-sequence alignment tools respectively. The experimental results are shown in Figure 2 , and it can be found that the bitpal tool takes the least time and has the fastest running speed, showing the best performance, and is suitable for applications in scenarios where a large amount of data needs to be processed quickly.

[0063] Experiment 2: The purpose is to evaluate the peak memory occupancy of different tools when processing long-read sequencing data. In the experiment, dual-sequence data with sequence lengths of 100bp and error rates of 5%, 10%, 15%, and 20% was generated, and 4 existing advanced methods - the edlib tool, the bitpal tool, the ksw2 tool, and the wfa2 tool were selected as dual-sequence alignment tools respectively. The experimental results are shown in Figure 3 , it can be found that the bitpal tool has the smallest peak memory occupancy and the best performance, and is suitable for use in scenarios with strict space requirements.

[0064] Experiment 3: The purpose is to evaluate the accuracy of different tools when processing long-read sequencing data. In the experiment, dual-sequence data with a sequence length of 100bp and an error rate of 5% was generated, and 4 existing advanced methods - the edlib tool, the bitpal tool, the ksw2 tool, and the wfa2 tool were selected as dual-sequence alignment tools respectively. The experimental results are shown in Figure 4 ( Figure 4 In which A, B, and C respectively represent ratio A, ratio B, and ratio C), it can be found that the edlib tool and the bitpal tool have a relatively high ratio A and a relatively good performance in terms of comparison accuracy, and are suitable for use in scenarios with strict space requirements.

[0065] This embodiment also provides a dual-sequence alignment tool evaluation system for implementing the above-mentioned dual-sequence alignment tool evaluation method, including:

[0066] A real query sequence existence judgment module for judging whether there is a real query sequence; if not, execute the first dual-sequence data generation module; if so, execute the second dual-sequence data generation module;

[0067] The first dual-sequence data generation module is used to obtain a number of sequencing sequences and simulation comparison results based on the reference genome, log-normal distribution parameters, and sequence information error rate; respectively extract the sequence information of a number of sequencing sequences to obtain a number of simulation query sequences; based on the positions and offsets of a number of simulation query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a number of reference sequences; a number of simulation query sequences and the corresponding reference sequences form dual-sequence data, and execute the evaluation result generation module;

[0068] The second dual-sequence data generation module is used to perform alignment processing on the real query sequence and the reference genome in terms of position, length, and sequence information to obtain an alignment result; based on the alignment result, extract the corresponding sequence from the reference genome as the reference sequence; the real query sequence and the reference sequence form dual-sequence data, and execute the evaluation result generation module;

[0069] The evaluation result generation module is used to input the double-sequence data into a double-sequence alignment tool for processing to obtain a target alignment result; and obtain an evaluation result based on the running time, and / or peak memory occupancy, and / or accuracy during the processing of the double-sequence alignment tool.

[0070] The data visualization module is used to visualize the evaluation results and supports multiple file formats (including SVG, PNG, PDF).

[0071] Among them, the data visualization module includes a bar chart display sub-module, a line chart display sub-module, and a box plot display sub-module.

[0072] The bar chart display sub-module is used to display the performance of different double-sequence alignment tools in terms of running time, peak memory occupancy, and accuracy under fixed set sequence lengths and error rates using bar charts. The horizontal axis of each bar chart represents different double-sequence alignment tools, and the vertical axis represents the metric values, making it clear at a glance the comparison between different tools.

[0073] The line chart display sub-module is used to display the performance changes of different double-sequence alignment tools under different parameter settings using line charts. Through the line chart, it can be observed how the alignment accuracy or running time of different tools changes when the error rate, sequence length, or other parameters change.

[0074] The box plot display sub-module is used to display the performance distribution of different double-sequence alignment tools on multiple test data using box plots, which is suitable for analyzing the distribution of peak memory occupancy and running time.

[0075] This embodiment also provides a double-sequence alignment tool evaluation device, including a processor and a memory. Among them, when the processor executes the computer program stored in the memory, the above-mentioned double-sequence alignment tool evaluation method is implemented.

[0076] This embodiment also provides a computer-readable storage medium for storing a computer program. Among them, when the computer program is executed by a processor, the above-mentioned double-sequence alignment tool evaluation method is implemented.

[0077] This embodiment provides a method for evaluating a pairwise alignment tool. First, different pairwise sequence data generation methods are formulated for the cases where the true query sequence is known and unknown, respectively, to construct diverse pairwise sequence data with different complexities. When the true query sequence is unknown, based on the reference genome, sequencing sequences that not only follow the read length characteristics of the lognormal distribution but also can truly simulate the error distribution characteristics of actual sequencing data are generated. The pairwise sequence data formed in this way has better quality score non-uniformity and homopolymer error rate, and can better reflect the comparison effect of the pairwise alignment tool. When the true query sequence is known, based on the alignment results of the reference genome and the true query sequence in terms of position, length, and sequence information, the reference sequence is determined, and pairwise sequence data is formed, which can more quickly reflect the comparison effect of the pairwise alignment tool. Finally, combining the running time, peak memory occupancy, and accuracy, the comparison effect of the pairwise alignment tool is comprehensively and accurately evaluated, and the performance of different pairwise alignment tools is fairly compared, which helps to select the pairwise alignment tool that performs best in their respective fields.

Claims

1. A method for evaluating a double sequence alignment tool, characterized in that Including the following operations: S1. Determine whether there is a real query sequence; if not, execute S2; if so, execute S3; S2. Based on the reference genome, lognormal distribution parameters, and sequence information error rate, obtain a number of sequencing sequences and a simulation comparison result; the operation of obtaining a number of sequencing sequences is as follows: according to the preset position parameters and data distribution dispersion degree, control the reference genome to generate a number of initial sequencing sequences with read length characteristics subject to a lognormal distribution; for each base position in each initial sequencing sequence, assign an insertion, deletion, or substitution error according to the corresponding base error type preference and base error expected probability to obtain a number of sequencing sequences; Extract the base information of a number of sequencing sequences respectively to obtain a number of simulated query sequences; based on the positions and offsets of a number of simulated query sequences on the reference genome, extract the corresponding sequence information on the reference genome respectively to obtain a number of reference sequences; The number of simulated query sequences and the corresponding reference sequences form double-sequence data, and execute S4; S3. Perform alignment processing on the real query sequence and the reference genome in terms of position, length, and sequence information to obtain an alignment result; based on the alignment result, extract the corresponding sequence from the reference genome as the reference sequence; the real query sequence and the reference sequence form double-sequence data, and execute S4; S4. Input the double-sequence data into a double-sequence alignment tool for processing to obtain a target alignment result; Based on the running time, and / or peak memory occupancy, and / or accuracy during the processing of the double-sequence alignment tool, obtain an evaluation result; The accuracy is the sequence score similarity between the target alignment result and the simulation comparison result.

2. The double-sequence alignment tool evaluation method according to claim 1, wherein In S4, the accuracy is the sequence score similarity between the target alignment result and the simulation comparison result, which is ratio A, or ratio B, or ratio C, or is obtained based on ratio A, ratio B, and ratio C; Ratio A is the ratio of the number of simulated query sequences with the same sequence alignment score in the target alignment result and the corresponding sequence in the simulation comparison result to the total number of simulated query sequences; Ratio B is the ratio of the number of simulated query sequences with the total sequence alignment score in the target alignment result greater than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulated query sequences; Ratio C is the ratio of the number of simulated query sequences with the total sequence alignment score in the target alignment result less than the total sequence alignment score of the corresponding sequence in the simulation comparison result to the total number of simulated query sequences.

3. The double sequence alignment tool evaluation method according to claim 2, wherein The sequence alignment score includes: sequence edit distance score, sequence linear gap penalty score, sequence affine gap penalty score, sequence two-stage affine gap penalty score.

4. The double-sequence alignment tool evaluation method according to claim 1, characterized in that In S2, the simulation comparison result is the alignment processing result of the sequencing sequence and the reference genome in terms of position, length, and sequence information.

5. A dual-sequence alignment tool evaluation system for implementing the dual-sequence alignment tool evaluation method described in claim 1, characterized in that, Including: A real query sequence existence judgment module for judging whether there is a real query sequence; If not, execute the first double-sequence data generation module; If so, execute the second double-sequence data generation module; The first dual-sequence data generation module is used to obtain a number of sequencing sequences and simulation comparison results based on a reference genome, lognormal distribution parameters, and sequence information error rate; respectively extract the sequence information of the number of sequencing sequences to obtain a number of simulated query sequences; based on the positions and offsets of the number of simulated query sequences on the reference genome, respectively extract the corresponding sequence information on the reference genome to obtain a number of reference sequences; the number of simulated query sequences and the corresponding reference sequences form dual-sequence data and execute the evaluation result generation module; The second dual-sequence data generation module is used to perform alignment processing on the real query sequence and the reference genome in terms of position, length, and sequence information to obtain an alignment result; based on the alignment result, extract the corresponding sequence from the reference genome as the reference sequence; the real query sequence and the reference sequence form dual-sequence data and execute the evaluation result generation module; The evaluation result generation module is used to input the dual-sequence data into a dual-sequence alignment tool for processing to obtain a target alignment result; Based on the running time, and / or peak memory occupancy, and / or accuracy during the processing of the dual-sequence alignment tool, obtain the evaluation result.

6. An evaluation device for a double sequence alignment tool, characterized in that, It includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the dual-sequence alignment tool evaluation method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, It is used to store a computer program. Among them, when the computer program is executed by the processor, it implements the dual-sequence alignment tool evaluation method according to any one of claims 1-4.