A method, device, equipment and medium for screening whole genome resequencing data of a species

By setting filtering conditions and base quantity thresholds in the NCBI SRA database, the system can automatically screen and download whole-genome sequencing data, solving the problems of low efficiency and lack of standardization in existing technologies. This enables efficient and accurate data acquisition, making it suitable for bioinformatics research.

CN122417166APending Publication Date: 2026-07-17JIANGXI AGRICULTURAL UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGXI AGRICULTURAL UNIVERSITY
Filing Date
2026-05-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for obtaining data from the NCBI SRA database for specific species and sequencing strategies are inefficient, have limited screening criteria, and are difficult to process in batches. In particular, there is a lack of simple and efficient solutions for automated screening and downloading of whole-genome sequencing data.

Method used

This paper provides a method for screening whole-genome resequencing data of a species. By setting filtering conditions, sequencing data of a specified species are retrieved from the NCBI SRA database. The SRA number, sequencing strategy, and base count are identified. A base count filtering threshold is introduced. Data filtering and downloading are integrated into the automated workflow. The prefetch tool is used for batch downloading, and failed samples are recorded.

Benefits of technology

It greatly reduces the operational threshold, improves data acquisition efficiency and accuracy, ensures that the filtered data meets the analysis requirements, supports batch processing and fault tolerance, and is suitable for research fields such as population genetics and molecular marker development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417166A_ABST
    Figure CN122417166A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for screening whole-genome resequencing data of a species, relating to the field of bioinformatics technology. The method includes setting filtering conditions and obtaining a user-inputted species name; searching the NCBI SRA database based on the specified species name to obtain all sequencing data for the specified species; identifying the SRA number, sequencing strategy, and base count based on all sequencing data; filtering the data based on the filtering conditions according to the sequencing strategy and the base count to obtain filtered samples; and determining the final selected samples based on the SRA number of the filtered samples. This application can improve the efficiency, comprehensiveness, and accuracy of screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and in particular to a method, apparatus, equipment and medium for screening whole genome resequencing data of a species. Background Technology

[0002] With the rapid development of high-throughput sequencing technology, the volume of biological data is growing exponentially. The Sequence Read Archive (SRA) database of the National Center for Biotechnology Information (NCBI) is one of the world's largest public sequencing data repositories, providing valuable data resources for researchers in life sciences, medicine, agriculture, and environmental protection. Efficiently, comprehensively, and accurately obtaining data on specific species and sequencing strategies from the SRA database is a crucial prerequisite for conducting subsequent bioinformatics analyses (such as population genetics, species identification, and molecular marker development). Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, equipment, and medium for screening whole-genome resequencing data of a species, which can improve the efficiency, comprehensiveness, and accuracy of screening.

[0004] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for screening whole-genome resequencing data of a species, including: Set filter criteria and obtain the scientific name of the species input by the user; Based on the defined species scientific name, the NCBI SRA database was used to retrieve all sequencing data for the defined species; Identify the SRA number, sequencing strategy, and base count based on all sequencing data; Based on the sequencing strategy and the number of bases, the data is filtered according to the filtering conditions to obtain the filtered sample. The final selected samples are determined based on the SRA number of the filtered samples.

[0005] In one embodiment, the filtering conditions include a target sequencing strategy; the target sequencing strategy is whole-genome sequencing.

[0006] In one embodiment, the filtration conditions further include a minimum base number filtration threshold; the minimum base number filtration threshold is 1 million base pairs.

[0007] In one embodiment, the designated species scientific name is the designated species Latin scientific name.

[0008] In one embodiment, identifying the SRA number, sequencing strategy, and base count based on all sequencing data specifically includes: Export a runinfo file containing complete source data from all sequencing data; The header information of the runinfo format file is parsed to identify and locate the SRA number column, sequencing strategy column, and base count column, and the location result is obtained; the location result includes SRA number, sequencing strategy, and base count.

[0009] In one embodiment, the species whole-genome resequencing data screening method further includes: Based on the total number of samples ultimately selected and the maximum number of downloads input by the user, a list of files to be downloaded is obtained; Download the files from the list to be downloaded to the target data storage path using the prefetch tool.

[0010] In one embodiment, downloading the files to be downloaded using the prefetch tool to the target data storage path based on the download list files specifically includes: Iterate through each SRA number in the download list file, and call the prefetch tool in turn to download the file corresponding to each SRA number to the output directory of the species set in the data storage path target; When a file corresponding to an SRA number fails to download, the SRA number that failed to download is recorded in the failure record file, and after the download is completed, the number of failed download samples is counted.

[0011] Secondly, this application provides a species whole-genome resequencing data screening device, comprising: The settings and retrieval modules are used to set filter conditions and retrieve the scientific names of species input by the user. The retrieval module is used to retrieve all sequencing data of the specified species using the NCBI SRA database based on the scientific name of the specified species. The identification module is used to identify the SRA number, sequencing strategy, and base count based on all sequencing data; The data filtering module is used to filter data based on the sequencing strategy and the number of bases according to the filtering conditions to obtain the filtered sample. The filtering module is used to determine the final filtered samples based on the SRA number of the filtered samples.

[0012] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the species whole-genome resequencing data screening method described above.

[0013] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the species whole-genome resequencing data screening method described above.

[0014] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, apparatus, device, and medium for screening whole-genome resequencing data of a species, which integrates retrieval, parsing, and filtering. Users only need to provide the scientific name of the species to complete the data screening, which greatly reduces the operational threshold and improves the efficiency of data acquisition. Data filtering is completed by sequencing strategy and base number. The introduction of base number as a filtering indicator can ensure that the filtered data meets the analysis requirements, thereby improving comprehensiveness and accuracy. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A schematic flowchart illustrating a method for screening whole-genome resequencing data of a species, provided as an embodiment of this application; Figure 2 A schematic diagram of the functional modules of a species whole genome resequencing data screening device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application; Figure 4 The image shows the test results. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Traditional methods of acquiring SRA data primarily rely on researchers manually accessing the NCBI website to search, filter, and download data. While these traditional methods can meet the needs to some extent, they have significant limitations. For example, when processing multiple species or large numbers of samples, manual operation is time-consuming, labor-intensive, and inefficient. Furthermore, the limited selection criteria make it difficult to perform batch and precise filtering based on the specific needs of research projects (such as sequencing depth and data volume). In addition, these traditional methods depend on network environments and manual operation, making it difficult to integrate them into automated workflows.

[0019] While existing command-line tools such as esearch, efetch, and prefetch provide interfaces for data retrieval and downloading, they typically require users to combine multiple commands and write complex parsing logic, posing a high barrier to entry for researchers without a bioinformatics background. In particular, for the specific type of "whole genome sequencing," and supplemented by key indicators such as "data volume," an integrated method for automated screening and batch downloading currently lacks a simple and efficient solution. Therefore, developing a rapid method for automating and standardizing the retrieval and screening of SRA data is of significant practical importance for improving research efficiency and promoting data sharing and utilization.

[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] In one exemplary embodiment, such as Figure 1 As shown, a method for screening whole-genome resequencing data of a species is provided. This method is executed by computer equipment, specifically by a computer device such as a terminal or server alone, or by a terminal and server together, and includes the following steps.

[0022] Step 101: Set filter conditions and obtain the scientific name of the species input by the user. The scientific name of the species is its Latin name.

[0023] Step 102: Based on the defined species name, search the NCBI SRA database to obtain all sequencing data for the defined species.

[0024] Step 103: Identify the SRA number, sequencing strategy, and base count based on all sequencing data.

[0025] Step 104: Filter the data based on the sequencing strategy and the number of bases according to the filtering conditions to obtain the filtered sample.

[0026] Step 105: Determine the final selected samples based on the SRA number of the filtered samples.

[0027] By integrating retrieval, parsing, and filtering, users only need to provide the scientific name of the species to complete data screening, which greatly reduces the operational threshold and improves data acquisition efficiency. Data filtering is completed through sequencing strategies and base count. Introducing base count as a filtering indicator can ensure that the filtered data meets the analysis requirements, thereby improving comprehensiveness and accuracy.

[0028] In one exemplary embodiment, the filtering conditions include a target sequencing strategy; the target sequencing strategy is whole genome sequencing (WGS).

[0029] In another exemplary embodiment, the filtering conditions further include a minimum base number filtering threshold; the minimum base number filtering threshold is 1 million base pairs. Setting the minimum base number filtering threshold to no less than 1,000,000 bp ensures that the data volume is sufficient to support subsequent population genome data analysis.

[0030] In an exemplary embodiment, identifying the SRA number, sequencing strategy, and base count based on all sequencing data specifically includes: exporting a runinfo format file containing complete raw data from all sequencing data; parsing the header information of the runinfo format file to identify and locate the positions of the SRA number column, sequencing strategy column, and base count column to obtain the location result; the location result includes the SRA number, sequencing strategy, and base count.

[0031] In one example embodiment, the species whole-genome resequencing data screening method further includes: obtaining a list file to be downloaded based on the total number of samples finally screened and the maximum download quantity input by the user; and downloading the data to the target data storage path using the prefetch tool based on the list file to be downloaded.

[0032] The process of downloading the SRA number to the target data storage path using the prefetch tool based on the list of files to be downloaded includes: iterating through each SRA number in the list of files to be downloaded, and sequentially calling the prefetch tool to download the file corresponding to each SRA number to the output directory of the species specified in the target data storage path; when a file corresponding to an SRA number fails to download, the SRA number that failed to download is recorded in the failure record file, and after the download is completed, the number of samples that failed to download is counted.

[0033] This application aims to address the technical problems of low efficiency, limited selection criteria, and difficulty in batch processing in existing technologies. The method for screening whole-genome resequencing data and the subsequent download process in practical applications include the following steps: S1. Set the data storage path directory, target sequencing strategy, and minimum base number filtering threshold.

[0034] S2. The user enters the Latin name of the given species and the maximum number of downloads. If the user does not enter a maximum number of downloads, the system will download the first 10 samples that meet the criteria by default.

[0035] S3. Based on the input species' Latin name, use the esearch and efetch commands in the Edirect tool to retrieve all sequencing data for that species from the NCBI SRA database and export a runinfo file containing complete metadata.

[0036] S4. Parse the runinfo file, automatically identify and locate the “Run” column (SRA number), “LibraryStrategy” column (sequencing strategy), and “Bases” column (base count) in the file.

[0037] S5. Based on the target sequencing strategy and minimum base number filtering conditions set in step S1, filter the data in the runinfo file, retain only samples with the target sequencing strategy and a base number greater than or equal to the set threshold, and extract their SRA numbers.

[0038] S6. Based on the maximum download quantity set in step S2, extract the first N SRA numbers that meet the conditions to form a list to be downloaded.

[0039] S7. Using NCBI's prefetch tool, iterate through the list of files to be downloaded, download the SRA data files to the species-specific path directory set in step S1, and record the numbers of any download failures. Enable the resume function in the prefetch tool to handle download interruptions caused by network fluctuations.

[0040] Compared with existing technologies, this application has the following advantages: (1) High degree of automation: The complex multi-step retrieval, parsing, filtering and download process is integrated into one process. Users only need to provide the species name to run it with one click, which greatly reduces the operation threshold and improves the data acquisition efficiency.

[0041] (2) Precise screening: It is not limited to the screening of sequencing strategies, but also innovatively introduces the key quality control indicator of "base quantity" for filtering, which can ensure that the amount of downloaded data meets the needs of subsequent analysis and avoids the download of invalid data.

[0042] (3) Batch processing and fault tolerance: It supports batch download and can automatically record samples that fail to download, making it convenient for users to download again and ensuring the integrity of data acquisition.

[0043] (4) Wide applicability: This method can be widely applied to bioinformatics research fields that require a large amount of WGS data, such as population genetics, molecular marker development, and species identification.

[0044] In another exemplary embodiment, a method is provided for automatically screening and downloading whole-genome resequencing data for a user-specified species. This method is implemented using a script called "seeksra" and specifically includes the following steps: Step 1: Environment configuration and parameter settings.

[0045] On a Linux server or system with Edirect and SRA Toolkit installed, configure the script. The core configuration section of the script defines: Define the root storage path for SRA data to clarify the overall storage location for sequencing data of all species.

[0046] The fixed-target sequencing strategy is whole-genome sequencing, which focuses on the core types of data to be screened.

[0047] The minimum base number filtering threshold is set to 1 million base pairs. This parameter can be flexibly adjusted according to actual research needs. If data does not need to be filtered by base number, this parameter can be set to unlimited.

[0048] Step 2: Input parameters and initialization.

[0049] Users launch the pre-built script in the system command line, passing in the relevant parameters: the Latin name of the target species is required, and the maximum number of samples to be downloaded is optional; if no maximum download number is specified, the script will automatically use the default value of 10. After receiving the parameters, the script will automatically create a dedicated output directory for the target species in the pre-defined root save path, based on the Latin name of the target species. All subsequent sequencing data, screening results, and log files for that species will be saved in this directory, achieving categorized data storage.

[0050] Step 3: Data retrieval and metadata acquisition.

[0051] The script will automatically invoke the relevant functions of the Edirect toolset, using the input species' Latin name as the search keyword, to comprehensively search the NCBI SRA database for all sequencing data entries for that species. After the search is complete, the script will organize and export all complete metadata information from the search results, including SRA number, sample information, base count, sequencing strategy, etc., into a runinfo file and save it to the current system directory, providing a complete data foundation for subsequent data analysis and screening.

[0052] Step 4: Analysis and Filtering.

[0053] The script automatically reads the header information of the runinfo format metadata file. By parsing the header, it accurately identifies and locates the SRA number column, sequencing strategy column, and base count column. Based on the location results, the script automatically filters all data in the file according to preset screening criteria, retaining only samples that simultaneously meet the criteria of "sequencing strategy is WGS" and "base count is greater than or equal to 1 million base pairs". It then extracts the corresponding SRA numbers from these valid samples, organizes all the numbers, and saves them into a dedicated number file to form a candidate download list.

[0054] Step 5: Generate the download list.

[0055] The script first counts the valid SRA numbers in the candidate download list to obtain the total number of samples of this species in the database that meet the screening criteria. Then, based on the maximum number of downloads provided by the user, it selects the top N SRA numbers from the candidate download list, organizes them, and generates the final download list file, thus precisely controlling the number of samples to be downloaded in this data download.

[0056] Step 6: Batch download and log recording.

[0057] The script automatically iterates through each SRA number in the final download list, sequentially calling the prefetch tool in the SRA Toolkit to download the SRA data file corresponding to each number to the species' dedicated output directory. During the download process, the system displays the download progress of each file in real time, and the download tool is configured with no file size limit and supports resume functionality to handle download interruptions caused by network fluctuations. If the download of a particular SRA number fails, the script records the failure number in a dedicated failure log file. After the download is complete, the script automatically calculates the number of successful and failed samples.

[0058] Step 7: Summarize the results.

[0059] After the script completes all download operations, it will automatically print a detailed summary report of the running results on the system's operating terminal. The report includes key information such as the target species name, the screening conditions used, the total number of samples of the species that meet the conditions in the database, the number of successful and failed samples downloaded, and the absolute path where the sequencing data of the species is saved, allowing users to clearly and comprehensively understand the overall situation of this data acquisition.

[0060] For example, in population genetics studies of snow leopard samples, the method provided in this application can be used to quickly and accurately screen and download all publicly available snow leopard WGS data, providing a solid data foundation for subsequent species identification marker development or population history analysis. Test results are as follows... Figure 4 .

[0061] #! / bin / bash # Function: Runs locally, filters and downloads WGS-type SRA data for a specified species (filtering only the number of bases). # Usage: seeksra "species Latin name" [maximum number of downloads (optional, default 10)] export PATH= / public / home / xiexionghui / bin / edirect: PATH export PATH= / public / home / xiexionghui / bin / prefetch: PATH # ---------------------- Core Configuration (Can be adjusted as needed) ---------------------- # SRA data is stored in the root directory (absolute path recommended). BASE_OUT_DIR=" HOME / ncbi_sra_data" # Default maximum number of downloads (users can override this using the second parameter) DEFAULT_MAX_SRA=10 # ===== WGS + Base Count Filter Configuration ===== # 1. Sequencing strategy to be selected: WGS (whole genome sequencing) TARGET_STRATEGY="WGS" # 2. Minimum number of bases (# of Bases): Only download samples with a base count ≥ this value (unit: bp; leave blank to avoid filtering). # Examples: 1000000 (1 million bps), 10000000 (10 million bps), "" (unfiltered) MIN_BASES="1000000" # ---------------------- Parameter Check and Initialization ---------------------- if [ -z " 1" ]; then echo "[Incorrect Usage] Please enter the Latin name of the species!" echo "Example: bash 0 'Arabidopsis thaliana' 15" echo "Parameter description:" echo "First parameter: Latin name of the species (required, must be enclosed in quotation marks containing spaces)" echo "The second parameter: maximum number of downloads (optional, default)" DEFAULT_MAX_SRA) exit 1 fi SPECIES=" 1" MAX_SRA= {2:- DEFAULT_MAX_SRA} SPECIES_DIR_NAME=" {SPECIES / / / _}" OUT_DIR=" BASE_OUT_DIR / SPECIES_DIR_NAME" # Dependency Checking Tool if ! command -v esearch&> / dev / null; then echo "[Error] edirect tool not found, please install it first!" exit 1 fi if ! command -v prefetch&> / dev / null; then echo "[Error] Prefetch tool not found, please install it first!" exit 1 fi # ---------------------- Data Retrieval and Filtering ---------------------- mkdir -p " OUT_DIR" cd " OUT_DIR" || { echo "[Error] Unable to enter directory"; OUT_DIR! "; exit 1;} echo "===== Start Search" {SPECIES} {TARGET_STRATEGY} type SRA data =====" # 1. Retrieve and export the complete runinfo data (including all metadata columns) esearch -db sra -query " {SPECIES}[Organism]" | \ efetch -format runinfo>sra_runinfo_full.txt # 2. Extract the column name row and determine the column number of each field (adapt to different versions of runinfo format) HEADER= (head -n 1 sra_runinfo_full.txt) # Get the column number (SRA number) of the Run column. RUN_COL= (echo " HEADER" | tr ',' '\n' | grep -n "^Run " | cut -d ':'-f 1) # Get the column number of the LibraryStrategy column (sequencing strategy) STRATEGY_COL= (echo " HEADER" | tr ',' '\n' | grep -n "^LibraryStrategy " | cut -d ':' -f 1) # Get the column number of the Bases column (number of bases, # of Bases) BASES_COL= (echo " HEADER" | tr ',' '\n' | grep -n "^Bases " | cut -d':' -f 1) # 3. Filtering logic: Only retain WGS type + minimum base count echo "===== Filtering conditions: {TARGET_STRATEGY} type + number of bases ≥ {MIN_BASES}bp=====" # Build awk filter conditions AWK_CONDITION="NR>1&&\ s_col == target_s" # Filtering based on the number of bases only if [ -n " MIN_BASES" ]; then AWK_CONDITION+="&&\ b_col>= min_bases" fi # Perform filtering awk -F ',' -v s_col=" STRATEGY_COL" -v target_s=" TARGET_STRATEGY" \ -v b_col=" BASES_COL" -v min_bases=" MIN_BASES" \ -v r_col=" RUN_COL {AWK_CONDITION} {print \ r_col} " sra_runinfo_full.txt>sra_accessions.txt # ---------------------- Download Logic ---------------------- # Check the filter results if [ ! -s sra_accessions.txt ]; then echo "[Error] No matching SRA data found ( {TARGET_STRATEGY} + Number of bases ≥ {MIN_BASES}bp)! exit 1 fi TOTAL_FOUND= (wc -l <sra_accessions.txt) echo "Total retrieved" {TOTAL_FOUND} eligible SRA samples will be downloaded this time. {MAX_SRA} ... # Extract the first N Accessions head -n " MAX_SRA" sra_accessions.txt>sra_accessions_selected.txt echo "SRA ID to be downloaded:" cat sra_accessions_selected.txt # Download SRA data (with resume capability) echo -e "\n===== Start downloading SRA data=====" rm -f sra_download_failed.txt while read -r acc; do echo "Downloading" {acc}..." prefetch -p -O " OUT_DIR" --max-size 0 " acc" if [ ? -eq 0 ]; then echo " {acc} Download complete! else echo " Download failed for {acc}, this has been logged! echo " acc">>sra_download_failed.txt fi done <sra_accessions_selected.txt # Download Summary echo -e "\n===== Download Summary=====" echo "Species: {SPECIES}" echo "Filtering conditions: {TARGET_STRATEGY} type + number of bases ≥ {MIN_BASES}bp" echo "Total number of samples meeting the criteria:" {TOTAL_FOUND}" echo "Number of samples downloaded this time: " {MAX_SRA}" SUCCESS_COUNT= (ls -1 " OUT_DIR" / .sra 2> / dev / null | wc -l) FAILED_COUNT= (cat sra_download_failed.txt 2> / dev / null | wc -l) echo "Successful downloads: " {SUCCESS_COUNT}" echo "Number of failed downloads: " {FAILED_COUNT}" if [ -f sra_download_failed.txt ]&&[ FAILED_COUNT -gt 0 ]; then echo "Failed SRA ID:" cat sra_download_failed.txt fi echo "Data storage directory: {OUT_DIR}" echo "========================" This application, by setting a target sequencing strategy and a minimum base number filtering threshold, and given the Latin name of the target species, integrates the Edirect tool to automatically retrieve and download SRA data for that species. The invention automates the complex SRA data retrieval, screening, and download process, innovatively introducing base number-based quality control filtering. This solves the problems of low efficiency, single screening criteria, and difficulty in batch processing in existing technologies, offering advantages of high efficiency, accuracy, and ease of use. It can be widely applied to data acquisition in research fields such as population genetics and molecular marker development.

[0062] Based on the same inventive concept, this application also provides a species whole-genome resequencing data screening device for implementing the species whole-genome resequencing data screening method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more species whole-genome resequencing data screening device embodiments provided below can be found in the limitations of the species whole-genome resequencing data screening method described above, and will not be repeated here.

[0063] In one exemplary embodiment, such as Figure 2 As shown, a species whole-genome resequencing data screening device is provided, comprising: The settings and retrieval modules are used to set filter conditions and retrieve the scientific names of species input by the user.

[0064] The retrieval module is used to search the NCBI SRA database based on the specified species scientific name to obtain all sequencing data of the specified species.

[0065] The identification module is used to identify the SRA number, sequencing strategy, and base count based on all sequencing data.

[0066] The data filtering module is used to filter data based on the sequencing strategy and the number of bases according to the filtering conditions to obtain filtered samples.

[0067] The filtering module is used to determine the final filtered samples based on the SRA number of the filtered samples.

[0068] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores species whole-genome resequencing data for screening. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a species whole-genome resequencing data screening method.

[0069] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment to which the present application is applied. Specific computer equipment may include, for example, [the following is a list of possible additional structures]. Figure 3 The embodiments show more or fewer components, combinations of certain components, or different component arrangements. In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the above-described method embodiments.

[0070] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method embodiments.

[0071] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method embodiments.

[0072] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0073] In this application, all actions to acquire signals, information, or data are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0074] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0075] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0076] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0077] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for screening whole-genome resequencing data of a species, characterized in that, include: Set filter criteria and obtain the scientific name of the species input by the user; Based on the defined species scientific name, the NCBI SRA database was used to retrieve all sequencing data for the defined species; Identify the SRA number, sequencing strategy, and base count based on all sequencing data; Based on the sequencing strategy and the number of bases, the data is filtered according to the filtering conditions to obtain the filtered sample. The final selected samples are determined based on the SRA number of the filtered samples.

2. The method for screening species whole-genome resequencing data according to claim 1, characterized in that, The filtering conditions include a target sequencing strategy; the target sequencing strategy is whole genome sequencing.

3. The method for screening species whole-genome resequencing data according to claim 2, characterized in that, The filtration conditions also include a minimum base number filtration threshold; the minimum base number filtration threshold is 1 million base pairs.

4. The method for screening species whole-genome resequencing data according to claim 1, characterized in that, The scientific name of the specified species is the Latin name of the specified species.

5. The method for screening species whole-genome resequencing data according to claim 1, characterized in that, Based on all sequencing data, the SRA number, sequencing strategy, and base count were identified, specifically including: Export a runinfo file containing complete source data from all sequencing data; The header information of the runinfo format file is parsed to identify and locate the SRA number column, sequencing strategy column, and base count column, and the location result is obtained; the location result includes SRA number, sequencing strategy, and base count.

6. The method for screening species whole-genome resequencing data according to claim 1, characterized in that, Also includes: Based on the total number of samples ultimately selected and the maximum number of downloads input by the user, a list of files to be downloaded is obtained; Download the files from the list to be downloaded to the target data storage path using the prefetch tool.

7. The method for screening species whole-genome resequencing data according to claim 6, characterized in that, Based on the list of files to be downloaded, use the prefetch tool to download them to the target data storage path. Specifically, this includes: Iterate through each SRA number in the download list file, and call the prefetch tool in turn to download the file corresponding to each SRA number to the output directory of the species set in the data storage path target; When a file corresponding to an SRA number fails to download, the SRA number that failed to download is recorded in the failure record file, and after the download is completed, the number of failed download samples is counted.

8. A device for screening whole-genome resequencing data of a species, characterized in that, include: The settings and retrieval modules are used to set filter conditions and retrieve the scientific names of species input by the user. The retrieval module is used to retrieve all sequencing data of the specified species using the NCBI SRA database based on the scientific name of the specified species. The identification module is used to identify the SRA number, sequencing strategy, and base count based on all sequencing data; The data filtering module is used to filter data based on the sequencing strategy and the number of bases according to the filtering conditions to obtain the filtered sample. The filtering module is used to determine the final filtered samples based on the SRA number of the filtered samples.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the species whole-genome resequencing data screening method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the species whole-genome resequencing data screening method according to any one of claims 1-7.