Pathogenic microorganism database and construction method thereof
By constructing a database containing pathogenic microbial phylogenetic relationships, genomic data, detection information and antibody information, the problem that existing databases cannot display phylogenetic tree maps and lack of information integration is solved, and the function of quickly obtaining and personalizing the processing of pathogenic microbial related data is realized, and the efficiency of scientific research is improved.
Patent Information
- Application Number
- CN202410571105.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-05-06
AI Technical Summary
The existing pathogenic microbial database cannot display the phylogenetic status tree map of microorganisms, lacks comparative information with similar pathogenic microorganisms, as well as the integration and visualization of pathogenic microorganisms, diseases caused by them, related antibodies, etc.
By obtaining data of target pathogenic microorganisms, a database containing a variety of pathogenic microorganisms and related extension information is constructed, phylogenetic relationship visualization, information collection and processing, genomic data, detection information and antibody information are integrated, and personalized information arrangement, combination and download functions are provided.
It has achieved rapid localization of the association between disease and pathogenic microorganisms, quickly understood the detection methods and verified sequences of pathogenic microorganisms, quickly obtained relevant antibodies, genome and annotation information, supported the arrangement, combination and download of custom information, and improved the efficiency of scientific research.
Smart Images

Figure CN119943165A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of biological information analysis, and discloses a method, device, equipment and medium for constructing a pathogenic microorganism database. Background Art
[0002] Pathogenic microorganisms refer to microorganisms that can cause diseases in animals and plants, including bacteria, viruses, fungi, parasites, etc. These microorganisms can grow and reproduce in the host body, and cause varying degrees of tissue damage and pathological reactions to the host, leading to the occurrence of various diseases. For humans, common diseases caused by pathogenic microorganisms include colds, pneumonia, tuberculosis, cholera, malaria, leukemia, etc. Pathogenic microorganisms can enter the human body through a variety of pathways, such as air droplets, food, water, contact transmission, sexual contact, etc. They may lurk in the human body for a long time without causing obvious symptoms, or they may cause acute infection in a short period of time, and even lead to death. In order to better prepare for disease prevention and control, pathogenic microorganism detection and new drug development, it is particularly important to understand the pathogenic microorganisms that cause these diseases.
[0003] At present, there are many databases related to pathogenic microorganisms in the world, including the domestic Global Pathogen List Database GCP, Pathogen Database PVD, etc., the New Coronavirus Database GISAID, RCoV19, etc., antibody-related abYsis, SAbDab, etc., and NCBI related to genome data, etc.
[0004] In the prior art, databases that cover a relatively comprehensive list of pathogenic microorganisms cannot display a tree diagram of their phylogenetic status when displaying the classification of a certain microorganism. Databases for a specific type of pathogenic microorganism lack comparative information with similar pathogenic microorganisms. There is no integrated and structured database of pathogenic microorganisms, diseases caused by them, literature on related antibodies, antibody sequences and annotation information, genomic data, etc. Although existing pathogenic microorganism databases can specify the type of data to be downloaded, such as genome sequences, etc., they cannot display and download customized information for self-selected species or individuals. Databases with antibody information are more focused on the operation of antibody sequence and structure information, such as BLAST searches, etc., while ignoring their association with pathogenic microorganisms and diseases caused by pathogenic microorganisms.
[0005] The present invention provides a method for establishing a database that can quickly locate the relationship between diseases and pathogenic microorganisms, quickly understand the reported detection methods and verified sequences of pathogenic microorganisms, quickly obtain antibodies, genomes, annotation information and other data related to a certain pathogenic microorganism, and allow users to perform personalized arrangement, combination and download of target data. Summary of the invention
[0006] Based on the above objectives, the present invention first provides a method for constructing a pathogenic microorganism database containing multiple pathogenic microorganisms and related extended information, the method comprising: Step 1. Data acquisition: Search for data information of pathogenic microorganisms in set literature sites and search engines according to the name of the target pathogenic microorganism; Step 2. Visualization of phylogenetic relationships: Determine the phylogenetic relationship of the target pathogenic microorganisms based on the data obtained in step 1 to obtain a phylogenetic tree and visualize it; Step 3. Information collection: Collect the genome data information of the target pathogenic microorganism, pathogenic microorganism detection information, and pathogenic microorganism related antibody information; Step 4. Information processing: First, classify the information to be processed. The information is classified into: Chinese and English scientific names and aliases of target microorganisms, diseases caused, reported detection methods of target microorganisms, literature citations, genome information of target microorganisms that have been sequenced, and antibody information related to target microorganisms. For example, enter a microorganism name in the name column to view the Chinese and English scientific names and aliases of the microorganism, related classifications (viruses, bacteria, fungi), belonging lists, subspecies branches, number of literature, number of antibodies, and number of genome sequences.
[0007] According to the above-mentioned classification of demand information, different subject pages (such as pathogenic microorganism home page, literature page, antibody page, genome information page) can be realized under different tags (different labels under different subject pages, pathogenic microorganism home page includes: pathogenic microorganism name, alias, Chinese name, related source list, branch information, taxonomic domain, number of related literature, number of antibodies, number of genome sequences; literature page includes: date, detection method, related sequence, PubMed number, DOI number, literature name; antibody page includes: antibody accession number, heavy and light chain, antibody name, antibody length, host species, references; genome information The page includes: genome ID, genome name, taxonomic ID, taxonomic levels, bacteria / strain name, country of isolation, genome sequencing status, GenBank accession number, sequencing center, sequencing platform, sequencing depth, assembly method, presence of plasmid, number of contigs, genome size, number of CDS). If quantifiable, the columns are sorted from large to small or from small to large. If not quantifiable, they are sorted and filtered by type. If filtering is performed, the filtering can be matched according to the input words. The filtering conditions include species classification, detection method, article publication date, species name, genome sequencing status, etc.
[0008] Step 5. Based on the screening results obtained in step 4, a pathogenic microorganism database is constructed.
[0009] In a specific embodiment of the present invention, in step 1, the search engine includes academic websites such as NCBI (National Biotechnology Center: https: / / www.ncbi.nlm.nih.gov / ), BV-BRC (Bacterial and Viral Bioinformation Resource Center: https: / / www.bv-brc.org / ), abYsis (http: / / www.abysis.org / abysis / ) and BING search (https: / / global.bing.com).
[0010] In a specific embodiment of the present invention, in step 1, the data information of the pathogenic microorganism includes scientific name, alias, abbreviation, and the name of the related disease caused by it.
[0011] In a preferred embodiment, the method for obtaining and visualizing the phylogenetic relationship in step 2 is to search in the database by the scientific name or alias of the pathogenic microorganism, enter the Taxonomy classification section, record the 8-level taxonomic status of the pathogenic microorganism by domain, phylum, class, order, family, genus and species, and classify it according to fungi, bacteria and viruses, and use JavaScript tree map dynamic rendering technology to convert the data in the NCBI database into a visualized phylogenetic tree; the JavaScript tree map dynamic rendering technology includes: creating an HTML page and a container, introducing a JavaScript file, adding data to be displayed in the tree map, writing JavaScript code to create a tree map, and adding the data to the tree map to display it in the form of a tree map.
[0012] In a preferred embodiment of the present invention, in step 3, the collection of the genome data information of the pathogenic microorganism is specifically as follows: search on the Bacteria and Virus Biological Information Resource Center BV-BRC (https: / / www.bv-brc.org / ) according to the name of the pathogenic microorganism, select the Genomes classification, enter the name of the pathogenic microorganism to be searched, first store all the data in the database, then extract the target data from the database, and then archive it. The target data includes sequenced genome information and spliced genome information. The spliced genome information includes Contigs, assembly method, size, sequencing center, platform, sequencing method, sequencing depth, host sample from which the pathogenic microorganism is isolated, geographic location, and ID number or Accession number that records the relevant information that can be found on NCBI.
[0013] In a preferred embodiment of the present invention, in step 3, the collection of pathogen detection information is specifically as follows: according to the alias of each pathogen and the name of the disease caused by it, the PubMed database is searched and crawled as keywords to obtain the document name, publication time, PubMed ID number, DOI number of the document and record them, and then "detect" and "microorganism name" are used to retrieve the documents that appear in the title and abstract at the same time, wherein "microorganism name" is searched as a word, and "detect" only needs to appear (for example, detect, detection, detected, etc. are all valid); then, according to the screened documents, the sequence data in the documents are extracted according to different detection methods: qPCR (including: real time RT-PCR, RT-qPCR), dPCR, LAMP (RT-LAMP), RPA, EXPAR, SDA, NASBA, HDA, RCA, immunoChromatography, ELISA, ELFA, TRFIA, CILA, MAFIA, HTRF, AlphaLISA, xMAP, Luminex x-TAG, Aptamer.
[0014] Batch search the detection methods and related sequence information involved in the literature for detecting the pathogenic microorganism and record them, archive all processed literature data, and display the processed data through UI rendering. The UI rendering uses three technologies: HTML, CSS, and JavaScript. First, the browser will parse the HTML / SVG / XHTML file, and a DOM Tree will be generated by parsing these three files. Next, the browser will parse the CSS file to generate a CSS rule tree. The last step is the execution of the script, which mainly uses the DOM API and CSSOM API to operate the DOM Tree and CSS Rule Tree for visualization and literature jump.
[0015] In a preferred embodiment of the present invention, in step 3, the collection of pathogen-related antibody information is specifically as follows: searching for relevant antibodies in the abYsis database according to the name of the pathogen, recording the species from which the antibody is isolated, the antibody accession number, the antibody name, the light and heavy chains, the amino acid sequence, the annotation information of different regions, and the relevant literature information, and classifying and arranging the data, and visualizing its content and function on the UI end: using HTML Tags to create a table, each cell uses Label, header use Tags, display the antibody name, light and heavy chains, species from which the antibody was isolated, and related literature information in a table. Tags are used to create cards, and then CSS is used for style design. The amino acid sequence and annotation information of different regions are displayed in the form of cards. Each card contains partial information, and users can click on the card to view more detailed information and download it.
[0016] The present invention also provides a pathogenic microorganism database constructed by the above method, the pathogenic microorganism database comprising: a data acquisition module, a phylogenetic relationship visualization module, an information collection module, an information processing module and a database construction module; The data acquisition module is used to search for data information of pathogenic microorganisms in set literature sites and search engines according to the names of target pathogenic microorganisms; The phylogenetic relationship visualization module is used to determine the phylogenetic relationship of pathogenic microorganisms based on the data acquired by the data acquisition module, obtain a phylogenetic tree and visualize it; The information collection module is used to collect genome data information of pathogenic microorganisms, pathogenic microorganism detection information and pathogenic microorganism related antibody information; The information processing module is used to classify the information to be processed, and the information classification is: Chinese and English scientific names and aliases of target microorganisms, diseases caused, reported detection methods of target microorganisms, literature citations, genome information of target microorganisms that have been sequenced, related antibody information, etc. For example, by entering a microorganism name in the name column, you can view the information of the microorganism's Chinese and English aliases, related classifications (viruses, bacteria, fungi), belonging lists, subspecies branches, number of literature, number of antibodies, and number of genome sequences.
[0017] According to the above-mentioned classification of demand information, different subject pages (such as pathogenic microorganism home page, literature page, antibody page, genome information page) can be realized under different tags (different labels under different subject pages, pathogenic microorganism home page includes: pathogenic microorganism name, alias, Chinese name, related source list, branch information, taxonomic domain, number of related literature, number of antibodies, number of genome sequences; literature page includes: date, detection method, related sequence, PubMed number, DOI number, literature name; antibody page includes: antibody accession number, heavy and light chain, antibody name, antibody length, host species, references; genome information The page includes: genome ID, genome name, taxonomic ID, taxonomic levels, bacteria / strain name, country of isolation, genome sequencing status, GenBank accession number, sequencing center, sequencing platform, sequencing depth, assembly method, presence of plasmid, number of contigs, genome size, number of CDS). If quantifiable, the columns are sorted from large to small or from small to large. If not quantifiable, they are sorted and filtered by type. If filtering is performed, the filtering can be matched according to the input words. The filtering conditions include species classification, detection method, article publication date, species name, genome sequencing status, etc.
[0018] The database construction module is used to construct a pathogenic microorganism database according to the screening results obtained by the information processing module.
[0019] The present invention also provides a gene database construction device, comprising: a memory and a processor; the memory is used to store a program; the processor is used to execute the program to implement each step of the method described in the present invention.
[0020] The present invention also provides a computer storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, each step of the method described in the present invention is implemented.
[0021] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: The present invention aims to solve the problem that a database with a relatively comprehensive list of pathogenic microorganisms does not display a phylogenetic tree diagram when displaying the classification of a certain microorganism, thereby failing to understand its evolutionary status. A phylogenetic tree diagram is attached to each pathogenic microorganism. The present invention can search and sort out the reported detection methods or means and the nucleic acid sequences used for a certain type of pathogenic microorganism, and provide the information required for biological control, detection and new drug development of the pathogenic microorganism in a targeted manner; The technical solution of the present invention combines the problem that the database of a certain type of pathogenic microorganism lacks comparative information with similar pathogenic microorganisms, summarizes the detection method of a certain pathogenic microorganism, and can compare and download the genomes and related information of multiple microorganisms according to any category in taxonomy; The present invention overcomes the problem of lack of comprehensive organization and visualization of information on pathogenic microorganisms, diseases caused by them, and antibodies related thereto, and the constructed database integrates and structures information on diseases, pathogenic microorganisms, related antibodies, antibody sequences and annotation information, genome data, sequencing information, and the like; In view of the technical problem that existing databases often cannot freely arrange and combine existing information and organize and download it, the database constructed by the present invention provides a relatively rich information category, and can be selected and arranged in sequence as needed, and supports customized arrangement and combination of specific pathogenic microorganisms and specific bacteria / strain information, and personalized organization and download of customized information, which provides convenience for scientific research; The present invention directly maps diseases to pathogenic microorganisms and then to antibody sequences, integrates information from multiple databases such as antibody databases, pathogenic microorganism databases, and literature databases, establishes a connection between antibody information and diseases, and lays a foundation for disease prevention, detection, and the development of new drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of the method for constructing a pathogenic microorganism database of the present invention; Figure 2 The taxonomic information of severe acute respiratory syndrome coronavirus collected in NCBI; Figure 3 A tree view of the taxonomy information of severe acute respiratory syndrome coronavirus in the database of the present invention is presented; Figure 4 A flow chart for the detection method and sequence acquisition of the present invention; Figure 5 Display of antibody information corresponding to pathogenic microorganisms in the database of the present invention; Figure 6 A flowchart of data acquisition and updating of the present invention; Figure 7 Display and download of customized information in the database of the present invention; Figure 8 A schematic diagram showing the flexibility of personalized arrangement and combination of genomic modules of the present invention (1); Fig. 9 A schematic diagram showing the flexibility of the personalized arrangement and combination of genomic modules of the present invention (2). DETAILED DESCRIPTION
[0023] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of protection defined by the claims of the present invention.
[0024] Example 1. Construction of a specific pathogenic microorganism database Figure 1 A flow chart of a method for constructing a database for a specific pathogenic microorganism of the present invention is given. In conjunction with the accompanying drawings of the present invention, in a specific embodiment of a method for constructing a pathogenic microorganism database provided by the present invention, the method includes: 1. Obtain taxonomic data (1) Determine the Chinese and English scientific names and aliases of the microorganisms to be searched According to the names of the pathogenic microorganisms in the list, search for the scientific name, alias, and abbreviation of the pathogenic microorganism, as well as the various names of the related diseases caused by it in the literature, academic websites such as NCBI, and search engines such as Bing. The specific operation method is: search for the pathogenic microorganism on an authoritative database (such as NCBI) or an authoritative search engine (such as Bing), determine the common name, scientific name, alias, abbreviation, etc. of the pathogenic microorganism, and lock the Chinese and English scientific name and alias of the disease caused by it through the detailed information of the pathogenic microorganism; For example: SARS-CoV, search for "SARS-CoV" in Bing, the corresponding infectious disease name is "Severe Acute Respiratory Syndrome", the corresponding Chinese name is "SARS coronavirus", and it has an alias "SARS-related coronavirus" in Wikipedia. Search for "SARS-CoV" in NCBI and find that its current scientific name is "Severe acute respiratory syndrome coronavirus", the common name is "SARS-CoV", and the informal name is "SARS coronavirus, SARS virus".
[0025] (2) Determine the eight-level taxonomic status and fungal, bacterial, or viral categories in the NCBI Taxonomy section and visualize the phylogenetic relationship diagram.
[0026] Determine the phylogenetic relationship of pathogenic microorganisms and obtain a phylogenetic tree. The specific operation method is: search the scientific name or alias of the pathogenic microorganism in authoritative databases such as NCBI, click on the information of the Taxonomy classification section, record its 8-level taxonomic units of domain, phylum, class, order, family, genus and species, and classify it according to fungi, bacteria and viruses, and then visualize its evolutionary taxonomic status in the form of a tree diagram; For example, SARS-CoV, search for "SARS-CoV" in the "TAXONOMY" category in NCBI to obtain the systematic classification of the microorganism. The retrieved results are as follows Figure 2 As shown, the phylogenetic tree (in both hierarchical and tree view styles) is rendered on the page using the data in the database.
[0027] In a specific embodiment of the present invention, JavaScript tree graph dynamic rendering technology is used, which can convert the data in the NCBI database into a visual phylogenetic tree, so that users can intuitively understand the systematic classification and evolutionary relationship of microorganisms. The following are the steps to create a JavaScript line tree graph: Create an HTML page and container: First, you need to create an HTML page and add a block-level element as a container for the line chart. This element needs to have a unique ID for subsequent reference.
[0028] HTML <title> Line Chart JS< / title> <style type="text / css">html, body, #container {width: 100%;height: 100%;margin: 0;padding: 0;}< / style> Import JavaScript files: Then, you need to import the corresponding JavaScript files.
[0029] HTML <script src="https: / / cdn.anychart.com / releases / 8.11.0 / js / anychart-base.min.js">< / script> Add data: Next, you need to add the data to be displayed in the line chart. This data is usually provided in the form of an array, and each array element represents a data point.
[0030] JavaScript var data = [ ["2003", 1, 0, 0], ["2004", 4, 0, 0], / / ... ["2022", 20, 22, 20] ]; Writing the visualization code: Finally, you need to write the JavaScript code to create the line chart and add the data to the chart.
[0031] JavaScript anychart.onDocumentReady (function () { / / Create a dataset var dataSet = anychart.data.set(data); / / Map all series var firstSeriesData = dataSet.mapAs({x: 0, value: 1}); var secondSeriesData = dataSet.mapAs({x: 0, value: 2}); var thirdSeriesData = dataSet.mapAs({x: 0, value: 3}); / / Create a line chart var chart = anychart.line(); / / Create series and name var firstSeries = chart.line(firstSeriesData); firstSeries.name("Roger Federer"); var secondSeries = chart.line(secondSeriesData); secondSeries.name("Rafael Nadal"); var thirdSeries = chart.line(thirdSeriesData); thirdSeries.name("Novak Djokovic"); }); Figure 2 The taxonomic information of coronaviruses collected in NCBI is given. Figure 3 gives a tree view of the taxonomic information in this database, and Figure 2 Compare, Figure 3 The phylogenetic tree diagram can clearly show the relationship and taxonomic status between the object of interest and related species or bacteria / strains, making the information more intuitive for users to receive.
[0032] 2. Genome data acquisition Collect the genome data of pathogenic microorganisms and related information such as sequencing and splicing. The specific operation method is: search on BV-BRC according to the name of the pathogenic microorganism, select the Genomes classification, enter the name of the pathogenic microorganism to be searched, and search for the genome number, genome name, classification number, taxonomic level, sequenced genome information, status, strain, GenBank accession number, spliced genome information, and evaluation indicators after assembly, including Contigs, assembly method, genome size, etc., the center where sequencing was carried out, platform, sequencing method, sequencing depth, plasmid, number of Contigs, number of CDS, host sample from which the pathogenic microorganism was isolated, geographical location of isolation, etc. (Genome ID, Genome Name, Taxon ID, Superkingdom, Kingdom, Phylum, Class, Order, Family, Genus, Species, Genome, Status, Strain, GenBank Accessions, Sequencing Center, Sequencing Platform, SequencingDepth, Assembly Method, Plasmids, Contigs, Size, CDS, Isolation Country, Host CommonName) records this part of information and stores it in the local storage. For example: "SARS-CoV", search on BV-BRC, select Genomes classification, enter "SARS-CoV" to search, click on one of them (such as the first data), and the results you can get include bacteria / strain name, GenBank Accession, genome size, number of contigs, Genome ID, GenomeName, taxonomic information, genome status, data completion date, sequencing platform, assembly method, GC content, ContigL50, Contig N50, CDS number, CDS ratio, assumed CDS number and ratio, bacteria / strain collection year and date, isolated country, and isolated substrate information.
[0033] 3. Pathogen detection methods and acquisition and visualization of application sequences Collect literature information related to pathogen detection. The specific operation method is: according to the alias of each pathogen and the name of the disease it causes, use this as a keyword to search and crawl in databases such as PubMed. Specific operation method: obtain the document name, publication time, PubMed ID number, DOI number and other information of the document and record them, use "detect" and "microorganism name" to search the content in the title and abstract, and only when they are retrieved at the same time, the document is considered to be retrieved, "microorganism name" is searched by word, and "detect" only needs to appear (for example, detect, detection, detected, etc. are all valid). Then, according to the selected literature, according to different detection methods (microbial detection methods recognized by the academic community), about 20 types (qPCR (including: real time RT-PCR, RT-qPCR), dPCR, LAMP (RT-LAMP), RPA, EXPAR, SDA, NASBA, HDA, RCA, immunoChromatography, ELISA, ELFA, TRFIA, CILA, MAFIA, HTRF, AlphaLISA, xMAP, Luminexx-TAG, Aptamer), the sequence data in the literature is extracted. The detection methods and related sequence information involved in the detection of the pathogenic microorganisms in the literature are searched in batches and recorded, and all processed literature data are archived and displayed through UI rendering (UI rendering involves three technologies: HTML, CSS and JavaScript. First, the browser parses HTML / SVG / XHTML files and generates a DOM Tree by parsing these three files. Next, the browser parses the CSS file to generate a CSS rule tree. The last step is the execution of the script, and the DOM Tree and CSS Rule Tree are operated through the DOM API and CSSOM API). Visualization and literature jump; For example: "SARS-CoV", search "detect SARS-CoV" in PUBMED, grab all the retrieved literature data URLs, parse the URLs, process the required data into json format and store them in the database, then parse the corresponding literature name, publication time, PubMed ID number, DOI number and other information from the downloaded literature data and save them in the database, and then further parse the saved data, for example, use PYTHON's ELEMENTTREE to parse, where PYTHON's FITZ is used to parse PDF, RE (regular expression) is used to parse strings, and XPATH is used to parse documents; for each literature entry, extract the required information (such as the name of the pathogenic microorganism, the detection method used), and store the extracted information in the database.
[0034] Filter out the detection methods related to the pathogenic microorganisms in the literature according to the keywords. The specific steps are: first, when parsing each literature entry, obtain the abstract and full text of the literature, use regular expressions to filter out the keywords related to the detection method in the obtained text, and then create a list containing the keywords of the required detection method; traverse the keyword list and search for matching keywords in the text; if matching keywords are found, it means that the literature involves the corresponding detection method, and the relevant information of the literature is stored in the database. Figure 4 Shown is a flowchart for obtaining reported detection methods and sequences used therein for a specific object.
[0035] Matching screening and sequence data extraction are performed according to the 20 methods provided in the early stage (rules: a. Generally, it only contains four letters ATCG, which are randomly arranged (other uppercase letters may also appear, and spaces may appear in the letters. After removing the spaces, the proportion of the four letters ATCG in the entire sequence should be controlled above 80%. The letters allowed to appear can be [ATCGRNX], which is case-insensitive. In addition, the sequence is allowed to contain no more than 20% of N or X or R, and only one of N, X or R is allowed to appear). B. The length of the sequence must be more than 10 letters. C. The sequence data generally contains key information such as "5'", "3'", "primer(s)", "probe(s)").
[0036] 4. Antibody data acquisition and visualization Collect and obtain information on antibodies related to pathogenic microorganisms. The specific operation method is: search for related antibodies in the abYsis database according to the name of the pathogenic microorganism, record the species from which the antibody is isolated, the antibody accession number, the antibody name, the light and heavy chains, the amino acid sequence, the annotation information of different regions, the relevant literature information, etc., and classify and organize the data. The UI end visualizes its content and functions: Use HTML Tags to create a table, each cell uses Label, header use Tags, display the antibody name, light and heavy chains, species from which the antibody was isolated, and related literature information in a table. Tags are used to create cards, and then CSS is used for style design. The amino acid sequence and annotation information of different regions are displayed in the form of cards. Each card contains partial information. Users can click on the card to view more detailed information and download it. For example: "SARS-CoV", search for "SARS-CoV" in AbYsis, grab the "heavy chain table" and "light chain table" data respectively, click on an antibody accession number in the table, such as "2dd8_H", and the data in the jump page includes the annotation information of the antibody, the sequence of the antibody, and the antibody sequence annotation method (including heavy chain and light chain information). Then save all the acquired data in the database, and through the rendering of the front and back ends, the data can be directly displayed and downloaded on the page. Specifically, JAVASCRIPT is used to parse the data structure and display the table fields in sequence, and the number of sequence rows and sequence positions are calculated according to the sequence length through MATH and other digital calculation functions and ARRAY operations, and then CSS is used to beautify and finally presented on the page. The rendering result is as follows: Figure 5 As shown, Figure 5 The antibody details page shown can display the antibody accession number, antibody amino acid sequence, sequence and length of each region in the antibody's Kabat annotation, antibody light and heavy chain information, antibody name, antibody source literature, antibody source host species, antibody length and other information.
[0037] 5. Update of data in the database All data retrieved from target pathogenic microorganisms need to be updated regularly. The data retrieval and update process diagram is as follows: Figure 6 As shown: Each time the target is searched in different databases through the retrieval unit, the obtained data is stored in the local storage unit after the aforementioned analysis, and is presented on the web page through later rendering. A regular update time is set through the program, such as a cycle of every 30 days, and the stored data is re-retrieved in the database for the target object, and the obtained data is stored instead of the original data.
[0038] 6. Implement sorting, filtering and exporting operations under different tags on different topic pages according to needs, and realize the flexibility of personalized arrangement and combination display of genome modules. The specific implementation methods are as follows: The data obtained above are stored in the local storage. Through UI rendering, the processed data are displayed, and the arrangement and combination of information can be customized according to needs. There are currently 24 information cards (Genome ID, Genome Name, Taxon ID, Superkingdom, Kingdom, Phylum, Class, Order, Family, Genus, Species, Genome Status, Strain, GenBank Accessions, Sequencing Center, Sequencing Platform, Sequencing Depth, Assembly Method, Plasmids, Contigs, Size, CDS, Isolation Country, Host Common Nane), which can be combined in any number and arranged in any order. And you can download the Excel version of all information according to the custom selected sample.
[0039] For example: "MYCOPLASMA PNEUMONIAE", search for "MYCOPLASMA PNEUMONIAE" in this database, click the content under "GENOME SEQUENCE COUNT", and jump to the details page. In the details on the right, select "Customize display columns" to select the column data you want to display; by sorting different columns, you can view the sorted data; select the corresponding data through the selection box (you can also select all with one click), after selecting the data, click "Export file" or "Download sequence", you can export or download the selected data (see Figure 7 ).
[0040] Figure 8 and Fig. 9 A schematic diagram showing the flexibility of personalized permutations and combinations of genomic modules of the present invention.
[0041] Figure 8 Displays a custom selection combination of data-related information cards: There are currently 24 optional information cards, including Genome ID, GenomeName, Taxon ID, Superkingdom, Kingdom, Phylum, Class, Order, Family, Genus, Species, Genome Status, Strain, GenBank Accessions, Sequencing Center, SequencingPlatform, Sequencing Depth, Assembly Method, Plasmids, Contigs, Size, CDS, IsolationCountry, and Host Common Nane. Select the required information card according to your needs. Figure 8 The chart shows seven types of information cards, including Genome ID, Genome Name, Genus, Species, Genome status, Strain, and Host commonname. The default arrangement of these seven types of information can be displayed in the table.
[0042] Fig. 9 Displays the custom sorting order of data-related information cards: After customizing and selecting the required information cards, you can customize the sorting order of different information. Fig. 9 The information card combination displayed is GENBANK ACCESSIONS, GENOMEID, GENUS, SPECIES, STRAIN, ISOLATION COUNTRY, and HOST COMMON NAME. By customizing the sorting order and putting the HOST COMMON NAME of the information card in the first place, the sorted results can be displayed in the table below.
[0043] Comparative Example: Take "SARS-CoV" as an example. The search and analysis path of the prior art is 1. To view the antibody information related to "SARS-CoV", you first need to open the abYsis official website, then enter "SARS-CoV" in the search box, click search, and you can see the data of antibodies related to the pathogen.
[0044] 2. To view the genome data related to "SARS-CoV", you first need to open the BV-BRC official website, select the "Genomes" category, enter "SARS-CoV" in the search bar, click search, and you can see the data related to the genome of this microorganism.
[0045] 3. To view the literature related to the detection methods involved in "SARS-CoV", you first need to enter the search keywords and the corresponding logical relationships in PUBMED (such as ["SARS-CoV" and "detect"] or ["SARS-CoV" and "detected"] or ["SARS-CoV" and "detection"], etc.), so that you can find the corresponding literature, and then manually browse and screen to obtain the detection methods and sequence information used.
[0046] 4. To view the scientific name and evolutionary classification of "SARS-CoV", you need to enter "SARS-CoV" in the "Taxonomy" category in NCBI and click on the details page to see the corresponding data.
[0047] Taking the database of the present invention as an example, to search for antibodies, genomes, literature on detection methods, detection methods and sequences used, and phylogenetic information of "SARS-CoV", one only needs to search for "SARS-CoV" in the "Name" column on the database homepage, and then directly jump to the corresponding page by clicking on the four labels of antibody count, genome sequence count, literature count, and branch to obtain antibody information, detection methods, primer sequences used, source literature, genome details, taxonomic status, etc., and add corresponding classification standards and Chinese names. All information can be obtained with one-click operation, which greatly improves the retrieval speed and efficiency.
[0048] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0049] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for constructing a pathogenic microorganism database, characterized in that: The method comprises: Step 1. Data acquisition: Search for data information of pathogenic microorganisms in a set literature site or search engine according to the name of the target pathogenic microorganism; Step 2. Visualization of phylogenetic relationships: Determine the phylogenetic relationship of the target pathogenic microorganisms based on the data obtained in step 1, obtain a phylogenetic tree and visualize it; Step 3. Information collection: Collect the genome data information of the target pathogenic microorganism, pathogenic microorganism detection information, and pathogenic microorganism related antibody information; Step 4. Information processing: First, classify the information to be processed. The information is classified into: Chinese and English scientific names and aliases of the target microorganisms, diseases caused, reported detection methods of the target microorganisms and literature citations, genome information of the target microorganisms that have been sequenced, and antibody information related to the target microorganisms; then sort the information under different subject pages and different tags. If quantifiable, sort by the number relationship from large to small or from small to large; if not quantifiable, sort and filter by type. If filtering is performed, it can be matched and filtered according to the input words. The filtering conditions include species classification, detection method, article publication date, species name, and genome sequencing status; Step 5. Based on the screening results obtained in step 4, a pathogenic microorganism database is constructed.
2. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 1, the search engine includes NCBI, BV-BRC, abYsis and / or BING search.
3. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 1, the data information of the pathogenic microorganism includes scientific name, alias, abbreviation, and the name of the related disease caused by it.
4. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: The method for obtaining and visualizing the phylogenetic tree in step 2 is to search the NCBI database by the scientific name or alias of the pathogenic microorganism, enter the Taxonomy classification section, record the 8-level taxonomic status of the pathogenic microorganism by domain, phylum, class, order, family, genus and species, and classify it according to fungi, bacteria and viruses, and use JAVASCRIPT tree map dynamic rendering technology to convert the data in the NCBI database into a visualized phylogenetic tree; The JAVASCRIPT tree map dynamic rendering technology includes: creating an HTML page and a container, introducing a JAVASCRIPT file, adding data to be displayed in the tree map, writing JAVASCRIPT code to create the tree map, and adding data to the tree map to display it in the form of a tree map.
5. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 3, the collection of the genome data information of the pathogenic microorganism is specifically as follows: search on BV-BRC according to the name of the pathogenic microorganism, select the Genomes classification, enter the name of the pathogenic microorganism to be searched, first store all the data in the database, then extract the target data from the database, and then archive it. The target data includes sequenced genome information and spliced genome information. The spliced genome information includes Contigs, assembly method, size, sequencing center, platform, sequencing method, sequencing depth, host sample and geographical location from which the pathogenic microorganism is isolated, and record the ID number or Accession number for finding relevant information on NCBI.
6. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 3, the collection of pathogen detection information is specifically as follows: according to the alias of each pathogen and the name of the disease caused by it, the PubMed database is searched and crawled as keywords to obtain the document name, publication time, PubMed ID number, DOI number of the document and record them, and then "detect" and "microorganism name" are used to retrieve the documents that appear in the title and abstract at the same time, wherein "microorganism name" is searched as a word, and "detect" only needs to appear; then, according to the screened documents, according to different detection methods: qPCR, dPCR, LAMP, RPA, EXPAR, SDA, NASBA, HDA, RCA, immunoChromatography, ELISA, ELFA, TRFIA, CILA, MAFIA, HTRF, AlphaLISA, xMAP, Luminexx-TAG, Aptamer, the sequence data in the document is extracted, and the processed data is displayed through UI rendering.
7. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 3, the collection of pathogen-related antibody information is as follows: searching for relevant antibodies in the abYsis database according to the name of the pathogen, recording the species from which the antibody is isolated, the antibody accession number, the antibody name, the light and heavy chains, the amino acid sequence, the annotation information of different regions, the antibody literature information, and classifying and arranging the data. The UI end visualizes its content and functions: using HTML Tags to create a table, each cell uses Label, header use Tags, the antibody name, light and heavy chains, species from which the antibody was isolated, and related literature information are displayed in a table format; at the same time, HTML Tags are used to create cards, and then CSS is used for style design. The amino acid sequence and annotation information of different regions are displayed in the form of cards. Each card contains partial information, and users can click on the card to view more detailed information and download it.
8. The method for constructing a pathogenic microorganism database according to claim 1, characterized in that: In step 4, the subject page includes a pathogenic microorganism home page, a literature page, an antibody page, and a genome information page, wherein the tags of the pathogenic microorganism page include: pathogenic microorganism name, alias, Chinese name, related source list, branch information, taxonomic domain, number of related literature, number of antibodies, and number of genome sequences; the tags of the literature page include: date, detection method, related sequence, PubMed number, DOI number, and literature name; the tags of the antibody page include: antibody accession number, heavy and light chains, antibody name, antibody length, host species, and references; the tags of the genome information page include: genome ID, genome name, taxonomic ID, various taxonomic levels, bacteria / strain name, country of isolated bacteria / strain, genome sequencing status, GenBank accession number, sequencing center, sequencing platform, sequencing depth, assembly method, presence or absence of plasmid, number of contigs, genome size, and number of CDS.
9. A pathogenic microorganism database constructed by using the method described in any one of claims 1 to 8, characterized in that: The pathogenic microorganism database includes the following modules: Data acquisition module: The data acquisition module searches for data information of pathogenic microorganisms in set literature sites and search engines according to the names of target pathogenic microorganisms; Phylogenetic relationship visualization module: the phylogenetic relationship visualization module determines the phylogenetic relationship of the target pathogenic microorganism based on the data acquired by the data acquisition module to obtain a phylogenetic tree and visualize it; Information collection module: The information collection module collects genome data information of target pathogenic microorganisms, pathogenic microorganism detection information, and pathogenic microorganism related antibody information; Information processing module: The information processing module first classifies the information to be processed, and the information classification is as follows: Chinese and English aliases of target microorganisms, diseases caused, reported detection methods of target microorganisms, literature citations, genome information of target microorganisms that have been sequenced, and antibody information related to target microorganisms; then the required information is sorted under different subject pages and different tags, and if quantifiable, it is sorted by the quantitative relationship from large to small or from small to large; If it is not quantifiable, sort and filter by type. If filtering is performed, it can be matched and filtered according to the input words. The filtering conditions include species classification, detection method, article publication date, species name, and genome sequencing status.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Preparation method of new antibiotic and platform system based on method
CN103159856A
Methods, systems and processes of determining transmission paths of infectious agents
CN108475297A
Porcine breed and canine brucella identification method based on specific sequence and SNP (Single Nucleotide Polymorphism) site
CN117904338A
Database for microbial investigations
WO2005036369A2
Systems and methods relating to bioinformatics
WO2023172923A2