A high-throughput sequencing data visualization method, device, medium and equipment

CN116153421BActive Publication Date: 2026-09-29陆燊
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310196131.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2026-09-29
Estimated Expiration
2043-03-02

AI Technical Summary

Technical Problem

这类传统的可视化方法的缺陷是:第一,由于高通量测序产生的reads数量巨大,把整个文件读进内存需要很长时间;第二,读入整个文件需要的内存远远超过了普通计算机的内存上限

Benefits of technology

[0039]本发明结合高通量测序数据特性,将不同类型的数据按照不同的规则进行处理和存储在不同的数据库中,基于此后续可以在对应的数据库中查找对应的数据,提高了高通量测序数据的存储和查询效率;且可实现随着测序数据量的增加,交互可视化效率依然保持稳定。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116153421B_ABST
    Figure CN116153421B_ABST
Patent Text Reader

Abstract

The application belongs to the field of biological information processing, and provides a high-throughput sequencing data visualization method, device, medium and equipment. The high-throughput sequencing data visualization method comprises the following steps: obtaining high-throughput sequencing data, and extracting visualization image data and corresponding descriptive information data from the high-throughput sequencing data; performing hierarchical fragmentation processing on the visualization image data based on a preset hierarchical fragmentation mechanism, and constructing a visualization image data query index; storing the hierarchical fragmented data of the visualization image data together with the query index in the nodes in a visualization image database for storage; storing the descriptive information data in a descriptive information database according to a descriptive information data query index; wherein the visualization image data query index and the descriptive information data query index have common unique identification information; rendering and displaying the visualization image data called based on the visualization image data query index, and displaying the matching descriptive information data based on the descriptive information data query index and the rendered visualization image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics processing, and in particular relates to a method, apparatus, medium and device for high-throughput sequencing data visualization. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In existing technologies, high-throughput sequencing sequence visualization typically involves acquiring the sequencing file output by the sequencer and reading it entirely into local memory. Then, it compares the data with a reference genome fragment and finally displays the comparison results graphically. The drawbacks of this traditional visualization method are: first, due to the massive number of reads generated by high-throughput sequencing, reading the entire file into memory takes a very long time; second, the memory required to read the entire file far exceeds the memory limit of a typical computer. For example, the existing IGV (Integrated Genome Viewer) stores all data locally, requires data to be read into memory during runtime (approximately 3-5 seconds startup time), and has significant memory requirements, exceeding 1GB during runtime, consuming substantial local resources and potentially causing computer lag.

[0004] Mainstream web-based genome browsers, such as the UCSC Genome Browser, store all data on the server side. Query requests are sent via a web browser, and images are rendered on the server and then sent back to the web browser for display. While storing data on the server side solves the storage limit problem to some extent, this type of visualization has several drawbacks: First, since both querying and image rendering are performed on the server side, the visualization is delayed and cannot be displayed automatically. Second, rendering images on the backend consumes excessive computing resources, increasing server load, resulting in slow visualization speeds and a poor user experience. JBrowse does not require a database and can be deployed locally or in the cloud. It extracts sequencing data for the corresponding display region directly from high-throughput data files using an HTTP chunked request mechanism, and performs visualization rendering on the client side after the data is returned, reducing the server's computing load. However, this method also has a drawback: since the data is directly obtained from the sequencing files, the speed at which the data is returned to the client becomes increasingly slower as the file size increases, leading to visualization delays and stuttering.

[0005] In summary, the inventors have found that existing methods for visualizing high-throughput sequencing sequences all suffer from problems such as increasingly slower data interaction between the client and server, delayed visualization response, and lag as the file data volume increases. Summary of the Invention

[0006] To address the technical problems existing in the background art, the present invention provides a method, apparatus, medium and device for high-throughput sequencing data visualization, which not only improves the storage and query efficiency of high-throughput sequencing data, but also maintains stable interactive visualization efficiency as the amount of sequencing data increases.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The first aspect of the present invention provides a method for visualizing high-throughput sequencing data.

[0009] A method for visualizing high-throughput sequencing data, comprising:

[0010] Acquire high-throughput sequencing data and extract visualization image data and corresponding descriptive information data from the high-throughput sequencing data;

[0011] The visualized image data is processed by hierarchical segmentation based on a preset hierarchical segmentation mechanism, and a visualized image data query index is constructed. The hierarchical segmented data of the visualized image data, along with its query index, is distributed to nodes in the visualized image database for storage. The descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier.

[0012] The visualized image data retrieved from the visualized image data query index is rendered and displayed, while the descriptive information data matching the rendered visualized image data is displayed based on the descriptive information data query index.

[0013] As one implementation method, in the process of performing hierarchical and segmented processing on the visualized image data based on a preset hierarchical and segmented mechanism, the calculation rule for the relationship between the hierarchical level and the number of segments is as follows: when the hierarchical level is n, the number of segments corresponding to each level is 2 raised to the power of n; where n is a natural number.

[0014] As one implementation method, in the process of performing hierarchical segmentation on the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1≤M*2^N / C<2; where C represents the chromosome length, i.e. the total number of chromosome bases, N represents the number of levels that this chromosome sequence needs to be divided into, and also represents the last level of the hierarchical segmentation, the Nth level; M represents the minimum segmentation sequence.

[0015] As one implementation method, in the process of performing hierarchical segmentation processing on the visualized image data based on a preset hierarchical segmentation mechanism, the minimum segmentation sequence refers to the length of the base sequence represented by each segment in the last level of the hierarchy, which is determined by the base resolution, the client image rendering speed, and the network transmission speed.

[0016] As one implementation method, in the process of performing hierarchical segmentation processing on the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the length of the base sequence represented by each segment in the nth level is: M*2^(Nn), where N represents the last level of the hierarchical process, n represents the nth level, that is, any level from 0 to N, and M represents the minimum segmentation sequence.

[0017] As one implementation method, in the process of performing hierarchical segmentation on the visualized image data based on a preset hierarchical segmentation mechanism, after the same chromosome is processed by hierarchical segmentation, the sum of the base sequence lengths represented by all segments in each level is M*2^N, where N represents the last level of the hierarchical segmentation and M represents the smallest segmentation sequence.

[0018] As one implementation method, the visual image data query index is composed of the base starting positions of each level and each segment, following the principle of first hierarchical and then segmented.

[0019] A second aspect of the present invention provides a high-throughput sequencing data visualization device.

[0020] A high-throughput sequencing data visualization device, comprising:

[0021] The data extraction module is used to acquire high-throughput sequencing data and extract visualization image data and its corresponding descriptive information data from the high-throughput sequencing data.

[0022] The data storage module is used to perform hierarchical and segmented processing on the visualized image data based on a preset hierarchical and segmented mechanism, and to construct a visualized image data query index. The hierarchical and segmented data of the visualized image data, along with its query index, are distributed to nodes in the visualized image database for storage. Descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image module is used to ensure that the data query index and the descriptive information data query index share a common unique identifier.

[0023] The data display module is used to render and display the visualized image data retrieved based on the visualized image data query index, and at the same time, display the descriptive information data that matches the rendered visualized image data based on the descriptive information data query index.

[0024] As one implementation method, in the process of performing hierarchical and segmented processing on the visualized image data based on a preset hierarchical and segmented mechanism, the calculation rule for the relationship between the hierarchical level and the number of segments is as follows: when the hierarchical level is n, the number of segments corresponding to each level is 2 raised to the power of n; where n is a natural number.

[0025] As one implementation, in the data storage module, during the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1≤M*2^N / C<2; where C represents the chromosome length, i.e., the total number of chromosome bases, N represents the number of levels that this chromosome sequence needs to be divided into, and also represents the last level of the hierarchical segmentation, the Nth level; M represents the minimum segmentation sequence.

[0026] In the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the minimum segmentation sequence refers to the length of the base sequence represented by each segment in the last level of the hierarchical segmentation, which is determined by the base resolution, the client image rendering speed and the network transmission speed.

[0027] As one implementation method, in the data storage module, during the hierarchical segmentation process of the visualized image data based on the preset hierarchical segmentation mechanism, the calculation rule for the length of the base sequence represented by each segment in the nth level is: M*2^(Nn), where N represents the last level of the hierarchical process, n represents the nth level, that is, any level from 0 to N, and M represents the minimum segmentation sequence.

[0028] As one implementation, in the data storage module, during the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, after the same chromosome is segmented, the sum of the base sequence lengths represented by all segments in each level is M*2^N, where N represents the last level of the hierarchical segmentation and M represents the smallest segmentation sequence.

[0029] In one implementation, the visualization image data query index in the data storage module is composed of the base starting positions of each level and each segment, following the principle of first hierarchical and then segmented.

[0030] A third aspect of the invention provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps in the high-throughput sequencing data visualization method described above.

[0031] A fourth aspect of the present invention provides a high-throughput sequencing data visualization device, characterized in that it includes a server and a client.

[0032] The server is configured as follows:

[0033] Acquire high-throughput sequencing data and extract visualization image data and corresponding descriptive information data from the high-throughput sequencing data;

[0034] The visualized image data is processed by hierarchical segmentation based on a preset hierarchical segmentation mechanism, and a visualized image data query index is constructed. The hierarchical segmented data of the visualized image data, along with its query index, is distributed to nodes in the visualized image database for storage. The descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier.

[0035] Render the visual image data retrieved from the visual image data query index;

[0036] The client is configured as follows:

[0037] It displays the rendered visual image data, and simultaneously displays descriptive information data that matches the rendered visual image data based on the descriptive information data query index.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] This invention combines the characteristics of high-throughput sequencing data, processes and stores different types of data according to different rules in different databases, and based on this, the corresponding data can be searched in the corresponding database, thus improving the storage and query efficiency of high-throughput sequencing data; and it can also ensure that the interactive visualization efficiency remains stable as the amount of sequencing data increases.

[0040] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0041] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0042] Figure 1 This is a schematic diagram showing the comparison of a short sequence (FASTQ file) to a reference genome (FASTA file);

[0043] Figure 2 This is a flowchart of the high-throughput sequencing data visualization method according to an embodiment of the present invention;

[0044] Figure 3 It is the process of visual display on the client side;

[0045] Figure 4A sample contains different data types, corresponding to different visualization image channels. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0048] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] Example 1

[0050] like Figure 2 As shown, this embodiment provides a method for visualizing high-throughput sequencing data, which includes:

[0051] Step 1: Acquire high-throughput sequencing data and extract visualization image data and its corresponding descriptive information data from the high-throughput sequencing data;

[0052] In step 1, the high-throughput sequencing data consists of:

[0053] Reference genome file: FASTA file. It contains the complete genome of a species, specifically the sequence information of each chromosome (ATGC…), where one letter represents one base. For example, humans have 23 chromosomes, approximately 3 billion bases.

[0054] Gene annotation files: GFF / GTF / BED files, etc. Gene annotation files contain the start and stop base positions of macromolecules such as genes, transcripts, introns, exons, and proteins.

[0055] Raw sequencing data files: FASTQ files, short sequences (200bp-300bp) generated by high-throughput sequencing, file data size = sequencing depth * reference genome length * coverage.

[0056] Comparison results files: The raw sequencing data (FASTQ) is compared (mapped) to the reference genome using FASTA to obtain details of how each short sequence is mapped onto the reference genome, such as start and end sites, orientation (positive or negative), etc. These are typically SAM and BAM files. The content of the BAM file is the same as that of the SAM file; only the data storage format differs. A SAM file stored in binary code becomes a BAM file.

[0057] Other files: such as VCF (Variant Call Format) files that record base mutation site information, can be downloaded from SNPdb to find the corresponding species' VCF file.

[0058] Coverage file: The coverage of short sequences in the FASTQ file for each base in the genome is calculated using software such as Samtools.

[0059] High-throughput sequencing data processing workflow and the relationships between individual files:

[0060] The data conversion device processes files 1, 2, 3, 4, and 5 to generate visual image data and descriptive information data suitable for direct display on the client side.

[0061] Data 1, 2, and 4 can be directly processed by the data conversion device. Data 3 needs to be compared with 1 (reference genome) to generate comparison result file 4, and then 5 is generated from file 4. All information from files 3 and 5 can be found in file 4.

[0062] The base position information in files 2, 3, 4, and 5 corresponds one-to-one with file 1 as the coordinate system. Essentially, all file contents correspond to a dataset of one-dimensional sequences (ATGC…).

[0063] Constant files: The size of the file data is related to the length of the reference genome. There will be differences between different species, but the data will not change between the same species. For example: 1, 2, 5.

[0064] Variable files: The data volume of these files is related to the sequencing depth and is unknown. For example, files 3 and 4, after analysis, show that constant files 1, 2, and 4, due to their stable data volume, will not affect the efficiency of visualization data in queries. Therefore, the factor causing unstable visualization efficiency lies in files 3 and 4. Since file 4 contains all the information from file 3, we only need to focus on file 4.

[0065] Step 2: Based on a preset hierarchical and segmented mechanism, the visualized image data is processed into hierarchical and segmented data, and a visualized image data query index is constructed. The hierarchical and segmented data of the visualized image data, along with its query index, are distributed to nodes in the visualized image database for storage. The descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier.

[0066] Specifically, the rules for hierarchical segmentation of each chromosome sequence in the genome are summarized as follows:

[0067]

Rule 1

[0068]

Rule 2

[0069]

Rule 3

Rule 2

[0070]

Rule 4

[0071]

Rule 5

[0072] The specific usage of the above rules is as follows:

[0073] First, based on base resolution, client image rendering speed, and network transmission speed, the minimum segmentation sequence length was determined to be M [Rule 3]. Given the chromosome length C, it was substituted into [Rule 2] to calculate: 1 ≤ M * 2^N / C < 2. This formula means that the total number of base sequences (length) contained in all segments of the last level is greater than or equal to the number of base sequences (length) of the current chromosome (if this restriction is not applied, N can be infinitely small), but less than twice the total number of base sequences (length) contained in the current chromosome (if this restriction is not applied, N can be infinitely large). Knowing M and C, a unique value N can be directly substituted into the formula to calculate, representing the chromosome sequence that needs to be divided into N levels.

[0074] Knowing the level N, we can calculate: a. The number of segments contained in the nth level (n represents any level from 0 to N): 2^n (2 to the power of n) [Rule 1] b. The length (or range) of the base sequence represented by each segment in the nth level: M(minimum segmentation sequence * 2^(Nn) [Rule 4]

[0075] The total number of bases in level n [Rule 5] = 2.a * 2.b = the number of segments in level n (where n represents any level from 0 to N) * the length (or range) of the base sequence represented by each segment in level n = 2^n * M * 2^(Nn) = M * 2^N. Therefore, it is clear that Rule 5, which states that the total number of bases in any level of the same chromosome is equal to M * 2^N, is entirely true. Thus, formulas 1, 2, 3, 4, and 5 above are logically consistent in their calculations.

[0076] storage:

[0077] Extract the ID and location information of each sequence from the comparison result file, and then use hierarchical tiling technology (similar to the vector tile technology of maps, the difference being that the vector tile technology processes two-dimensional data, while this technology processes one-dimensional data) to generate a visualization image that can be directly rendered on the client side. Store it in the visualization information database according to certain indexing rules for displaying frequently interactive visualization information.

[0078] The complete textual descriptive information in the comparison results file is stored in a distributed database, and the index is the ID of each sequence.

[0079] There is a one-to-one correspondence between the ID of each sequence in the visual image database and the ID stored in the descriptive information database.

[0080] Step 3: Render and display the visualized image data retrieved from the visualized image data query index, and at the same time display the descriptive information data that matches the rendered visualized image data based on the descriptive information data query index.

[0081] Client-side interaction:

[0082] The client is responsible for rendering the visualization information into a visualization image, while the server is responsible for transmitting information between the client and the database over the network. The database is only involved in searching for visualization images and descriptive information data (preprocessing high-throughput sequencing file data).

[0083] When a user's interaction causes a change in the actual base range displayed on the reference genome corresponding to the visualized image, the client calculates the storage level of the visualized image data in the database and the actual base range on the reference sequence based on the currently selected window pixel range and translation scaling scale. The client then sends a request to the server with these parameters. The server builds an index, finds the visualized image data according to the index, returns it, and then renders it into the corresponding visualized image through the client.

[0084] Before the display scale or window position is about to trigger a change in the visualization level, the client will return the visualization image data of the next base range to be displayed to the client for rendering in advance, so that the display range can be smoothly transitioned.

[0085] When a user's interaction (hover, click, touch, etc.) triggers the display of descriptive information, the client calculates the ID information corresponding to the visualization selected by the user based on the position of the clicked viewport and the zoom level. This ID information is then used as an index to search the descriptive information database, and the client is then returned to display the detailed information.

[0086] Query:

[0087] Data is retrieved from databases of visualized image data and descriptive information via network requests.

[0088] When user interaction changes the visualization scale and range (zooming, panning, or directly determining the visualization range through the search box, etc.), the client calculates the layering of the vector information corresponding to the current visualization area in the database and the base range of the actual sequence, and then sends this information to the server. The server combines this information with the database indexing rules to find the corresponding visualization image data and then sends it back to the client for rendering.

[0089] When a user's interaction triggers the display of descriptive information, the client calculates the target visualization image corresponding to the specific pixel position triggered by the user, obtains the vector information for rendering this part of the image, and sends the ID contained therein to the server. The server searches for an index that matches this ID in the distributed storage of descriptive information, obtains the complete descriptive information of the corresponding sequence, and returns it to the client for display.

[0090] Specifically, in estimating the data volume of FASTQ and SAM files, sequencing depth refers to the ratio of the total number of bases (bp) obtained from sequencing to the genome size, i.e., sequencing depth = data volume / reference genome size. Alternatively, it can be understood as the average number of times each base in the genome is sequenced.

[0091] Sequencing coverage refers to the proportion of the entire genome obtained through sequencing. Alternatively, it can be understood as the proportion of regions (or bases) on the genome that have been detected at least once.

[0092] Data volume calculation formula: Data volume = Sequencing depth * Genome size * Sequencing coverage.

[0093] Taking the human genome as an example, it contains approximately 3 billion bases, with each base occupying one byte. The human genome sequence data is approximately 3GB. Assuming a sequencing depth of 50x (50x is already very deep) and a coverage of 98%, it would generate approximately 150GB of FASTQ data. The SAM file generated after comparison is approximately 2-3 times larger than the FASTQ file. The text data from 4 samples could exceed 1TB (1024GB). Therefore, fast and efficient querying of this data must be achieved through database retrieval, such as... Figure 1 As shown.

[0094] Interactive visualization efficiency is affected by three factors: database query efficiency, network transmission speed, and client rendering efficiency.

[0095] Interactive visualization requires querying two types of data: visualization image data, which is related to the formation of the visualization image, mainly the positional information of bases in the genome; if the client's interactive behavior causes the base sequence range to be visualized to change (therefore not limited to translation and scaling), making it impossible to draw the visualization image based on the existing visualization image data, the task of querying the corresponding visualization image data will be triggered, and the new visualization image data will be returned to the client for rendering.

[0096] Descriptive information data: Information displayed when annotating and describing the content of a visual image during interaction (generally through clicking, hovering, or touch screen).

[0097] Visual image data and descriptive information data are separated during the interaction process. For example, when rendering a visual image, only the visual image data is used, and detailed descriptive data is not displayed. Only after the image rendering is completed, and the interaction determines whether to display a visual image with detailed description information, will the specific descriptive information be queried and returned to the front end for display.

[0098] In short, to maintain query efficiency, it's crucial to ensure efficient retrieval of visual images and descriptive information data related to the four documents. Simultaneously, segmenting queries according to the interaction logic will reduce the server's data query load.

[0099] To ensure efficient querying of visualized graphical data, the following design is adopted:

[0100] Visualized Graphical Data Storage and Query Design:

[0101] A hierarchical segmentation strategy is developed based on the reference genome sequence length of the current species (the mechanism is inspired by the vector tile technology used in map scenes);

[0102] Extract the information used for visualization in file 4 (short sequence ID, start base, stop base, positive and negative strand information, mutation site information), and generate vector information according to the hierarchical segmentation mechanism;

[0103] Build an index according to hierarchical sharding rules and distribute the sharding information evenly to different database nodes for storage;

[0104] The information displayed varies at different scaling scales. The visualized image data only retains information related to image generation at different scaling scales, which can effectively reduce the storage space occupied by the visualized image data and improve query and network transmission efficiency.

[0105] Visual image hierarchical and segmentation mechanism:

[0106] Each chromosome sequence (file 1) of the reference genome is segmented into hierarchical pieces (similar to vector tile technology). The segmentation rules are shown in Table 1, which provides the specific segmentation rules:

[0107] Table 1 Rules for Hierarchical Segmentation

[0108] 0 2^0=1 M*2^N 1 2^1=2 M*2^(N-2) 2 2^2=4 M*2^(N-4) 3 2^3=8 M*2^(N-8) 4 2^4=16 M*2^(N-16) ... ... ... N 2^N M

[0109] The final level of segmentation follows the rule: 1 <= M * 2^N / C < 2, where C is the chromosome length, M is the minimum segmentation sequence, and n represents the level (n is a natural number). In Table 1, level 0 has 2^0 = 1 segment, and the next level after level 0 is level 1, which is then divided into 2^1 = 2 segments. Therefore, by determining the length of the minimum segmentation sequence in the final level, we can deduce the maximum number of levels a chromosome sequence can be divided into. For example, the longest human chromosome, chromosome 1, has approximately 245,000,000 bases. If the minimum segmentation sequence is set to 512, the number of segments in the next level is twice that of the previous level, resulting in a total of 19 levels (starting from 0). The length of each sequence in the final level (level 19) is 512 (2^9) bases. For the human genome, following a division method where the number of fragments in each subsequent level is twice that of the previous level, a chromosome can be divided into a maximum of 19 levels. The last level has the most fragments, approximately 500,000, and the total number of fragments will not exceed 1 million (because 2^0 + 2^1 + 2^2 + 2^(n-1) < 2^n, so the total number of fragments in the first n-1 levels is less than that in the last level; that is, each additional level doubles the number of fragments). In other words, for the human reference genome, the number of database indexes built based on the number of visualization image fragments and the number of visualization image data under each index remain essentially constant, ensuring stable query efficiency. Therefore, the longer the length of M, the fewer fragments are in the last level, and the less storage space is required. Thus, without affecting client rendering efficiency, a longer M is better.

[0110] The optimal length of M is determined by the base resolution (i.e., the maximum number of bases that can be clearly seen within the actual observation range of the rendered visual image data on the user's device terminal), the client's image rendering speed, and the network transmission speed. In other words, it's the maximum number of bases that can be clearly seen through the viewport area. For example, the maximum display area of ​​a local computer viewport is approximately 1600 pixels. Clearly displaying each base requires at least 4 pixels, and a 1600-pixel wide viewport can display a maximum of 400 bases. Therefore, calculations show that when the client queries the last level, uploading a fragment of bases with a range of 512 is sufficient to cover the entire viewport area (512 * 4 = 2048 > 1600), which is more than enough to display base details on the screen. Because the data volume of 512 bases is small enough, displaying 10 visual tracks would not exceed 1MB of data, and would not put significant pressure on network transmission. Even if a smaller base range needs to be displayed later, such as increasing the magnification (reducing the number of bases displayed in the viewport) or switching the client from a computer to a mobile phone, there's no need to request new visualization data. Rendering can be performed directly based on the last-level data (small data volume, minimal computation, and no impact on response speed). Therefore, when the base resolution is reached, further magnification will be performed, and the client will use the returned last-level data to calculate and render the completed visualization. This reduces data redundancy caused by unnecessary hierarchical segmentation and improves data storage and retrieval efficiency. Therefore, the value of M needs to be greater than the base resolution (e.g., 400), but cannot exceed the maximum base value for network transmission and client calculation (e.g., 1024 bases).

[0111] Table 2 shows the storage rules for SAM / BAM files in the visualization images, and the schematic diagram of the short sequence dataset (using chromosome 1 as an example, levels 9-17).

[0112]

[0113] The minimum segmentation sequence (M) for each chromosome is determined using a reference genome. Files 2, 3, 4, and 5 are then segmented and partitioned according to this determined M. However, the visualization details recorded for each file differ. For example, when segmenting and recording SAM files into a visualization image database, in addition to assigning them to different levels and partitions based on the start and end positions of short sequence bases, positive and negative strand information is also recorded. Color or arrows are added later to differentiate between these. Furthermore, the differences in image display at different levels must be considered. For instance, if the maximum viewport width is 1600 pixels, and the view is currently being performed at level 0, one pixel represents 245,000,000 / 1600 = 153,125. If each pixel represents a 200bp short sequence, and calculations show 2^9 < 153125 / 200 < 2^10, levels 0-9 can only display the results of short sequence alignment (mapping) to the reference genome through coverage. At this scale, each slice stores only one coverage value, and the visualization of coverage in the reference genome fragment can be represented by a histogram or line graph. From level 10 onwards, the specific location of the short sequence alignment (mapping) to the reference genome can be displayed through a short sequence diagram. Each slice stores the start base position of the short sequence corresponding to the reference genome and the short sequence length, creating a diagram composed of rectangles.

[0114] When a short sequence spans two segments, it is split into two parts. If the start position of the short sequence is in segment A and the end position is in segment B, then the short sequence in segment A is recorded with its start base position and segment end position. The short sequence in segment B is recorded with its segment start base position and short sequence end position. Other location-independent information is also retained on each segment, such as an ID that identifies each short sequence.

[0115] Hierarchical and segmented storage mechanism for visualized image data:

[0116] Distributed storage is used: reference genomes and annotation files for different species are stored in a public database (files 1, 2, and 5a), with a fixed data volume; sample information is stored in a sample database, which can be expanded or reduced as needed (files 3, 4, and 5b). A single sample contains different data types, corresponding to different visualization image tracks, such as... Figure 4 As shown, each data type is processed according to a hierarchical and fragmented mechanism. Queries are uniformly calculated based on the actual observed base range to determine the corresponding reference genome's hierarchy and fragmentation. The selected data is then retrieved through an index.

[0117] After the hierarchical sharding is completed, the resulting data will be distributed and stored. The query index consists of the starting position of the hierarchical and shard bases (or can be obtained by calculating the hash), and the shards will be evenly stored on different nodes.

[0118] For example, a 102,400-byte genome is divided into levels 0-7, with each level having 1, 2, 4, 8, 16, 32, 64, and 128 fragments. If this data is distributed across four databases, based on the formula 2^0 + 2^1 + 2^2 + 2^(n-1) < 2^n, levels 0-5 would be placed in one database (31 fragments); level 6 would be placed in a separate database (32 fragments); and level 7 fragments would be distributed across two databases, each containing 32 fragments. Therefore, the corresponding level's dataset location can be quickly determined by base intervals. Since the data volume in each dataset is almost identical, stable query speeds are ensured. Taking level 2 as an example, with four fragments and a minimum segmentation sequence (M) of 100, the calculated indices are 2-0, 2-25600, 2-51200, and 2-102400. If the visualized base range is between 10,000 and 20,000 bases, the complete visualized image data can be obtained through a 2-0 partition. If the visualized base range is between 10,000 and 30,000 bases, the complete visualized image data can be obtained through two partitions: 2-0 and 2-25,600. If the visualized base range is between 10,000 and 60,000 bases, a level 1 query is required, and the complete visualized image data can be obtained through two partitions: 1-51,200 and 1-102,400. Under this hierarchical distributed storage mechanism with equal base range partitioning, the fluctuation in the query efficiency of visualized image data due to changes in the base range is negligible.

[0119] Therefore, it is evident that querying one or two slices for each visualization image type (corresponding to one visualization image track) is sufficient to meet the necessary information for the current viewport's visualization display. Assuming four tracks are to be displayed, four to eight slices would be required. The small data volume of a single slice reduces network transmission and client rendering pressure; distributed data storage ensures controllable query data volume; and precise indexing ensures high and controllable query efficiency.

[0120] To ensure efficient querying of descriptive information data, the following design is adopted:

[0121] Descriptive information data storage and query design:

[0122] The descriptive information contains all the information for files 2, 3, 4, and 5. These file data can be converted into a format that can be displayed on the client.

[0123] The ID of the visualized object (short sequence, gene annotation) is used as the index; according to the base range, all short sequence information is evenly distributed to different database nodes for storage;

[0124] There is a one-to-one correspondence between the ID of each sequence in the visual image database and the ID stored in the descriptive information database;

[0125] This design features simple expansion rules and strong scalability. Storage nodes are located by base range, and queries are performed using IDs as indexes, resulting in high and controllable query speeds. Taking human chromosome 1 as an example, with 245,000,000 bases, descriptive information is stored in the database in increments of one million bases, resulting in 245 nodes. Assuming a short sequence length of 200bp and a sequencing depth of 50x, one million bases would contain approximately 50 * 1,000,000 / 200 = 250,000 short sequence data. If the client sends an ID starting at position 100,000, the server will locate the first storage node and then perform a query using that ID as an index. For hundreds of thousands of data points, the location time is negligible, and the query time can be controlled within 0.2 seconds (200ms). The query efficiency is consistent and stable across all nodes.

[0126] Query design:

[0127] Data is retrieved from databases of visual image data and descriptive information via AJAX network requests. AJAX stands for Asynchronous JavaScript and XML, meaning that JavaScript performs asynchronous network requests. This request mechanism allows the user to remain on the current page while a new HTTP request is sent to retrieve data from the server, which is then displayed on the page. Advantages include reduced bandwidth transfer between the client and server, faster response times, and a better user experience without reloading the entire page. AJAX is an asynchronous request pattern that can be implemented in various ways. For example, modern browsers primarily rely on the XMLHttpRequest object for AJAX implementations. Other clients have alternative implementation methods, as long as they conform to the asynchronous request pattern of AJAX.

[0128] Visualized image data is retrieved via GET requests, while descriptive information data is retrieved via POST requests (other methods are also possible; this is just one example). The advantage of using GET requests for visualized images is that the client's calculated hierarchy and segmentation information is passed to the server as parameters. The server then calculates the index value and retrieves the data based on that index, returning it to the client. GET requests also allow the URL to be copied and sent to others, who can then reproduce the results, facilitating data writing. Descriptive information is obtained via POST. Firstly, this allows the server to distinguish the type of data to retrieve based on the network request type. Secondly, since viewing specific information requires considering the visualized results, it's unnecessary to include this information in the GET request, resulting in an excessively long URL.

[0129] When user interaction changes the visualization scale and range (zooming, panning, or directly determining the visualization range through the search box, etc.), the client calculates the actual base range corresponding to the current viewport area, calculates the hierarchical and segmented position of the visualization image information in the database based on the actual base range, and then sends this information to the server. The server combines this information with the database index rules to find the corresponding visualization image data, and then sends this data back to the client for rendering.

[0130] When a user's interaction triggers the display of descriptive information, the client calculates the target visualization image corresponding to the specific pixel position triggered by the user, obtains the image data for rendering, and sends the ID contained therein to the server. The server searches for an index that matches this ID in the distributed storage of descriptive information, obtains the complete descriptive information of the corresponding sequence, and returns it to the client for display.

[0131] Window area: The range of the actual rendered result of the visualized image data.

[0132] Viewport area: The actual viewing range of the rendered results of the visualized image data.

[0133] Each time the requested visualization image fragment data is sent to the client, it is rendered in the window area. The user can only see part of the visualization image in the window area through the viewport area. However, by zooming, panning, and moving the viewport, the overflowing parts can be seen. If the viewport display area is about to reach the window rendering boundary after zooming or panning, for example, if it is 50px away from the window rendering image boundary, a new request will be triggered, returning the visualization image fragment data that will be displayed in the viewport area and rendering it into an image, making panning and zooming display smoother and without lag.

[0134] Client-side interaction design:

[0135] like Figure 3 As shown, the client is responsible for rendering the visualization information into a visualization image, the server is responsible for transmitting information between the client and the database over the network, and the database is only involved in searching for visualization images and descriptive information data (preprocessing high-throughput sequencing file data).

[0136] When a user's interaction causes a change in the actual base display range of the reference genome corresponding to the visualized image, the client calculates the level at which the visualized image data is stored in the database and the actual base range on the reference sequence according to the currently selected viewport pixel range and pan and zoom scale, then the client sends a request carrying these parameters to the server. The server constructs an index, finds the visualized image data according to the index and returns it, and then the client renders it into a corresponding visualized image. For example, when a user inputs a given base range to search in the search box, and clicks on the pan window area of the reference genome schematic diagram.

[0137] Before the display scale or window position is about to trigger a change in the visualization level, the client will pre-fetch the visualization image data of the base interval range to be displayed next to the client for rendering, so that a smooth transition can be achieved between display ranges.

[0138] When a user's interaction (hovering, clicking, touching the screen, etc.) triggers the display of descriptive information, the client calculates according to the position of the viewport clicked by the user and the zoom scale, finds the id information corresponding to the visualization image selected by the user, uses this as an index to search in the descriptive information database, and then returns it to the client to display the detailed information.

[0139] Expression of the linear mapping function:

[0140] y=scale(x)

[0141] Domain interval range:

[0142] domain:[d0,d1]

[0143] Range interval range:

[0144] range:[r0,r1]

[0145] Linear mapping function:

[0146]

[0147] The meaning of the linear mapping function is: given a domain from d0 to d1, and a range from r0 to r1. There is currently an x between d0 and d1 in the domain, denoted as d0<x<d1. Substituting into the formula scale(x)=(r0*d1-r0*x+r1*x-r1*d0) / (d1-d0) gives a y value, and this value is between r0 and r1, that is r0<y<r1. What is the use of this? For example, if x=2, the domain range is [1,3], the range is [1,9], substituting into the formula gives 5. Since 2 is exactly in the middle of [1,3], correspondingly, it should also be 5 in the range [1,9].

[0148] The function of a linear mapping function is to map a position x in an interval ([d0,d1]) to a position y in another interval (r0,r1).

[0149] For example, the observed pixel range (start and end points) is scaled on the client side, and then converted into the range (start and end points) of the observed chromatin through a linear mapping function. The database is then searched based on the converted real position interval.

[0150] For example:

[0151] Assuming the base range of human chromosome x is [0, 200, 000, 000] and the display pixel range is [0, 2, 000], and now we select the pixel interval [100, 500], what is the actual base range of the chromosome?

[0152] First, we need to determine which is the range and which is the domain. Here, we are calculating the range of base pairs. The pixel range ([0,2,000]) is the range, while the range of base x on chromosome ([0,200,000,000]) is the domain. Then, substituting [100,500] into the linear mapping formula yields:

[0153] d0 = 0, d1 = 2,000

[0154] r0 = 0, r1 = 200,000,000

[0155]

[0156]

[0157] Calculations show that the pixel range [100, 500] displayed on the screen corresponds to the actual base range [10, 000, 000, 50, 000, 000]. If the pixel range is inferred from the actual base range, then the domain definition can be swapped for calculation.

[0158] Front-end and back-end interaction logic: Visual image part:

[0159] 1. Determine the actual range (pixel interval) based on the sliding window.

[0160] 2. The actual domain (base range) on the genome is calculated using a linear mapping function.

[0161] 3. The client calculates the level and segment number based on the actual base range.

[0162] 4. Send a request for the corresponding level and segment number.

[0163] For example, GET https: / / example.com / query?level=10&block=100, this request means sending a Get request with a level of 10 and a block number of 100 to the backend.

[0164] 5. After the backend receives the request, it locates the corresponding database node (or database table) by level, and finds the data information of the corresponding block number in the corresponding database node (or database table).

[0165] 6. From the returned data, extract the corresponding base interval data from the returned data through calculation according to the domain base interval saved at the front end.

[0166] 7. Convert the actual base interval into a pixel interval through a linear mapping formula for the extracted data, and render it on the client side via SVG / Canvas / WebGL / WebGPU:

[0167] Explanation of storage and query mechanism for descriptive information database:

[0168] The descriptive information database is divided into storage nodes (several databases) or divided into tables (several tables in one database) according to the base range. For example: humans have 23 pairs of chromosomes, each chromosome is assigned to one database, and the database node is located by the chromosome number when viewing a certain chromosome. Assuming that chromosome 1 has approximately 245,000,000 bases, every 1,000,000 bases is stored in one table, so there are 245 tables stored in the chromosome 1 database. According to the base range, the time for locating the database and querying the corresponding table is almost negligible.

[0169] Descriptive information database:

[0170] 1. After the client completes visual image rendering, when the user interacts with the visualized image, for example, clicking the visualized image with a mouse;

[0171] 2. The client obtains the ID information of the corresponding image according to the image position where the click is triggered, and sends a POST request (with the base range and image ID of the corresponding image) to the backend;

[0172] 3. After the backend server receives the request, it extracts the base range information. For example, when querying image information on chromosome 1, the ID range is between [1500000, 1510000]. It can be seen from the above example that according to the index comparison rule ( (n-1)*1000000 <1500000> < n*1000000; (n-1)*1000000 <1510000> < n*1000000 ), it is obtained that n=2. Therefore, the ID information must be in table 2;

[0173] 4. By using the ID as an index, the corresponding image information in the table can be found and displayed on the client.

[0174] Example 2

[0175] This embodiment provides a high-throughput sequencing data visualization device, which includes:

[0176] The data extraction module is used to acquire high-throughput sequencing data and extract visualization image data and its corresponding descriptive information data from the high-throughput sequencing data.

[0177] The data storage module is used to perform hierarchical and segmented processing on the visualized image data based on a preset hierarchical and segmented mechanism, and to construct a visualized image data query index. The hierarchical and segmented data of the visualized image data, along with its query index, are distributed to nodes in the visualized image database for storage. Descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image module is used to ensure that the data query index and the descriptive information data query index share a common unique identifier.

[0178] The data display module is used to render and display the visualized image data retrieved based on the visualized image data query index, and at the same time, display the descriptive information data that matches the rendered visualized image data based on the descriptive information data query index.

[0179] In the process of performing hierarchical and segmented processing on the visualized image data based on a preset hierarchical and segmented mechanism, the calculation rule for the relationship between the hierarchical level and the number of segments is as follows: when the hierarchical level is n, the number of segments corresponding to each level is 2 raised to the power of n; where n is a natural number.

[0180] In the data storage module, during the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1≤M*2^N / C<2; where C represents the chromosome length, i.e. the total number of chromosome bases, N represents the number of levels that this chromosome sequence needs to be divided into, and also represents the last level of the hierarchical segmentation, the Nth level; M represents the minimum segmentation sequence.

[0181] In the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the minimum segmentation sequence refers to the length of the base sequence represented by each segment in the last level of the hierarchical segmentation, which is determined by the base resolution, the client image rendering speed and the network transmission speed.

[0182] In the data storage module, during the hierarchical segmentation process of the visualized image data based on the preset hierarchical segmentation mechanism, the calculation rule for the length of the base sequence represented by each segment in the nth level is: M*2^(Nn), where N represents the last level of the hierarchical process, n represents the nth level, that is, any level from 0 to N, and M represents the minimum segmentation sequence.

[0183] In the data storage module, during the hierarchical segmentation process of the visualized image data based on the preset hierarchical segmentation mechanism, after the same chromosome is segmented, the sum of the base sequence lengths represented by all segments in each level is M*2^N, where N represents the last level of the hierarchical segmentation and M represents the smallest segmentation sequence.

[0184] In the data storage module, the visualization image data query index is constructed according to the principle of first hierarchical and then segmented, and is composed of the base starting positions of each segment at each level.

[0185] It should be noted that the modules in this embodiment do not correspond to the steps in Embodiment 1, and will not be repeated here.

[0186] Example 3

[0187] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the high-throughput sequencing data visualization method described above.

[0188] It should be noted that the steps in the high-throughput sequencing data visualization method are the same as in Example 1, and will not be repeated here.

[0189] Example 4

[0190] This embodiment provides a high-throughput sequencing data visualization device, which includes a server and a client.

[0191] The server is configured as follows:

[0192] Acquire high-throughput sequencing data and extract visualization image data and corresponding descriptive information data from the high-throughput sequencing data;

[0193] The visualized image data is processed by hierarchical segmentation based on a preset hierarchical segmentation mechanism, and a visualized image data query index is constructed. The hierarchical segmented data of the visualized image data, along with its query index, is distributed to nodes in the visualized image database for storage. The descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier.

[0194] Render the visual image data retrieved from the visual image data query index;

[0195] The client is configured as follows:

[0196] It displays the rendered visual image data, and simultaneously displays descriptive information data that matches the rendered visual image data based on the descriptive information data query index.

[0197] It should be noted that the steps in this embodiment are the same as those in Embodiment 1, and will not be repeated here.

[0198] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0199] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for visualizing high-throughput sequencing data, characterized in that, include: Acquire high-throughput sequencing data and extract visualization image data and corresponding descriptive information data from the high-throughput sequencing data; The visualized image data is processed by hierarchical segmentation based on a preset hierarchical segmentation mechanism, and a visualized image data query index is constructed. The hierarchical segmented data of the visualized image data, along with its query index, is distributed to nodes in the visualized image database for storage. The descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier. In the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1≤M 2^N / C < 2; where C represents chromosome length, i.e., the total number of chromosome bases; N represents the number of levels this chromosome sequence needs to be divided into, which is also the last level of the division, the Nth level; M represents the minimum segmentation sequence, which refers to the length of the base sequence represented by each segment in the last level of the division, and is determined by the base resolution, the client image rendering speed, and the network transmission speed; the value of M is greater than the base resolution, but not greater than the maximum base value of network transmission and client calculation; The visualization image data retrieved from the visualization image data query index is rendered and displayed. At the same time, the descriptive information data matching the rendered visualization image data is displayed based on the descriptive information data query index. During the rendering and display of the visualization image data retrieved from the visualization image data query index, when the base resolution is reached, it is further magnified. The client uses the returned last level data to calculate and complete the visualization image display.

2. The high-throughput sequencing data visualization method as described in claim 1, characterized in that, In the process of performing hierarchical and segmented processing on the visualized image data based on the preset hierarchical and segmented mechanism, the calculation rule for the relationship between the hierarchical level and the number of segments is as follows: when the hierarchical level is n, the number of segments corresponding to each level is 2 raised to the power of n; where n is a natural number.

3. The high-throughput sequencing data visualization method as described in claim 1, characterized in that, In the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the length of the base sequence represented by each segment in the nth level is as follows: M 2^(Nn), where N represents the last level of the hierarchy, n represents the nth level, that is, any level from 0 to N, and M represents the minimum segmentation sequence.

4. The high-throughput sequencing data visualization method as described in claim 1, characterized in that, During the hierarchical segmentation process of the visualized image data based on a preset hierarchical segmentation mechanism, after the same chromosome is segmented, the sum of the base sequence lengths represented by all segments in each level is M. 2^N, where N represents the last level of the hierarchy and M represents the minimum segmentation sequence.

5. The high-throughput sequencing data visualization method as described in claim 1, characterized in that, The visualized image data query index follows the principle of first hierarchical and then segmented, and is composed of the base starting positions of each level and segment.

6. A high-throughput sequencing data visualization device, characterized in that, include: The data extraction module is used to acquire high-throughput sequencing data and extract visualization image data and its corresponding descriptive information data from the high-throughput sequencing data. The data storage module is used to perform hierarchical segmentation processing on the visualized image data based on a preset hierarchical segmentation mechanism, and to construct a visualized image data query index. The hierarchical segmented data of the visualized image data, along with its query index, is distributed and stored in nodes of the visualized image database. Descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image module ensures that the data query index and the descriptive information data query index share a common unique identifier. During the hierarchical segmentation processing of the visualized image data based on the preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1 ≤ M. 2^N / C < 2; where C represents chromosome length, i.e., the total number of chromosome bases; N represents the number of levels this chromosome sequence needs to be divided into, which is also the last level of the division, the Nth level; M represents the minimum segmentation sequence, which refers to the length of the base sequence represented by each segment in the last level of the division, and is determined by the base resolution, the client image rendering speed, and the network transmission speed; the value of M is greater than the base resolution, but not greater than the maximum base value of network transmission and client calculation; The data display module is used to render and display the visualized image data retrieved based on the visualized image data query index, and at the same time, it displays the descriptive information data that matches the rendered visualized image data based on the descriptive information data query index. During the process of rendering and displaying the visualized image data retrieved based on the visualized image data query index, when the base resolution is reached, it continues to zoom in, and the client uses the returned last level data to calculate and complete the visualization image display.

7. The high-throughput sequencing data visualization device as described in claim 6, characterized in that, In the data storage module, during the hierarchical and segmented processing of the visualized image data based on a preset hierarchical and segmented mechanism, the calculation rule for the relationship between the hierarchical level and the number of segments is as follows: when the hierarchical level is n, the number of segments corresponding to each level is 2 raised to the power of n; where n is a natural number. or In the data storage module, during the hierarchical segmentation process of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1 ≤ M. 2^N / C<2; where C represents chromosome length, i.e., the total number of chromosome bases; N represents the number of levels this chromosome sequence needs to be divided into, which is also the last level of the division, the Nth level; M represents the minimum segmentation sequence; In the process of hierarchical segmentation of the visualized image data based on a preset hierarchical segmentation mechanism, the minimum segmentation sequence refers to the length of the base sequence represented by each segment in the last level of the hierarchical segmentation, which is determined by the base resolution, the client image rendering speed and the network transmission speed. or In the data storage module, during the hierarchical segmentation process of the visualized image data based on a preset hierarchical segmentation mechanism, the calculation rule for the length of the base sequence represented by each segment in the nth level is: M 2^(Nn), where N represents the last level of the hierarchy, n represents the nth level, that is, any level from 0 to N, and M represents the minimum segmentation sequence; or In the data storage module, during the hierarchical segmentation process of the visualized image data based on a preset hierarchical segmentation mechanism, after the same chromosome undergoes hierarchical segmentation, the sum of the base sequence lengths represented by all segments in each level is M. 2^N, where N represents the last level of the hierarchy and M represents the minimum segmentation sequence; or In the data storage module, the visualization image data query index is constructed according to the principle of first hierarchical and then segmented, and is composed of the base starting positions of each segment at each level.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the high-throughput sequencing data visualization method as described in any one of claims 1-4.

9. A high-throughput sequencing data visualization device, characterized in that, Includes server-side and client-side components; The server is configured as follows: Acquire high-throughput sequencing data and extract visualization image data and corresponding descriptive information data from the high-throughput sequencing data; The visualized image data is processed by hierarchical segmentation based on a preset hierarchical segmentation mechanism, and a visualized image data query index is constructed. The hierarchical segmented data of the visualized image data, along with its query index, is distributed and stored in nodes of the visualized image database. Descriptive information data is stored in the descriptive information database according to the descriptive information data query index. The visualized image data query index and the descriptive information data query index share a common unique identifier. During the hierarchical segmentation of the visualized image data based on the preset hierarchical segmentation mechanism, the calculation rule for the hierarchical level N of each chromosome is: 1 ≤ M. 2^N / C < 2; where C represents chromosome length, i.e., the total number of chromosome bases; N represents the number of levels this chromosome sequence needs to be divided into, which is also the last level of the division, the Nth level; M represents the minimum segmentation sequence, which refers to the length of the base sequence represented by each segment in the last level of the division, and is determined by the base resolution, the client image rendering speed, and the network transmission speed; the value of M is greater than the base resolution, but not greater than the maximum base value of network transmission and client calculation; Render the visual image data retrieved from the visual image data query index; The client is configured as follows: The system displays the rendered visual image data, and simultaneously displays descriptive information data that matches the rendered visual image data based on the descriptive information data query index. During the rendering and display of the visual image data retrieved based on the visual image data query index, when the base resolution is reached, it continues to zoom in, and the client uses the returned last level data to calculate and complete the rendering of the visual image display.

Citation Information

Patent Citations

  • Storage method and query method for high-throughput sequencing sequences

    CN107506618A