Gene Sequence Alignment and Mutation Detection System and Method Based on Cloud Computing
Through a cloud-based gene sequence alignment and variant detection system, user gene sequence data is collected and processed in real time, and global alignment algorithms are used to perform comparison and analysis and mutation site location, which solves the problem of poor gene sequence alignment and variant detection in the existing technology, and achieves more efficient gene sequence alignment and variant detection.
Patent Information
- Application Number
- CN202510098265.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing gene sequence alignment and variant detection technologies cannot effectively perform gene sequence alignment and variant detection localization, resulting in poor results.
A cloud-based gene sequence alignment and variant detection system is adopted, including a gene acquisition terminal, a data communication module and a cloud server terminal. By collecting, processing and comparing user gene sequence data in real time, a global comparison algorithm is used for comparison and analysis, and a location of mutation sites is found.
Effective alignment and variation detection and localization of gene sequences are achieved, and the effect of gene sequence alignment and variation detection is improved.
Smart Images

Figure CN119541644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene detection, and specifically to a gene sequence alignment and variation detection system and method based on cloud computing. Background Art
[0002] With the development of genomics and the continuous progress of gene sequencing technology, more and more bioinformatics data need to be processed and analyzed; among them, gene sequence alignment and variation detection are two important tasks in bioinformatics.
[0003] Chinese Patent No. CN107633158B discloses a method and device for compressing and decompressing gene sequences. The method for compressing gene sequences includes: generating a variant reference sequence according to high-frequency variation information and a standard reference sequence; compressing the gene sequence to be processed according to the matching result between the gene sequence to be processed and the variant reference sequence to obtain a compressed gene sequence; which can improve the compression rate of gene sequences, thereby reducing the storage space of gene sequences and facilitating the copying and transmission of gene sequences; however, this patent has the following defects:
[0004] The existing ones cannot effectively align gene sequences and cannot effectively detect and locate gene sequence variation situations, resulting in poor gene sequence alignment and variation detection effects. Summary of the Invention
[0005] The purpose of the present invention is to provide a gene sequence alignment and variation detection system and method based on cloud computing, which can effectively align gene sequences and can effectively detect and locate gene sequence variation situations, and can improve the gene sequence alignment and variation detection effects, solving the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A gene sequence alignment and variation detection system based on cloud computing, comprising:
[0008] A gene collection terminal, used for managing users and collecting real-time data of user gene sequences;
[0009] A data communication module, used for transmitting real-time data of user gene sequences and establishing data communication between the gene collection terminal and the cloud server terminal;
[0010] A cloud server terminal, used for processing real-time data of user gene sequences, performing gene sequence alignment and variation detection, and determining gene sequence alignment results and gene sequence alignment and variation detection results.
[0011] Preferably, the gene collection terminal includes:
[0012] A user management unit for managing user identity registration, login, and user permissions;
[0013] Among them, the user management interface prompts the user to input a gene sequence, and the user inputs the gene sequence according to the prompts of the user management interface;
[0014] A gene collection unit for real-time collection of the gene sequence input by the user;
[0015] After the user inputs the gene sequence, the user management interface performs real-time collection on the gene sequence, and automatically converts the gene sequence into a format recognizable by a computer to determine the real-time data of the user gene sequence.
[0016] Preferably, the gene collection terminal further includes:
[0017] A response duration real-time monitoring module for real-time monitoring of the collection response duration of the gene collection unit after the user inputs the gene sequence;
[0018] A duration comparison module for comparing the collection response duration of the gene collection unit after the user inputs the gene sequence with a preset collection response duration threshold;
[0019] A collection monitoring quantity setting module for, when the collection response duration of the gene collection unit after the user inputs the gene sequence exceeds the preset collection response duration threshold, setting the collection monitoring quantity by using the collection response duration that exceeds the preset collection response duration threshold;
[0020] Among them, the collection monitoring quantity is obtained through the following formula:
[0021] ; where Q represents the collection monitoring quantity; INT[] represents rounding up the function inside the parentheses; T represents the collection response duration that exceeds the preset collection response duration threshold; T th represents the preset collection response duration threshold; T z represents the maximum allowable collection response duration corresponding to the minimum sensitivity for normal operation of data collection; λ represents an adjustment coefficient, and the value range of the adjustment coefficient is 0.13 - 0.36; t d represents the average time interval of data collection of the gene collection unit;
[0022] A timeliness anomaly determination module for, after the collection response duration exceeds the preset collection response duration threshold, determining whether there is an anomaly in the data collection timeliness of the gene collection unit by using the collection response duration of the real-time collection times of the gene collection unit corresponding to the collection monitoring quantity.
[0023] Preferably, the timeliness anomaly determination module includes;
[0024] A collection parameter monitoring module, configured to monitor the real-time collection times of the gene collection unit according to the collection monitoring quantity after the collection response duration exceeds a preset collection response duration threshold;
[0025] A collection response duration retrieval module, configured to retrieve the collection response duration corresponding to each data collection in the real-time collection times of the gene collection unit when the real-time collection times of the gene collection unit reach the collection monitoring quantity;
[0026] A data collection sensitivity coefficient acquisition module, configured to acquire the data collection sensitivity coefficient corresponding to the gene collection unit by using the collection response duration corresponding to each data collection in the real-time collection times of the gene collection unit;
[0027] Wherein, the data collection sensitivity coefficient corresponding to the gene collection unit is obtained through the following formula:
[0028] ;
[0029] Wherein, K represents the data collection sensitivity coefficient corresponding to the gene collection unit; Q represents the collection monitoring quantity; T represents the collection response duration exceeding the preset collection response duration threshold; T th represents the preset collection response duration threshold; T z represents the maximum allowable collection response duration corresponding to the minimum sensitivity for the normal operation of data collection; T max and T min represent the maximum and minimum values of the collection response duration corresponding to Q data collections of the gene collection unit; T i represents the collection response duration corresponding to the i-th data collection of the gene collection unit;
[0030] A coefficient comparison module, configured to compare the data collection sensitivity coefficient with a preset sensitivity coefficient threshold;
[0031] An anomaly alarm determination and module, configured to determine that there is an anomaly in the timeliness of data collection of the gene collection unit and perform an anomaly alarm when the data collection sensitivity coefficient is lower than the preset sensitivity coefficient threshold.
[0032] Preferably, the data communication module includes:
[0033] A data sending unit, configured to send real-time user gene sequence data;
[0034] A data receiving unit, configured to receive real-time user gene sequence data;
[0035] According to the requirements of gene sequence alignment and mutation detection based on cloud computing, establish a data communication link between the data sending unit and the data receiving unit;
[0036] Among them, the data sending unit receives the real-time user gene sequence data transmitted from the gene collection terminal, and the data sending unit sends the received real-time user gene sequence data to the data receiving unit;
[0037] Among them, the data receiving unit receives the real-time user gene sequence data transmitted from the data sending unit, and the data receiving unit sends the received real-time user gene sequence data to the cloud server terminal.
[0038] Preferably, the cloud server terminal includes:
[0039] A data processing module for processing the real-time user gene sequence data to determine the user gene sequence characteristic data;
[0040] A comparison and analysis module for comparing and analyzing the user gene sequence characteristic data to determine the gene sequence alignment result based on cloud computing;
[0041] A mutation detection module for searching and locating gene sequence mutation sites to determine the gene sequence alignment and mutation detection result based on cloud computing;
[0042] A storage and display module for storing and displaying the gene sequence alignment and mutation detection report based on cloud computing.
[0043] Preferably, the data processing module includes:
[0044] A data cleaning unit for cleaning the real-time user gene sequence data;
[0045] Obtain the real-time user gene sequence data;
[0046] Clean the real-time user gene sequence data;
[0047] Remove inconsistent data, invalid values, and missing values from the real-time user gene sequence data;
[0048] Determine the real-time user gene sequence data useful for gene sequence alignment and mutation detection;
[0049] A feature extraction unit for extracting features from the real-time user gene sequence data;
[0050] Obtain the real-time user gene sequence data useful for gene sequence alignment and mutation detection after cleaning;
[0051] Extract features from the real-time user gene sequence data useful for gene sequence alignment and mutation detection;
[0052] Extract the features that can reflect gene sequence alignment and mutation detection;
[0053] Determine the user's gene sequence feature data.
[0054] Preferably, the alignment analysis module includes:
[0055] A standard setting unit for setting the standard data of the user's gene sequence;
[0056] According to the requirements of gene sequence alignment and mutation detection based on cloud computing, pre-set the standard data of the user's gene sequence, and the standard data of the user's gene sequence is used to provide a reference basis for gene sequence alignment;
[0057] An alignment analysis unit for performing alignment analysis on the user's gene sequence feature data;
[0058] Obtain the user's gene sequence feature data and the standard data of the user's gene sequence;
[0059] Based on the global alignment algorithm, perform alignment analysis on the user's gene sequence feature data and the standard data of the user's gene sequence, judge the similarity and difference between the user's gene sequence feature data and the standard data of the user's gene sequence, and determine the gene sequence alignment result based on cloud computing;
[0060] Among them, if the user's gene sequence feature data is within the range of the standard data of the user's gene sequence, the gene sequence alignment result based on cloud computing is that the user's gene sequence is normal;
[0061] Among them, if the user's gene sequence feature data is not within the range of the standard data of the user's gene sequence, the gene sequence alignment result based on cloud computing is that the user's gene sequence has a mutation.
[0062] Preferably, the mutation detection module includes:
[0063] A search and positioning unit for searching and positioning gene sequence mutation sites;
[0064] Obtain the gene sequence alignment result based on cloud computing;
[0065] According to the gene sequence alignment result based on cloud computing, perform mining analysis on the user's gene sequence feature data to find the mutation region of the user's gene sequence;
[0066] Based on the mutation region of the user's gene sequence, search for and locate the gene sequence mutation sites, and locate the gene sequence mutation sites;
[0067] Determine the gene sequence alignment and mutation detection result based on cloud computing.
[0068] Preferably, the storage and display module includes:
[0069] A distributed database for storing gene sequence alignment and mutation detection reports based on cloud computing;
[0070] Obtain user gene sequence feature data, gene sequence alignment results, and gene sequence alignment and mutation detection results;
[0071] Analyze and integrate the user gene sequence feature data, gene sequence alignment results, and gene sequence alignment and mutation detection results to form a gene sequence alignment and mutation detection report based on cloud computing;
[0072] Store the gene sequence alignment and mutation detection report based on cloud computing;
[0073] A display and sharing unit for visually displaying the gene sequence alignment and mutation detection report based on cloud computing and sharing the gene sequence alignment and mutation detection report based on cloud computing.
[0074] According to another aspect of the present invention, there is provided a gene sequence alignment and mutation detection method based on cloud computing, which is implemented based on the gene sequence alignment and mutation detection system based on cloud computing as described above, and includes the following steps:
[0075] S1: Manage user identity registration, login, and user permissions through a gene collection terminal, collect real-time user gene sequence data, transmit the real-time user gene sequence data through a data communication module, and establish data communication between the gene collection terminal and the cloud server terminal;
[0076] S2: Process the real-time user gene sequence data through the cloud server terminal, perform alignment analysis on the user gene sequence feature data and the user gene sequence standard data based on the global alignment algorithm to determine the gene sequence alignment result, search and locate the gene sequence mutation sites, and determine the gene sequence alignment and mutation detection result.
[0077] Compared with the prior art, the beneficial effects of the present invention are:
[0078] The present invention manages users, collects real-time data of user gene sequences, processes the real-time data of user gene sequences to determine user gene sequence feature data, and based on a global alignment algorithm, compares and analyzes the user gene sequence feature data and user gene sequence standard data to determine the similarity and difference between the user gene sequence feature data and the user gene sequence standard data, determines a gene sequence alignment result based on cloud computing, and locates gene sequence variation sites to determine a gene sequence alignment and variation detection result based on cloud computing, which can effectively align gene sequences and effectively detect and locate gene sequence variation conditions, and can improve the gene sequence alignment and variation detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a framework diagram of the gene sequence alignment and variation detection system based on cloud computing of the present invention;
[0080] Figure 2 is a flowchart of the gene sequence alignment and variation detection method based on cloud computing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0082] To solve the problem that the existing technology cannot effectively align gene sequences and cannot effectively detect and locate gene sequence variation conditions, resulting in poor gene sequence alignment and variation detection effects, please refer to Figure 1 - Figure 2 , the following technical solutions are provided in this embodiment:
[0083] A gene sequence alignment and variation detection system based on cloud computing includes:
[0084] A gene collection terminal for managing users and collecting real-time data of user gene sequences.
[0085] In this embodiment, the gene collection terminal includes:
[0086] A user management unit for managing user identity registration, login, and user permissions;
[0087] Specifically, according to the gene sequence alignment and variation detection requirements based on cloud computing, a user registration requirement is displayed on the user management interface, and the user inputs user identity information according to the user registration requirement to complete user identity registration;
[0088] Among them, the user identity information includes name, gender, age, address, phone number, account number and password;
[0089] Specifically, after the user identity registration is completed, the user logs in based on the account number and password, and automatically authenticates the account number and password entered by the user. After the authentication is qualified, the user logs in successfully;
[0090] Specifically, after the user logs in successfully, the user exercises the authority of gene sequence alignment and mutation detection. The user management interface prompts the user to input the gene sequence, and the user inputs the gene sequence according to the prompts of the user management interface;
[0091] The gene collection unit is used to collect the gene sequence input by the user in real time;
[0092] After the user inputs the gene sequence, the user management interface collects the gene sequence in real time, and automatically converts the gene sequence into a format that can be recognized by a computer, and determines the real-time data of the user gene sequence, providing data support for subsequent gene sequence alignment and mutation detection.
[0093] It should be noted that the real-time collection of the user identity information and the gene sequence input by the user has obtained the prior consent of the user. After obtaining the user's consent, the user identity information and the real-time data of the user gene sequence are collected, analyzed and used; the user identity information and the real-time data of the user gene sequence are legally collected and legally used.
[0094] The data communication module is used to transmit the real-time data of the user gene sequence and establish data communication between the gene collection terminal and the cloud server terminal.
[0095] Specifically, the gene collection terminal further includes:
[0096] The response duration real-time monitoring module is used to monitor in real time the collection response duration of the gene collection unit after the gene sequence input by the user;
[0097] The duration comparison module is used to compare the collection response duration of the gene collection unit after the gene sequence input by the user with a preset collection response duration threshold;
[0098] The collection monitoring quantity setting module is used to, when the collection response duration of the gene collection unit after the gene sequence input by the user exceeds the preset collection response duration threshold, set the collection monitoring quantity by using the collection response duration that exceeds the preset collection response duration threshold;
[0099] Among them, the collection monitoring quantity is obtained through the following formula:
[0100] ; where Q represents the number of acquisition monitors; INT[] represents rounding up the function inside the parentheses; T represents the acquisition response duration exceeding the preset acquisition response duration threshold; T th represents the preset acquisition response duration threshold; T z represents the maximum allowable acquisition response duration corresponding to the minimum sensitivity for normal operation of data acquisition; λ represents the adjustment coefficient, and the value range of the adjustment coefficient is 0.13 - 0.36; t d represents the average time interval of data acquisition of the gene acquisition unit;
[0101] The timeliness anomaly determination module is used to determine whether there is an anomaly in the data acquisition timeliness of the gene acquisition unit by using the acquisition response duration of the real-time acquisition times of the gene acquisition unit corresponding to the acquisition monitor number after the acquisition response duration exceeds the preset acquisition response duration threshold.
[0102] The technical effects of the above technical solution are as follows: Through the response duration real-time monitoring module, the gene acquisition terminal can monitor the acquisition response duration of the gene acquisition unit in real time after the user inputs the gene sequence. This helps to promptly detect delays or lags in the acquisition process and ensure the timeliness of data acquisition. The duration comparison module compares the real-time monitored acquisition response duration with the preset acquisition response duration threshold. Once the threshold is exceeded, the abnormal response mechanism is triggered, which helps to promptly identify and handle potential acquisition problems. When the acquisition response duration exceeds the preset threshold, the acquisition monitor number setting module dynamically sets the acquisition monitor number according to the acquisition response duration exceeding the threshold. This mechanism is calculated by a formula, considering various factors (such as the acquisition response duration exceeding the threshold, the preset threshold, the maximum allowable acquisition response duration, the adjustment coefficient, and the average time interval of data acquisition), making the setting of the monitor number more reasonable and scientific. The timeliness anomaly determination module uses the acquisition response duration of the real-time acquisition times of the gene acquisition unit corresponding to the acquisition monitor number to determine whether there is an anomaly in the data acquisition timeliness after the acquisition response duration exceeds the preset threshold. This determination helps to further confirm whether there are problems in the acquisition process and take corresponding measures for correction. Overall, this technical solution acts on the gene acquisition process through multiple links such as real-time monitoring, intelligent comparison, dynamic adjustment, and anomaly determination, which helps to improve the efficiency and accuracy of data acquisition. At the same time, by dynamically adjusting the acquisition monitor number, it can more flexibly respond to different acquisition scenarios and requirements. This technical solution can promptly detect and handle problems in the acquisition process through real-time monitoring and the anomaly handling mechanism, thereby enhancing the reliability and stability of the entire gene acquisition system. This is of great significance for ensuring the integrity and accuracy of gene data.
[0103] In summary, through a series of intelligent monitoring and processing mechanisms, this technical solution realizes the comprehensive monitoring and management of the gene collection process, improves the efficiency and accuracy of data collection, and enhances the reliability and stability of the system.
[0104] Specifically, the timeliness anomaly determination module includes:
[0105] The acquisition parameter monitoring module is used to monitor the real-time acquisition times of the gene acquisition unit according to the acquisition monitoring quantity after the acquisition response duration exceeds the preset acquisition response duration threshold.
[0106] The acquisition response duration retrieval module is used to retrieve the acquisition response duration corresponding to each data acquisition in the real-time acquisition times of the gene acquisition unit when the real-time acquisition times of the gene acquisition unit reach the acquisition monitoring quantity.
[0107] The data acquisition sensitivity coefficient acquisition module is used to obtain the data acquisition sensitivity coefficient corresponding to the gene acquisition unit by using the acquisition response duration corresponding to each data acquisition in the real-time acquisition times of the gene acquisition unit.
[0108] Among them, the data acquisition sensitivity coefficient corresponding to the gene acquisition unit is obtained through the following formula:
[0109] ; where K represents the data acquisition sensitivity coefficient corresponding to the gene acquisition unit; Q represents the acquisition monitoring quantity; T represents the acquisition response duration exceeding the preset acquisition response duration threshold; T th represents the preset acquisition response duration threshold; c represents the maximum allowable acquisition response duration corresponding to the minimum sensitivity for normal operation of data acquisition; T max and T min represent the maximum and minimum values of the acquisition response duration corresponding to Q data acquisitions of the gene acquisition unit; T i represents the acquisition response duration corresponding to the i-th data acquisition of the gene acquisition unit;
[0110] The coefficient comparison module is used to compare the data acquisition sensitivity coefficient with the preset sensitivity coefficient threshold.
[0111] The anomaly alarm determination and module is used to determine that there is an anomaly in the timeliness of data acquisition of the gene acquisition unit and perform an anomaly alarm when the data acquisition sensitivity coefficient is lower than the preset sensitivity coefficient threshold.
[0112] The technical effects of the above technical solution are as follows: Through the acquisition parameter monitoring module, this module can, after the acquisition response duration exceeds the preset threshold, monitor the real-time acquisition times of the gene acquisition unit in real time according to the set acquisition monitoring quantity. This ensures the continuity and integrity of data acquisition and provides a basis for subsequent data analysis. The acquisition response duration retrieval module can accurately retrieve the acquisition response duration corresponding to each data acquisition when the real-time acquisition times of the gene acquisition unit reach the monitoring quantity. This step is crucial for calculating the data acquisition sensitivity coefficient subsequently. The data acquisition sensitivity coefficient acquisition module calculates the data acquisition sensitivity coefficient of the gene acquisition unit using the collected response durations through the above formula. This coefficient synthesizes multiple factors, such as the acquisition monitoring quantity, the acquisition response duration exceeding the threshold, the preset threshold, the maximum allowable acquisition response duration, and the maximum and minimum values of the actual acquisition response duration, etc., so as to more accurately reflect the sensitivity and response speed of the gene acquisition unit. The coefficient comparison module compares the calculated data acquisition sensitivity coefficient with the preset sensitivity coefficient threshold. Once the sensitivity coefficient is lower than the threshold, the abnormal alarm determination module will immediately determine that there is an abnormality in the timeliness of data acquisition of the gene acquisition unit and trigger an abnormal alarm. This mechanism ensures that problems can be discovered and processed in a timely manner, avoiding delays or failures in data acquisition. Overall, through multiple links such as real-time monitoring, precise analysis, scientific calculation, and intelligent determination, this technical solution acts on the gene acquisition process together, significantly improving the quality and efficiency of data acquisition. It can not only discover and handle problems in the acquisition process in a timely manner, but also ensure the accuracy and integrity of data, providing strong support for subsequent gene analysis and research. This technical solution enhances the reliability and stability of the gene acquisition system through a series of intelligent monitoring and determination mechanisms. It can operate stably in various complex and changeable environments, ensuring the continuity and accuracy of data acquisition, thereby improving the performance and usability of the entire system.
[0113] To sum up, through multiple aspects such as precise monitoring, scientific calculation, and intelligent determination, this technical solution realizes the comprehensive monitoring and management of the gene acquisition process, significantly improving the quality and efficiency of data acquisition and enhancing the reliability and stability of the system.
[0114] In this embodiment, the data communication module includes:
[0115] A data sending unit for sending real-time user gene sequence data;
[0116] A data receiving unit for receiving real-time user gene sequence data;
[0117] According to the requirements of gene sequence alignment and mutation detection based on cloud computing, establish a data communication link between the data sending unit and the data receiving unit;
[0118] Wherein, the data sending unit transmits an instruction requesting to establish a data communication link to the data receiving unit;
[0119] The data receiving unit receives the instruction transmitted by the data sending unit requesting to establish a data communication link, and the data receiving unit checks the data communication channel according to the data communication requirement to check whether there is an idle data communication channel for transmitting the real-time data of the user gene sequence;
[0120] When the data receiving unit does not have an idle data communication channel for transmitting the real-time data of the user's gene sequence, the data receiving unit transmits an instruction to the data sending unit not to agree to establish a data communication link, and at this time, the data sending unit queues and waits;
[0121] When the data receiving unit has an idle data communication channel for transmitting the real-time data of the user's gene sequence, the data receiving unit transmits an instruction to agree to establish a data communication link to the data sending unit;
[0122] The data sending unit receives the instruction of agreeing to establish the data communication link transmitted by the data receiving unit, and the data sending unit establishes the data communication link with the data receiving unit according to the instruction of agreeing to establish the data communication link transmitted by the data receiving unit;
[0123] Specifically, the data sending unit receives the user gene sequence real-time data transmitted from the gene collection terminal, and the data sending unit sends the received user gene sequence real-time data to the data receiving unit;
[0124] Specifically, the data receiving unit receives the user gene sequence real-time data transmitted from the data sending unit, and the data receiving unit sends the received user gene sequence real-time data to the cloud server terminal.
[0125] The cloud server terminal is used to process the user's real-time gene sequence data, perform gene sequence comparison and variation detection, and determine the gene sequence comparison results and gene sequence comparison and variation detection results.
[0126] In this embodiment, the cloud server terminal includes:
[0127] The data processing module is used to process the real-time data of the user's gene sequence and determine the user's gene sequence feature data.
[0128] In this embodiment, the data processing module includes:
[0129] A data cleaning unit, used to clean the user's gene sequence real-time data;
[0130] Obtain real-time data of user gene sequences;
[0131] Clean the real-time data of the user's gene sequence, including:
[0132] Check the consistency of the real-time data of the user's gene sequence;
[0133] According to the reasonable value ranges and mutual relationships of each parameter in the real-time data of the user's gene sequence, check whether the real-time data of the user's gene sequence meets the requirements;
[0134] Remove the inconsistent data in the real-time data of the user's gene sequence that exceeds the normal range, is logically unreasonable or contradictory;
[0135] Process the invalid values and missing values in the real-time data of the user's gene sequence;
[0136] According to the requirements of data validity and integrity, check whether the real-time data of the user's gene sequence contains invalid values and missing values;
[0137] Remove the invalid values and missing values in the real-time data of the user's gene sequence that are useless for gene sequence alignment and variant detection;
[0138] Determine the real-time data of the user's gene sequence that is useful for gene sequence alignment and variant detection;
[0139] It should be noted that data cleaning refers to the process of processing and organizing the real-time data of the user's gene sequence during the data analysis process to improve the quality and usability of the real-time data of the user's gene sequence.
[0140] Among them, data cleaning includes checking data consistency, processing invalid values and missing values. When checking the consistency of the real-time data of the user's gene sequence, according to the reasonable value ranges and mutual relationships of each parameter in the real-time data of the user's gene sequence, check whether the real-time data of the user's gene sequence meets the requirements, and remove the inconsistent data in the real-time data of the user's gene sequence that exceeds the normal range, is logically unreasonable or contradictory; for example, if a variable measured on a 1-7 scale shows a value of 0, or if the weight shows a negative number, it should be regarded as exceeding the normal value range; computer software such as SPSS, SAS, and Excel can automatically identify each variable value that exceeds the defined range; answers with logical inconsistencies may appear in various forms: for example, many respondents say they drive to work but also report not having a car; or respondents report being heavy purchasers and users of a certain brand but at the same time give a very low score on the familiarity scale; when inconsistencies are found, the questionnaire serial number, record serial number, variable name, error category, etc. should be listed for further verification and correction.
[0141] Meanwhile, due to investigation, coding, and input errors, there may be some invalid values and missing values in the real-time user gene sequence data. It is necessary to process the invalid values and missing values to improve the subsequent processing accuracy and efficiency of the real-time user gene sequence data.
[0142] A feature extraction unit for extracting features from the real-time user gene sequence data;
[0143] Obtain the real-time user gene sequence data that is useful for gene sequence alignment and mutation detection after cleaning;
[0144] Extract features from the real-time user gene sequence data that is useful for gene sequence alignment and mutation detection;
[0145] Extract features that can reflect gene sequence alignment and mutation detection;
[0146] Determine the user gene sequence feature data.
[0147] A comparison and analysis module for comparing and analyzing the user gene sequence feature data to determine the gene sequence alignment result based on cloud computing.
[0148] In this embodiment, the comparison and analysis module includes:
[0149] A standard setting unit for setting the user gene sequence standard data;
[0150] According to the requirements of gene sequence alignment and mutation detection based on cloud computing, the user gene sequence standard data is preset in advance. The user gene sequence standard data is used to provide a reference basis for gene sequence alignment;
[0151] A comparison and analysis unit for comparing and analyzing the user gene sequence feature data;
[0152] Obtain the user gene sequence feature data and the user gene sequence standard data;
[0153] Based on the global alignment algorithm, compare and analyze the user gene sequence feature data and the user gene sequence standard data, judge the similarity and difference between the user gene sequence feature data and the user gene sequence standard data, and determine the gene sequence alignment result based on cloud computing;
[0154] Among them, if the user gene sequence feature data is within the range of the user gene sequence standard data, the gene sequence alignment result based on cloud computing is that the user gene sequence is normal;
[0155] Among them, if the user gene sequence feature data is not within the range of the user gene sequence standard data, the gene sequence alignment result based on cloud computing is that the user gene sequence has a mutation.
[0156] It should be noted that the global alignment algorithm is a method of comparing a given sequence from beginning to end to make as many characters as possible match in the same column; this algorithm is applicable to sequences with relatively high similarity and similar lengths. The global alignment algorithm usually uses dynamic programming algorithms, such as the Needleman-Wunsch algorithm and the Smith-Waterman algorithm.
[0157] A mutation detection module, which is used to search and locate the mutation sites of gene sequences and determine the gene sequence alignment and mutation detection results based on cloud computing.
[0158] In this embodiment, the mutation detection module includes:
[0159] A search and location unit, which is used to search and locate the mutation sites of gene sequences;
[0160] Obtain the gene sequence alignment results based on cloud computing;
[0161] According to the gene sequence alignment results based on cloud computing, mine and analyze the user's gene sequence characteristic data to find the mutation region of the user's gene sequence;
[0162] Based on the mutation region of the user's gene sequence, search for and locate the mutation sites of the gene sequence to locate the mutation sites of the gene sequence;
[0163] Determine the gene sequence alignment and mutation detection results based on cloud computing.
[0164] A storage and display module, which is used to store and display the gene sequence alignment and mutation detection report based on cloud computing.
[0165] In this embodiment, the storage and display module includes:
[0166] A distributed database, which is used to store the gene sequence alignment and mutation detection report based on cloud computing;
[0167] Obtain the user's gene sequence characteristic data, gene sequence alignment results, and gene sequence alignment and mutation detection results;
[0168] Analyze and integrate the user's gene sequence characteristic data, gene sequence alignment results, and gene sequence alignment and mutation detection results to form a gene sequence alignment and mutation detection report based on cloud computing;
[0169] Store the gene sequence alignment and mutation detection report based on cloud computing;
[0170] A display and sharing unit, which is used to display the gene sequence alignment and mutation detection report based on cloud computing in a visual form and share the gene sequence alignment and mutation detection report based on cloud computing.
[0171] To better demonstrate the gene sequence alignment and variant detection process based on cloud computing, this embodiment now provides a gene sequence alignment and variant detection method based on cloud computing, which is implemented based on the above-mentioned gene sequence alignment and variant detection system based on cloud computing, and includes the following steps:
[0172] S1: Manage user identity registration, login, and user permissions through a gene collection terminal, collect real-time data of the user's gene sequence, transmit the real-time data of the user's gene sequence through a data communication module, and establish data communication between the gene collection terminal and the cloud server terminal;
[0173] S2: Process the real-time data of the user's gene sequence through the cloud server terminal, perform alignment analysis on the feature data and standard data of the user's gene sequence based on the global alignment algorithm, determine the gene sequence alignment result, search and locate the gene sequence variant sites, and determine the gene sequence alignment and variant detection result.
[0174] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0175] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A gene sequence alignment and variation detection system based on cloud computing, characterized in that: include: Gene collection terminal, used to manage users and collect real-time data of user gene sequences; Data communication module, used to transmit user gene sequence real-time data and establish data communication between gene collection terminal and cloud server terminal; The cloud server terminal is used to process the user's real-time gene sequence data, perform gene sequence comparison and variation detection, and determine the gene sequence comparison results and gene sequence comparison and variation detection results; The gene collection terminal comprises: User management unit, used to manage user identity registration, login and user rights; The user management interface prompts the user to input the gene sequence, and the user inputs the gene sequence according to the prompt of the user management interface; A gene collection unit, used for real-time collection of gene sequences input by users; After the user inputs the gene sequence, the user management interface collects the gene sequence in real time and automatically converts the gene sequence into a format that can be recognized by the computer to determine the real-time data of the user's gene sequence; The gene collection terminal further includes: A response time real-time monitoring module is used to monitor in real time the collection response time of the gene collection unit after the gene sequence is input by the user; A duration comparison module, used to compare the collection response duration of the gene collection unit after the gene sequence input by the user with a preset collection response duration threshold; A collection monitoring quantity setting module, used for setting the collection monitoring quantity using the collection response time exceeding the preset collection response time threshold when the collection response time of the gene collection unit after the gene sequence input by the user exceeds the preset collection response time threshold; The collected monitoring quantity is obtained by the following formula: Where, Q represents the number of collected monitoring; INT[] represents rounding up the function inside the brackets; T represents the collection response time that exceeds the preset collection response time threshold; T th Indicates the preset collection response time threshold; T z represents the maximum allowable acquisition response time corresponding to the minimum sensitivity that satisfies the normal operation of data acquisition; λ represents the adjustment coefficient, and the value range of the adjustment coefficient is 0.13-0.36; t d represents the average time interval for data collection of the gene collection unit; A timeliness abnormality determination module is used to determine whether there is a data collection timeliness abnormality in the gene collection unit by using the collection response time of the real-time collection times of the gene collection unit corresponding to the collection monitoring quantity after the collection response time exceeds a preset collection response time threshold; The timeliness abnormality determination module includes: A collection parameter monitoring module, used to monitor the real-time collection times of the gene collection unit according to the collection monitoring quantity after the collection response time exceeds a preset collection response time threshold; A collection response duration calling module is used to call the collection response duration corresponding to each data collection in the real-time collection times of the gene collection unit when the real-time collection times of the gene collection unit reaches the collection monitoring number; A data collection sensitivity coefficient acquisition module is used to obtain the data collection sensitivity coefficient corresponding to the gene collection unit by using the collection response time corresponding to each data collection in the real-time collection times of the gene collection unit; The data collection sensitivity coefficient corresponding to the gene collection unit is obtained by the following formula: Wherein, K represents the data collection sensitivity coefficient corresponding to the gene collection unit; Q represents the number of collection monitoring; T represents the collection response time exceeding the preset collection response time threshold; T th Indicates the preset collection response time threshold; T z Indicates the maximum allowable acquisition response time corresponding to the minimum sensitivity that satisfies the normal operation of data acquisition; T max and T min represents the maximum and minimum collection response time corresponding to Q data collections of the gene collection unit; T i represents the collection response time corresponding to the i-th data collection of the gene collection unit; A coefficient comparison module, used for comparing the data acquisition sensitivity coefficient with a preset sensitivity coefficient threshold; An abnormal alarm determination and module, used to determine that the gene collection unit has abnormal data collection timeliness when the data collection sensitivity coefficient is lower than a preset sensitivity coefficient threshold, and to issue an abnormal alarm; The cloud server terminal comprises: A data processing module is used to process the real-time data of the user's gene sequence and determine the characteristic data of the user's gene sequence; The comparison and analysis module is used to compare and analyze the user's gene sequence feature data and determine the gene sequence comparison results based on cloud computing; The mutation detection module is used to locate the mutation sites of gene sequences and determine the gene sequence comparison and mutation detection results based on cloud computing; The storage and display module is used to store and display gene sequence comparison and variation detection reports based on cloud computing.
2. The cloud computing-based gene sequence alignment and variation detection system according to claim 1, characterized in that: The data communication module comprises: A data sending unit, used to send user gene sequence real-time data; A data receiving unit, used to receive real-time gene sequence data from users; According to the requirements of gene sequence comparison and variation detection based on cloud computing, a data communication link is established between the data sending unit and the data receiving unit; The data sending unit receives the user gene sequence real-time data transmitted from the gene collection terminal, and the data sending unit sends the received user gene sequence real-time data to the data receiving unit; The data receiving unit receives the user gene sequence real-time data transmitted by the data sending unit, and the data receiving unit sends the received user gene sequence real-time data to the cloud server terminal.
3. The cloud computing-based gene sequence alignment and variation detection system according to claim 2, characterized in that: The data processing module comprises: A data cleaning unit, used to clean the user's gene sequence real-time data; Obtain real-time data of user gene sequences; Clean the user's real-time gene sequence data; Remove inconsistent data, invalid values and missing values in the user's real-time gene sequence data; Determine the user's real-time gene sequence data that is useful for gene sequence alignment and variation detection; A feature extraction unit, used to extract features from user gene sequence real-time data; Obtain real-time user gene sequence data that is useful for gene sequence alignment and variation detection after cleaning; Extract features from real-time user gene sequence data useful for gene sequence alignment and variation detection; Extract features that can reflect gene sequence alignment and variation detection; Determine the user's gene sequence feature data.
4. The cloud computing-based gene sequence alignment and variation detection system according to claim 3, characterized in that: The comparison and analysis module comprises: A standard setting unit, used to set user gene sequence standard data; According to the needs of gene sequence comparison and variation detection based on cloud computing, user gene sequence standard data is pre-set, and the user gene sequence standard data is used to provide a reference for gene sequence comparison; Comparison and analysis unit, used to compare and analyze user gene sequence feature data; Obtain user gene sequence characteristic data and user gene sequence standard data; Based on the global comparison algorithm, the user's gene sequence feature data and the user's gene sequence standard data are compared and analyzed to determine the similarities and differences between the user's gene sequence feature data and the user's gene sequence standard data, and determine the gene sequence comparison result based on cloud computing; Wherein, if the user's gene sequence characteristic data is within the user's gene sequence standard data range, the gene sequence comparison result based on cloud computing is that the user's gene sequence is normal; Among them, if the user's gene sequence feature data is not within the user's gene sequence standard data range, the gene sequence comparison result based on cloud computing is the user's gene sequence variation.
5. The cloud computing-based gene sequence alignment and variation detection system according to claim 4, characterized in that: The variation detection module comprises: A search and positioning unit is used to search and locate the gene sequence variation site; Obtain gene sequence comparison results based on cloud computing; According to the gene sequence comparison results based on cloud computing, the user's gene sequence feature data is mined and analyzed to find the user's gene sequence variation area; Based on the user's gene sequence variation region, the gene sequence variation site is searched and located to locate the gene sequence variation site; Determine the results of gene sequence alignment and variation detection based on cloud computing.
6. The cloud computing-based gene sequence alignment and variation detection system according to claim 5, characterized in that: The storage display module includes: A distributed database for storing cloud-based gene sequence alignment and mutation detection reports; Obtain user gene sequence feature data, gene sequence comparison results, and gene sequence comparison and variation detection results; Analyze and integrate user gene sequence feature data, gene sequence comparison results, and gene sequence comparison and variation detection results to form a gene sequence comparison and variation detection report based on cloud computing; Store cloud-based gene sequence alignment and mutation detection reports; The display sharing unit is used to display the gene sequence alignment and variation detection report based on cloud computing in a visual form, and to share the gene sequence alignment and variation detection report based on cloud computing.
7. A method for gene sequence alignment and variation detection based on cloud computing, implemented based on the gene sequence alignment and variation detection system based on cloud computing according to claim 6, characterized in that: The steps include: S1: Manage user identity registration, login and user rights through the gene collection terminal, collect user gene sequence real-time data, transmit user gene sequence real-time data through the data communication module, and establish data communication between the gene collection terminal and the cloud server terminal; S2: Process the user's gene sequence real-time data through the cloud server terminal, compare and analyze the user's gene sequence feature data and the user's gene sequence standard data based on the global comparison algorithm, determine the gene sequence comparison result, search and locate the gene sequence variation site, and determine the gene sequence comparison and variation detection results.
Citation Information
Patent Citations
Method and device for compressing and decompressing gene sequences
CN107633158B
System used for intelligently analyzing individualized tumor gene detection
CN107437004A
Cloud computing-based method for quality control and management of gene sequence data
CN109584958A