Method for tamper-proofing encryption of brain glioma gene high-throughput sequencing result

By generating the first key and the second key, double encryption of the brain glioma gene sequence is solved, and the problem of poor encryption effect in the prior art is achieved, achieving more efficient data encryption and privacy protection.

CN120030607AInactive Publication Date: 2025-05-23JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510487307.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120030607A_ABST
    Figure CN120030607A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data encryption, in particular to a tamper-proof encryption method for a brain glioma gene high-throughput sequencing result. The method comprises the following steps: firstly, generating a first key of a brain glioma gene sequence; according to the repetition condition of a brain glioma gene sequence and a normal brain gene sequence, obtaining each segment to-be-encrypted gene sequence and a segment normal gene sequence; according to the difference condition between each segment of the to-be-encrypted gene sequence and the corresponding segment of the normal gene sequence, obtaining the important measure of each segment of the to-be-encrypted gene sequence; acquiring a second key of the brain glioma gene sequence according to the important measurement of all segmented to-be-encrypted gene sequences of the brain glioma gene sequence; and performing data encryption on the brain glioma gene sequence according to the first key and the second key of the brain glioma gene sequence. According to the method, the difference between the brain glioma gene sequence and the normal brain gene sequence is fully considered, and dual encryption is performed, so that the encryption effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data encryption, and in particular to a method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes. Background Art

[0002] High-throughput sequencing technology is a technology that can perform parallel sequencing on a large number of nucleic acid molecules at a time. It has the advantages of high sensitivity, high resolution and high throughput. Glioma is one of the common primary malignant tumors of the central nervous system. High-throughput sequencing can be used to comprehensively analyze the genes in brain glioma samples, providing a theoretical basis for the clinical diagnosis and treatment of glioma.

[0003] The high-throughput sequencing results of glioma genes contain a large amount of personal health information of patients, which is very sensitive. During the data transmission process, the use of tamper-proof encryption transmission methods can improve the security of data and ensure data sharing between different institutions while protecting the privacy of patients. The existing technology usually directly uses a symmetric encryption algorithm to generate a key to encrypt the glioma gene sequence. When using the existing technology to encrypt the glioma gene sequence, the difference between the glioma gene sequence and the normal brain gene sequence is not fully considered, making it difficult to ensure the encryption effect. Summary of the invention

[0004] In order to solve the technical problem that the prior art has poor encryption effect on brain glioma gene sequences, the purpose of the present invention is to provide a method for tamper-proof encryption of brain glioma gene high-throughput sequencing results. The technical solution adopted is as follows: A method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes, the method comprising: Obtaining gene sequences of brain glioma and normal brain; Generate a first key of the glioma gene sequence; divide the glioma gene sequence and the normal brain gene sequence according to the repetition of the glioma gene sequence and the normal brain gene sequence to obtain each segmented gene sequence to be encrypted and a segmented normal gene sequence; obtain an important metric of each segmented gene sequence to be encrypted according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; obtain a second key of the glioma gene sequence according to the important metrics of all segmented gene sequences to be encrypted of the glioma gene sequence; Data of the brain glioma gene sequence is encrypted according to the first key and the second key of the brain glioma gene sequence.

[0005] Furthermore, the method for obtaining the first key includes: The first key of the brain glioma gene sequence is generated by using the AES algorithm.

[0006] Furthermore, the method for obtaining each segmented gene sequence to be encrypted and each segmented normal gene sequence includes: Obtaining each characteristic substring according to the repetition of the brain glioma gene sequence and the normal brain gene sequence; Acquire characteristic parameters of the characteristic substring according to the sequence length and repetition frequency of the characteristic substring; The largest characteristic parameter corresponds to the characteristic substring as a partition substring; The brain glioma gene sequence and the normal brain gene sequence are divided respectively by using the divided substrings to obtain each segmented gene sequence to be encrypted and a segmented normal gene sequence.

[0007] Furthermore, the method for acquiring the characteristic parameters includes: For any characteristic substring, the repetition frequency of the characteristic substring in the glioma gene sequence is taken as the first frequency; the repetition frequency of the characteristic substring in the normal brain gene sequence is taken as the second frequency; the sum of the first frequency and the second frequency is calculated as the repetition parameter of the characteristic substring; the sum of the repetition parameter of the characteristic substring and the sequence length is calculated and normalized to obtain the characteristic parameter of the characteristic substring.

[0008] Furthermore, the method for obtaining each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence includes: The brain glioma gene sequence is divided by using the division substring to obtain each segmented gene sequence to be encrypted; the normal brain gene sequence is divided by using the division substring to obtain each segmented normal gene sequence.

[0009] Furthermore, the method for obtaining the important metrics includes: Obtaining a first important metric of the segmented gene sequence to be encrypted according to the difference in element proportions between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; According to the difference in sequence length between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence, obtaining a second important metric of the segmented gene sequence to be encrypted; Obtaining a third important metric of the segmented gene sequence to be encrypted according to the difference in the characteristic parameters included in the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; The first important metric, the second important metric and the third important metric are forwardly integrated to obtain the important metric of the segmented gene sequence to be encrypted.

[0010] Furthermore, the method for obtaining the first important metric includes: In the segmented gene sequence to be encrypted, the proportion of each element in the segmented gene sequence to be encrypted is taken as the first proportion of each element; in the segmented normal gene sequence, the proportion of each element in the segmented normal gene sequence is taken as the second proportion of each element; The absolute value of the difference between the first proportion and the second proportion of the same element is calculated to obtain a local difference value of the element; and the local difference values ​​of all elements of the segmented gene sequence to be encrypted are averaged to obtain a first important metric.

[0011] Furthermore, the method for obtaining the third important metric includes: In the segmented gene sequence to be encrypted, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a first characteristic parameter; in the segmented normal gene sequence, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a second characteristic parameter; the absolute value of the difference between the first characteristic parameter and the second characteristic parameter is calculated to obtain a third important metric of the segmented gene sequence to be encrypted.

[0012] Furthermore, the method for obtaining the second key includes: The segmented gene sequence to be encrypted whose importance metric is less than a preset importance threshold is used as a non-critical segmented sequence; for the suffix tree of the glioma gene sequence, the nodes corresponding to the non-critical segmented sequence are removed to obtain a simplified suffix tree; the segmented gene sequence to be encrypted is encrypted using a hash algorithm to obtain a hash value of the segmented gene sequence to be encrypted; the hash values ​​of the segmented gene sequence to be encrypted are sorted according to the simplified suffix tree structure to obtain the hash value of the glioma gene sequence, and the hash value is used as the second key of the glioma gene sequence.

[0013] Furthermore, the method for encrypting data of the brain glioma gene sequence includes: The first key is used to encrypt the brain glioma gene sequence to generate a ciphertext; the second key of the brain glioma gene sequence is used to generate a digital signature; and the ciphertext and the digital signature are used as encrypted data.

[0014] The present invention has the following beneficial effects: In order to encrypt data, the first key of the glioma gene sequence is first generated; in order to divide the gene sequence into smaller units, the glioma gene sequence and the normal brain gene sequence are divided to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence; according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence, the important metric of each segmented gene sequence to be encrypted is obtained; the important metric evaluates the importance of the segmented gene sequence to be encrypted to the onset of glioma. Combining the important metrics of all segmented gene sequences to be encrypted of the glioma gene sequence, the second key can be generated. Double data encryption is performed using the first key and the second key to improve the encryption effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 A flowchart of a method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes provided by one embodiment of the present invention; Figure 2 A flow chart of a method for obtaining segmented gene sequences to be encrypted and segmented normal gene sequences provided by one embodiment of the present invention; Figure 3 A flow chart of a method for obtaining important metrics provided by an embodiment of the present invention; Figure 4 A structural diagram of a system for tamper-proof encryption of high-throughput sequencing results of glioma genes provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes proposed by the present invention, its specific implementation method, structure, characteristics and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0018] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0019] The specific scheme of the method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes provided by the present invention is described in detail below with reference to the accompanying drawings.

[0020] The present invention provides a method for encrypting the results of high-throughput sequencing of glioma genes to prevent tampering. Figure 1 , which shows a flow chart of a method for tamper-proof encryption of brain glioma gene high-throughput sequencing results provided by an embodiment of the present invention, the method comprising the following steps: Step S1: Obtaining glioma gene sequences and normal brain gene sequences.

[0021] In order to improve confidentiality, the difference between glioma gene sequences and normal brain gene sequences needs to be fully considered. First, the glioma gene sequences and normal brain gene sequences need to be obtained.

[0022] The data acquisition process must strictly follow the user authorization principle to ensure that the glioma gene sequence and normal brain gene sequence are legally obtained from the monitoring system. The specific acquisition process includes: First, tissue sections are obtained from brain tissues of glioma patients and healthy individuals, and DNA samples are extracted. Subsequently, a suitable high-throughput sequencing system is selected for sequencing according to the specific needs of the study. The Illumina sequencing system is a high-throughput sequencing system with efficient and accurate sequencing capabilities. The present invention uses the Illumina sequencing system to introduce the constructed DNA sample into the sequencing flow pool, and start the automated sequencing process to accurately capture and parse the gene sequence of the sample. The gene sequence of glioma patients is used as the glioma gene sequence for subsequent encryption of the gene sequence to protect the privacy and information security of the patient; the gene sequence of healthy individuals is used as the normal brain gene sequence, and the normal brain gene sequence can be used as a control group for comparing and analyzing the differences between the glioma gene sequence and the normal brain gene sequence. It should be noted that the data collection in the present invention is authorized by the user, does not violate relevant laws and regulations, and does not violate public order and good customs.

[0023] Since the prior art only uses one key to encrypt the glioma gene sequence to ensure that only the sender and receiver can decrypt and read the transmitted data during the data transmission process, the risk of a single key being cracked is high, making it difficult to ensure the encryption effect. The present invention constructs a dual encryption scheme to effectively ensure the encryption effect.

[0024] Step S2: Generate the first key for the glioma gene sequence; divide the glioma gene sequence and the normal brain gene sequence according to the repetition situation of the glioma gene sequence and the normal brain gene sequence to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence; obtain the importance measure of each segmented gene sequence to be encrypted according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; obtain the second key of the glioma gene sequence according to the importance measures of all the segmented gene sequences to be encrypted of the glioma gene sequence.

[0025] To perform data encryption, first generate the first key for the glioma gene sequence; to divide the gene sequence into smaller units, divide the glioma gene sequence and the normal brain gene sequence to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence; obtain the importance measure of each segmented gene sequence to be encrypted according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; the importance measure evaluates the importance of the segmented gene sequence to be encrypted for the occurrence of glioma. Combining the importance measures of all the segmented gene sequences to be encrypted of the glioma gene sequence, the second key can be generated.

[0026] To ensure the encryption effect, first, the first key for the glioma gene sequence needs to be generated. Preferably, in an embodiment of the present invention, the method for obtaining the first key includes: Use the AES (Advanced Encryption Standard) algorithm to generate the first key for the glioma gene sequence. It should be noted that the AES algorithm is a symmetric encryption algorithm that uses the same key to encrypt and decrypt data. Using the AES algorithm to generate the first key for the glioma gene sequence is a well-known prior art to those skilled in the art. Here, it is only briefly described: use the AES algorithm to generate the AES key for the glioma gene sequence, and use the AES key as the first key.

[0027] Considering that the glioma gene sequence has specific biological significance, glioma is a complex type of tumor, and its occurrence and development involve abnormalities of multiple genes. By segmenting the glioma gene sequence, gene segments related to the tumor can be more accurately identified. Please refer to Figure 2 , which shows a flowchart of a method for obtaining each segmented gene sequence to be encrypted and the segmented normal gene sequence in an embodiment of the present invention. Preferably, in an embodiment of the present invention, the method for obtaining each segmented gene sequence to be encrypted and the segmented normal gene sequence includes: Step S201: Obtain each characteristic substring according to the repetition situation of the glioma gene sequence and the normal brain gene sequence.

[0028] To identify substrings that appear repeatedly in glioma gene sequences and normal brain gene sequences.

[0029] Preferably, in one embodiment of the present invention, the method for obtaining the characteristic substring includes: The Ukkonen algorithm (On-line construction of suffix trees) is used to construct suffix trees for glioma gene sequences and normal brain gene sequences respectively; repeated substrings in the suffix trees of both glioma gene sequences and normal brain gene sequences are used as characteristic substrings. It should be noted that the Ukkonen algorithm is an online algorithm for constructing suffix trees, and the Ukkonen algorithm is a technical means well known to those skilled in the art, and will not be described in detail here.

[0030] In view of the above steps, considering that the repetitive substrings in the gene sequence have sequences with key biological significance, such as the promoter region or specific repetitive sequences in the gene sequence, the suffix tree is used to compare the gene sequence segments of tumor tissue with those of normal tissue, and the repetitive substrings in the suffix tree of both the glioma gene sequence and the normal brain gene sequence are used as characteristic substrings. The characteristic substrings represent conserved sequences that are important for maintaining gene function.

[0031] Step S202: Acquire characteristic parameters of the characteristic substring according to the sequence length and repetition frequency of the characteristic substring.

[0032] The expression intensity of the characteristic substring in the gene sequence is characterized by constructing characteristic parameters.

[0033] Preferably, in one embodiment of the present invention, the method for acquiring characteristic parameters includes: For any characteristic substring, the repetition frequency of the characteristic substring in the glioma gene sequence is taken as the first frequency; the repetition frequency of the characteristic substring in the normal brain gene sequence is taken as the second frequency; the sum of the first frequency and the second frequency is calculated as the repetition parameter of the characteristic substring; the sum of the repetition parameter of the characteristic substring and the sequence length is calculated and normalized to obtain the characteristic parameter of the characteristic substring. Normalization is a technical means well known to those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.

[0034] For the above steps, the frequency of repetition of the characteristic substring in the gene sequence is a direct reflection of the importance of expression. For example, the characteristic substring that may play a key role in gene replication, transcription or regulation has a higher frequency of occurrence; considering that the sequence length of the characteristic substring is also an indicator of its importance. Longer characteristic substrings may contain more genetic information and represent that the characteristic substring is more conservative, with stronger functional constraints and stability. The repetition parameter reflects the overall repetition of the characteristic substring in the two gene sequences. The characteristic parameter combines the repetition parameter and the sequence length to more comprehensively evaluate the expression strength of the characteristic substring in the gene sequence.

[0035] Step S203: The feature substring corresponding to the largest feature parameter is used as the partition substring.

[0036] The characteristic parameter reflects the expression strength of the characteristic substring in the gene sequence. The characteristic substring with the strongest expression strength is used as the partition substring for the subsequent partitioning of the gene sequence.

[0037] Step S204: using the partitioned substrings, the glioma gene sequence and the normal brain gene sequence are partitioned respectively to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence.

[0038] The selected partition substrings are used to partition the glioma gene sequence and the normal brain gene sequence to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence.

[0039] Preferably, in one embodiment of the present invention, the glioma gene sequence is divided by dividing substrings to obtain each segmented gene sequence to be encrypted; the normal brain gene sequence is divided by dividing substrings to obtain each segmented normal gene sequence. Among them, the specific process of dividing the glioma gene sequence is described: in the glioma gene sequence, each division substring is identified, and the first data and the last data of the division substring are used as division points to divide the glioma gene sequence to obtain each segmented gene sequence to be encrypted. It should be noted that the division point belongs to each division substring. It should be noted that the means for dividing the normal brain gene sequence and the glioma gene sequence are the same and will not be repeated here. In order to compare in sequence, the segmented gene sequence to be encrypted is numbered according to the order of the segmented gene sequence to be encrypted in the glioma gene sequence, and the sequence number of the segmented gene sequence to be encrypted is obtained; the segmented normal gene sequence is numbered according to the order of the segmented normal gene sequence in the normal brain gene sequence, and the number of the segmented normal gene sequence is obtained. The segmented gene sequence to be encrypted and the segmented normal gene sequence with the same label are regarded as the segmented gene sequence to be encrypted and the segmented normal gene sequence corresponding to each other.

[0040] Considering the differences between glioma gene sequences and normal brain gene sequences, this difference is very important in medicine and bioinformatics, because glioma gene sequences usually contain specific mutation sites, which are the key to revealing the disease mechanism. An important metric is constructed to quantify the importance of the segmented gene sequence to be encrypted to the onset of glioma. Figure 3 , which shows a flow chart of a method for obtaining important metrics in one embodiment of the present invention, the method for obtaining important metrics includes: Step S211: obtaining a first important metric of the segmented gene sequence to be encrypted according to the difference in element proportions between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence.

[0041] Taking into account that the segmented gene sequences to be encrypted due to gene mutations caused by the onset of brain glioma are different from the segmented normal gene sequences, in order to compare the difference in the element proportions of the segmented gene sequences to be encrypted and the corresponding segmented normal gene sequences, the difference in element composition between the two is measured.

[0042] In one embodiment of the present invention, in a segmented gene sequence to be encrypted, the proportion of each element in the segmented gene sequence to be encrypted is taken as the first proportion of each element; in a segmented normal gene sequence, the proportion of each element in the segmented normal gene sequence is taken as the second proportion of each element; the absolute value of the difference between the first proportion and the second proportion of the same element is calculated to obtain a local difference value of the element; the local difference values ​​of all elements of the segmented gene sequence to be encrypted are averaged to obtain a first important metric.

[0043] For the above steps, the first important metric reflects the difference in elemental composition between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence. The larger the first important metric, the greater the difference in elemental composition, and the more important the segmented gene sequence to be encrypted is to the onset of brain glioma.

[0044] Step S212: obtaining a second important metric of the segmented gene sequence to be encrypted according to the sequence length difference between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence.

[0045] The difference in length between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence is compared to measure the difference in length between the two.

[0046] In one embodiment of the present invention, the absolute value of the difference between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence is calculated to obtain the second important metric of the segmented gene sequence to be encrypted.

[0047] For the above steps, the second important metric reflects the difference in sequence length between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence. The larger the second important metric, the greater the difference in sequence length, and the more important the segmented gene sequence to be encrypted is to the onset of brain glioma.

[0048] Step S213: obtaining a third important metric of the segmented gene sequence to be encrypted according to the difference in characteristic parameters between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence.

[0049] By comparing the differences in characteristic parameters included in the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence, the difference in characteristic parameters between the two can be measured.

[0050] In one embodiment of the present invention, in a segmented gene sequence to be encrypted, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a first characteristic parameter; in a segmented normal gene sequence, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a second characteristic parameter; and the absolute value of the difference between the first characteristic parameter and the second characteristic parameter is calculated to obtain a third important metric of the segmented gene sequence to be encrypted.

[0051] For the above steps, the third important metric reflects the degree of difference between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence at the feature level. The larger the third important metric is, the greater the difference in the feature parameters is, and the more important the segmented gene sequence to be encrypted is to the onset of brain glioma.

[0052] Step S214: forwardly fuse the first important metric, the second important metric and the third important metric to obtain the important metric of the segmented gene sequence to be encrypted.

[0053] Combining the above three important metrics, we get an important metric that can fully reflect the difference between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence. The larger the important metric, the greater the comprehensive difference between the segmented gene sequence to be encrypted and the corresponding segmented normal gene sequence, and the more important the segmented gene sequence to be encrypted is to the onset of brain glioma.

[0054] It should be noted that forward fusion is an existing technology well known to those skilled in the art, and forward fusion can adopt simple product, arithmetic mean or other suitable fusion methods. In one embodiment of the present invention, the product of the first important metric, the second important metric and the third important metric is calculated and normalized to obtain the important metric of the segmented gene sequence to be encrypted. Among them, the present invention adopts linear normalization method for normalization.

[0055] As sensitive information in the biomedical field, the glioma gene sequence needs to be fully protected. By constructing a second key, access control to the glioma gene sequence data can be increased to prevent unauthorized access and leakage. Preferably, in one embodiment of the present invention, the method for obtaining the second key includes: The segmented gene sequence to be encrypted whose importance metric is less than the preset importance threshold is used as a non-critical segmented sequence; for the suffix tree of the glioma gene sequence, the nodes corresponding to the non-critical segmented sequence are removed to obtain a simplified suffix tree; the segmented gene sequence to be encrypted is encrypted using a hash algorithm to obtain the hash value of the segmented gene sequence to be encrypted; the hash values ​​of the segmented gene sequence to be encrypted are sorted according to the simplified suffix tree structure to obtain the hash value of the glioma gene sequence, and the hash value is used as the second key of the glioma gene sequence. In one embodiment of the present invention, the preset importance threshold is set to 0.57, and the implementer can set it according to the implementation scenario. It should be noted that the hash algorithm is a prior art well known to those skilled in the art and will not be described in detail here.

[0056] For the above steps, the construction of the second key is based on the importance metric. The importance metric reflects the importance of the segmented gene sequence to be encrypted in the pathogenesis of glioma. For the suffix tree of the glioma gene sequence, by removing the non-critical segmented sequences with low importance metrics, the suffix tree structure can be simplified, the complexity of the hash calculation can be reduced, and the security and efficiency of the key can be improved. The segmented gene sequence to be encrypted is encrypted using a hash algorithm to obtain the hash value of each segmented gene sequence to be encrypted. These hash values ​​are sorted according to the simplified suffix tree structure to retain the structural information. The sorted hash value sequence can be regarded as a characteristic representation of the glioma gene sequence. The sorted hash value sequence is used as the second key of the glioma gene sequence. The second key not only contains the key information of the gene sequence, but also enhances security through the hash algorithm and the suffix tree structure.

[0057] Step S3: Encrypt the data of the glioma gene sequence according to the first key and the second key of the glioma gene sequence.

[0058] Double data encryption is performed using a first key and a second key to improve encryption effect.

[0059] Preferably, in one embodiment of the present invention, the method for encrypting data of a glioma gene sequence comprises: Based on the AES algorithm, the first key is used to encrypt the brain glioma gene sequence to generate a ciphertext; the second key of the brain glioma gene sequence is used to generate a digital signature; and the ciphertext and the digital signature are used as encrypted data.

[0060] It should be noted that the AES algorithm is a well-known technical means in the art. Here, a brief description is given of how to encrypt the brain glioma gene sequence using the first key to generate a ciphertext: Set the block length to 128 bits, process the brain glioma gene sequence in blocks, and use padding technology to fill in data blocks with a length of less than 128 bits to ensure that the size of each data block is consistent. Use the first key to encrypt each data block to obtain the ciphertext block corresponding to the data block, and connect these ciphertext blocks in series in the order of the original sequence to obtain the encrypted data sequence, and use the encrypted data sequence as the ciphertext. Use the second key generated previously to generate a digital signature. The digital signature can be the result of a hash operation on the second key. The purpose of the digital signature is to verify the integrity and authenticity of the data and ensure that the ciphertext has not been tampered with during transmission. Use the ciphertext and digital signature as encrypted data. When transmitting or storing, these two parts of data are usually processed together.

[0061] Furthermore, the first key, the second key and the encrypted data are packaged and uploaded. During subsequent use, the authorized recipient first uses the same hash function to generate a new hash code and verify the second key. When the second key verification is successful, the encrypted data is decrypted according to the comparison result, and the first key is further verified using the public key provided by the sender, and the data summary is calculated at the same time. When the double keys match at the same time, it means that the original data has not been tampered with and the integrity is good. When the signature verification is successful, the original data is restored and decryption access is allowed.

[0062] In summary, the embodiment of the present invention provides a method for tamper-proof encryption of high-throughput sequencing results of glioma genes. First, a first key of a glioma gene sequence is generated; according to the repetition of the glioma gene sequence and the normal brain gene sequence, the glioma gene sequence and the normal brain gene sequence are divided to obtain each segmented gene sequence to be encrypted and the segmented normal gene sequence; according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence, the important metric of each segmented gene sequence to be encrypted is obtained; according to the important metric of all segmented gene sequences to be encrypted of the glioma gene sequence, the second key of the glioma gene sequence is obtained; according to the first key and the second key of the glioma gene sequence, the glioma gene sequence is encrypted. In the embodiment of the present invention, by fully considering the difference between the glioma gene sequence and the normal brain gene sequence, double encryption is performed to improve the encryption effect.

[0063] The present invention also proposes a tamper-proof encryption system for high-throughput sequencing results of brain glioma genes, please refer to Figure 4, which shows a system structure diagram of tamper-proof encryption of brain glioma gene high-throughput sequencing results provided by an embodiment of the present invention. The system includes: a data acquisition module 101, a key analysis module 102 and an encryption module 103.

[0064] The data acquisition module 101 is used to acquire the gene sequence of brain glioma and the gene sequence of normal brain.

[0065] The key analysis module 102 is used to generate a first key for a glioma gene sequence; divide the glioma gene sequence and the normal brain gene sequence according to their repetitions to obtain each segmented gene sequence to be encrypted and a segmented normal gene sequence; obtain an important metric for each segmented gene sequence to be encrypted according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; obtain a second key for the glioma gene sequence according to the important metrics of all segmented gene sequences to be encrypted of the glioma gene sequence.

[0066] The encryption module 103 is used to encrypt data of the glioma gene sequence according to the first key and the second key of the glioma gene sequence.

[0067] It should be noted that: the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the system for tamper-proof encryption of high-throughput sequencing results of glioma genes provided in the above embodiment and the method embodiment for tamper-proof encryption of high-throughput sequencing results of glioma genes belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0068] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0069] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. A method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes, characterized in that: The method comprises: Obtaining gene sequences of brain glioma and normal brain; Generate a first key of the glioma gene sequence; divide the glioma gene sequence and the normal brain gene sequence according to the repetition of the glioma gene sequence and the normal brain gene sequence to obtain each segmented gene sequence to be encrypted and a segmented normal gene sequence; obtain an important metric of each segmented gene sequence to be encrypted according to the difference between each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; obtain a second key of the glioma gene sequence according to the important metrics of all segmented gene sequences to be encrypted of the glioma gene sequence; Data of the brain glioma gene sequence is encrypted according to the first key and the second key of the brain glioma gene sequence.

2. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 1, characterized in that: The method for obtaining the first key includes: The first key of the brain glioma gene sequence is generated by using the AES algorithm.

3. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 1, characterized in that: The method for obtaining each segmented gene sequence to be encrypted and each segmented normal gene sequence comprises: Obtaining each characteristic substring according to the repetition of the brain glioma gene sequence and the normal brain gene sequence; Acquire characteristic parameters of the characteristic substring according to the sequence length and repetition frequency of the characteristic substring; The largest characteristic parameter corresponds to the characteristic substring as a partition substring; The brain glioma gene sequence and the normal brain gene sequence are divided respectively by using the divided substrings to obtain each segmented gene sequence to be encrypted and a segmented normal gene sequence.

4. A method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes according to claim 3, characterized in that: The method for obtaining the characteristic parameters includes: For any characteristic substring, the repetition frequency of the characteristic substring in the glioma gene sequence is taken as the first frequency; the repetition frequency of the characteristic substring in the normal brain gene sequence is taken as the second frequency; the sum of the first frequency and the second frequency is calculated as the repetition parameter of the characteristic substring; the sum of the repetition parameter of the characteristic substring and the sequence length is calculated and normalized to obtain the characteristic parameter of the characteristic substring.

5. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 2, characterized in that: The method for obtaining each segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence comprises: The brain glioma gene sequence is divided by using the division substring to obtain each segmented gene sequence to be encrypted; the normal brain gene sequence is divided by using the division substring to obtain each segmented normal gene sequence.

6. The method for tamper-proof encryption of high-throughput sequencing results of brain glioma genes according to claim 3, characterized in that: The method for obtaining the important metrics includes: Obtaining a first important metric of the segmented gene sequence to be encrypted according to the difference in element proportions between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; According to the difference in sequence length between the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence, obtaining a second important metric of the segmented gene sequence to be encrypted; Obtaining a third important metric of the segmented gene sequence to be encrypted according to the difference in the characteristic parameters included in the segmented gene sequence to be encrypted and its corresponding segmented normal gene sequence; The first important metric, the second important metric and the third important metric are forwardly integrated to obtain the important metric of the segmented gene sequence to be encrypted.

7. A method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 6, characterized in that: The method for obtaining the first important metric includes: In the segmented gene sequence to be encrypted, the proportion of each element in the segmented gene sequence to be encrypted is taken as the first proportion of each element; in the segmented normal gene sequence, the proportion of each element in the segmented normal gene sequence is taken as the second proportion of each element; The absolute value of the difference between the first proportion and the second proportion of the same element is calculated to obtain a local difference value of the element; and the local difference values ​​of all elements of the segmented gene sequence to be encrypted are averaged to obtain a first important metric.

8. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 6, characterized in that: The method for obtaining the third important metric includes: In the segmented gene sequence to be encrypted, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a first characteristic parameter; in the segmented normal gene sequence, the mean of the characteristic parameters of all characteristic substrings is calculated to obtain a second characteristic parameter; the absolute value of the difference between the first characteristic parameter and the second characteristic parameter is calculated to obtain a third important metric of the segmented gene sequence to be encrypted.

9. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 1, characterized in that: The method for obtaining the second key includes: The segmented gene sequence to be encrypted whose importance metric is less than a preset importance threshold is used as a non-critical segmented sequence; for the suffix tree of the glioma gene sequence, the nodes corresponding to the non-critical segmented sequence are removed to obtain a simplified suffix tree; the segmented gene sequence to be encrypted is encrypted using a hash algorithm to obtain a hash value of the segmented gene sequence to be encrypted; the hash values ​​of the segmented gene sequence to be encrypted are sorted according to the simplified suffix tree structure to obtain the hash value of the glioma gene sequence, and the hash value is used as the second key of the glioma gene sequence.

10. The method for tamper-proof encryption of high-throughput sequencing results of glioma genes according to claim 1, characterized in that: The method for encrypting data of brain glioma gene sequence includes: Based on the AES algorithm, the glioma gene sequence is encrypted using a first key to generate a ciphertext; a digital signature is generated using a second key of the glioma gene sequence; and the ciphertext and the digital signature are used as encrypted data.

Citation Information

Patent Citations

  • Text privacy data dual encryption protection method, device and equipment

    CN116418481A

  • Encryption and tamper-proof transmission method for dialysis data

    CN118487750A

  • Gene sequence screening method and system for tumor markers

    CN119170100A

  • Gene data effective protection method and system based on computer encryption technology

    CN119227119A