Sequencing data processing method and system based on methylation analysis of target region

By obtaining and analyzing cfDNA sequencing data of target patients, combining methylation calculation rules and patient information, the methylation concern region was screened, which solved the problems of low methylation feature recognition accuracy and insufficient personalized screening in the prior art, and improved the specificity and biological correlation of methylation analysis.

CN120048351AInactive Publication Date: 2025-05-27SHENZHEN RAPHA BIOTECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510504370.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing methylation analysis methods have insufficient detection sensitivity and coverage, resulting in a decrease in the recognition accuracy of methylation characteristics, and fail to fully combine the specific information of the target patient, so personalized screening cannot be carried out.

Method used

By obtaining the cfDNA sequencing data of the target patient, methylation site information is determined based on the methylation calculation rules, and methylation region of concern is screened based on the patient information to finally determine the methylation characteristics.

Benefits of technology

Improves the specificity and biological relevance of methylation analysis, providing a reliable data basis for disease prediction or risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048351A_ABST
    Figure CN120048351A_ABST
Patent Text Reader

Abstract

The invention discloses a sequencing data processing method and system based on methylation analysis of a target region. The method comprises the following steps: acquiring cfDNA sequencing data of a target patient to be processed; based on a methylation calculation rule, methylation point location information in the sequencing data is determined; based on patient information of the target patient, determining a methylation region of interest of the sequencing data; and according to the methylation point location information and the methylation interest area, determining methylation characteristics corresponding to the sequencing data. Therefore, the specificity and biological correlation of methylation analysis can be improved, and a reliable data basis is provided for subsequent disease prediction or risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a sequencing data processing method and system based on target region methylation analysis. Background Art

[0002] Circulating DNA (cfDNA) is free DNA fragments existing in body fluids such as blood and urine. It usually originates from normal cells, cancer cells or fetuses in the body. It is non-invasive and has a wide range of applications in tumor treatment, prenatal diagnosis and organ transplantation. Existing methylation detection technologies mainly rely on high-throughput sequencing technology to analyze its methylation pattern by sequencing cfDNA. These methods have made significant progress in detection sensitivity and coverage, making cfDNA-based methylation analysis gradually become an important tool for early disease screening and risk assessment. Current methylation analysis methods usually directly extract methylation information from sequencing data and perform statistical analysis. However, due to the complex source of cfDNA fragments, sequencing data often contains a lot of background noise and individual differences, which may lead to reduced recognition accuracy of methylation features. In addition, existing methods usually do not fully combine the specific information of the target patients, and cannot perform personalized screening of methylation focus areas based on the biological characteristics of different individuals, resulting in some key methylation features not being effectively extracted. These deficiencies affect the specificity and biological relevance of methylation analysis, thereby limiting its application value in disease detection and risk assessment. It can be seen that the existing technology has defects that need to be solved urgently. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a sequencing data processing method and system based on target region methylation analysis, which can improve the specificity and biological relevance of methylation analysis and provide a reliable data basis for subsequent disease prediction or risk assessment.

[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a sequencing data processing method based on target region methylation analysis, the method comprising: Obtaining sequencing data of cfDNA of a target patient to be processed; Based on the methylation calculation rules, determining the methylation site information in the sequencing data; Based on the patient information of the target patient, determining the methylation focus region of the sequencing data; The methylation features corresponding to the sequencing data are determined according to the methylation site information and the methylation focus region.

[0005] As an optional embodiment, in the first aspect of the present invention, the methylation site information includes at least one of chromosome information, chromosome site information, methylation read information, unmethylated read information and methylation rate information.

[0006] As an optional embodiment, in the first aspect of the present invention, determining the methylation site information in the sequencing data based on the methylation calculation rule includes: Performing a quality assessment on each data portion of the sequencing data to obtain an assessment parameter corresponding to each data portion; Determine the data portion whose evaluation parameter is greater than a first parameter threshold as data to be calculated; Based on the MethyDackel software, the methylation sites of the data to be calculated are identified to obtain the methylation site information.

[0007] As an optional embodiment, in the first aspect of the present invention, performing quality assessment on each data portion of the sequencing data to obtain an assessment parameter corresponding to each data portion includes: Each data portion of the sequencing data is input into a trained quality assessment neural network to obtain an assessment parameter corresponding to each data portion; the quality assessment neural network is trained by a training data set including a plurality of training sequencing data and corresponding quality assessment annotations.

[0008] As an optional embodiment, in the first aspect of the present invention, the patient information includes at least one of the patient's medical history, the patient's treatment history, the patient's medication history, the patient's physiological characteristics, and the patient's real-time sensor data.

[0009] As an optional embodiment, in the first aspect of the present invention, determining the methylation focus region of the sequencing data based on the patient information of the target patient includes: For each historical patient in the historical patient database, calculating a first similarity between the patient information of the historical patient and the patient information of the target patient; Calculating a second similarity between the methylation feature data of the historical patient and the methylation point information; Calculate a weighted average of the first similarity and the second similarity to obtain a similarity parameter of the historical patient; Screening out all the historical patients whose similarity parameters are greater than a second parameter threshold to obtain a plurality of similar patients; The methylation focus region of the sequencing data is determined based on the methylation feature data of all the similar patients.

[0010] As an optional embodiment, in the first aspect of the present invention, determining the methylation focus region of the sequencing data according to the methylation feature data of all the similar patients includes: Determine the methylation rate of each region in the methylation feature data of each of the similar patients, and determine the region where the methylation rate is greater than a third parameter threshold as an abnormal region; The intersection of the abnormal regions of all the similar patients is calculated to determine the methylation focus region of the sequencing data.

[0011] As an optional embodiment, in the first aspect of the present invention, determining the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region includes: In any of the methylation focus regions, according to the methylation site information, a regional feature corresponding to the methylation focus region is calculated; the regional feature includes at least one of a methylation rate, a methylation site continuity feature, a methylation site change feature, and a methylation site map feature; The regional features corresponding to at least two of the methylation focus regions are determined as the methylation features corresponding to the sequencing data, so as to be used for subsequent model training or disease prediction and evaluation of the target patient.

[0012] A second aspect of an embodiment of the present invention discloses a sequencing data processing system based on target region methylation analysis, the system comprising: An acquisition module, used to acquire sequencing data of cfDNA of a target patient to be processed; A first determination module, used to determine the methylation site information in the sequencing data based on the methylation calculation rule; A second determination module is used to determine the methylation focus region of the sequencing data based on the patient information of the target patient; The third determination module is used to determine the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region.

[0013] As an optional embodiment, in the second aspect of the present invention, the methylation site information includes at least one of chromosome information, chromosome site information, methylation read information, unmethylated read information and methylation rate information.

[0014] As an optional embodiment, in the second aspect of the present invention, the first determination module determines the specific manner of the methylation site information in the sequencing data based on the methylation calculation rule, including: Performing a quality assessment on each data portion of the sequencing data to obtain an assessment parameter corresponding to each data portion; Determine the data portion whose evaluation parameter is greater than a first parameter threshold as data to be calculated; Based on the MethyDackel software, the methylation sites of the data to be calculated are identified to obtain the methylation site information.

[0015] As an optional embodiment, in the second aspect of the present invention, the first determination module performs quality assessment on each data portion of the sequencing data to obtain a specific method of the assessment parameter corresponding to each data portion, including: Each data portion of the sequencing data is input into a trained quality assessment neural network to obtain an assessment parameter corresponding to each data portion; the quality assessment neural network is trained by a training data set including a plurality of training sequencing data and corresponding quality assessment annotations.

[0016] As an optional embodiment, in the second aspect of the present invention, the patient information includes at least one of the patient's medical history, the patient's treatment history, the patient's medication history, the patient's physiological characteristics, and the patient's real-time sensor data.

[0017] As an optional embodiment, in the second aspect of the present invention, the second determination module determines the specific manner of the methylation focus region of the sequencing data based on the patient information of the target patient, including: For each historical patient in the historical patient database, calculating a first similarity between the patient information of the historical patient and the patient information of the target patient; Calculating a second similarity between the methylation feature data of the historical patient and the methylation point information; Calculate a weighted average of the first similarity and the second similarity to obtain a similarity parameter of the historical patient; Screening out all the historical patients whose similarity parameters are greater than a second parameter threshold to obtain a plurality of similar patients; The methylation focus region of the sequencing data is determined based on the methylation feature data of all the similar patients.

[0018] As an optional embodiment, in the second aspect of the present invention, the second determination module determines the specific manner of the methylation focus region of the sequencing data according to the methylation feature data of all the similar patients, including: Determine the methylation rate of each region in the methylation feature data of each of the similar patients, and determine the region where the methylation rate is greater than a third parameter threshold as an abnormal region; The intersection of the abnormal regions of all the similar patients is calculated to determine the methylation focus region of the sequencing data.

[0019] As an optional embodiment, in the second aspect of the present invention, the third determination module determines the specific manner in which the methylation feature corresponding to the sequencing data is determined according to the methylation site information and the methylation focus region, including: In any of the methylation focus regions, according to the methylation site information, a regional feature corresponding to the methylation focus region is calculated; the regional feature includes at least one of a methylation rate, a methylation site continuity feature, a methylation site change feature, and a methylation site map feature; The regional features corresponding to at least two of the methylation focus regions are determined as the methylation features corresponding to the sequencing data, so as to be used for subsequent model training or disease prediction and evaluation of the target patient.

[0020] The third aspect of the present invention discloses another sequencing data processing system based on target region methylation analysis, the system comprising: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute part or all of the steps in the sequencing data processing method based on target region methylation analysis disclosed in the first aspect of the present invention.

[0021] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the sequencing data processing method based on target region methylation analysis disclosed in the first aspect of the present invention.

[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention extracts methylation site information in sequencing data based on methylation calculation rules, and screens methylation focus areas in combination with the patient information of the target patient to comprehensively determine the precise methylation characteristics, thereby improving the specificity and biological relevance of methylation analysis and providing a reliable data basis for subsequent disease prediction or risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1It is a flow chart of a sequencing data processing method based on target region methylation analysis disclosed in an embodiment of the present invention.

[0025] Figure 2 It is a structural schematic diagram of a sequencing data processing system based on target region methylation analysis disclosed in an embodiment of the present invention.

[0026] Figure 3 It is a schematic diagram of the structure of another sequencing data processing system based on target region methylation analysis disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or equipment.

[0029] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] The present invention discloses a sequencing data processing method and system based on target region methylation analysis, extracts methylation site information in sequencing data based on methylation calculation rules, and screens methylation focus regions in combination with patient information of target patients to comprehensively determine accurate methylation features, thereby improving the specificity and biological relevance of methylation analysis and providing a reliable data basis for subsequent disease prediction or risk assessment. The following are detailed descriptions.

[0031] Embodiment 1 See also Figure 1 , Figure 1 : is a flow chart of a sequencing data processing method based on target region methylation analysis disclosed in an embodiment of the present invention. Figure 1 The described sequencing data processing method based on target region methylation analysis can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 1 As shown, the sequencing data processing method based on target region methylation analysis may include the following operations: 101. Obtain sequencing data of cfDNA of the target patient to be processed.

[0032] 102. Based on the methylation calculation rules, determine the methylation site information in the sequencing data. 103. Based on the patient information of the target patient, determine the methylation focus areas of the sequencing data. 104. Determine the methylation features corresponding to the sequencing data based on the methylation site information and the methylation focus area.

[0033] It can be seen that the above-mentioned embodiment of the invention extracts methylation site information in sequencing data based on methylation calculation rules, and screens methylation focus areas in combination with the patient information of the target patient to comprehensively determine the precise methylation characteristics, thereby improving the specificity and biological relevance of methylation analysis, and providing a reliable data basis for subsequent disease prediction or risk assessment.

[0034] As an optional embodiment, in the above steps, the methylation site information includes at least one of chromosome information, chromosome site information, methylation read information, unmethylated read information and methylation rate information.

[0035] It can be seen that through the above optional embodiments, the content of the methylation site information is limited to comprehensively characterize the methylation-related features of the sequencing data, so as to facilitate subsequent accurate feature calculations, assist in improving the specificity and biological relevance of methylation analysis, and provide a reliable data basis for subsequent disease prediction or risk assessment.

[0036] As an optional embodiment, in the above step, determining the methylation site information in the sequencing data based on the methylation calculation rule includes: Performing quality assessment on each data portion of the sequencing data to obtain assessment parameters corresponding to each data portion; Determine the data portion whose evaluation parameter is greater than the first parameter threshold as the data to be calculated; Based on MethyDackel software, methylation site identification is performed on the data to be calculated to obtain methylation site information.

[0037] It can be seen that through the above optional embodiments, the quality of each data part of the sequencing data is evaluated, and the data to be calculated that meet the quality requirements are screened out based on the first parameter threshold, and then the MethyDackel software is called to perform methylation site recognition operations on the screened data, thereby effectively eliminating the interference caused by low-quality sequencing fragments, improving the accuracy of methylation site recognition and the reliability of data analysis, and providing higher quality basic data support for subsequent methylation feature extraction and disease association analysis.

[0038] As an optional embodiment, in the above steps, performing quality assessment on each data portion of the sequencing data to obtain assessment parameters corresponding to each data portion includes: Each data portion of the sequencing data is input into a trained quality assessment neural network to obtain an assessment parameter corresponding to each data portion; the quality assessment neural network is trained by a training data set including a plurality of training sequencing data and corresponding quality assessment annotations.

[0039] It can be seen that through the above optional embodiments, each data part of the sequencing data is input into the trained quality assessment neural network to output the corresponding assessment parameters. The neural network is trained by a large amount of training sequencing data and its corresponding quality assessment annotations, and can accurately reflect the quality level of the sequencing data according to its characteristics, thereby improving the intelligence and accuracy of the quality assessment, and providing efficient and reliable quality judgment basis for subsequent data screening and methylation analysis.

[0040] As an optional embodiment, in the above steps, the patient information includes at least one of the patient's medical history, the patient's treatment history, the patient's medication history, the patient's physiological characteristics and the patient's real-time sensor data.

[0041] It can be seen that through the above optional embodiments, the content of the patient information is limited to comprehensively characterize the relevant characteristics of the patient, so as to facilitate the subsequent accurate region determination and feature calculation, assist in improving the specificity and biological relevance of methylation analysis, and provide a reliable data basis for subsequent disease prediction or risk assessment.

[0042] As an optional embodiment, in the above step, determining the methylation focus region of the sequencing data based on the patient information of the target patient includes: For each historical patient in the historical patient database, calculating a first similarity between the patient information of the historical patient and the patient information of the target patient; Calculating a second similarity between the methylation feature data and the methylation point information of the historical patient; Calculate the weighted average of the first similarity and the second similarity to obtain a similarity parameter of the historical patient; Screen out all historical patients whose similarity parameters are greater than the second parameter threshold to obtain multiple similar patients; Based on the methylation signature data of all similar patients, the methylation regions of interest in the sequencing data were determined.

[0043] It can be seen that through the above optional embodiments, for each historical patient in the historical patient database, the first similarity between its patient information and the target patient information, as well as the second similarity between its methylation feature data and the methylation point information in the target sequencing data are calculated respectively, and the similarity parameter is determined based on the weighted sum average value, and then similar patients with similarity parameters higher than the set threshold are screened out, and finally the methylation focus area of ​​the target sequencing data is determined in combination with the methylation feature data of these similar patients, thereby realizing personalized regional screening based on individual characteristics and molecular characteristics of the population, effectively improving the pertinence and biological relevance of methylation analysis, and laying a more reliable foundation for subsequent accurate diagnosis or disease risk prediction.

[0044] As an optional embodiment, in the above step, determining the methylation focus region of the sequencing data according to the methylation feature data of all similar patients includes: Determine the methylation rate of each region in the methylation feature data of each similar patient, and determine the region whose methylation rate is greater than the third parameter threshold as an abnormal region; The intersection of abnormal regions of all similar patients was calculated to determine the methylation focus regions of the sequencing data.

[0045] It can be seen that through the above optional embodiments, the methylation rate is calculated region by region for the methylation feature data of each similar patient, and the regions with methylation rates exceeding the set threshold are screened out as abnormal regions. Then, the intersection operation is performed on the abnormal regions of all similar patients, and finally the methylation focus regions of the target sequencing data are determined, thereby ensuring the significance of the methylation characteristics while enhancing the stability of the analysis and the ability to identify commonalities, effectively improving the accuracy and reliability of the screening of methylation focus regions, and providing more biologically meaningful candidate regions for subsequent disease correlation analysis.

[0046] As an optional embodiment, in the above step, determining the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region includes: In any methylation focus region, according to the methylation site information, the regional feature corresponding to the methylation focus region is calculated; optionally, the regional feature includes at least one of the methylation rate, the methylation site continuity feature, the methylation site change feature and the methylation site map feature; The regional features corresponding to at least two methylation focus regions are determined as methylation features corresponding to the sequencing data, so as to be used for subsequent model training or disease prediction and evaluation of target patients.

[0047] It can be seen that through the above optional embodiments, in any methylation focus area, multidimensional regional features including methylation rate, continuity, change trend and graphic structure are extracted based on methylation point information, and regional features corresponding to multiple focus areas are integrated to construct a complete methylation feature, which can fully explore the regional methylation patterns in the sequencing data, improve the richness and discrimination of feature expression, and provide a more accurate and effective data basis for subsequent model training and disease prediction and evaluation of target patients.

[0048] Embodiment 2 See also Figure 2 , Figure 2 : is a schematic diagram of a sequencing data processing system based on target region methylation analysis disclosed in an embodiment of the present invention. Figure 2 The described sequencing data processing system based on target region methylation analysis can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 2 As shown, the sequencing data processing system based on target region methylation analysis may include: The acquisition module 201 is used to obtain the sequencing data of the cfDNA of the target patient to be processed.

[0049] The first determination module 202 is used to determine the methylation site information in the sequencing data based on the methylation calculation rule. The second determination module 203 is used to determine the methylation focus region of the sequencing data based on the patient information of the target patient. The third determination module 204 is used to determine the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region.

[0050] It can be seen that the above-mentioned embodiment of the invention extracts methylation site information in sequencing data based on methylation calculation rules, and screens methylation focus areas in combination with the patient information of the target patient to comprehensively determine the precise methylation characteristics, thereby improving the specificity and biological relevance of methylation analysis, and providing a reliable data basis for subsequent disease prediction or risk assessment.

[0051] As an optional embodiment, the methylation site information includes at least one of chromosome information, chromosome site information, methylation read information, unmethylated read information and methylation rate information.

[0052] It can be seen that through the above optional embodiments, the content of the methylation site information is limited to comprehensively characterize the methylation-related features of the sequencing data, so as to facilitate subsequent accurate feature calculations, assist in improving the specificity and biological relevance of methylation analysis, and provide a reliable data basis for subsequent disease prediction or risk assessment.

[0053] As an optional embodiment, the first determination module determines the specific manner of the methylation site information in the sequencing data based on the methylation calculation rule, including: Performing quality assessment on each data portion of the sequencing data to obtain assessment parameters corresponding to each data portion; Determine the data portion whose evaluation parameter is greater than the first parameter threshold as the data to be calculated; Based on MethyDackel software, methylation site identification is performed on the data to be calculated to obtain methylation site information.

[0054] It can be seen that through the above optional embodiments, the quality of each data part of the sequencing data is evaluated, and the data to be calculated that meet the quality requirements are screened out based on the first parameter threshold, and then the MethyDackel software is called to perform methylation site recognition operations on the screened data, thereby effectively eliminating the interference caused by low-quality sequencing fragments, improving the accuracy of methylation site recognition and the reliability of data analysis, and providing higher quality basic data support for subsequent methylation feature extraction and disease association analysis.

[0055] As an optional embodiment, the first determination module performs quality assessment on each data portion of the sequencing data to obtain a specific method of the assessment parameter corresponding to each data portion, including: Each data portion of the sequencing data is input into a trained quality assessment neural network to obtain an assessment parameter corresponding to each data portion; the quality assessment neural network is trained by a training data set including a plurality of training sequencing data and corresponding quality assessment annotations.

[0056] It can be seen that through the above optional embodiments, each data part of the sequencing data is input into the trained quality assessment neural network to output the corresponding assessment parameters. The neural network is trained by a large amount of training sequencing data and its corresponding quality assessment annotations, and can accurately reflect the quality level of the sequencing data according to its characteristics, thereby improving the intelligence and accuracy of the quality assessment, and providing efficient and reliable quality judgment basis for subsequent data screening and methylation analysis.

[0057] As an optional embodiment, the patient information includes at least one of the patient's medical history, the patient's treatment history, the patient's medication history, the patient's physiological characteristics, and the patient's real-time sensor data.

[0058] It can be seen that through the above optional embodiments, the content of the patient information is limited to comprehensively characterize the relevant characteristics of the patient, so as to facilitate the subsequent accurate region determination and feature calculation, assist in improving the specificity and biological relevance of methylation analysis, and provide a reliable data basis for subsequent disease prediction or risk assessment.

[0059] As an optional embodiment, the second determination module determines the specific manner of the methylation focus region of the sequencing data based on the patient information of the target patient, including: For each historical patient in the historical patient database, calculating a first similarity between the patient information of the historical patient and the patient information of the target patient; Calculating a second similarity between the methylation feature data and the methylation point information of the historical patient; Calculate the weighted average of the first similarity and the second similarity to obtain a similarity parameter of the historical patient; Screen out all historical patients whose similarity parameters are greater than the second parameter threshold to obtain multiple similar patients; Based on the methylation signature data of all similar patients, the methylation regions of interest in the sequencing data were determined.

[0060] It can be seen that through the above optional embodiments, for each historical patient in the historical patient database, the first similarity between its patient information and the target patient information, as well as the second similarity between its methylation feature data and the methylation point information in the target sequencing data are calculated respectively, and the similarity parameter is determined based on the weighted sum average value, and then similar patients with similarity parameters higher than the set threshold are screened out, and finally the methylation focus area of ​​the target sequencing data is determined in combination with the methylation feature data of these similar patients, thereby realizing personalized regional screening based on individual characteristics and molecular characteristics of the population, effectively improving the pertinence and biological relevance of methylation analysis, and laying a more reliable foundation for subsequent accurate diagnosis or disease risk prediction.

[0061] As an optional embodiment, the second determination module determines the specific manner of the methylation focus region of the sequencing data according to the methylation feature data of all similar patients, including: Determine the methylation rate of each region in the methylation feature data of each similar patient, and determine the region whose methylation rate is greater than the third parameter threshold as an abnormal region; The intersection of abnormal regions of all similar patients was calculated to determine the methylation focus regions of the sequencing data.

[0062] It can be seen that through the above optional embodiments, the methylation rate is calculated region by region for the methylation feature data of each similar patient, and the regions with methylation rates exceeding the set threshold are screened out as abnormal regions. Then, the intersection operation is performed on the abnormal regions of all similar patients, and finally the methylation focus regions of the target sequencing data are determined, thereby ensuring the significance of the methylation characteristics while enhancing the stability of the analysis and the ability to identify commonalities, effectively improving the accuracy and reliability of the screening of methylation focus regions, and providing more biologically meaningful candidate regions for subsequent disease correlation analysis.

[0063] As an optional embodiment, the third determination module determines the specific manner in which the methylation feature corresponding to the sequencing data is determined according to the methylation site information and the methylation focus region, including: In any methylation focus region, according to the methylation site information, the regional feature corresponding to the methylation focus region is calculated; optionally, the regional feature includes at least one of a methylation rate, a methylation site continuity feature, a methylation site change feature, and a methylation site map feature; The regional features corresponding to at least two methylation focus regions are determined as methylation features corresponding to the sequencing data, so as to be used for subsequent model training or disease prediction and evaluation of target patients.

[0064] It can be seen that through the above optional embodiments, in any methylation focus area, multidimensional regional features including methylation rate, continuity, change trend and graphic structure are extracted based on methylation point information, and regional features corresponding to multiple focus areas are integrated to construct a complete methylation feature, which can fully explore the regional methylation patterns in the sequencing data, improve the richness and discrimination of feature expression, and provide a more accurate and effective data basis for subsequent model training and disease prediction and evaluation of target patients.

[0065] Embodiment 3 See also Figure 3 , Figure 3 It is another sequencing data processing system based on target region methylation analysis disclosed in an embodiment of the present invention. Figure 3 The described sequencing data processing system based on target region methylation analysis is applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 3 As shown, the sequencing data processing system based on target region methylation analysis may include: A memory 301 storing executable program codes; a processor 302 coupled to the memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the sequencing data processing method based on target region methylation analysis described in the first embodiment.

[0066] Embodiment 4 An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the sequencing data processing method based on target region methylation analysis described in the first embodiment.

[0067] Embodiment 5 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the sequencing data processing method based on target region methylation analysis described in Example 1.

[0068] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0069] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0070] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0071] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may be in the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of this specification may be in the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0072] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0073] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0076] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0077] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0078] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0079] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0080] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0081] Finally, it should be noted that the sequencing data processing method and system based on target region methylation analysis disclosed in the embodiment of the present invention only discloses the preferred embodiment of the present invention, which is only used to illustrate the technical scheme of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical schemes described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical schemes from the spirit and scope of the technical schemes of the embodiments of the present invention.

Claims

1. A sequencing data processing method based on target region methylation analysis, characterized in that: The method comprises: Obtaining sequencing data of cfDNA of a target patient to be processed; Based on the methylation calculation rules, determining the methylation site information in the sequencing data; Based on the patient information of the target patient, determining the methylation focus region of the sequencing data; The methylation features corresponding to the sequencing data are determined according to the methylation site information and the methylation focus region.

2. The method for processing sequencing data based on target region methylation analysis according to claim 1, characterized in that: The methylation site information includes at least one of chromosome information, chromosome site information, methylation read information, unmethylated read information and methylation rate information.

3. The sequencing data processing method based on target region methylation analysis according to claim 1, characterized in that: The step of determining the methylation site information in the sequencing data based on the methylation calculation rule includes: Performing a quality assessment on each data portion of the sequencing data to obtain an assessment parameter corresponding to each data portion; Determine the data portion whose evaluation parameter is greater than a first parameter threshold as data to be calculated; Based on the MethyDackel software, the methylation sites of the data to be calculated are identified to obtain the methylation site information.

4. The method for processing sequencing data based on target region methylation analysis according to claim 3, characterized in that: The step of performing quality assessment on each data portion of the sequencing data to obtain an assessment parameter corresponding to each data portion includes: Each data portion of the sequencing data is input into a trained quality assessment neural network to obtain an assessment parameter corresponding to each data portion; the quality assessment neural network is trained by a training data set including a plurality of training sequencing data and corresponding quality assessment annotations.

5. The method for processing sequencing data based on target region methylation analysis according to claim 1, characterized in that: The patient information includes at least one of the patient's medical history, the patient's treatment history, the patient's medication history, the patient's physiological characteristics, and the patient's real-time sensor data.

6. The method for processing sequencing data based on target region methylation analysis according to claim 1, characterized in that: The determining of the methylation focus region of the sequencing data based on the patient information of the target patient includes: For each historical patient in the historical patient database, calculating a first similarity between the patient information of the historical patient and the patient information of the target patient; Calculating a second similarity between the methylation feature data of the historical patient and the methylation point information; Calculate a weighted average of the first similarity and the second similarity to obtain a similarity parameter of the historical patient; Screening out all the historical patients whose similarity parameters are greater than a second parameter threshold to obtain a plurality of similar patients; The methylation focus region of the sequencing data is determined based on the methylation feature data of all the similar patients.

7. The method for processing sequencing data based on target region methylation analysis according to claim 6, characterized in that: Determining the methylation focus region of the sequencing data according to the methylation feature data of all the similar patients includes: Determine the methylation rate of each region in the methylation feature data of each of the similar patients, and determine the region where the methylation rate is greater than a third parameter threshold as an abnormal region; The intersection of the abnormal regions of all the similar patients is calculated to determine the methylation focus region of the sequencing data.

8. The method for processing sequencing data based on target region methylation analysis according to claim 1, characterized in that: Determining the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region includes: In any of the methylation focus regions, according to the methylation site information, a regional feature corresponding to the methylation focus region is calculated; the regional feature includes at least one of a methylation rate, a methylation site continuity feature, a methylation site change feature, and a methylation site map feature; The regional features corresponding to at least two of the methylation focus regions are determined as the methylation features corresponding to the sequencing data, so as to be used for subsequent model training or disease prediction and evaluation of the target patient.

9. A sequencing data processing system based on target region methylation analysis, characterized in that: The system comprises: An acquisition module, used to acquire sequencing data of cfDNA of a target patient to be processed; A first determination module, used to determine the methylation site information in the sequencing data based on the methylation calculation rule; A second determination module is used to determine the methylation focus region of the sequencing data based on the patient information of the target patient; The third determination module is used to determine the methylation features corresponding to the sequencing data according to the methylation site information and the methylation focus region.

10. A sequencing data processing system based on target region methylation analysis, characterized in that: The system comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the sequencing data processing method based on target region methylation analysis according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Differential methylation region screening method and device

    CN114171115A

  • DNA methylation level spectrum prediction method and system combined with pathological phenotypic characteristics

    CN116246702A

  • Enhancement of cancer screening using cell-free viral nucleic acids

    US20190032145A1