Method, device and equipment for training model, medium and product

By selecting appropriate marker regions and combining computational methods, and training the model using sample data, the problem of poor disease identification accuracy in existing technologies has been solved, achieving higher disease identification accuracy and sensitivity.

CN121279488APending Publication Date: 2026-01-06SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511437214.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing technologies, when using a uniform method for calculating methylation signals, ignore the differences between different biomarker regions, resulting in poor accuracy of disease identification models.

Method used

By selecting the most suitable combination from multiple biomarker regions and calculation methods, and training the model using sample data, the accuracy and sensitivity of disease identification can be improved.

Benefits of technology

This improved the sensitivity and accuracy of the disease identification model, enhancing its ability to identify diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121279488A_ABST
    Figure CN121279488A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method, device and equipment for training a model, a medium and a product. The method comprises the following steps: selecting at least one marker region and at least one calculation method from a plurality of marker regions and a plurality of first calculation methods based on second sample data corresponding to the plurality of marker regions and the plurality of first calculation methods in first sample data, wherein the plurality of marker regions correspond to a first plurality of calculation methods, the calculation methods are used to distinguish between a disease and a non-disease, and the first sample data includes sample data for the disease and the non-disease. The method further includes training a model for identifying a disease based on third sample data of the first sample data corresponding to the at least one marker region and the at least one computing method. Through the method, the model is trained by using the marker area and the corresponding calculation method, so that the accuracy of identifying diseases by the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure generally relate to the field of machine learning for disease identification, and more specifically to methods, apparatus, devices, media, and program products for training models. Background Technology

[0002] With technological advancements, machine learning is developing at an increasingly rapid pace. Currently, machine learning plays a significant role in daily life and work, and can be applied across various industries. For example, it can be used to recommend products on e-commerce platforms and songs on music platforms. Furthermore, machine learning can be used in natural language processing, image recognition, and autonomous driving. In addition, machine learning can be applied to assist in disease diagnosis and drug development.

[0003] Currently, machine learning technology is being applied more and more widely in the healthcare field. As a crucial technology in healthcare, machine learning has achieved significant results in multiple application scenarios. Through in-depth data mining and continuous model optimization, machine learning technology has played a vital role in disease prediction, diagnosis, and treatment, driving the development of healthcare. Therefore, given the current rapid development of artificial intelligence, research on machine learning technology in the field of early disease screening is becoming increasingly important. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, device, medium, and program product for training a model.

[0005] According to a first aspect of this disclosure, a method for training a model is provided. The method further includes selecting at least one marker region and at least one calculation method from the plurality of marker regions and the plurality of calculation methods based on second sample data corresponding to a plurality of marker regions and a plurality of calculation methods in first sample data, wherein the plurality of marker regions correspond to the plurality of calculation methods, and the calculation methods are used to distinguish between diseases and non-diseases, and the first sample data includes sample data for both diseases and non-diseases. The method further includes training a model for identifying diseases based on third sample data corresponding to at least one marker region and at least one calculation method in the first sample data.

[0006] In a second aspect of this disclosure, an apparatus for training a model is provided. The apparatus includes a data selection module configured to select at least one marker region and at least one calculation method from the plurality of marker regions and the plurality of calculation methods based on second sample data corresponding to a plurality of marker regions and a plurality of calculation methods in first sample data, wherein the plurality of marker regions correspond to the plurality of calculation methods, and the calculation methods are used to distinguish between diseases and non-diseases, the first sample data including sample data for diseases and non-diseases; and a model training module configured to train a model for identifying diseases based on third sample data corresponding to the at least one marker region and the at least one calculation method in the first sample data.

[0007] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a memory for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to a first aspect of this disclosure.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0010] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0012] Figure 1 The illustration shows a schematic diagram of an example environment in which some embodiments of the present disclosure may be implemented;

[0013] Figure 2 The illustration shows a schematic diagram of an example method for training a model according to some embodiments of the present disclosure;

[0014] Figure 3 The illustration shows a schematic diagram of an example architecture for training a model according to some embodiments of the present disclosure;

[0015] Figure 4 The illustration shows a schematic diagram of an example of a methylated region according to some embodiments of the present disclosure;

[0016] Figure 5 The illustration shows a schematic diagram of an example of determining a verification set according to some embodiments of the present disclosure;

[0017] Figure 6 The illustration shows an example of a bar chart illustrating model performance according to some embodiments of the present disclosure;

[0018] Figure 7 The illustration shows an example of a line graph illustrating model performance according to some embodiments of the present disclosure;

[0019] Figures 8A to 8B The illustration is a schematic diagram of an example of a bar chart showing the distribution of different marker regions in different cancer types according to some embodiments of the present disclosure.

[0020] Figure 9 The illustration shows a schematic block diagram of an apparatus for training a model according to some embodiments of the present disclosure;

[0021] Figure 10 A schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure is illustrated.

[0022] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] For example, when acquiring sample data related to the medical field, prompts can be provided to users to clearly indicate whether their data can be used. This allows users to autonomously choose whether to provide the corresponding data as sample data based on the prompts.

[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0029] With the rapid development of machine learning technology, it plays a significant role in daily life and work, and can be applied to various industries, especially in the healthcare field. For example, machine learning can be used to assist in disease prediction and diagnosis, drug development, and other areas. In this context, improving the accuracy of machine learning in predicting and diagnosing diseases has become an increasingly important research question.

[0030] In the field of early disease screening, when modeling using deoxyribonucleic acid (DNA) methylation signals, different biomarker regions often employ a single and consistent method for calculating methylation signals, such as uniformly using average methylation level, methylation haplotype load (MHL), or the number of aberrant fragments. However, different biomarker regions exhibit differences in noise levels / noise tolerance due to variations in cytosine-guanine dinucleotide (CpG) density, base composition, and background methylation levels in healthy individuals. For example, regions with high CpG density can use a more lenient threshold when statistically analyzing principal differential haplotypes (fully methylated / fully demethylated). Therefore, using a uniform calculation method ignores the differences between different biomarker regions, resulting in trained models that perform less accurately than expected in disease identification.

[0031] To address at least the aforementioned and other potential problems, embodiments of this disclosure propose a method for training a model. In this method, a computing device first selects at least one computational method corresponding to a marker region from a second plurality of computational methods for distinguishing between disease and non-disease in a sample dataset, based on first sample data for disease and non-disease. Then, a first plurality of computational methods are determined based on the at least one computational method. In one example, the disease described herein is cancer. In another example, the disease described herein can be any suitable symptom. After determining the first plurality of computational methods, the computing device further selects at least one marker region and at least one method from the plurality of marker regions and the first plurality of computational methods based on second sample data corresponding to the plurality of marker regions and the first plurality of computational methods in the first sample data. The plurality of marker regions correspond to the first plurality of computational methods, and the computational methods are used to distinguish between disease and non-disease. Finally, after determining the at least one marker region and at least one method, the computing device trains a model for disease identification based on third sample data corresponding to the at least one marker region and at least one method in the first sample data. This method utilizes sample data for both diseases and non-diseases to select appropriate marker regions and calculation methods from multiple marker regions. Furthermore, it selects at least one marker region and at least one method from multiple marker regions and multiple calculation methods based on second sample data from the first sample data. Finally, it combines third sample data from the first sample data to train a model capable of identifying diseases, thereby improving the sensitivity and accuracy of the model in identifying diseases.

[0032] It is understood that this disclosure is only described in conjunction with relevant embodiments and is not intended to limit the scope of protection of this disclosure. Any technical solutions in other disclosures that fall within the scope of protection of this disclosure should be protected.

[0033] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 The illustration shows an example environment in which the devices and / or methods of embodiments of this disclosure can be implemented. In environment 100, computing device 102 may be a dedicated processing device with data processing capabilities, such as a Graphics Processing Unit (GPU). Computing device 102 may also have the ability to store data. Computing device 102 may also have the ability to receive and send data. Computing device 102 can process local data or online data. Computing device 102 may also include the ability to convert non-textual information into textual information; for example, the computing device may run various suitable machine learning models.

[0034] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0035] like Figure 1 As shown, the computing device 102 can first determine first sample data 106 from the sample dataset 104 for training a model to identify diseases. Next, based on the first sample data 106, at least one calculation method corresponding to marker region 116 among multiple marker regions 114 for DNA can be selected from a second set of multiple calculation methods 120. The first calculation method in the second set of multiple calculation methods 120 can be used to distinguish between diseases and non-diseases. The multiple marker regions 114 include multiple marker regions such as marker region 116 and marker region 118, and marker region 116 corresponds to the second set of multiple calculation methods 120, while marker region 118 corresponds to the second set of multiple calculation methods 122. Additionally, the second set of multiple calculation methods 120 and 122 can be the same second set of multiple calculation methods; for example, the calculation methods in the second set of multiple calculation methods 120 and 122 are completely identical.

[0036] The computing device 102 can use sample data to determine at least one calculation method 126 corresponding to the marker region 116 from a second plurality of calculation methods 120. For example, at least one calculation method 126 can be selected from the second plurality of calculation methods 120 based on the classification performance of the first calculation method in the second plurality of calculation methods 120 in distinguishing between diseased and non-diseased patients from the sample data, such as selecting the calculation method with the best or highest classification performance. Similarly, for other marker regions, corresponding calculation methods can be selected from the corresponding second plurality of calculation methods. For example, for marker region 118, at least one calculation method 128 can be selected from the second plurality of calculation methods 122. After determining at least one calculation method 126 for marker region 116 and at least one calculation method 128 for marker region 118, corresponding to the plurality of marker regions 114, the computing device 102 further selects at least one marker region and at least one method 130 from the plurality of marker regions 114 and the first plurality of calculation methods 124 based on second sample data 108 in the first sample data 106. The second sample data 108 includes values ​​calculated by the corresponding calculation method for the corresponding marker region in the first sample of the first sample data 106.

[0037] Finally, the computing device 102 can train the model 112 based on the third sample data 110 in the first sample data 106, which corresponds to at least one marker region and at least one method 130. For example, for the first marker region in the at least one marker region, the value calculated by the corresponding computing method for the first sample in the sample data in this region can be determined. Additionally, identification information as to whether the first sample is a patient with a disease can also be obtained. The trained model 112 can then be used to identify diseases.

[0038] Additionally, above Figure 1 The disease described is the first disease, and the sample data in sample dataset 104 is sample data for the first disease. Similarly, sample dataset 104 may also include sample data for other diseases, based on... Figure 1 The method for obtaining the third sample data 110 is described in the same way as obtaining the fourth sample data for other diseases. Next, a fifth sample data corresponding to at least one second biomarker region and at least one second calculation method is selected from the fourth sample data. Then, the model is trained based on the third sample data 110 and the fifth sample data. If the third sample data 110 and the fifth sample data have overlapping biomarker regions and methods, the overlapping biomarker regions and methods can be merged. The first calculation method corresponding to biomarker region 116 can also form a corresponding identifier for binding the corresponding biomarker region 116 and the first calculation method.

[0039] This method utilizes sample data for both diseases and non-diseases to select appropriate marker regions and calculation methods from multiple marker regions. Furthermore, it selects at least one marker region and at least one method from multiple marker regions and multiple calculation methods based on second sample data from the first sample data. Finally, it combines third sample data from the first sample data to train a model capable of identifying diseases, thereby improving the sensitivity and accuracy of the model in identifying diseases.

[0040] The above combination Figure 1 The following is a schematic diagram illustrating an example environment in which some embodiments of this disclosure may be implemented, in conjunction with... Figure 2 A schematic diagram illustrating an example method for training a model according to some embodiments of the present disclosure. This example method can be provided by... Figure 1 The computing device 102 or any suitable computing device in the system shall execute the test.

[0041] like Figure 2As shown, in example method 200, at block 202, based on second sample data corresponding to multiple marker regions and multiple calculation methods in the first sample data, at least one marker region and at least one calculation method are selected from the multiple marker regions and the multiple calculation methods, wherein the multiple marker regions correspond to the multiple calculation methods, and the calculation methods are used to distinguish between diseases and non-diseases. The first sample data includes sample data for diseases and non-diseases. Furthermore, the first calculation method in the first multiple calculation methods 124 can be used to distinguish between diseases and non-diseases. The first sample data 106 can be sample data for a specific disease, such as sample data for a specific type of cancer. For example, the first sample data 106 can be sample data for colorectal cancer (CRC), or it can be a sample dataset for other types of cancer; this application does not impose any limitations here. Additionally, the sample dataset 104 can include sample data for multiple diseases.

[0042] In some embodiments, the first sample in the first sample data 106 includes a first value for a combination of a marker region 116 among a plurality of marker regions and a first calculation method among a plurality of calculation methods 120. The first value is calculated by applying the corresponding first calculation method to the corresponding marker region 116. Additionally, in determining the first plurality of calculation methods 124, the computing device 102 first determines a plurality of values ​​corresponding to the plurality of samples in the first sample data 106. These plurality of values ​​are values ​​for the plurality of samples for the plurality of marker regions and the plurality of calculation methods, and each of the plurality of values ​​in the first sample data 106 includes a value for any combination of the plurality of marker regions among the plurality of marker regions and the plurality of calculation methods among the plurality of calculation methods. Alternatively, for a marker region in a sample of the sample data, if the region is not covered by a segment, there may be no value corresponding to the marker region and the corresponding method. In this case, adjustments can be made in a subsequent process, for example, using the mean of non-disease samples in the training set under the marker region and the corresponding method as the value of the sample for the marker region and the corresponding method. It should be noted that non-disease samples can be used to indicate samples that do not correspond to any existing known disease type, or samples that do not correspond to a specific disease type (such as CRC cancer).

[0043] In some embodiments, the first sample data 106 in the sample dataset 104 may include disease samples or non-disease samples, and disease samples and non-disease samples can be distinguished by marker regions of markers in DNA. In one example, disease samples and non-disease samples can be distinguished by calculating features of marker regions (e.g., calculating the value of methylation features of marker regions) using calculation methods in a second plurality of calculation methods 120 or 122.

[0044] In some embodiments, the computing device 102 can, for a marker region 116 among multiple marker regions, determine a first value calculated using a first calculation method among a second plurality of calculation methods 120 in a first sample from multiple samples in the first sample data 106, and then determine the classification performance of the first calculation method for distinguishing between diseases and non-diseases based on the first value. For example, for the first calculation method, the area under the ROC curve (AUC) for distinguishing between diseases and non-diseases is determined using the value obtained using the calculation method in the sample and the sample label. This area under the ROC curve can be referred to as the first classification performance corresponding to the first calculation method. In one example, for the marker region, after determining the value calculated by the first calculation method among the second plurality of calculation methods 120, the classification performance of the first calculation method for distinguishing between diseases and non-diseases can be determined based on the size of the area under the ROC curve. Generally, the value range of the area under the ROC curve is [0.5, 1]. The closer the value of the area under the ROC curve is to 1, the better the classification performance; conversely, the closer the value of the area under the ROC curve is to 0.5, the worse the classification performance.

[0045] In some embodiments, after determining the first value of the first calculation method in the second plurality of calculation methods 120 corresponding to the marker region 116 corresponding to the plurality of samples, the computing device 102 may further determine a plurality of values ​​corresponding to the plurality of samples, thereby further determining a plurality of classification performances corresponding to the plurality of calculation methods, wherein the calculation methods in the plurality of calculation methods have corresponding classification performances, and the determined plurality of classification performances include the aforementioned first classification performance. Then, the computing device determines a target classification performance from the determined plurality of classification performances, for example, determining the highest classification performance as the target classification performance, and subsequently determining the calculation method corresponding to the target classification performance as the calculation method corresponding to the marker region 116. In one example, among the plurality of classification performances corresponding to the plurality of calculation methods in the second plurality of calculation methods 120, at least one calculation method 126 has the highest classification performance, and the computing device 102 may determine at least one calculation method 126 as the calculation method corresponding to the marker region 116. In another example, among the plurality of classification performances corresponding to the plurality of calculation methods in the second plurality of calculation methods 122, at least one calculation method 128 has the highest classification performance, and the computing device 102 may determine at least one calculation method 128 as the calculation method corresponding to the marker region 118. Therefore, a first plurality of calculation methods 124 can be determined for the plurality of marker regions 114.

[0046] In some embodiments, the computing device 102 first obtains second sample data 108 from the first sample data 106, corresponding to a plurality of marker regions 114 and corresponding first plurality of calculation methods 124. For example, each row in the sample data matrix represents a sample, and each column corresponds to a marker region and a calculation method. Therefore, a column identifier can be formed by a marker region and a calculation method, and the data in that column is a value generated by the calculation method based on the information in the marker region. Thus, for N marker regions and M methods, N... M columns. For the column identifier formed by the marker regions and corresponding calculation methods in the multiple marker regions 114, a corresponding column can be found from the sample data matrix. Therefore, multiple columns corresponding to the multiple marker regions can be found from the sample data matrix to form the second sample data 108. After obtaining the second sample data 108, from the first multiple calculation methods 124 corresponding to the multiple marker regions 114, a first importance score for distinguishing between disease and non-disease is determined for the first marker region and its corresponding first calculation method. For example, the importance score of calculation method 126 corresponding to marker region 116 is determined. For example, the second sample data 108 is modeled according to the Boruta algorithm to determine multiple importance scores corresponding to the first multiple calculation methods 124 corresponding to the multiple marker regions 114. Then, based on the determined multiple importance scores, at least one marker region and at least one method 130 are selected from the multiple marker regions 114 and the first multiple calculation methods 124. The first marker region and the first calculation method can form a marker region and method pair. Therefore, a predetermined number of marker regions and method pairs can be selected from multiple pairs of marker regions and methods to complete feature filtering and dimensionality reduction. The predetermined number can be set according to requirements, and this application does not impose any restrictions on it.

[0047] In some embodiments, based on multiple importance scores determined by multiple marker regions 114 and a first plurality of calculation methods 124, marker regions and calculation methods with importance scores greater than or equal to a threshold score are selected as at least one marker region and at least one calculation method 130. For example, if the threshold score is 0.5, the importance score of marker region 116 and calculation method 126 is 0.9, and the importance score of marker region 118 and calculation method 128 is 0.8, then marker region 116 and calculation method 126, as well as marker region 118 and calculation method 128, can be determined as at least one marker region and at least one calculation method 130. The threshold score values ​​are merely examples and can be set according to actual needs; this application does not impose limitations on them.

[0048] In some embodiments, the determined importance scores determined by the plurality of marker regions 114 and the first plurality of calculation methods 124 can be sorted in descending order, and a specified number of marker regions and calculation methods can be selected (e.g., the top 200 marker regions and calculation method pairs).

[0049] Finally, at box 204, a model for disease identification is trained based on third sample data corresponding to the at least one marker region and the at least one calculation method in the first sample data. In some embodiments, after the computing device 102 determines the at least one marker region and the at least one calculation method 130, the determined at least one marker region and at least one calculation method 130, and third sample data 110 corresponding to the at least one marker region and at least one calculation method 130 in the sample data, can be used to train model 112. The process of acquiring the third sample data 110 is similar to the process of acquiring the second sample data 108, which is to query the data corresponding to the column identifier formed by the marker region and the calculation method from the sample data. The trained model 112 can be used to identify diseases, for example, it can be used to identify CRC disease.

[0050] In some embodiments, the sample dataset 104 may include sample data for multiple types of diseases, wherein the disease corresponding to the first sample data 106 is a first disease, and marker regions such as marker region 116 and marker region 118 are first marker regions. Furthermore, the sample dataset 104 may also include second target sample data for a second disease.

[0051] In some embodiments, the computing device 102 can also select at least one second marker region and at least one second calculation method for distinguishing the second disease from a plurality of marker regions 114 and a first plurality of calculation methods 124 for the plurality of marker regions 114, based on fourth sample data for the second disease. This acquisition of at least one second marker region and at least one second calculation method is similar to the acquisition of at least one marker region and at least one calculation method corresponding to the first disease described above. Then, the computing device can determine a fifth sample data corresponding to at least one second marker and at least one second calculation method from the fourth sample data, and finally train model 112 using the third and fifth sample data. The trained model 112 has the ability to identify the first disease and the second disease. Additionally, data for other diseases can also be input to train the model. Therefore, the model can be a model for identifying multiple diseases. It should be noted that model 112 can be a neural network model, model 112 can be a gradient boosting model based on a decision tree, or model 112 can be other types of models capable of implementing the scheme of this application; this application does not impose any limitations here.

[0052] Additionally, the trained model 112 can be validated. The computing device 102 first determines a validation dataset based on the sample dataset. This determination can be achieved by splitting the sample dataset to identify the validation dataset for each sample dataset. The validated dataset is then used to validate the performance of model 112. Furthermore, the splitting operation on the sample dataset 104 can include at least one of the following: splitting based on an independent sampling center, splitting based on sampling time, splitting based on experimental batches, or random splitting. Splitting based on sampling time can be based on the timing of the sampling, and splitting based on experimental batches can be based on the different experimental batches. Additionally, when splitting the sample dataset 104, the sampling centers are distinguished only during the splitting process for independent sampling centers; they are not distinguished during other types of splitting. It should be noted that the validation dataset includes sampling centers and disease types. Sampling centers can be different hospitals, and disease types can include known diseases such as CRC or hepatocellular carcinoma (HCC).

[0053] Additionally, the sampling center, sampling time, or experimental batch in the validation dataset determined from the sample dataset 104 does not appear in the corresponding training set, and the ratio of the validation dataset to the training dataset in the sample dataset 104 can be adjusted according to user needs, for example, the sample dataset 104 can be used to determine the validation dataset and the training dataset in a 1:4 ratio.

[0054] In some embodiments, the computing device 102 can perform a predetermined number of splitting operations on the sample dataset to obtain corresponding validation datasets. For example, when it is necessary to split the sample dataset to obtain 5 validation datasets, 5 splitting operations can be performed on these sample datasets. The predetermined multiple splitting operation yields multiple candidate validation datasets; this predetermined multiple can be defined by the user. In one example, when the predetermined multiple is 20, the sample dataset needs to be repeatedly split 5 times starting with different random seeds. 20 = 100 times, resulting in 100 candidate validation datasets. Then, computing device 102 determines 5 validation datasets from these 100 candidate validation datasets. It should be noted that there may be overlap in sampling centers and disease types among the 100 determined candidate validation datasets.

[0055] In some embodiments, when splitting to determine the candidate validation dataset, the computing device 102 may first acquire data from a sampling center and then determine whether the data from the sampling center covers all target diseases (e.g., if there are 15 target diseases). If not fully covered, a center may be randomly selected from other centers that include the missing diseases until all diseases are covered. Furthermore, it is further determined whether the number of samples for all target diseases in the candidate validation dataset is greater than a predetermined number (e.g., a predetermined number of 10). If the number of samples is less than the predetermined number, random sampling is performed from the remaining centers containing the disease until all diseases have at least the predetermined number of samples in the test set.

[0056] In some embodiments, after determining that the 100 split candidate validation datasets satisfy all predetermined conditions, a target candidate validation dataset (5 target candidate validation datasets) corresponding to a predetermined number of splits can be determined based on the degree of mutual exclusion between the candidate validation datasets in the 100 candidate validation datasets. For example, the degree of mutual exclusion between candidate validation datasets can be determined based on the number of intersections of the sampling centers between the candidate validation datasets. For example, the smaller the number of intersections of the sampling centers, the greater the degree of mutual exclusion between the candidate validation datasets. In one example, the computing device 102 can first calculate the number of intersections of the sampling centers of any two candidate validation datasets in the 100 candidate validation datasets and determine the difference matrix for the number of intersections. Then, a candidate validation dataset R1 is randomly selected from the 100 datasets. Then, the candidate validation dataset R2 with the smallest number of intersections with R1 is selected from the remaining 99 candidate validation datasets. Subsequently, the candidate validation dataset R3 with the smallest maximum number of intersections with R1 and R2 is selected from the remaining 98 datasets. This step is repeated until the required 5 target candidate validation datasets are selected from the 100 candidate validation datasets. This method ensures that the sampling centers in the five selected target candidate validation datasets are as mutually exclusive as possible, thereby improving the reliability and accuracy of the validation model performance.

[0057] This method utilizes sample data for both diseases and non-diseases to select appropriate marker regions and calculation methods from multiple marker regions. Furthermore, it selects at least one marker region and at least one method from multiple marker regions and multiple calculation methods based on second sample data from the first sample data. Finally, it combines third sample data from the first sample data to train a model capable of identifying diseases, thereby improving the sensitivity and accuracy of the model in identifying diseases.

[0058] The above combination Figure 2 A schematic diagram illustrating example methods for training models according to some disclosed embodiments is described below. Figure 3A schematic diagram illustrating examples of architectures for training models according to some embodiments of the present disclosure.

[0059] like Figure 3 As shown in Example 300, during the calculation of features for marker regions, the representation method for calculating features is first optimized in box 302. That is, the most suitable calculation method is selected for different marker regions to calculate the methylation features used to identify diseases. For ease of understanding, marker regions can be referred to as regions. Among them, multiple marker regions include regions 1 (304), 2 (306), 3 (308), and n (310), etc. Region 1 is used as an example for illustration.

[0060] For region 1 304, the second plurality of calculation methods includes multiple calculation methods from calculation method 1 to calculation method m. The calculation device 102 first determines, at position 314, the calculation method 8 316 corresponding to region 1 304 from the second plurality of calculation methods 312, based on sample data for a single disease (such as CRC) in the sample data. Similarly, for multiple regions from region 2 306 to region n 310, corresponding calculation methods are selected. For ease of understanding, the following table can be used for explanation:

[0061] Here, X represents the input samples, with samples 1 to 1000 as an example. Of these 1000 samples, 500 are non-disease samples and 500 are disease samples (e.g., CRC samples). Each sample needs to be calculated in regions 1 304 to n 310 using calculation methods 1 to m to obtain multiple values ​​for all samples. Then, for multiple regions, the classification performance of multiple calculation methods for differentiating between diseases and non-diseases is determined based on the calculated values. For example, the classification performance of a single calculation method is determined using the area under the receiver operating characteristic curve (AUC) or the Boruta algorithm. Subsequently, multiple classification performances for multiple calculation methods are determined from the multiple values ​​for the 1000 samples. Finally, the calculation method corresponding to the highest classification performance can be determined. Thus, the computing device 102 can determine regions 1 304 and calculation methods 8 316 to n and calculation method m. Additionally, in determining the classification performance of a single computational method, other metrics besides AUC or other algorithms besides the Boruta algorithm may be used, and this application makes no limitation herein. Furthermore, the data in the tables above and other tables in this application are merely illustrative and are not intended to limit the number of samples.

[0062] Subsequently, global feature filtering is performed at box 318. At box 320, all regions and calculation methods are merged. Then, at box 322, the top-ranked features are selected. The computing device 102 can use the sample data to determine the importance scores for differentiating between disease and non-disease regions for multiple marker regions and methods. For example, the Boruta algorithm can be used for modeling. For ease of understanding, the following table can be used for explanation:

[0063] The top features can be selected based on their importance scores, for example, from highest to lowest. In one example, the top 800 features with the highest importance scores are selected according to their corresponding regions and calculation methods, thus identifying the top 800 calculated features for a single disease at position 324. The number of top features can be defined by the user.

[0064] Finally, the identified 800 features were modeled at position 328 to train a model that can be used to distinguish between diseases and non-diseases. Additionally, the leading features at position 326 corresponding to other diseases can be input into the model for training, resulting in a model capable of recognizing multiple disease types, thereby improving the accuracy and classification of disease identification.

[0065] The above combination Figure 3 A schematic diagram illustrating examples of architectures for training models according to some disclosed embodiments is shown below. Figure 4 A schematic diagram illustrating examples of methylated regions according to some embodiments of the present disclosure.

[0066] like Figure 4 As shown in Example 400, at 402, a region is shown to be relatively more methylated in diseased individuals compared to non-diseased individuals. In this marker region, individual CpG sites are methylated in the non-disease state, while the remaining CpG sites are unmethylated. At 404, the methylation status of CpG sites in a DNA methylation sequencing fragment covering this region from a sample is shown. Fragments included in the statistics are required to cover at least three CpG sites.

[0067] In one example, each fragment is categorized into anomalous and normal fragments based on the degree of methylation. For instance, in regions where disease patients have higher methylation levels than non-disease patients, fragments with a methylated CpG site ratio reaching a threshold t are considered anomalous. Conversely, in regions where disease patients have lower methylation levels than non-disease patients, fragments with an unmethylated CpG site ratio reaching a threshold t are considered anomalous. The number of anomalous fragments can be termed UMcount, where U represents unmethylated sites and M represents methylated sites. The total number of fragments is total_count, and the number of normal fragments (Xcount) is (total_count - UMcount). The threshold t can be set with multiple gradients, meaning different levels of strictness are used when defining anomalous fragments to tolerate different levels of noise (e.g., error rate per unit point). In calculating the methylation characteristics of marker regions, the number of anomalous fragments can be used directly, or the anomalous fragment ratio (UMratio) can be obtained by standardizing the number of anomalous fragments. The standardization calculation of the number of anomalous fragments can be performed using the following four formulas: (1) (2) (3) (4)

[0068] Methylation features are calculated using various methods, including but not limited to those mentioned above, targeting marker regions. This allows for a multi-faceted and multi-dimensional characterization of the marker regions, thereby improving the accuracy of the trained model in disease identification.

[0069] The above combination Figure 4 A schematic diagram illustrating examples of cancer-to-non-disease methylated regions according to some disclosed embodiments is shown below. Figure 5 The illustration depicts examples of determining a verification set according to some embodiments of the present disclosure. For ease of description, the diseases described herein may also be referred to as cancer types.

[0070] like Figure 5As shown in Example 500, during the process of validating model performance, it is necessary to first split the sample set into a challenging validation set at point 502. For example, the sample dataset can be split into an independent center validation set 504, a sampling time validation set 506, an experimental batch validation set 508, and a randomly split validation set 508, based on different sampling centers. Specifically, at point 512, the independent center validation set covers all target cancer types, with at least 10 samples per cancer type. The sampling centers in the validation set do not appear in the training set, and the sampling centers in the multiple split validation sets are as mutually exclusive as possible. At point 514, samples from the last 20% to 30% of the sampling time are used as the sampling validation set 506, and the sampling time in the validation set does not appear in the training set. At point 516, samples from the last 20% to 30% of the experimental batches are used as the experimental batch validation set 508, and the experimental batches in the validation set do not appear in the training set. At point 518, the training set and validation set are randomly split proportionally to generate a randomly split validation set 510. It should be noted that during the validation of model performance, any one or more of the following validation sets can be used: independent center validation set 504, sampling time validation set 506, experimental batch validation set 508, or random split validation set 510.

[0071] In determining the independent center validation set 504, the validation dataset obtained by splitting according to the independent sampling centers is determined as the validation dataset. The computing device 102 can also perform a predetermined number of splitting operations on the dataset, for example, splitting the sample dataset into 10 candidate validation datasets. The training set samples corresponding to multiple validation datasets in the 10 candidate validation datasets are the total sample dataset minus the samples of the validation datasets. The predetermined multiple can be defined by the user. In one example, when the predetermined multiple is 30, the dataset needs to be split repeatedly 300 times with different random seeds to obtain 300 candidate validation datasets. Then, the computing device 102 determines 10 datasets from these 300 candidate validation datasets. It should be noted that there can be overlap in the sampling centers and disease types among the determined 300 candidate validation datasets. When obtaining a single candidate validation dataset, the following steps can be taken: Randomly select a center and check whether the cancer types covered by that center cover all target cancer types. If not, randomly select a center from other centers that contain the missing cancer types until all cancer types are covered. Check whether the number of each cancer type in the validation set is less than 10. If it is less than 10, randomly sample the remaining centers that contain that cancer type until all cancer types have at least 10 samples in the test set.

[0072] After determining 300 candidate validation datasets using the above method, computing device 102 further identifies 10 target candidate validation datasets from the 300 candidate validation datasets.

[0073] After confirming that the 300 candidate validation datasets meet all predetermined conditions, the computing device 102 first calculates the number of intersections of the sampling centers of any two candidate validation datasets from the 300 datasets and determines the difference matrix for the number of intersections. Then, it randomly selects a candidate validation dataset R1 from the 300 datasets. Next, it selects the candidate validation dataset R2 with the smallest number of intersections with R1 from the remaining 299 datasets. Then, it further selects the candidate validation dataset R3 with the smallest maximum number of intersections with R1 and R2 from the remaining 298 datasets. This process is repeated until the desired 10 target candidate validation datasets are selected from the 300 datasets. This method ensures that the sampling centers in the selected 10 datasets are as mutually exclusive as possible, improving the reliability and accuracy of the validation model's performance and more accurately evaluating the model's generalization ability. For ease of understanding, the following table explains this:

[0074] After identifying candidate validation dataset R2 with the smallest number of intersections with candidate validation dataset R1, the remaining 298 candidate validation datasets are further selected from those with the smallest number of intersections with both R1 and R2. For example, if R3 has 4 intersections with R1, 2 with R2, 5 with R3, and 4 with R2, then the maximum intersections between R3 and R4 with R1 and R2 are 4 and 5 respectively. Therefore, when comparing only R3 and R4, R3 is selected as the next candidate validation dataset. This process continues from the remaining 297 candidate validation datasets until 10 target candidate validation datasets, including R1, R2, and R3, are chosen. This method ensures that the sampling centers in the selected target candidate validation datasets are as mutually exclusive as possible, resulting in better performance of the validation model.

[0075] This approach allows for the verification of model performance from multiple perspectives, enabling a more accurate assessment of the model's generalization ability.

[0076] The above combination Figure 5 A schematic diagram illustrating an example of determining a verification set according to some disclosed embodiments is described below. Figure 6 A schematic diagram illustrating an example of a histogram of model performance according to some embodiments of the present disclosure.

[0077] like Figure 6As shown in Example 600, combined with the previous Example 500, the ordinate of bar chart 602 represents the mean difference in sensitivity under 99 specificity across multiple sets of model parameters (Adaptive Feature Optimization Two-Layer Frame Algorithm (AFOA) - Traditional Method), and the ordinate of bar chart 604 represents the mean difference in area under the curve across multiple sets of model parameters (AFOA - Traditional Method). The x-axis of both bar charts 602 and 604 represents the validation dataset from rep1 to rep10. Bar charts 602 and 604 demonstrate that the model trained using AFOA shows a significant improvement in both sensitivity and AUC compared to the traditional model trained using a single computational method.

[0078] The above combination Figure 6 A schematic diagram illustrating an example of a histogram of model performance according to some disclosed embodiments is provided below. Figure 7 A schematic diagram illustrating an example of a line graph depicting model performance according to some embodiments of the present disclosure.

[0079] like Figure 7 In Example 702 of Example 700, lines 704, 708, 712, 716, 720, 724, 728, 732, 736, and 740 correspond to the performance of the model trained using AFOA in 10 splits, while lines 706, 710, 714, 718, 722, 726, 730, 734, 738, and 742 correspond to the performance of the model trained using traditional methods in 10 splits. The horizontal axis represents the model under different parameter settings. It can be seen that, when comparing aligned model parameters, the model trained using AFOA shows a systematic improvement in the Area Under the Curve (AUC) compared to the traditional method. This line graph indicates that the model trained using AFOA has improved disease identification performance compared to the model trained using traditional methods.

[0080] The above combination Figure 7 A schematic diagram illustrating an example of a model performance line graph according to some disclosed embodiments is provided below. Figures 8A to 8B A schematic diagram of an example of a bar chart illustrating the distribution of different marker regions in different cancer types according to some embodiments of the present disclosure.

[0081] like Figures 8A to 8BAs shown in bar charts 802 and 804 of Examples 800A and 800B, the horizontal axis of the bar chart represents the optimal method (e.g., calculation method 1 to calculation method 10) among at least one calculation method corresponding to the biomarker regions for different cancer types, and the vertical axis represents the number of biomarker regions. It can be seen that the optimal calculation methods for selecting different biomarker regions differ; and the distribution of the optimal calculation methods for biomarker region selection varies under different clinical states of different cancer types, such as early-stage CRC, late-stage CRC, early-stage HCC, and late-stage HCC.

[0082] like Figure 9 As shown, it illustrates a device 900 for training a model, which can be implemented in Figure 1 The device 900 includes a data selection module 902 configured to select at least one marker region and at least one calculation method from the plurality of marker regions and the plurality of calculation methods based on second sample data corresponding to the plurality of marker regions and the plurality of calculation methods in the first sample data, wherein the plurality of marker regions correspond to the plurality of calculation methods, and the calculation methods are used to distinguish between diseases and non-diseases, and the first sample data includes sample data for diseases and non-diseases; and a model training module 904 configured to train a model for identifying diseases based on third sample data corresponding to the at least one marker region and the at least one calculation method in the first sample data.

[0083] In some embodiments, the apparatus 900 further includes: a first plurality of calculation methods selection module configured to select a first plurality of calculation methods; and the first plurality of calculation methods selection module includes: a value determination module configured to determine a plurality of values ​​corresponding to a plurality of samples in the first sample data, the plurality of values ​​being values ​​of the plurality of samples for a plurality of marker regions and a second plurality of calculation methods; a classification performance determination module configured to determine, based on the plurality of values ​​corresponding to the plurality of samples, the classification performance of the second plurality of calculation methods for distinguishing between diseases and non-diseases in the plurality of marker regions; and a first calculation method determination module configured to determine, based on the classification performance, a first plurality of calculation methods corresponding to the plurality of marker regions from the second plurality of calculation methods.

[0084] In some embodiments, the first calculation method determination module includes: a target classification performance selection module configured to determine a target classification performance that meets predetermined conditions from the classification performance; and a second calculation method determination module configured to determine the calculation method corresponding to the target classification performance as the first plurality of calculation methods corresponding to the plurality of marker regions.

[0085] In some embodiments, the data selection module 902 includes: a first data acquisition module configured to acquire second sample data corresponding to a plurality of marker regions and a plurality of calculation methods from the first sample data; a score determination module configured to determine, based on the second sample data, a plurality of importance scores for distinguishing between diseases and non-diseases corresponding to the plurality of marker regions and the plurality of calculation methods; and a third data selection module configured to select at least one marker region and at least one calculation method from the plurality of marker regions and the plurality of calculation methods based on the plurality of importance scores.

[0086] In some embodiments, the disease is a first disease, the first sample data further includes fourth sample data for a second disease, at least one biomarker region is at least one first biomarker region, and at least one calculation method is at least one first calculation method.

[0087] In some embodiments, the model training module 904 includes: a fourth data selection module configured to determine at least one second biomarker region and at least one second calculation method for distinguishing the second disease from a first plurality of calculation methods for a plurality of biomarker regions based on fourth sample data for the second disease; a fifth data selection module configured to select fifth sample data corresponding to the at least one second biomarker region and at least one second calculation method from the fourth sample data; and a first model training module configured to train a model based on third sample data and fifth sample data.

[0088] In some embodiments, the model is one of the following: a neural network model and a decision tree-based gradient boosting model.

[0089] In some embodiments, the apparatus 900 further includes: a validation dataset determination module configured to determine a validation dataset for the model based on a sample dataset; and a performance verification module configured to verify the performance of the model based on the validation dataset.

[0090] In some embodiments, the validation dataset determination module includes a splitting module configured to determine multiple validation data subsets by splitting the sample dataset.

[0091] In some embodiments, the splitting module includes: a candidate dataset determination module configured to determine a plurality of candidate independent center validation datasets by performing a predetermined number of splitting operations on the sample dataset, wherein the sampling centers in the candidate independent center validation datasets are different from the sampling centers of the data used to train the model; a mutual exclusion determination module configured to obtain a set of candidate independent center validation datasets from the plurality of candidate independent center validation datasets where the mutual exclusion degree of the sampling centers is greater than a threshold; and a validation dataset determination module configured to use the candidate independent center validation datasets in the set of candidate independent center validation datasets as validation datasets for validating the model.

[0092] In some embodiments, the splitting operation for the sample dataset may include at least one of the following: splitting based on an individual sampling center, splitting based on the sampling time, splitting based on the experimental batch, or random splitting.

[0093] Figure 10 A schematic block diagram of an example device 1000 that can be used to implement embodiments of the present disclosure is shown. Figure 1 The computing device 102 can be implemented using device 1000. As shown, device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1002 or loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 can also store various programs and data required for the operation of device 1000. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0094] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0095] The various processes and handling described above, such as method 200, can be executed by processing unit 1001. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more actions of the example method 200 described above can be performed.

[0096] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0097] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0098] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. Network adapter cards or network interfaces in all computing / processing devices receive the computer-readable program instructions from the network and forward them to computer-readable storage media in the various computing / processing devices.

[0099] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0100] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that all blocks of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0101] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0102] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, all blocks in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that all blocks in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0104] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for training a model, comprising: selecting, based on second sample data in first sample data corresponding to a plurality of marker regions and a first plurality of computing methods, at least one marker region and at least one computing method from the plurality of marker regions and the first plurality of computing methods, wherein the plurality of marker regions correspond to the first plurality of computing methods, the computing methods are for distinguishing between a disease and a non-disease, and the first sample data comprises sample data for the disease and the non-disease; and training a model for identifying the disease based on third sample data in the first sample data corresponding to the at least one marker region and the at least one computing method.

2. The method of claim 1, further comprising: selecting a first plurality of computing methods; and wherein selecting the first plurality of computing methods comprises: determining a plurality of values corresponding to a plurality of samples in the first sample data, the plurality of values being values of the plurality of samples for the plurality of marker regions and a second plurality of computing methods; determining, based on the plurality of values corresponding to the plurality of samples, classification performances of the second plurality of computing methods for distinguishing between the disease and the non-disease in the plurality of marker regions; and determining, based on the classification performances, the first plurality of computing methods corresponding to the plurality of marker regions from the second plurality of computing methods.

3. The method of claim 2, wherein determining, based on the classification performances, the first plurality of computing methods corresponding to the plurality of marker regions from the second plurality of computing methods comprises: determining, from the classification performances, a target classification performance satisfying a predetermined condition; and determining, as the first plurality of computing methods corresponding to the plurality of marker regions, a computing method corresponding to the target classification performance.

4. The method of claim 1, wherein selecting, from the plurality of marker regions and the first plurality of computing methods, the at least one marker region and the at least one computing method comprises: obtaining, from the first sample data, second sample data corresponding to the plurality of marker regions and the first plurality of computing methods; determining, based on the second sample data, a plurality of importance scores corresponding to the plurality of marker regions and the first plurality of computing methods for distinguishing between the disease and the non-disease; and selecting, based on the plurality of importance scores, the at least one marker region and the at least one computing method from the plurality of marker regions and the first plurality of computing methods.

5. The method of claim 1, wherein the disease is a first disease, the first sample data further comprises fourth sample data for a second disease, the at least one marker region is at least one first marker region, and the at least one computing method is at least one first computing method.

6. The method of claim 5, wherein training a model for identifying the disease based on third sample data in the first sample data corresponding to the at least one marker region and the at least one computing method comprises: ​ ​ determining, from the first plurality of computing methods for the plurality of marker regions, at least one second marker region and at least one second computing method for distinguishing the second disease based on the fourth sample data for the second disease; selecting, from the fourth sample data, fifth sample data corresponding to the at least one second marker region and the at least one second computing method; and training the model based on the third sample data and the fifth sample data.

7. The method of claim 1, wherein the model is one of: a neural network model and a gradient boosting model based on decision trees.

8. The method of claim 1, further comprising: determining, based on the sample data set, a validation data set for the model; and validating, based on the validation data set, a performance of the model.

9. The method of claim 8, wherein determining a validation data set for the sample data set comprises: determining a validation data set by a splitting operation on the sample data set.

10. The method of claim 9, wherein determining a validation data set by a splitting operation on the sample data set comprises: determining a plurality of candidate independent center validation data sets by a predetermined number of splitting operations on the sample data set, wherein a sampling center in a candidate independent center validation data set is different from a sampling center of data used for training the model; obtaining, from the plurality of candidate independent center validation data sets, a group of candidate independent center validation data sets whose sampling centers are mutually exclusive to a threshold degree; and using a candidate independent center validation data set in the group of candidate independent center validation data sets as a validation data set for validating the model.

11. The method of claim 9, wherein the splitting operation on the sample data set can comprise at least one of: a split for independent sampling centers, a split according to sampling time, a split according to experimental batches, or a random split.

12. An apparatus for training a model, comprising: a data selection module configured to select, based on second sample data in first sample data corresponding to a plurality of marker regions and a first plurality of computing methods, at least one marker region and at least one computing method from the plurality of marker regions and the first plurality of computing methods, wherein the plurality of marker regions correspond to the first plurality of computing methods, and the computing method is used to distinguish a disease and a non-disease, and the first sample data comprises sample data for the disease and the non-disease; and a model training module configured to train a model for identifying the disease based on third sample data in the first sample data corresponding to the at least one marker region and the at least one computing method.

13. An electronic device, comprising: at least one processor; and memory storing at least one program, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-11. ​ ​ 14. A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method according to any one of claims 1-11.

15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11.