Colorectum prognostic marker screening method and system
By collecting and analyzing colorectal prognostic markers and population diagnostic data, and classifying and screening population data, the problem of difficulty in determining colorectal cancer prognostic markers in the prior art is solved, and efficient prognostic marker screening and analysis and judgment are achieved.
Patent Information
- Application Number
- CN202510162747.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively determine the prognostic markers of colorectal cancer, which leads to difficulty in prognosis analysis and judgment.
By collecting potential colorectal prognostic marker data sets and population diagnostic data, performing screening analysis, and different classifications screening population data to form a reasonable combination of prognostic marker.
Efficient screening of colorectal prognostic markers is achieved, and a combination of markers can be formed that can accurately analyze and judge colorectal prognostic markers.
Smart Images

Figure CN120148816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for screening colorectal prognosis markers. Background Art
[0002] Colorectal cancer is a common type of cancer and has a high incidence. Clinically, there is currently no good indicator for analyzing and judging the prognosis of colorectal cancer. This is because the prognosis of colorectal cancer is closely related to individual factors, and at the same time, it is also because it is difficult to select its prognosis markers. Currently, efforts are also being made to search for colorectal prognosis markers to determine a stable marker for accurate analysis and judgment of colorectal prognosis.
[0003] Currently, the determination of colorectal prognosis markers mainly involves in-depth analysis and exploration of new marker objects. The determination of markers in this way is often time-consuming and it is not easy to obtain suitable markers. The confirmation of a single marker indeed cannot accurately analyze the prognosis in different individuals.
[0004] Therefore, designing a method and system for screening colorectal prognosis markers, collecting current potential markers and using a reasonable population of diagnoses for rational analysis to screen a combination of prognosis markers formed by multiple markers, and forming a basis for reasonable analysis and judgment of colorectal prognosis is an urgent problem to be solved currently. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for screening colorectal prognosis markers. By collecting the current data set of markers recognized as potential colorectal prognosis markers, and at the same time collecting the reasonable population diagnosis data for screening and analyzing the potential prognosis markers in the colorectal area, and then screening and judging them based on the diagnostic results of different potential colorectal prognosis markers in the population diagnosis data. Of course, since the influence of different diagnostic result individuals in the population diagnosis data on the diagnostic results of the markers is different, when using the population diagnosis data to screen the potential markers, it is necessary to first make a reasonable division of the population data to provide a data reference for mutual comparison in the screening analysis. The colorectal prognosis markers formed through the screening analysis can form a reasonable and accurate prognosis analysis and judgment of colorectal cancer in the form of a combination, providing a new analysis and processing method for the prognosis analysis and judgment of colorectal cancer.
[0006] The object of the present invention is also to provide a screening system for colorectal prognosis markers. The system effectively and reasonably realizes the screening of colorectal prognosis markers by configuring an overall system capable of collecting colorectal potential prognosis markers and screening population data, which is the material basis for ensuring the correct and efficient screening of colorectal prognosis markers.
[0007] In a first aspect, the present invention provides a method for screening colorectal prognosis markers, including: collecting colorectal potential prognosis markers to form a colorectal potential prognosis set; collecting screening population data and performing classification based on population sources to form different classified screening population data; for different classified screening population data, screening and analyzing all colorectal potential prognosis markers in the colorectal potential prognosis set according to the diagnosis results to form prognosis marker screening data.
[0008] In the present invention, the method forms a data set of markers by collecting currently recognized potential colorectal prognosis markers, and at the same time collects reasonable population diagnosis data for screening and analyzing the markers in the colorectal potential prognosis set. Then, it screens and judges the markers based on the performance of the diagnosis results of different colorectal potential prognosis markers in the population diagnosis data. Of course, since the influence of different diagnostic result individuals in the population diagnosis data on the diagnosis results of the markers is different, it is necessary to first perform reasonable division of the population data when using the population diagnosis data to screen potential markers, so as to provide mutually comparable data references for screening and analysis. The colorectal prognosis markers formed through screening and analysis can form a reasonable and accurate prognosis analysis and judgment of colorectal cancer in a combined form, providing a new analysis and processing method for the prognosis analysis and judgment of colorectal cancer.
[0009] As a possible implementation, collecting colorectal potential prognosis markers to form a colorectal potential prognosis set includes: determining the marker collection range, performing semantic-based extraction of colorectal-related markers to form a colorectal-related marker set; setting the target screening range of the markers, and screening all colorectal-related markers in the colorectal-related marker set to determine all colorectal potential prognosis markers and form a colorectal potential prognosis set.
[0010] In the present invention, the extraction of potential prognostic markers for colorectal cancer mainly involves large-scale data screening. Of course, to obtain potential prognostic markers corresponding to colorectal cancer, semantic analysis of colorectal cancer prognostic markers is required to extract possible prognostic markers. Here, it should be noted that the scope of marker collection can be reasonably divided according to the actual situation. For example, the collection scope can be a resource database shared in the medical field or a database of frontier medical journal papers, which are all appropriate data sources for obtaining potential prognostic markers for colorectal cancer. For the target screening scope of markers, it can be determined according to actual needs. The setting of the screening scope mainly limits the possible types of potential prognostic markers, such as limiting to markers at the gene level, markers at the cellular level, markers of body-related pheromones, etc., so as to make the screening more targeted and at the same time provide guidance for the extraction of population data, and screen prognostic markers in a better and more reasonable way.
[0011] As a possible implementation method, collect and screen population data, and conduct classification based on the population source to form different classified and screened population data, including: determining the scope of population information collection, collecting individual prognostic index data corresponding to the potential prognostic markers of colorectal cancer in the potential prognosis of colorectal cancer, and forming screened population data; based on the collection source of population information, classify and divide the screened population data to form different classified and screened population data.
[0012] In the present invention, the screening of population data mainly involves establishing an assessment of whether the prognostic marker has a prognostic analysis effect. It can be understood that for prognostic markers, the expression levels of the same marker in the population without colorectal cancer and in the population with colorectal cancer are different. At the same time, the expression levels of the same marker in patients who have achieved recovery and patients whose condition has deteriorated are also different. Therefore, when collecting and screening population data, on the one hand, it is necessary to obtain population data under different diagnostic conditions to ensure the sufficiency of the data, and thus ensure that the analysis is more reasonable when screening and analyzing prognostic markers. On the other hand, it is also necessary to ensure that there is sufficient population data in different classifications to avoid deviation of the analysis results due to insufficient data and inability to correctly judge whether the potential prognostic marker is a true prognostic marker.
[0013] As a possible implementation, based on the sources of population information collection, the screened population data is classified and divided to form different classified and screened population data, including: determining the sources of different individual prognostic index data in the screened population data, clustering all individual prognostic index data with the same source to form different source-screened population data groups; extracting the colorectal cancer diagnosis results of all individuals in different source-screened population data groups, and determining the prevalence ratio of the population groups diagnosed with colorectal cancer in the corresponding source-screened population data groups; classifying and dividing according to the prevalence ratios of the population groups corresponding to different source-screened population data groups to form different classified and screened population data.
[0014] In the present invention, for the classification of population screening, it is considered that the classification of population data is mainly carried out based on the diagnosis results. Therefore, in different classified data, the prevalence ratio of the diagnosis results of colorectal cancer, especially positive results, is different. Through this ratio, the division of different types of population data can be accurately achieved. Of course, it should be noted that among the different population data formed by the division based on the diagnosis result ratio, there will also be some individuals with diagnosis results different from this type. For example, in the population data of the rehabilitation category, there may be individual data of deterioration. After all, most of the data collection is carried out according to the diagnosis types of medical institutions. For example, directly collect all colorectal population data in the rehabilitation department and directly collect all colorectal population data in the ICU department. However, this situation will not have a substantial impact on the screening results. After all, the screening analysis is carried out by extracting the corresponding same diagnosis result data in the population data.
[0015] As a possible implementation, according to the prevalence ratios of the population groups corresponding to different source-screened population data groups, classify and divide to form different classified and screened population data, including: clustering the source-screened population data groups with a prevalence ratio of zero for all population groups to form undiagnosed screened population data; classifying and dividing the remaining different source-screened population data groups based on the ratio difference to form different classified and screened population data.
[0016] In the present invention, the division of population data considers the performance of potential prognostic markers under different diagnosis results. Therefore, the population data with a negative diagnosis result is a necessary type of classified data. This type of data provides the most basic reference for the screening of prognostic markers.
[0017] As a possible implementation, classify and divide the remaining screened population data groups from different sources based on the proportion gap to form different classified screened population data, including: obtaining the disease prevalence proportions of the population groups corresponding to the remaining screened population data groups from different sources, arranging the disease prevalence proportions of different population groups in ascending order, and determining the differences between the disease prevalence proportions of adjacent population groups; locating the largest difference, and determining the smaller disease prevalence proportion that forms the largest difference as the classification proportion; clustering the source-screened population data groups corresponding to the classification proportion and the source-screened population data groups corresponding to the disease prevalence proportions of all population groups smaller than the classification proportion to form general diagnostic screened population data; clustering the source-screened population data groups corresponding to the disease prevalence proportions of all population groups larger than the classification proportion to form enriched screened population data.
[0018] In the present invention, for the classified population data of the rehabilitation type and the deterioration type, since the prognosis results are in two completely opposite directions, there will be a large difference in the proportions. Therefore, when screening to form classified population data, this feature can be used to take the size of the proportion gap as the basis for judging the classification type of the population data.
[0019] As a possible implementation, for different classified screened population data, screen and analyze all colorectal potential prognostic markers in the colorectal potential prognosis set according to the diagnostic results to form prognostic marker screening data, including: extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the undiagnosed screened population data; extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the general diagnostic screened population data; extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the enriched screened population data; screening out non-prognostic markers from the colorectal potential prognostic markers in the colorectal potential prognosis set according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data to form a colorectal prognosis set; performing differential-based prognostic marker screening analysis on the colorectal potential prognostic markers in the colorectal prognosis set according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data to determine the colorectal prognostic markers.
[0020] In the present invention, after classifying the population data, different population data can be used to screen and analyze potential prognostic markers. Here, when screening, first, the result values of different potential prognostic markers need to be extracted from the three screened population data, and then comparative analysis is carried out based on the result values to determine the prognostic markers. Of course, there are also two obvious situations for potential prognostic markers, that is, some potential prognostic markers cannot be used as the judgment for prognostic analysis, and the other type is that some potential prognostic markers have reference significance for prognostic analysis judgment.
[0021] As a possible implementation, according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data, the colorectal potential prognostic markers in the colorectal potential prognosis set are screened out as non-prognostic markers to form a colorectal prognosis set, including: for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the prognostic result values corresponding thereto in all different screened population data coincide, the corresponding colorectal potential prognostic marker is screened out and labeled; for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the prognostic result values corresponding thereto in any two different screened population data coincide, the corresponding colorectal potential prognostic marker is screened out and labeled; the colorectal potential prognostic markers labeled as screened out in the colorectal potential prognosis set are excluded to form a colorectal prognosis set.
[0022] In the present invention, for a potential prognostic marker, if it can be used as reference data for prognostic analysis, then the performance of the potential prognostic marker in three different screened population data should be different. Through this method of analysis and screening, potential prognostic markers that have no impact on prognostic analysis can be labeled, thereby completing a relatively accurate screening analysis.
[0023] As a possible implementation, according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data, a differential-based screening analysis of the colorectal potential prognostic markers in the colorectal prognosis set is performed to determine colorectal prognostic markers, including: for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the prognostic result values corresponding thereto in all different screened population data do not coincide, and the minimum value of the difference between the prognostic result values corresponding thereto in any two screened population data does not exceed the significant difference threshold, the corresponding colorectal potential prognostic marker is labeled as a general prognostic marker; for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the prognostic result values corresponding thereto in all different screened population data do not coincide, and the minimum value of the difference between the prognostic result values corresponding thereto in any two screened population data exceeds the significant difference threshold, the corresponding colorectal potential prognostic marker is labeled as a significant prognostic marker.
[0024] In the present invention, of course, for potential prognostic markers that can truly provide a basis for prognostic analysis and judgment, there are also certain differences in the significance of their expressions. Therefore, when performing differential analysis, it is necessary to consider whether the potential prognostic markers are obvious, and then provide a classification of the influence degree of the prognostic analysis of these potential prognostic markers, so that the effect of prognostic analysis is more obvious and effective.
[0025] In a second aspect, the present invention provides a colorectal prognosis marker screening system, comprising: a data acquisition unit for collecting potential colorectal prognosis markers to form a set of potential colorectal prognoses, and for collecting data of the screening population; a screening and analysis unit for classifying and dividing the data of the screening population collected by the data acquisition unit to form different classified screening population data, and for obtaining the set of potential colorectal prognoses formed by the data acquisition unit to perform screening and analysis to form prognosis marker screening data.
[0026] In the present invention, the system is configured as a whole capable of collecting potential colorectal prognosis markers and data of the screening population for screening colorectal prognosis markers, effectively and reasonably realizing the screening of colorectal prognosis markers, and is the material basis for ensuring the correct and efficient screening of colorectal prognosis markers.
[0027] The beneficial effects of a colorectal prognosis marker screening method and system provided by the present invention are as follows:
[0028] This method forms a data set of markers by collecting currently recognized potential colorectal prognosis markers, and at the same time collects reasonable population diagnosis data for screening and analyzing the markers in the set of potential colorectal prognoses. Then, it screens and judges them based on the diagnostic result performances of different potential colorectal prognosis markers in the population diagnosis data. Of course, since the impacts of different diagnostic result individuals in the population diagnosis data on the diagnostic results of the markers are different, when using the population diagnosis data to screen potential markers, it is necessary to first perform reasonable division of the population data to provide mutually comparable data references for screening and analysis. The colorectal prognosis markers formed through screening and analysis can form a reasonable and accurate prognosis analysis and judgment for colorectal cancer in a combined form, providing a new analysis and processing method for the prognosis analysis and judgment of colorectal cancer.
[0029] The system is configured as a whole capable of collecting potential colorectal prognosis markers and data of the screening population for screening colorectal prognosis markers, effectively and reasonably realizing the screening of colorectal prognosis markers, and is the material basis for ensuring the correct and efficient screening of colorectal prognosis markers. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 It is a step diagram of the colorectal prognosis marker screening method provided by the embodiment of the present invention;
[0032] Figure 2 Schematic structural diagram of the colorectal prognosis marker screening system provided by the embodiment of the present invention;
[0033] Figure 3 Classification and screening population data type diagram of the colorectal prognosis marker screening method provided by the embodiment of the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.
[0035] Colorectal cancer is a common type of cancer and has a high incidence. Clinically, there is currently no good indicator for analyzing and judging the prognosis of colorectal cancer. This is because the prognosis of colorectal cancer is closely related to individual factors, and at the same time, it is also because the selection of its prognostic markers is quite difficult. Currently, efforts are also being actively made to search for colorectal prognostic markers to determine a stable marker for accurate analysis and judgment of colorectal prognosis.
[0036] Currently, the determination of colorectal prognostic markers mainly involves in-depth analysis and exploration of new marker objects. This way of determining markers is often time-consuming and it is not easy to obtain suitable markers. The confirmation of a single marker indeed cannot perform accurate prognosis analysis on different individuals.
[0037] Reference Figures 1 to 3 , the embodiment of the present invention provides a colorectal prognosis marker screening method. This method collects the currently recognized potential colorectal prognosis markers to form a data set of markers, and at the same time collects the reasonable population diagnosis data for screening and analyzing the markers in the potential colorectal prognosis set. Then, it screens and judges them based on the diagnostic result performances of different potential colorectal prognosis markers in the population diagnosis data. Of course, since the impacts of different diagnostic result individuals in the population diagnosis data on the diagnostic results of the markers are different, when using the population diagnosis data to screen the potential markers, it is necessary to first make a reasonable division of the population data to provide mutually comparable data references for the screening analysis. The colorectal prognosis markers formed through the screening analysis can form a reasonable and accurate prognosis analysis and judgment of colorectal cancer in a combined form, providing a new analysis and processing method for the prognosis analysis and judgment of colorectal cancer.
[0038] The colorectal prognosis marker screening method specifically includes the following steps:
[0039] S1: Collect colorectal potential prognosis markers to form a colorectal potential prognosis set.
[0040] Collect potential prognostic markers for colorectal cancer to form a potential prognostic set for colorectal cancer, including: determining the marker collection scope, extracting colorectal cancer-related markers based on semantics to form a colorectal cancer-related marker set; setting the target screening scope for markers, and screening all colorectal cancer-related markers in the colorectal cancer-related marker set to determine all potential prognostic markers for colorectal cancer and form a potential prognostic set for colorectal cancer.
[0041] The extraction of potential prognostic markers for colorectal cancer mainly involves data screening within a relatively large range. Of course, to obtain potential prognostic markers corresponding to colorectal cancer, semantic analysis of colorectal cancer prognostic markers is required to extract possible prognostic markers. Here, it should be noted that for determining the marker collection scope, it can be reasonably divided according to the actual situation. For example, the collection scope can be a shared resource database in the medical field or a database of cutting-edge medical journal papers, which are all appropriate data sources for obtaining potential prognostic markers for colorectal cancer. For the target screening scope of markers, it can be determined according to actual needs. The setting of the screening scope mainly limits the possible types of potential prognostic markers, such as limiting to markers at the gene level, markers at the cellular level, markers of body-related pheromones, etc., so that the screening is more targeted and also provides guidance for the extraction of population data to screen prognostic markers in a better and more reasonable way.
[0042] S2: Collect and screen population data, and conduct classification based on population sources to form different classified and screened population data.
[0043] Collect and screen population data, and conduct classification based on population sources to form different classified and screened population data, including: determining the population information collection scope, collecting individual prognostic index data corresponding to the potential prognostic markers for colorectal cancer in the potential prognostic set for colorectal cancer to form screened population data; classifying and dividing the screened population data based on the collection source of population information to form different classified and screened population data.
[0044] Screening of population data mainly involves establishing an assessment of whether a prognostic marker has a prognostic analysis effect. It can be understood that for a prognostic marker, its performance level in the population without colorectal cancer is different from that in the population with colorectal cancer. At the same time, the performance levels of the same marker in patients who have achieved recovery and those whose condition has deteriorated are also different. Therefore, when collecting and screening population data, on the one hand, it is necessary to obtain population data of different diagnostic situations to ensure the sufficiency of the data, thereby ensuring more rationality in the screening and analysis of prognostic markers. On the other hand, it is also necessary to ensure that there is sufficient population data in different classifications to avoid deviation of the analysis results due to insufficient data and unable to correctly determine whether a potential prognostic marker is a true prognostic marker.
[0045] Based on the source of population information collection, classify and divide the screened population data to form different classified screened population data, including: determining the sources of different individual prognostic index data in the screened population data, and clustering all individual prognostic index data of the same source to form different source-screened population data groups; extracting the colorectal cancer diagnosis results of all individuals in different source-screened population data groups, and determining the disease prevalence ratio of the population group diagnosed with colorectal cancer in the corresponding source-screened population data group; classifying and dividing according to the disease prevalence ratio of the population group corresponding to different source-screened population data groups to form different classified screened population data.
[0046] Regarding the classification of population screening, it is considered that the classification of population data is mainly based on the diagnosis results. Therefore, in different classified data, the proportion of positive diagnosis results for colorectal cancer, especially, is different. Through this proportion, the division of population data of different categories can be accurately achieved. Of course, it should be noted that among the different population data formed by the division using the diagnosis result proportion, there will also be some individuals with diagnosis results different from this type. For example, there may be individual data of deterioration in the population data of the recovery category. After all, most of the data collection is carried out according to the medical treatment types of medical institutions. For example, directly collect all colorectal population data in the rehabilitation department and directly collect all colorectal population data in the ICU department. However, this situation will not have a substantial impact on the screening results. After all, the screening analysis is carried out by extracting the corresponding same diagnosis result data in the population data.
[0047] Classify and divide according to the disease prevalence ratio of the population group corresponding to different source-screened population data groups to form different classified screened population data, including: clustering all source-screened population data groups with a disease prevalence ratio of zero in the population group to form undiagnosed screened population data; classifying and dividing the remaining different source-screened population data groups based on the ratio difference to form different classified screened population data.
[0048] The classification of population data takes into account the performance of potential prognostic markers under different diagnostic results. Therefore, the population data with negative diagnostic results is a necessary type of classification data. This type of data provides the most basic reference for the screening of prognostic markers.
[0049] Classify and divide the remaining screened population data groups from different sources based on the proportion gap to form different classified screened population data, including: obtaining the disease prevalence proportions of the population groups corresponding to the remaining screened population data groups from different sources, arranging the different disease prevalence proportions of the population groups in ascending order, and determining the differences between the disease prevalence proportions of adjacent population groups; locating the largest difference, and determining the classified proportion as the smaller disease prevalence proportion that forms the largest difference; clustering the source screened population data groups corresponding to the classified proportion and the source screened population data groups corresponding to all disease prevalence proportions of population groups smaller than the classified proportion to form general diagnostic screened population data; clustering the source screened population data groups corresponding to the disease prevalence proportions of all population groups larger than the classified proportion to form enriched screened population data.
[0050] For the classified population data of the recovery type and the deterioration type, since the prognostic results are in two completely opposite directions, there will be a large difference in the proportions. Therefore, when screening to form classified population data, this characteristic can be used to take the size of the proportion gap as the basis for judging the classification type of population data.
[0051] S3: For different classified screened population data, screen and analyze all colorectal potential prognostic markers in the colorectal potential prognosis set according to the diagnostic results to form prognostic marker screening data.
[0052] For different classified screened population data, screen and analyze all colorectal potential prognostic markers in the colorectal potential prognosis set according to the diagnostic results to form prognostic marker screening data, including: extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the undiagnosed screened population data; extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the general diagnostic screened population data; extracting the prognostic result values of all colorectal potential prognostic markers in the colorectal potential prognosis set existing in the enriched screened population data; screening out non-prognostic markers from the colorectal potential prognostic markers in the colorectal potential prognosis set according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data to form a colorectal prognosis set; performing differential-based prognostic marker screening analysis on the colorectal potential prognostic markers in the colorectal prognosis set according to the prognostic result values of all colorectal potential prognostic markers extracted from different screened population data to determine the colorectal prognostic markers.
[0053] After classifying the population data, different population data can be used to screen and analyze potential prognostic markers. Here, when conducting the screening, it is first necessary to extract the result values for different potential prognostic markers from three different screening population data, and then conduct a comparative analysis based on the result values to determine the prognostic markers. Of course, there are two obvious situations for potential prognostic markers, that is, the potential prognostic markers cannot be used as a judgment for prognostic analysis, and the other is that the potential prognostic markers have reference significance for prognostic analysis and judgment.
[0054] Based on the prognostic result values of all colorectal potential prognostic markers extracted from different screening population data, non-prognostic markers are screened out from the colorectal potential prognostic markers in the colorectal potential prognosis set to form a colorectal prognosis set, including: for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the corresponding prognostic result values in all different screening population data coincide, the corresponding colorectal potential prognostic marker is screened out and marked; for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the corresponding prognostic result values in any two different screening population data coincide, the corresponding colorectal potential prognostic marker is screened out and marked; the colorectal potential prognostic markers marked as screened out in the colorectal potential prognosis set are excluded to form a colorectal prognosis set.
[0055] For potential prognostic markers, if they can be used as reference data for prognostic analysis, their performance in three different screening population data should be different. Through this method of analysis and screening, potential prognostic markers that have no impact on prognostic analysis can be marked, and thus a relatively accurate screening analysis can be completed.
[0056] Based on the prognostic result values of all colorectal potential prognostic markers extracted from different screening population data, a differential-based screening analysis of prognostic markers is conducted on the colorectal potential prognostic markers in the colorectal prognosis set to determine the colorectal prognostic markers, including: for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the corresponding prognostic result values in all different screening population data do not coincide, and the minimum value of the difference between the corresponding prognostic result values in any two screening population data does not exceed the significant difference threshold, the corresponding colorectal potential prognostic marker is marked as a general prognostic marker; for any colorectal potential prognostic marker in the colorectal potential prognosis set, if the corresponding prognostic result values in all different screening population data do not coincide, and the minimum value of the difference between the corresponding prognostic result values in any two screening population data exceeds the significant difference threshold, the corresponding colorectal potential prognostic marker is marked as a significant prognostic marker.
[0057] Of course, for potential prognostic markers that can truly provide a basis for prognostic analysis and judgment, there are also certain differences in the significance of their expression. Therefore, when performing differential analysis, it is necessary to consider whether the potential prognostic markers are significant, and then provide a classification of the impact size of the prognostic analysis and treatment of these potential prognostic markers, so as to make the effect of prognostic analysis and treatment more obvious and effective.
[0058] The present invention also provides a colorectal prognostic marker screening system, which includes: a data acquisition unit for collecting colorectal potential prognostic markers to form a colorectal potential prognosis set and for collecting screening population data; a screening and analysis unit for classifying and dividing the screening population data collected by the data acquisition unit to form different classified screening population data, and obtaining the colorectal potential prognosis set formed by the data acquisition unit for screening and analysis to form prognostic marker screening data.
[0059] By being configured as an overall unit capable of collecting colorectal potential prognostic markers and screening population data for colorectal prognostic marker screening, the system effectively and reasonably realizes the screening of colorectal prognostic markers, which is the material basis for ensuring the correct and efficient screening of colorectal prognostic markers.
[0060] In summary, the beneficial effects of the colorectal prognostic marker screening method and system provided by the embodiments of the present invention are as follows:
[0061] This method forms a data set of markers by collecting currently recognized potential colorectal prognostic markers, and at the same time collects reasonable population diagnosis data for screening and analyzing the markers in the colorectal potential prognosis set. Then, it screens and judges them based on the diagnostic results of different colorectal potential prognostic markers in the population diagnosis data. Of course, since the impacts of different diagnostic result individuals in the population diagnosis data on the diagnostic results of markers are different, when using the population diagnosis data to screen potential markers, it is necessary to first perform a reasonable division of the population data to provide a data reference for mutual comparison in the screening analysis. The colorectal prognostic markers formed through screening and analysis can form a reasonable and accurate prognostic analysis and judgment of colorectal cancer in a combined form, providing a new analysis and treatment method for the prognostic analysis and judgment of colorectal cancer.
[0062] By being configured as an overall unit capable of collecting colorectal potential prognostic markers and screening population data for colorectal prognostic marker screening, the system effectively and reasonably realizes the screening of colorectal prognostic markers, which is the material basis for ensuring the correct and efficient screening of colorectal prognostic markers.
[0063] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. If the information indicated by a certain piece of information is called the information to be indicated, then in the specific implementation process, there are many ways to indicate the information to be indicated. For example, but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated, etc. It is also possible to indirectly indicate the information to be indicated by indicating other information, where there is an association relationship between the other information and the information to be indicated. It is also possible to only indicate a part of the information to be indicated, while the other parts of the information to be indicated are known or pre-agreed. For example, it is also possible to use the arrangement order of each piece of information pre-agreed (such as stipulated in the protocol) to achieve the indication of specific information, thereby reducing the indication overhead to a certain extent. At the same time, it is also possible to identify the common parts of each piece of information and indicate them uniformly to reduce the indication overhead caused by separately indicating the same information.
[0064] In addition, the specific indication method can also be various existing indication methods, such as, but not limited to, the above-mentioned indication methods and their various combinations, etc. The specific details of various indication methods can refer to the prior art and will not be elaborated herein. As can be seen from the above, for example, when it is necessary to indicate multiple pieces of information of the same type, it is possible that the indication methods of different pieces of information are different. In the specific implementation process, the required indication method can be selected according to specific needs. The embodiments of the present application do not limit the selected indication method. In this way, the indication methods involved in the embodiments of the present application should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.
[0065] It should be understood that the information to be indicated can be sent as a whole, or can be divided into multiple sub-information and sent separately, and the sending periods and / or sending times of these sub-information can be the same or different. The specific sending method is not limited in the embodiments of the present application. Among them, the sending periods and / or sending times of these sub-information can be predefined, such as predefined according to the protocol, or can be configured by the sending device by sending configuration information to the receiving device.
[0066] "Predefined" or "pre-configured" can be implemented by pre-saving the corresponding codes, tables or other ways that can be used to indicate relevant information in the device. The embodiments of the present application do not limit its specific implementation method. Among them, "saving" can mean saving in one or more memories. The one or more memories can be separately provided, or can be integrated in an encoder or decoder, a processor, or a communication device. The one or more memories can also be partly separately provided and partly integrated in a decoder, a processor, or a communication device. The type of the memory can be any form of storage medium, which is not limited in the embodiments of the present application.
[0067] The "protocol" involved in the embodiments of this application may refer to a protocol family in the communication field, a standard protocol with a frame structure similar to that of a protocol family, or a related protocol applied to future communication systems. The embodiments of this application do not make specific limitations on this.
[0068] In the embodiments of this application, descriptions such as "when...", "in the case of...", "if", and "when" all mean that the device will perform corresponding processing under certain objective circumstances, not limited to time, and it is not required that the device must have a judgment action when implemented, nor does it mean that there are other limitations.
[0069] In the description of the embodiments of this application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; the "and / or" in the embodiments of this application is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. Also, in the description of the embodiments of this application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression means any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple. Additionally, for the convenience of clearly describing the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit to be different. At the same time, in the embodiments of this application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0070] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0071] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM).
[0072] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0073] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0074] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0075] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0076] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0077] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0078] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0079] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0080] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0081] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0082] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for screening colorectal prognostic markers, characterized in that: include: Collect colorectal potential prognostic markers to form a colorectal potential prognostic set; Collect and screen population data, and classify and divide the population based on their sources to form different classified screening population data; For the different classified screening population data, all the colorectal potential prognostic markers in the colorectal potential prognostic set are screened and analyzed according to the diagnosis results to form prognostic marker screening data.
2. The method for screening colorectal prognostic markers according to claim 1, characterized in that: The colorectal potential prognostic markers are collected to form a colorectal potential prognostic set, including: Determine the scope of marker collection, perform semantic-based colorectal-related marker extraction, and form a colorectal-related marker set; A target screening range of markers is set, and all the colorectal-related markers in the colorectal-related marker set are screened to determine all colorectal potential prognostic markers to form the colorectal potential prognostic set.
3. The method for screening colorectal prognostic markers according to claim 2, characterized in that: The collected and screened population data is classified and divided based on the population source to form different classified and screened population data, including: Determining the scope of population information collection, collecting individual prognostic indicator data corresponding to the colorectal potential prognostic markers in the colorectal potential prognostic set, and forming screening population data; Based on the collection source of the crowd information, the screened crowd data is classified and divided to form different classified screened crowd data.
4. The method for screening colorectal prognostic markers according to claim 3, characterized in that: Based on the collection source of the crowd information, the screening crowd data is classified and divided to form different classified screening crowd data, including: Determining the sources of different individual prognostic indicator data in the screening population data, and clustering all the individual prognostic indicator data from the same source to form different source screening population data groups; Extracting the colorectal cancer diagnosis results of all individuals in the different source screening population data groups, and determining the proportion of the population group diagnosed with colorectal cancer in the corresponding source screening population data groups; According to the disease proportion of the population groups corresponding to the different source screening population data groups, classification is performed to form different classified screening population data.
5. The method for screening colorectal prognostic markers according to claim 4, characterized in that: The population group disease proportions corresponding to the population data groups screened from different sources are classified and divided to form different classified screening population data, including: Clustering the source screening population data group with a disease proportion of zero in all population groups to form undiagnosed screening population data; The remaining data groups of the population screened from different sources are classified and divided based on the difference in proportion to form different classified screening population data.
6. The method for screening colorectal prognostic markers according to claim 5, characterized in that: The remaining data groups of the screening population from different sources are classified and divided based on the difference in proportion to form different classified screening population data, including: Obtain the remaining disease proportions of the population groups corresponding to the different source screening population data groups, arrange the disease proportions of the different population groups in order from small to large, and determine the difference between the disease proportions of adjacent population groups; Locate the largest difference, and determine the disease proportion of the smaller population group that forms the largest difference as the classification proportion; Clustering the source screening population data group corresponding to the classification proportion and the source screening population data group corresponding to all the population group disease proportions smaller than the classification proportion to form general diagnosis screening population data; Clustering the source screening population data groups corresponding to all the disease proportions of the population group that are greater than the classification proportion to form enriched screening population data.
7. The method for screening colorectal prognostic markers according to claim 6, characterized in that: The different classified screening population data are screened and analyzed for all the colorectal potential prognostic markers in the colorectal potential prognostic set according to the diagnosis results to form prognostic marker screening data, including: Extracting the prognostic result values of all the colorectal potential prognostic markers in the colorectal potential prognostic set present in the undiagnosed screening population data; Extracting the prognostic result values of all the colorectal potential prognostic markers in the colorectal potential prognostic set present in the general diagnosis screening population data; Extracting the prognostic result values of all the colorectal potential prognostic markers in the colorectal potential prognostic set present in the enriched screening population data; According to the prognostic result values of all the colorectal potential prognostic markers extracted from the data of different screening populations, the colorectal potential prognostic markers in the colorectal potential prognostic set are screened out as non-prognostic markers to form a colorectal prognostic set; According to the prognostic result values of all the colorectal potential prognostic markers extracted from the data of different screening populations, the colorectal potential prognostic markers in the colorectal prognosis set are subjected to difference-based prognostic marker screening analysis to determine the colorectal prognostic markers.
8. The method for screening colorectal prognostic markers according to claim 7, characterized in that: The prognostic result values of all the colorectal potential prognostic markers extracted from the data of different screening populations are used to screen out the colorectal potential prognostic markers in the colorectal potential prognostic set as non-prognostic markers to form a colorectal prognostic set, including: For any colorectal potential prognostic marker in the colorectal potential prognostic set, if the corresponding prognostic result values in all different screening population data overlap, the corresponding colorectal potential prognostic marker is screened out and calibrated; For any colorectal potential prognostic marker in the colorectal potential prognostic set, if the corresponding prognostic result values in any two different screening population data overlap, the corresponding colorectal potential prognostic marker is screened out and calibrated; The colorectal potential prognostic markers marked as screened out in the colorectal potential prognostic set are excluded to form the colorectal prognostic set.
9. The method for screening colorectal prognostic markers according to claim 8, characterized in that: The method of performing a prognostic marker screening analysis based on differences on the colorectal potential prognostic markers in the colorectal prognosis set according to the prognostic result values of all the colorectal potential prognostic markers extracted from the data of different screening populations, and determining the colorectal prognostic markers, comprises: For any of the colorectal potential prognostic markers in the colorectal potential prognostic set, if the corresponding prognostic result values in all different screening population data do not overlap, and the minimum value of the difference between the corresponding prognostic result values in any two screening population data does not exceed the significant difference threshold, the corresponding colorectal potential prognostic marker is marked as a general prognostic marker; For any colorectal potential prognostic marker in the colorectal potential prognostic set, if the corresponding prognostic result values in all different screening population data do not overlap, and the minimum value of the difference between the corresponding prognostic result values in any two screening population data exceeds the significant difference threshold, then the corresponding colorectal potential prognostic marker will be marked as a significant prognostic marker.
10. A colorectal prognostic marker screening system, characterized in that: include: A data collection unit, used for collecting colorectal potential prognostic markers to form a colorectal potential prognostic set for collecting screening population data; The screening and analysis unit is used to classify and divide the screening population data collected by the data collection unit to form different classified screening population data, and to obtain the colorectal potential prognosis set formed by the data collection unit, perform screening and analysis, and form prognostic marker screening data.