A human brain tumor prediction model training system and training method

By selecting data selection methods based on heterogeneous crossover abundance and ROI feature coefficients in the human brain tumor prediction model, and optimizing data screening by combining training difficulty and judgment conditions, the problem of insufficient data balance in different treatment cycles of the model was solved, and the accuracy and stability of the model were improved.

CN120452770BActive Publication Date: 2025-12-05THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510517577.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-12-05
Estimated Expiration
2045-04-24

Smart Images

  • Figure CN120452770B_ABST
    Figure CN120452770B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of tumor prediction, and particularly relates to a human brain tumor prediction model training system and a training method, the system comprising: a data acquisition unit; a selection analysis unit configured to determine a data source state according to heterogeneous cross abundance and ROI characteristic coefficients, and determine a data selection mode according to the data source state; an extension selection unit configured to determine a processing mode of each data source according to a tumor data category in multi-data source extension selection; a supplementary selection unit configured to determine a representative data source based on a data source evaluation index in representative data source supplementary selection, determine a representative combination according to ROI similarity and effective similarity, and determine a supplementary selection mode of each representative combination according to a combination comparison coefficient; and a model training unit; the present application can improve the stability of model prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tumor prediction technology, and in particular to a training system and method for a human brain tumor prediction model. Background Technology

[0002] In oncology research and clinical practice, establishing accurate tumor prediction models is crucial for developing personalized treatment plans and evaluating treatment outcomes. However, due to the complexity and heterogeneity of tumors, a single data source often cannot provide comprehensive and accurate information to meet the needs of model training. Therefore, how to select training data to improve the stability of tumor prediction models is a technical problem that urgently needs to be solved by those skilled in the art.

[0003] Chinese Patent Publication No. CN119361146A discloses a data analysis-based tumor recurrence risk assessment and early warning system, including: a multi-data collection module, a big data processing platform, a model architecture system, and an individual data processing system. This invention addresses the problems of existing technologies, such as lack of big data support, limited personalized prediction capabilities, and insufficient real-time and continuous monitoring. By integrating large-scale historical case data and using these rich datasets for model training, this invention can significantly improve the accuracy and generalization ability of the prediction model. It also integrates multiple data sources such as genomic information, lifestyle habits, and environmental factors to provide personalized risk assessments for each patient. Furthermore, it monitors changes in patients' daily physiological parameters in real time through interfaces connected to various intelligent medical devices, providing continuous health management recommendations for post-tumor surgery patients. However, the above-mentioned technical solutions have the following problems: they do not screen the training data and do not consider the poor balance of data across different treatment cycles, leading to the model's inability to fully learn the characteristics of minority classes of data, resulting in poor stability of the model's prediction results. Summary of the Invention

[0004] To address this issue, the present invention provides a training system and method for a human brain tumor prediction model, which overcomes the problem in the prior art where the model cannot fully learn the features of minority class data when the balance of data at different treatment stages is poor, resulting in poor stability of the model prediction results.

[0005] To achieve the above objectives, the present invention provides a training system for a human brain tumor prediction model, comprising:

[0006] The data acquisition unit is used to acquire tumor data from several data sources;

[0007] An analysis unit is selected, which is connected to the data acquisition unit, to determine the data source status based on the heterogeneous cross abundance and ROI characteristic coefficients, and to determine the data selection method based on the data source status. The data selection method is either extended selection from multiple data sources or supplementary selection from representative data sources.

[0008] An extended selection unit, which is connected to the data acquisition unit and the selection analysis unit respectively, is used to determine the processing method of each data source according to the tumor data category in the extended selection of multiple data sources. The processing method includes determining a first screening method according to the training difficulty coefficient and determining a second screening method according to the judgment conditions.

[0009] The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values.

[0010] The supplementary selection unit is connected to the data acquisition unit and the selection analysis unit respectively. It is used to determine the representative data source based on the data source evaluation index, determine the representative combination based on ROI similarity and effective similarity, and determine the supplementary selection method of each representative combination based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient.

[0011] The model training unit is connected to the extended selection unit and the supplementary selection unit respectively, and is used to train the tumor prediction model using the selected tumor data as training data.

[0012] Furthermore, the selection analysis unit responds to the data source status to determine the data selection method, including:

[0013] If the data source status of the selected analysis unit response is that the heterogeneous cross abundance is less than the preset heterogeneous cross abundance or the ROI feature coefficient is less than the preset ROI feature coefficient, then the data selection method is determined to be multi-data source extended selection.

[0014] If the data source status of the selected analysis unit response is heterogeneous cross abundance greater than or equal to the preset heterogeneous cross abundance and ROI feature coefficient greater than or equal to the preset ROI feature coefficient, then the data selection method is determined to be supplementary selection of representative data sources.

[0015] Furthermore, the extended selection unit determines the tumor data category based on the treatment effectiveness and treatment fluctuation value. The tumor data categories include:

[0016] Tumor data that has a treatment effectiveness greater than or equal to the preset treatment effectiveness and a treatment fluctuation value greater than or equal to the preset treatment fluctuation value;

[0017] Type II tumor data: treatment effectiveness less than the preset treatment effectiveness or treatment fluctuation value less than the preset treatment fluctuation value.

[0018] Furthermore, the extended selection unit determines the processing method for each data source based on the tumor data category, including:

[0019] For each type of tumor data in a single data source, the processing method is to determine the first screening method based on the training difficulty coefficient;

[0020] For each type of tumor data in a single data source, the processing method is to determine the second screening method based on the judgment criteria.

[0021] Furthermore, the extended selection unit determines the first screening method based on the training difficulty coefficient, including:

[0022] If the training difficulty coefficient is greater than or equal to the preset training difficulty coefficient, the first screening method is determined to be to select tumor data based on the angiography feature value and the rhythm risk degree.

[0023] If the training difficulty coefficient is less than the preset training difficulty coefficient, then the first screening method is determined to be selecting tumor data based on tumor reference values.

[0024] Furthermore, the method for confirming the training difficulty coefficient includes:

[0025] If the heterogeneity correlation is greater than or equal to the preset heterogeneity correlation, the training difficulty coefficient is determined based on the heterogeneity interaction difficulty.

[0026] If the heterogeneity correlation is less than the preset heterogeneity correlation, the training difficulty coefficient is determined based on the homogeneity correlation coefficient.

[0027] Furthermore, the extended selection unit response determination condition for determining the second filtering method includes:

[0028] The extended selection unit response determination condition is that the edge fuzziness coefficient is greater than or equal to the preset edge fuzziness coefficient and the ROI comparison coefficient is less than the preset ROI comparison coefficient. Then the second screening method is to select tumor data based on the associated ROI feature value and the mediating diffusion index.

[0029] The extended selection criteria for unit response are: if the edge fuzziness coefficient is less than the preset edge fuzziness coefficient or the ROI comparison coefficient is greater than or equal to the preset ROI comparison coefficient, then the second screening method is to select tumor data based on the data evaluation value.

[0030] Furthermore, the method for confirming the associated ROI feature values ​​includes:

[0031] The initial characterization interval is determined based on the image characterization value, and the area of ​​the initial characterization interval is increased according to the specificity index of the initial characterization interval. The increased and adjusted initial characterization interval is recorded as the characterization interval.

[0032] The feature value of the associated ROI is determined by the ratio of the mutation coefficient corresponding to the characterization interval to the area of ​​the characterization interval.

[0033] Furthermore, the supplementary selection unit determines representative data sources based on the data source evaluation index, determines representative combinations based on ROI similarity and effective similarity, and determines the supplementary selection method for each representative combination based on the combination comparison coefficient.

[0034] For a single representative combination,

[0035] If the combined comparison coefficient is less than the preset combined comparison coefficient, the supplementary selection method is determined to be to select tumor data based on feature keywords;

[0036] If the combined comparison coefficient is greater than or equal to the preset combined comparison coefficient, then the supplementary selection method is determined to be to select tumor data based on the dynamic difference coefficient.

[0037] This invention also provides a method for training a human brain tumor prediction model, comprising:

[0038] Tumor data are acquired from several data sources. The data source status is determined based on heterogeneous crossover abundance and ROI characteristic coefficients. The data selection method is determined based on the data source status, which can be either extended selection from multiple data sources or supplementary selection from representative data sources.

[0039] In the extended selection of multiple data sources, the processing method of each data source is determined according to the tumor data category. The processing methods include determining the first screening method based on the training difficulty coefficient and determining the second screening method based on the judgment criteria.

[0040] The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values.

[0041] In the supplementary selection of representative data sources, representative data sources are determined based on the data source evaluation index, representative combinations are determined based on ROI similarity and effective similarity, and the supplementary selection method for each representative combination is determined based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient.

[0042] The selected tumor data was used as training data to train the tumor prediction model.

[0043] Compared with the prior art, the beneficial effects of the present invention are that, in the technical solution of the present invention, the data source status is determined according to the heterogeneous cross-abundance and ROI characteristic coefficient. The heterogeneous cross-abundance and ROI characteristic coefficient effectively reflect the richness of the data source, and then different data selection methods are adaptively selected according to the data source status, so that the selection of data selection methods is more in line with the actual application scenario, which can increase the diversity of tumor data and improve the accuracy of model prediction.

[0044] Furthermore, in this invention, tumor data categories are determined based on treatment effectiveness and treatment fluctuation values. Treatment effectiveness and treatment fluctuation values ​​effectively reflect the treatment effect of the treatment cycle. Then, the processing method of each data source is determined according to the tumor data category. The processing method can be dynamically adjusted to improve the balance of data in different treatment cycles, thereby improving the stability of the tumor prediction model.

[0045] Furthermore, this invention effectively reflects the training difficulty of tumor data through a training difficulty coefficient, and then determines the first screening method based on the training difficulty coefficient. This avoids the problem in the prior art where the model prediction results are biased when tumors with different treatment cycles have similar characteristics, which is not considered when selecting training data. It can select tumor data based on contrast feature values ​​and rhythm risk, which can reveal the hidden differences in tumor data, improve the richness of data selection, and thus improve the accuracy and generalization ability of the tumor prediction model.

[0046] Furthermore, the present invention effectively reflects the similarity between the two types of tumor data and the ambiguity of the ROI region through the judgment conditions. Then, different second screening methods are adaptively selected according to the judgment conditions, which can more comprehensively consider the multi-faceted characteristics of tumors, avoid the problem of unbalanced model training data selection due to insufficient similarity of a single indicator, take into account the impact of individual differences on data selection, optimize the model training effect, and thus improve the prediction accuracy and stability of the model. Attached Figure Description

[0047] Figure 1 This is a unit connection diagram of the human brain tumor prediction model training system of the present invention;

[0048] Figure 2 This is a flowchart illustrating how the present invention determines the data selection method based on the data source status;

[0049] Figure 3 This is a flowchart illustrating how the first selection method is determined based on the training difficulty coefficient according to the present invention.

[0050] Figure 4 This is a schematic diagram of the training method for the human brain tumor prediction model of the present invention. Detailed Implementation

[0051] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0052] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0053] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0054] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0055] Please see Figures 1 to 3 As shown, the present invention provides a training system for a human brain tumor prediction model, comprising:

[0056] The data acquisition unit is used to acquire tumor data from several data sources;

[0057] An analysis unit is selected, which is connected to the data acquisition unit, to determine the data source status based on the heterogeneous cross abundance and ROI characteristic coefficients, and to determine the data selection method based on the data source status. The data selection method is either extended selection from multiple data sources or supplementary selection from representative data sources.

[0058] An extended selection unit, which is connected to the data acquisition unit and the selection analysis unit respectively, is used to determine the processing method of each data source according to the tumor data category in the extended selection of multiple data sources. The processing method includes determining a first screening method according to the training difficulty coefficient and determining a second screening method according to the judgment conditions.

[0059] The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values.

[0060] The supplementary selection unit is connected to the data acquisition unit and the selection analysis unit respectively. It is used to determine the representative data source based on the data source evaluation index, determine the representative combination based on ROI similarity and effective similarity, and determine the supplementary selection method of each representative combination based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient.

[0061] The model training unit is connected to the extended selection unit and the supplementary selection unit respectively, and is used to train the tumor prediction model using the selected tumor data as training data.

[0062] The application scenario of this invention is the training of a prognostic prediction model for human brain tumors. This invention includes several data sources, each storing several tumor data. Each tumor data includes treatment data corresponding to different consultation times for a single patient. The treatment data corresponding to a single consultation time includes examination information, medication records, disease description, and treatment plan. Examination information includes, but is not limited to, human brain X-ray images, blood test information, and ultrasound contrast imaging information. Blood test information includes, but is not limited to, CA15-3 concentration value and red blood cell count. Each human brain X-ray image corresponds to a Region of Interest (ROI), which is the tumor area manually circled by the doctor using a delineation tool. Ultrasound contrast imaging information includes, but is not limited to, TIC images and ultrasound contrast imaging videos. TIC images are two-dimensional images with time as the horizontal axis and the signal intensity of the contrast agent as the vertical axis. This is content that is easily understood by those skilled in the art and will not be elaborated further.

[0063] This invention includes several historical records, each recording at least one instance of the human brain tumor prediction model training process, including heterogeneous crossover abundance, ROI feature coefficients, diagnostic effectiveness, diagnostic fluctuation value, training difficulty coefficient, and edge blur coefficient. Each historical record also has a corresponding qualification marker, which indicates whether the human brain tumor prediction model training process meets the user's needs. The qualification marker can be manually recorded. It is understood that the user can determine whether the human brain tumor prediction model training process meets the requirements based on self-defined indicators. These self-defined indicators can be, but are not limited to, the misjudgment rate, which will not be elaborated here. The misjudgment rate is the number of times the tumor prediction model incorrectly predicts the tumor treatment plan.

[0064] Using selected tumor data as training data to train a tumor prediction model can make the prediction results of the tumor prediction model more accurate. The specific training process is easy for those skilled in the art to understand and will not be described in detail.

[0065] Specifically, the selection analysis unit responds to the data source status to determine the data selection method, including:

[0066] If the data source status of the selected analysis unit response is that the heterogeneous cross abundance is less than the preset heterogeneous cross abundance or the ROI feature coefficient is less than the preset ROI feature coefficient, then the data selection method is determined to be multi-data source extended selection.

[0067] If the data source status of the selected analysis unit response is heterogeneous cross abundance greater than or equal to the preset heterogeneous cross abundance and ROI feature coefficient greater than or equal to the preset ROI feature coefficient, then the data selection method is determined to be supplementary selection of representative data sources.

[0068] The data source status includes a first data source status and a second data source status. The first data source status is when the heterogeneous cross abundance is less than the preset heterogeneous cross abundance or the ROI feature coefficient is less than the preset ROI feature coefficient. The second data source status is when the heterogeneous cross abundance is greater than or equal to the preset heterogeneous cross abundance and the ROI feature coefficient is greater than or equal to the preset ROI feature coefficient.

[0069] Heterogeneous cross abundance is the maximum value among the sub-heterogeneous cross abundances corresponding to each data source. For a single data source, this data source is denoted as the target data source, and other data sources other than the target data source are denoted as reference data sources. The sub-heterogeneous cross abundance corresponding to the target data source = the number of identical keywords / the total number of keywords appearing in the treatment plans corresponding to each tumor data in each data source. The treatment plans corresponding to each tumor data in each reference data source are denoted as reference records, and the treatment plans corresponding to each tumor data in the target data source are denoted as target records. The number of identical keywords is the number of keywords that appear in both the target record and the reference record.

[0070] The ROI feature coefficient is the maximum value among the sub-ROI feature coefficients corresponding to each data source. The sub-ROI feature coefficient corresponding to a single data source = ROI fluctuation coefficient + treatment fluctuation coefficient.

[0071] The ROI fluctuation coefficient is the standard deviation of the ROI reference value for each tumor data in a single data source. The ROI reference value for a single tumor data is the area of ​​the ROI region in the human brain X-ray image taken at the patient's first consultation time in that tumor data. The treatment fluctuation coefficient is the standard deviation of the treatment effectiveness for each tumor data in a single data source. The treatment effectiveness for a single tumor data is equal to the CA15-3 concentration value corresponding to the earliest treatment data of the patient in that tumor data minus the CA15-3 concentration value corresponding to the latest treatment data of the patient in that tumor data. The CA15-3 concentration value is usually measured by enzyme-linked immunosorbent assay or chemiluminescent immunoassay, which is easily understood by those skilled in the art and will not be elaborated further.

[0072] Users can determine the preset values ​​of heterogeneous cross abundance and preset ROI feature coefficients according to their actual application scenarios. The larger the preset values ​​of heterogeneous cross abundance and preset ROI feature coefficients, the greater the user's need for extended selection of multiple data sources. A preset value of heterogeneous cross abundance and preset ROI feature coefficients is provided. Historical records that represent supplementary selection of data sources are detected. The average value of the heterogeneous cross abundance corresponding to the historical records that meet the user's needs is recorded as the preset heterogeneous cross abundance, and the average value of the ROI feature coefficients corresponding to the historical records that meet the user's needs is recorded as the preset ROI feature coefficients.

[0073] Specifically, the extended selection unit determines the tumor data category based on the effectiveness of diagnosis and treatment and the fluctuation value of diagnosis and treatment. The tumor data categories include:

[0074] Tumor data that has a treatment effectiveness greater than or equal to the preset treatment effectiveness and a treatment fluctuation value greater than or equal to the preset treatment fluctuation value;

[0075] Type II tumor data: treatment effectiveness less than the preset treatment effectiveness or treatment fluctuation value less than the preset treatment fluctuation value.

[0076] Wherein, the treatment fluctuation value = standard deviation of CA15-3 concentration value corresponding to each consultation time in a single tumor data / percentage of decrease frequency, and percentage of decrease frequency = number of consultation times with decrease / (number of consultation times in a single tumor data - 1). For each consultation time in a single tumor data, the consultation times other than the earliest consultation time are recorded as reference consultation times. The method for confirming the decrease in consultation time is as follows: for a single reference consultation time, if the CA15-3 concentration value of the reference consultation time is greater than the CA15-3 concentration value of the adjacent consultation time corresponding to the reference consultation time, then the reference consultation time is recorded as the decrease in consultation time. The adjacent consultation time corresponding to a single reference consultation time is the consultation time that is adjacent to the reference consultation time and earlier than the reference consultation time.

[0077] Users can determine the preset treatment effectiveness and preset treatment fluctuation values ​​based on the actual application scenario. The smaller the preset treatment effectiveness and preset treatment fluctuation values, the greater the user's need to determine the first screening method based on the training difficulty coefficient. A preset treatment effectiveness and preset treatment fluctuation value are provided. The historical records of determining the first screening method based on the training difficulty coefficient are detected. The average treatment effectiveness corresponding to the historical records that meet the user's needs is recorded as the preset treatment effectiveness, and the average treatment fluctuation value corresponding to the historical records that meet the user's needs is recorded as the preset treatment fluctuation value.

[0078] Specifically, the extended selection unit determines the processing method for each data source based on the tumor data category, including:

[0079] For each type of tumor data in a single data source, the processing method is to determine the first screening method based on the training difficulty coefficient;

[0080] For each type of tumor data in a single data source, the processing method is to determine the second screening method based on the judgment criteria.

[0081] Specifically, the extended selection unit determines the first screening method based on the training difficulty coefficient, including:

[0082] If the training difficulty coefficient is greater than or equal to the preset training difficulty coefficient, the first screening method is determined to be to select tumor data based on the angiography feature value and the rhythm risk degree.

[0083] If the training difficulty coefficient is less than the preset training difficulty coefficient, then the first screening method is determined to be selecting tumor data based on tumor reference values.

[0084] The user can determine the value of the preset training difficulty coefficient according to the actual application scenario. The larger the value of the preset training difficulty coefficient, the greater the user's need to select tumor data based on tumor reference values. The system provides a preset training difficulty coefficient value, detects the historical records of selecting tumor data based on tumor reference values, and records the average value of the training difficulty coefficients corresponding to the historical records that meet the user's needs as the preset training difficulty coefficient.

[0085] When selecting tumor data based on imaging feature values ​​or rhythm risk, the first type 1 tumor data in the first reference sequence is taken as the starting point, and type 1 tumor data in the first reference sequence are selected sequentially at preset intervals as the selected training data. The first preset interval is the number of type 1 tumor data between two adjacent type 1 tumor data in the selected first reference sequence. The first preset interval is positively correlated with the total amount of type 1 tumor data in the data source.

[0086] The method for confirming the first reference sequence is as follows:

[0087] When selecting tumor data based on contrast feature values ​​and rhythm risk, the first reference sequence is a sequence of tumor data of each type in a single data source, sorted in descending order of comprehensive evaluation value. Comprehensive evaluation value = contrast feature value + rhythm risk.

[0088] When selecting tumor data based on tumor reference values, the first reference sequence is a sequence of tumor data of each type in a single data source, sorted in descending order according to the tumor reference values ​​corresponding to the time of first visit.

[0089] Rhythm risk level = Tumor reference value corresponding to the initial consultation time of a single type of tumor data / Rhythm disorder coefficient, Rhythm disorder coefficient = |Peak coefficient corresponding to the initial consultation time of a single type of tumor data - Preset peak coefficient|, The preset peak coefficient is determined by: for the tumor reference value corresponding to the initial consultation time of a single type of tumor data, the tumor reference value is recorded as a1, and the average value of the peak coefficients corresponding to the type of tumor data with the initial consultation time corresponding to a1 in the historical records is recorded as the preset peak coefficient;

[0090] Contrast feature value = peak coefficient + diffusion coefficient. Peak coefficient = signal intensity at peak time / time interval from initial time to peak time. For a single tumor data, the time point corresponding to the maximum signal intensity in the TIC image corresponding to the initial consultation time of the tumor data is recorded as the peak time. The termination time is the maximum time point corresponding to the TIC curve in the TIC image corresponding to the initial consultation time. The initial time is the minimum time point corresponding to the TIC curve in the TIC image corresponding to the initial consultation time. Diffusion coefficient = (area of ​​the first region / time interval from peak time to termination time) - (area of ​​the second region / time interval from initial time to peak time).

[0091] The line that passes through the peak time and is perpendicular to the x-axis is designated as the first line, the line that passes through the termination time and is perpendicular to the x-axis is designated as the second line, and the line that passes through the initial time and is perpendicular to the x-axis is designated as the third line. The area of ​​the first region is the area of ​​the closed region between the first line, the second line, the x-axis, and the TIC curve. The area of ​​the second region is the area of ​​the closed region between the first line, the third line, the x-axis, and the TIC curve.

[0092] Specifically, the methods for determining the training difficulty coefficient include:

[0093] If the heterogeneity correlation is greater than or equal to the preset heterogeneity correlation, the training difficulty coefficient is determined based on the heterogeneity interaction difficulty.

[0094] If the heterogeneity correlation coefficient is less than the preset heterogeneity correlation coefficient, the training difficulty coefficient is determined based on the homogeneity correlation coefficient.

[0095] If the heterogeneity correlation is greater than or equal to the preset heterogeneity correlation, the training difficulty coefficient is the average value of the heterogeneity interaction difficulty corresponding to each type of tumor data in a single data source.

[0096] If the heterogeneity correlation is less than the preset heterogeneity correlation, then the training difficulty coefficient and the homogeneity correlation coefficient are positively correlated.

[0097] The heterogeneous interaction difficulty for a single type I tumor data point from a single data source is the maximum value among the difficulty thresholds corresponding to that type I tumor data point and each type II tumor data point in the same data source. The difficulty threshold for a single type I tumor data point and a single type II tumor data point is calculated as: Hidden Difference Coefficient / ROI Similarity. The Hidden Difference Coefficient is the time interval between the target consultation time and the initial consultation time in the single type I tumor data point. For a single type II tumor data point, the tumor reference value corresponding to the earliest consultation time of that type II tumor data point is denoted as 'a', and the tumor reference value corresponding to the target consultation time of the single type I tumor data point is denoted as 'b'. The ROI Similarity is 1 / |ab|. The method for determining the target consultation time is as follows: for a single type I tumor data point... For tumor data, the deviation coefficient corresponding to each consultation time is detected. The consultation time with the smallest deviation coefficient is recorded as the target consultation time. The deviation coefficient corresponding to a single consultation time is = |tumor reference value corresponding to the consultation time - a|. The tumor reference value corresponding to a single consultation time is = pixel fluctuation coefficient / pixel mean + texture coefficient. The texture coefficient is = (4π×S) / L2, where S is the area of ​​the ROI region in the human brain X-ray image corresponding to a single consultation time, L is the perimeter of the ROI region in the human brain X-ray image corresponding to a single consultation time, the pixel mean is the average value of the pixel values ​​corresponding to each pixel in the ROI region, and the pixel fluctuation coefficient is the standard deviation of the pixel values ​​corresponding to each pixel in the ROI region.

[0098] The correlation coefficient is calculated as 1 / the standard deviation of the tumor reference value corresponding to the first visit time for each type of tumor data in a single data source.

[0099] The method for confirming heterogeneous relevance is as follows: For a single data source, each treatment plan corresponding to each type of tumor data in the data source is recorded as the first record, and each treatment plan corresponding to each type of tumor data in the data source is recorded as the second record. The distribution of the number of keywords in the first and second records is recorded as m1 and m2. Heterogeneous relevance = (the number of identical keywords in the first and second records) / (the larger value between m1 and m2). The preset value of heterogeneous relevance can be determined by the user according to the actual application scenario. The smaller the preset value of heterogeneous relevance, the greater the user's need to determine the training difficulty coefficient based on the heterogeneous interaction difficulty. One preset value of heterogeneous relevance is provided, with a preset heterogeneous relevance of 60%.

[0100] Specifically, the extended selection unit response determination conditions for determining the second filtering method include:

[0101] The extended selection unit response determination condition is that the edge fuzziness coefficient is greater than or equal to the preset edge fuzziness coefficient and the ROI comparison coefficient is less than the preset ROI comparison coefficient. Then the second screening method is to select tumor data based on the associated ROI feature value and the mediating diffusion index.

[0102] The extended selection criteria for unit response are: if the edge fuzziness coefficient is less than the preset edge fuzziness coefficient or the ROI comparison coefficient is greater than or equal to the preset ROI comparison coefficient, then the second screening method is to select tumor data based on the data evaluation value.

[0103] The judgment conditions include a first judgment condition and a second judgment condition. The first judgment condition is that the edge blur coefficient is greater than or equal to the preset edge blur coefficient and the ROI comparison coefficient is less than the preset ROI comparison coefficient. The second judgment condition is that the edge blur coefficient is less than the preset edge blur coefficient or the ROI comparison coefficient is greater than or equal to the preset ROI comparison coefficient.

[0104] The edge blurring coefficient is the average of the sub-blurring coefficients corresponding to each type II tumor data in a single data source. The sub-blurring coefficient is determined as follows: for a single type II tumor data, the human brain X-ray image corresponding to the initial consultation time of the type II tumor data is recorded as the target image, the pixel corresponding to the edge of the ROI region in the target image is recorded as the target pixel, and the pixel with the same pixel value as the target pixel and located outside the ROI region is recorded as the reference pixel. The sub-blurring coefficient = the number of reference pixels / the average distance, where the average distance is the average of the shortest distances from each reference pixel to the ROI region.

[0105] The ROI comparison coefficient is the maximum value of the comparison reference values ​​corresponding to each type II tumor data in a single data source, denoted as d1, and the minimum value, denoted as d2. The ROI comparison coefficient = (d1-d2) / d2. The comparison reference value corresponding to a single type II tumor data is = the tumor reference value corresponding to the first visit time of the type II tumor data / the ROI reference value corresponding to the type II tumor data.

[0106] Users can determine the values ​​of the preset edge blur coefficient and the preset ROI comparison coefficient according to the actual application scenario. The larger the value of the preset edge blur coefficient and the smaller the value of the preset ROI comparison coefficient, the greater the user's need to select tumor data based on the data evaluation value. The system provides a set of preset edge blur coefficient and preset ROI comparison coefficient values, detects the historical records of tumor data selected based on the data evaluation value, and records the average edge blur coefficient corresponding to the historical records that meet the user's needs as the preset edge blur coefficient, and records the average ROI comparison coefficient corresponding to the historical records that meet the user's needs as the preset ROI comparison coefficient.

[0107] The method for confirming the diffusion index and data evaluation value is as follows: for a single type II tumor data, the CA15-3 concentration value at the time of the first visit of the type II tumor data is recorded as k. The type II tumor data with the CA15-3 concentration value at the time of the first visit selected from the historical records is recorded as the analysis data. The diffusion index = 1 / (the average value of the treatment effectiveness corresponding to each analysis data). The data evaluation value = the comparison reference value corresponding to the type II tumor data + the sub-fuzzy coefficient corresponding to the type II tumor data.

[0108] When selecting tumor data based on associated ROI feature values ​​and mediating diffusion index or based on data evaluation values, the first type II tumor data in the second reference sequence is used as the starting point. Type II tumor data in the second reference sequence are selected sequentially at a second preset interval as the selected training data. The second preset interval is the number of type II tumor data between two adjacent type II tumor data in the selected second reference sequence. The second preset interval is positively correlated with the total amount of type II tumor data in the data source.

[0109] The confirmation method for the second reference sequence is as follows:

[0110] When selecting tumor data based on the associated ROI feature value and the mediating diffusion index, the second reference sequence is a sequence of tumor data of each type in a single data source sorted in descending order of feature index. Feature index = associated ROI feature value + mediating diffusion index.

[0111] When selecting tumor data based on data evaluation values, the second reference sequence is a sequence of sorted data for each type II tumor from a single data source in descending order of data evaluation values.

[0112] Specifically, the methods for confirming the associated ROI feature values ​​include:

[0113] The initial characterization interval is determined based on the image characterization value, and the area of ​​the initial characterization interval is increased according to the specificity index of the initial characterization interval. The increased and adjusted initial characterization interval is recorded as the characterization interval.

[0114] The feature value of the associated ROI is determined by the ratio of the mutation coefficient corresponding to the characterization interval to the area of ​​the characterization interval.

[0115] Wherein, image representation value = sub-fuzziness coefficient / preset sub-fuzziness coefficient + ROI reference value / preset ROI reference value; the preset sub-fuzziness coefficient and the preset ROI reference value are confirmed as follows: for a single type II tumor data, the average of the sub-fuzziness coefficients corresponding to each type II tumor data in the data source where the type II tumor data is located is recorded as the preset sub-fuzziness coefficient, and the average of the ROI reference values ​​corresponding to each type II tumor data in the data source where the type II tumor data is located is recorded as the preset ROI reference value;

[0116] The initial characterization interval is determined based on the image characterization value. This includes: recording the human brain X-ray image corresponding to the first visit time of a single tumor data as the analysis image; the initial characterization interval is the area that is larger than the area of ​​the ROI region in the analysis image and similar to the ROI region in the analysis image; the area of ​​the initial characterization interval is positively correlated with the image characterization value; it should be noted that the reference length corresponding to any point on the outline of the initial characterization interval is the same.

[0117] For a single point on the contour of the initially selected representation interval, this point is recorded as the target point. The reference length corresponding to the target point is the length of the line connecting the target point and the reference intersection point. The reference intersection point is the intersection point of the reference line segment and the ROI region. The reference line segment is the line connecting the target point and the reference point. The reference point is the center of the outer circle of the ROI region in the analysis image.

[0118] The increase in the area of ​​a single initial characterization interval is negatively correlated with the specificity index of that initial characterization interval;

[0119] For a single tumor data point, the specificity index of the initial characterization interval is the standard deviation of the modulated region coefficients corresponding to each type of tumor data point in the data source containing that tumor data point; the modulated region coefficient corresponding to a single type of tumor data point is the average pixel value corresponding to each pixel point in the initial characterization interval corresponding to that type of tumor data point.

[0120] For a single tumor dataset, the associated ROI feature value = mutation coefficient corresponding to the representation interval of the tumor dataset / area of ​​the representation interval of the tumor dataset. The mutation coefficient corresponding to the representation interval = average pixel value of each pixel in the representation interval / standard deviation of pixel value of each pixel in the representation interval.

[0121] Specifically, the supplementary selection unit determines representative data sources based on the data source evaluation index, determines representative combinations based on ROI similarity and effective similarity, and determines the supplementary selection method for each representative combination based on the combination comparison coefficient.

[0122] For a single representative combination,

[0123] If the combined comparison coefficient is less than the preset combined comparison coefficient, the supplementary selection method is determined to be to select tumor data based on feature keywords;

[0124] If the combined comparison coefficient is greater than or equal to the preset combined comparison coefficient, then the supplementary selection method is determined to be to select tumor data based on the dynamic difference coefficient.

[0125] Among them, determining the representative data source based on the data source evaluation index includes: selecting the data source with the largest data evaluation index as the representative data source; data source evaluation index = sub-ROI feature coefficient corresponding to a single data source + sub-heterogeneous cross abundance corresponding to a single data source;

[0126] The representative combination is determined based on ROI similarity and effective similarity, including: performing association analysis on each tumor data in the representative data source; when performing association analysis on a single tumor data, the tumor data is recorded as the target data; other tumor data in the representative data source other than the target data are recorded as reference data; the set of reference data with ROI similarity greater than the preset ROI similarity and effective similarity greater than the preset effective similarity, together with the target data, is recorded as a representative combination; and the association analysis continues on the tumor data not recorded in the representative combination until all tumor data are recorded in the representative combination.

[0127] The method for confirming ROI similarity and effective similarity is as follows: for any two tumor data, the larger value of the ROI reference values ​​corresponding to the two tumor data is recorded as r1, and the smaller value is recorded as r2. The ROI similarity = 1 - (r1 - r2) / r1. The larger value of the diagnostic and treatment effectiveness corresponding to the two tumor data is recorded as e1, and the smaller value is recorded as e2. The effective similarity = 1 - (e1 - e2) / e1.

[0128] Users can determine the preset ROI similarity and preset effective similarity values ​​according to the actual application scenario. The greater the user's need for similarity of tumor data in a single representative combination, the greater the preset ROI similarity and preset effective similarity values ​​will be. One preset ROI similarity is 75% and the preset effective similarity is 70%.

[0129] The method for confirming the combination comparison coefficient is as follows: for a single representative combination, the representative combination is recorded as the target representative combination, the treatment plan corresponding to each tumor data in the target representative combination is recorded as the first record, and the treatment plan corresponding to each tumor data in each representative combination other than the target representative combination is recorded as the second record. The combination comparison coefficient = (the number of keywords that appear in both the first and second records) / (the number of different keywords in the first record).

[0130] Tumor data from other data sources besides the representative data source are recorded as candidate data, and candidate data whose similarity threshold with the target representative is greater than the preset similarity threshold are recorded as matching data.

[0131] The similarity threshold between a single candidate data and the target representative combination is 1 / |ROI reference value corresponding to the candidate data - ROI mean| + 1 / |treatment effectiveness corresponding to the candidate data - treatment effectiveness mean|. The average ROI reference value corresponding to each tumor data in the target representative combination is recorded as the ROI mean, and the average treatment effectiveness corresponding to each tumor data in the target representative combination is recorded as the treatment effectiveness mean.

[0132] The user can determine the value of the preset combination comparison coefficient according to the actual application scenario. The smaller the value of the preset combination comparison coefficient, the greater the user's need to select tumor data based on the dynamic difference coefficient. One preset value of the combination comparison coefficient is provided, with the preset combination comparison coefficient being 50%.

[0133] Tumor data is selected based on feature keywords, including: for a single representative combination, the matching data of the treatment plan containing feature keywords are recorded as a type of selected data, and the type of selected data is selected in descending order of the number of feature keywords, until the number of selected data in a type reaches the preset number corresponding to the representative combination; the number of feature keywords is the total number of feature keywords contained in each treatment plan corresponding to a single type of selected data.

[0134] Tumor data is selected based on the dynamic difference coefficient, including: matching data whose dynamic difference coefficient with the target representative combination is greater than the preset dynamic difference coefficient is recorded as second-class selection data, and second-class selection data is selected in descending order of dynamic difference coefficient until the number of second-class selection data selected reaches the preset number corresponding to the representative combination.

[0135] The preset quantity corresponding to a single representative combination = the pre-selected quantity corresponding to the representative combination - the number of tumor data in the representative combination. The pre-selected quantity corresponding to a single representative combination is positively correlated with the proportion coefficient corresponding to the representative combination. The proportion coefficient corresponding to a single representative combination = the number of tumor data in the representative combination / the total amount of tumor data in the representative data source.

[0136] The dynamic difference coefficient between a single matched data point and the target representative combination is the average of the fluctuation differences corresponding to each tumor data point in the matched data point and the target representative combination. The fluctuation difference corresponding to a single tumor data point in the target representative combination is equal to |the treatment fluctuation value corresponding to the matched data point - the treatment fluctuation value corresponding to a single tumor data point in the target representative combination|.

[0137] Users can determine the values ​​of the preset similarity threshold and the preset dynamic difference coefficient according to the actual application scenario. The greater the user's need to improve the correlation of data selection, the larger the value of the preset similarity threshold and the smaller the value of the preset dynamic difference coefficient. One preset similarity threshold and preset dynamic difference coefficient are provided. The preset similarity threshold is 85%. The historical records of tumor data selected based on the dynamic difference coefficient are detected, and the average value of the dynamic difference coefficients corresponding to each type of selected data in the historical records that can meet the user's needs is recorded as the preset dynamic difference coefficient.

[0138] Please see Figure 4 As shown, this is a schematic diagram of the human brain tumor prediction model training method of the present invention. The present invention also provides a human brain tumor prediction model training method, comprising:

[0139] Tumor data are acquired from several data sources. The data source status is determined based on heterogeneous crossover abundance and ROI characteristic coefficients. The data selection method is determined based on the data source status, which can be either extended selection from multiple data sources or supplementary selection from representative data sources.

[0140] In the extended selection of multiple data sources, the processing method of each data source is determined according to the tumor data category. The processing methods include determining the first screening method based on the training difficulty coefficient and determining the second screening method based on the judgment criteria.

[0141] The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values.

[0142] In the supplementary selection of representative data sources, representative data sources are determined based on the data source evaluation index, representative combinations are determined based on ROI similarity and effective similarity, and the supplementary selection method for each representative combination is determined based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient.

[0143] The selected tumor data was used as training data to train the tumor prediction model.

[0144] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0145] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A training system for a human brain tumor prediction model, characterized in that, include: The data acquisition unit is used to acquire tumor data from several data sources; An analysis unit is selected, which is connected to the data acquisition unit, to determine the data source status based on the heterogeneous cross abundance and ROI characteristic coefficients, and to determine the data selection method based on the data source status. The data selection method is either extended selection from multiple data sources or supplementary selection from representative data sources. An extended selection unit, which is connected to the data acquisition unit and the selection analysis unit respectively, is used to determine the processing method of each data source according to the tumor data category in the extended selection of multiple data sources. The processing method includes determining a first screening method according to the training difficulty coefficient and determining a second screening method according to the judgment conditions. The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values. The supplementary selection unit is connected to the data acquisition unit and the selection analysis unit respectively. It is used to determine the representative data source based on the data source evaluation index, determine the representative combination based on ROI similarity and effective similarity, and determine the supplementary selection method of each representative combination based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient. The model training unit is connected to the extended selection unit and the supplementary selection unit respectively, and is used to train the tumor prediction model using the selected tumor data as training data. Heterogeneous cross abundance is the maximum value among the sub-heterogeneous cross abundances corresponding to each data source. The sub-heterogeneous cross abundance corresponding to the target data source = the number of identical keywords / the total number of keywords appearing in the treatment plan corresponding to each tumor data in each data source. The treatment plan corresponding to each tumor data in each reference data source is recorded as a reference record, and the treatment plan corresponding to each tumor data in the target data source is recorded as a target record. The number of identical keywords is the number of keywords that appear in both the target record and the reference record. The ROI feature coefficient is the maximum value among the sub-ROI feature coefficients corresponding to each data source. The sub-ROI feature coefficient corresponding to a single data source = ROI fluctuation coefficient + treatment fluctuation coefficient. The ROI fluctuation coefficient is the standard deviation of the ROI reference value corresponding to each tumor data in a single data source. The ROI reference value corresponding to a single tumor data is the area of ​​the ROI region in the human brain X-ray image taken at the patient's first consultation time in the tumor data. The treatment fluctuation coefficient is the standard deviation of the treatment effectiveness corresponding to each tumor data in a single data source. The treatment effectiveness corresponding to a single tumor data = CA15-3 concentration value corresponding to the earliest treatment data of the patient in the tumor data - CA15-3 concentration value corresponding to the latest treatment data of the patient in the tumor data. Contrast characteristic value = peak coefficient + diffusion coefficient, peak coefficient = signal strength at peak time / time interval from initial time to peak time, diffusion coefficient = (area of ​​first region / time interval from peak time to termination time) - (area of ​​second region / time interval from initial time to peak time). For a single tumor data point, the time point corresponding to the maximum signal intensity in the TIC image at the initial consultation time of the tumor data is recorded as the peak time. The termination time is the maximum time point of the TIC curve in the TIC image at the initial consultation time. The initial time is the minimum time point of the TIC curve in the TIC image at the initial consultation time. The straight line that passes through the peak time and is perpendicular to the x-axis is recorded as the first straight line. The straight line that passes through the termination time and is perpendicular to the x-axis is recorded as the second straight line. The straight line that passes through the initial time and is perpendicular to the x-axis is recorded as the third straight line. The area of ​​the first region is the area of ​​the closed region between the first straight line, the second straight line, the x-axis, and the TIC curve. The area of ​​the second region is the area of ​​the closed region between the first straight line, the third straight line, the x-axis, and the TIC curve. Rhythm risk score = Tumor reference value corresponding to the initial consultation time of a single type of tumor data / Rhythm disorder coefficient, Rhythm disorder coefficient = |Peak coefficient corresponding to the initial consultation time of a single type of tumor data - Preset peak coefficient|; The methods for confirming the associated ROI feature values ​​include: The initial characterization interval is determined based on the image characterization value, and the area of ​​the initial characterization interval is increased according to the specificity index of the initial characterization interval. The increased and adjusted initial characterization interval is recorded as the characterization interval. The feature value of the associated ROI is determined by the ratio of the mutation coefficient corresponding to the characterization interval to the area of ​​the characterization interval. Mediated diffusion index = 1 / (average of the diagnostic and treatment effectiveness corresponding to each analysis data); The method for confirming ROI similarity and effective similarity is as follows: for any two tumor data, the larger value of the ROI reference values ​​corresponding to the two tumor data is recorded as r1, and the smaller value is recorded as r2. The ROI similarity = 1 - (r1 - r2) / r1. The larger value of the diagnostic and treatment effectiveness corresponding to the two tumor data is recorded as e1, and the smaller value is recorded as e2. The effective similarity = 1 - (e1 - e2) / e1. For a single representative combination, the representative combination is recorded as the target representative combination. The treatment plan corresponding to each tumor data in the target representative combination is recorded as the first record. The treatment plan corresponding to each tumor data in each representative combination other than the target representative combination is recorded as the second record. The combination comparison coefficient = (the number of keywords that appear in both the first and second records) / (the number of different keywords in the first record). The dynamic difference coefficient between a single matched data point and the target representative combination is the average of the fluctuation differences corresponding to each tumor data point in the matched data point and the target representative combination. The fluctuation difference corresponding to a single tumor data point in the matched data point and the target representative combination is equal to |the treatment fluctuation value corresponding to the matched data point - the treatment fluctuation value corresponding to a single tumor data point in the target representative combination|.

2. The human brain tumor prediction model training system according to claim 1, characterized in that, The selection analysis unit responds to the data source status to determine the data selection method, including: If the data source status of the selected analysis unit response is that the heterogeneous cross abundance is less than the preset heterogeneous cross abundance or the ROI feature coefficient is less than the preset ROI feature coefficient, then the data selection method is determined to be multi-data source extended selection. If the data source status of the selected analysis unit response is heterogeneous cross abundance greater than or equal to the preset heterogeneous cross abundance and ROI feature coefficient greater than or equal to the preset ROI feature coefficient, then the data selection method is determined to be supplementary selection of representative data sources.

3. The human brain tumor prediction model training system according to claim 2, characterized in that, The extended selection unit determines the tumor data category based on the effectiveness of diagnosis and treatment and the fluctuation value of diagnosis and treatment. The tumor data categories include: Tumor data that has a treatment effectiveness greater than or equal to the preset treatment effectiveness and a treatment fluctuation value greater than or equal to the preset treatment fluctuation value; Type II tumor data: treatment effectiveness less than the preset treatment effectiveness or treatment fluctuation value less than the preset treatment fluctuation value.

4. The human brain tumor prediction model training system according to claim 3, characterized in that, The extended selection unit determines the processing method for each data source based on the tumor data category, including: For each type of tumor data in a single data source, the processing method is to determine the first screening method based on the training difficulty coefficient; For each type of tumor data in a single data source, the processing method is to determine the second screening method based on the judgment criteria.

5. The human brain tumor prediction model training system according to claim 4, characterized in that, The extended selection unit determines the first screening method based on the training difficulty coefficient, including: If the training difficulty coefficient is greater than or equal to the preset training difficulty coefficient, the first screening method is determined to be to select tumor data based on the angiography feature value and the rhythm risk degree. If the training difficulty coefficient is less than the preset training difficulty coefficient, then the first screening method is determined to be selecting tumor data based on tumor reference values.

6. The human brain tumor prediction model training system according to claim 5, characterized in that, The methods for confirming the training difficulty coefficient include: If the heterogeneity correlation is greater than or equal to the preset heterogeneity correlation, the training difficulty coefficient is determined based on the heterogeneity interaction difficulty. If the heterogeneity correlation is less than the preset heterogeneity correlation, the training difficulty coefficient is determined based on the homogeneity correlation coefficient. Heterogeneous relevance = (number of identical keywords in the first and second records) / (larger value between m1 and m2), where m1 and m2 are the number of keywords in the first and second records, respectively. The heterogeneous interaction difficulty for a single type 1 tumor data in a single data source is the maximum value among the difficulty thresholds corresponding to the type 1 tumor data and each type 2 tumor data in the same data source. The difficulty threshold corresponding to a single type 1 tumor data and a single type 2 tumor data is = hidden difference coefficient / ROI similarity. The hidden difference coefficient is the time interval between the target consultation time and the initial consultation time in a single type 1 tumor data. The correlation coefficient is calculated as 1 / the standard deviation of the tumor reference value corresponding to the initial consultation time of each type of tumor data in a single data source.

7. The human brain tumor prediction model training system according to claim 4, characterized in that, The extended selection unit response determination conditions for determining the second filtering method include: The extended selection unit response determination condition is that the edge fuzziness coefficient is greater than or equal to the preset edge fuzziness coefficient and the ROI comparison coefficient is less than the preset ROI comparison coefficient. Then the second screening method is to select tumor data based on the associated ROI feature value and the mediating diffusion index. The extended selection criteria for unit response are: if the edge fuzziness coefficient is less than the preset edge fuzziness coefficient or the ROI comparison coefficient is greater than or equal to the preset ROI comparison coefficient, then the second screening method is to select tumor data based on the data evaluation value. The edge blur coefficient is the average of the sub-blur coefficients corresponding to each type of tumor data in a single data source. The sub-blur coefficient = number of reference pixels / average distance, where the average distance is the average of the shortest distances from each reference pixel to the ROI region. ROI comparison coefficient = (d1-d2) / d2, the comparison reference value corresponding to a single type II tumor data = the tumor reference value corresponding to the initial consultation time of the type II tumor data / the ROI reference value corresponding to the type II tumor data, where d1 and d2 are the maximum and minimum values ​​of the comparison reference values ​​corresponding to each type II tumor data in a single data source.

8. The human brain tumor prediction model training system according to claim 2, characterized in that, The supplementary selection unit determines representative data sources based on the data source evaluation index, determines representative combinations based on ROI similarity and effective similarity, and determines the supplementary selection method for each representative combination based on the combination comparison coefficient. For a single representative combination, If the combined comparison coefficient is less than the preset combined comparison coefficient, the supplementary selection method is determined to be to select tumor data based on feature keywords; If the combined comparison coefficient is greater than or equal to the preset combined comparison coefficient, then the supplementary selection method is determined to be to select tumor data based on the dynamic difference coefficient.

9. A training method for a human brain tumor prediction model training system according to any one of claims 1 to 8, characterized in that, include: Tumor data are acquired from several data sources. The data source status is determined based on heterogeneous crossover abundance and ROI characteristic coefficients. The data selection method is determined based on the data source status, which can be either extended selection from multiple data sources or supplementary selection from representative data sources. In the extended selection of multiple data sources, the processing method of each data source is determined according to the tumor data category. The processing methods include determining the first screening method based on the training difficulty coefficient and determining the second screening method based on the judgment criteria. The first screening method is to select tumor data based on contrast feature values ​​and rhythm risk, or to select tumor data based on tumor reference values. The second screening method is to select tumor data based on associated ROI feature values ​​and mediating diffusion index, or to select tumor data based on data evaluation values. In the supplementary selection of representative data sources, representative data sources are determined based on the data source evaluation index, representative combinations are determined based on ROI similarity and effective similarity, and the supplementary selection method for each representative combination is determined based on the combination comparison coefficient. The supplementary selection method is to select tumor data based on feature keywords or dynamic difference coefficient. The selected tumor data was used as training data to train the tumor prediction model.

Citation Information

Patent Citations

  • Tumor chemotherapy effect prediction model construction method and system based on artificial intelligence

    CN119069124A

  • Tumor recurrence risk assessment and early warning system based on data analysis

    CN119361146A