A data processing method and device, electronic equipment and storage medium
By updating the decision tree model and optimizing the assessment of lesion imaging features using sample datasets, the problem of inaccurate lesion feature description in existing models is solved, thus improving the efficiency of early cancer screening.
Patent Information
- Application Number
- CN202310564606.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-05-18
AI Technical Summary
Existing CT image-based models for assessing the probability of malignancy of lesions lack accurate descriptions of lesion characteristics and cannot effectively assist in early cancer screening.
By updating the decision tree model and optimizing the first decision tree model using the sample dataset, a more accurate assessment of lesion imaging features can be obtained, until assessment data superior to the original model is obtained.
It enables a more accurate description of the imaging characteristics of lesions, improving the efficiency of early cancer screening.
Smart Images

Figure CN116563670B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a data processing method, apparatus, electronic device and storage medium. Background Technology
[0002] Currently, lesion malignancy probability assessment models based on computed tomography (CT) images have become one of the main methods for early cancer screening; however, the models in related technologies only obtain relevant features of lesions from CT images and lack more accurate descriptive models of lesion features, which cannot more efficiently assist doctors in early cancer screening. Summary of the Invention
[0003] This disclosure provides a data processing method, apparatus, electronic device, and storage medium, and provides a method for updating a decision tree model. The decision tree model obtained based on the method provided in this disclosure can more accurately describe the epidemiological characteristics of lesions and more efficiently assist doctors in early cancer screening.
[0004] According to a first aspect of this disclosure, a data processing method is provided, comprising:
[0005] The lesion imaging features corresponding to at least one CT image result of each object in the first dataset are input into the first decision tree model to obtain the first prediction score result;
[0006] In response to the first predicted score result including at least one negative sample, or the preset update condition being met, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model.
[0007] Based on the sample dataset, obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model respectively;
[0008] If the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, then the second decision tree model is confirmed as the target decision tree model; or, if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, then the first decision tree model is updated again based on the at least one negative sample and the sample dataset, and evaluation data is obtained until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model.
[0009] According to a second aspect of this disclosure, a data processing apparatus is provided, comprising:
[0010] The acquisition unit is used to input the lesion imaging features corresponding to at least one CT image result of each object in the first dataset into the first decision tree model to obtain the first prediction score result;
[0011] The update unit is used to update the first decision tree model based on the sample dataset to obtain a second decision tree model in response to the first prediction score result including at least one negative sample or meeting a preset update condition.
[0012] The evaluation unit is used to obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model based on the sample dataset.
[0013] The iteration unit is used to confirm the second decision tree model as the target decision tree model if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model; or, if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, it is used to update the first decision tree model again based on the at least one negative sample and the sample dataset, and obtain evaluation data until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model.
[0014] According to a third aspect of this disclosure, an electronic device is provided, comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0018] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described in this disclosure.
[0019] The data processing method disclosed herein involves inputting the lesion imaging features corresponding to at least one CT image result for each object in a first dataset into a first decision tree model to obtain a first predicted score result. In response to the first predicted score result including at least one negative sample, the first decision tree model is updated based on the at least one negative sample and the sample dataset to obtain a second decision tree model. Evaluation data corresponding to the first decision tree model and the second decision tree model are obtained based on the sample dataset. If the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, the second decision tree model is confirmed as the target decision tree model. Alternatively, if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, the first decision tree model is updated again based on the at least one negative sample and the sample dataset, and evaluation data is obtained, until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model. In this way, a decision tree model can be obtained, which can obtain a more accurate description of the lesion based on the lesion imaging features obtained from CT images, i.e., lesion-related scoring results, and can more efficiently assist doctors in early cancer screening.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0022] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0023] Figure 1 This illustration shows an optional flowchart of a data processing method provided in an embodiment of the present disclosure;
[0024] Figure 2 A schematic diagram of another optional flow of the data processing method provided in an embodiment of this disclosure is shown;
[0025] Figure 3 An optional schematic diagram of the decision tree model provided in an embodiment of this disclosure is shown;
[0026] Figure 4 A schematic diagram of another optional flow of the data processing method provided in an embodiment of this disclosure is shown;
[0027] Figure 5 Another optional schematic diagram of the decision tree model provided in this disclosure embodiment is shown;
[0028] Figure 6 A schematic diagram of an optional structure of the data processing apparatus provided in an embodiment of this disclosure is shown;
[0029] Figure 7 A schematic diagram of another optional structure of the data processing apparatus provided in this disclosure embodiment is shown;
[0030] Figure 8 An optional schematic diagram of a user interaction page provided in an embodiment of this disclosure is shown;
[0031] Figure 9 Another optional schematic diagram of the user interaction page provided in the embodiments of this disclosure is shown;
[0032] Figure 10 A schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0033] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0034] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0035] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0036] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.
[0037] It should be understood that in the various embodiments of this disclosure, the sequence number of each implementation process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.
[0038] Figure 1 An optional flowchart of the data processing method provided in an embodiment of this disclosure is shown, and the steps will be described accordingly.
[0039] Step S101: Input the lesion imaging features corresponding to at least one CT image result of each object in the first dataset into the first decision tree model to obtain the first prediction score result.
[0040] In some embodiments, the first dataset may be an online dataset, i.e., CT images acquired in real time. It should be noted that the CT images in the first dataset are all of the same type, such as all lung CT images, all brain CT images, all chest CT images, or all liver CT images, etc.
[0041] In some embodiments, the first dataset may include a single CT image of an object, or at least two CT images of an object, wherein the at least two CT images were acquired at different times. That is, the images in the first dataset include CT images of multiple objects, with a one-to-one correspondence between CT images and objects (i.e., one object corresponds to only one CT image); or, there may be a many-to-one correspondence between CT images and objects (i.e., one object corresponds to multiple CT images).
[0042] In some embodiments, the predicted scoring result may include a predicted classification result and a predicted score for the CT image. The predicted classification result includes the classification of the lesion corresponding to the CT image; for example, if the CT image is a liver CT image, the classification of the corresponding lesion may include one of the following: negative and positive; or, if the CT image is a chest CT image, the classification of the corresponding lesion may include the BIRADS classification. The predicted score may include a score for the lesion, and the branch is used to characterize the degree of the lesion.
[0043] In some embodiments, the decision tree model (including the first decision tree model, the second decision tree model, and other updated models) includes multiple branches, each branch including different lesion imaging feature classifications and lesion imaging feature values; each branch may include a lesion imaging feature classification and corresponding lesion imaging feature values, and the branches may be connected. For example, taking lung CT images as an example, the lesion long diameter branch may be followed by the lesion volume branch, and the lesion volume branch may be followed by the lesion density branch.
[0044] In some embodiments, the decision tree model can be built based on clinical guidelines (NCCN) and lesion imaging features. During the prediction phase, the model will proceed to the corresponding decision branch of the scoring decision tree according to the importance of each lesion imaging feature, and then the nodule features will be weighted to calculate the score. If multiple CT scans are performed (e.g., two CT scans), features such as changes in lesion volume, changes in the proportion of solid components, and lesion volume doubling time will also be added to the calculation of the decision tree.
[0045] In step S102, in response to the first prediction score result including at least one negative sample, or the preset update condition being met, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model.
[0046] In some embodiments, the preset update conditions may include satisfying a first period, that is, updating the first decision tree model based on the sample dataset every first period.
[0047] In some embodiments, the negative samples include data from the first predicted scoring results where the predicted scoring result does not match the lesion imaging features. The negative samples represent the lesion imaging features, i.e., the lesion imaging features corresponding to the mismatched predicted scoring result, or CT images. The mismatch between the predicted scoring result and the lesion imaging features may include: the predicted scoring result cannot be obtained based on the lesion imaging features. For example, the result corresponding to the lesion imaging features is healthy, but the predicted scoring result is unhealthy.
[0048] In some embodiments, if the first predicted score includes at least one negative sample, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model.
[0049] In some embodiments, the data in the sample dataset includes pre-collected sample data and negative samples generated during the period from the last update of the decision tree model to the update of the first decision tree model.
[0050] For example, if the decision tree model is updated according to preset update conditions, then the negative samples generated in the first period after the last update of the decision tree model are added to the sample dataset to update the decision tree model; or, if the decision tree model is updated based on the number of negative samples, then the newly added negative samples are added to the sample dataset to update the decision tree model.
[0051] Step S103: Based on the sample dataset, obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model.
[0052] In some embodiments, after obtaining the second decision tree model, in order to confirm the update effect, the data processing method implementation carrier (hereinafter referred to as the carrier) evaluates the first decision tree model and the second decision tree model respectively, obtains the corresponding evaluation data, and confirms the subsequent process based on the comparison results of the evaluation data.
[0053] In specific implementation, the evaluation data includes at least one of sensitivity data, specificity data, and Youden index; the carrier inputs the sample dataset into a first decision tree model to obtain the evaluation data corresponding to the first decision tree model; inputs the sample dataset into a second decision tree model to obtain the evaluation data corresponding to the second decision tree model; and compares at least one of sensitivity data, specificity data, and Youden index to confirm the superiority of the first decision tree model and the second decision tree model.
[0054] In some embodiments, if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, then step S104 is executed; if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, then the first decision tree model is updated again based on the at least one negative sample and the sample dataset, i.e., step S102 is executed again, and evaluation data is obtained until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model.
[0055] Step S104: Confirm that the second decision tree model is the target decision tree model.
[0056] In some embodiments, if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, it indicates that the update operation is effective. Subsequently, the imaging features of the lesions corresponding to the CT impact are processed based on the second decision tree model to obtain the prediction score result. The prediction score result is used to assist doctors in early screening and cannot be used as a direct diagnosis result or treatment suggestion.
[0057] Thus, according to the data processing method provided in this disclosure, a method for updating a decision tree model is provided. During the decision tree model update process, a first dataset (data collected online) is first input into the first decision tree model for verification to obtain a first prediction score. If negative samples exist or a first cycle is met, the first decision tree model is updated based on the sample dataset (which may include negative samples). After the update, the first and second decision tree models are evaluated to determine the optimal model for use. The final decision tree model can obtain a more accurate description of lesions based on the lesion imaging features obtained from CT images, i.e., lesion-related scoring results, which can more efficiently assist doctors in early cancer screening.
[0058] Figure 2 A schematic diagram of another optional flow of the data processing method provided in the embodiments of this disclosure is shown, and will be described according to each step.
[0059] Step S201: Obtain the lesion imaging features corresponding to the CT images in the first dataset based on the lesion prediction model.
[0060] In some embodiments, the first dataset may be an online dataset, i.e., CT images acquired in real time. It should be noted that the CT images in the first dataset are all of the same type, such as all lung CT images, all brain CT images, all chest CT images, or all liver CT images, etc.
[0061] In some embodiments, the first dataset may include a single CT image of an object, i.e., the images in the first dataset include CT images of multiple objects, and there is a one-to-one correspondence between the CT images and the objects (i.e., one object corresponds to only one CT image).
[0062] In some embodiments, the lesion imaging model takes CT images as input and outputs lesion imaging features corresponding to the CT images. Different types of lesions (i.e., different organs) correspond to different lesion imaging models. Taking lung CT images as an example, the lesion imaging features may include the Mayo lung cancer model, the Brock lung cancer model, or other models in related technologies that can obtain lesion imaging features based on CT images. This disclosure does not make any specific limitations.
[0063] In some embodiments, the output of the lesion imaging model may also include lesion phenotype probabilities (or model scoring results). The model scoring results are used to characterize the severity of the lesion; for example, a low score indicates less severe lesion than a high score.
[0064] Step S202: Input the lesion imaging features corresponding to at least one CT image result of each object in the first dataset into the first decision tree model to obtain the first prediction score result.
[0065] In some embodiments, the predicted scoring result may include a predicted classification result and a predicted score for the CT image. The predicted classification result includes the classification of the lesion corresponding to the CT image; for example, if the CT image is a liver CT image, the classification of the corresponding lesion may include one of the following: negative and positive; or, if the CT image is a chest CT image, the classification of the corresponding lesion may include the BIRADS classification. The predicted score may include a score for the lesion, and the branch is used to characterize the degree of the lesion.
[0066] In some embodiments, the decision tree model (including the first decision tree model, the second decision tree model, and other updated models) includes multiple branches, each branch including different lesion imaging feature classifications and lesion imaging feature values; each branch may include a lesion imaging feature classification and corresponding lesion imaging feature values, and the branches may be connected. For example, taking lung CT images as an example, the lesion long diameter branch may be followed by the lesion volume branch, and the lesion volume branch may be followed by the lesion density branch.
[0067] Figure 3 An optional schematic diagram of the decision tree model provided in an embodiment of this disclosure is shown. It should be noted that... Figure 3 The decision tree model provided is merely an example and is not intended to limit the scope of protection of this disclosure. Any decision tree model can be updated using the data processing methods provided in this disclosure.
[0068] like Figure 3 As shown, the decision tree model can be used for a single CT image. The first-level branch includes the model score result, as well as the age and model score result. The second-level branch (i.e., the branch of the first-level branch) includes the lesion size value. The output of the second-level branch is the predicted score result. The lesion classification includes positive and negative, and the predicted score includes the predicted grade and corresponding score. For example, a negative score is 0 to 30, and a positive grade 1 has scores of 30 to 40, 30 to 50, 50 to 60, etc.
[0069] In some embodiments, the carrier inputs the lesion imaging features corresponding to a single CT image result of each object in the first dataset into a first decision tree model. Based on each branch of the first decision tree model, the score and weight of each feature in the lesion imaging features are determined. The carrier performs a weighted summation on each feature included in the lesion imaging features to obtain a first prediction score result corresponding to each object in the first dataset.
[0070] Alternatively, in some embodiments, the carrier inputs the lesion imaging features corresponding to a single CT image result of each object in the first dataset into the first decision tree model. Based on different branch conditions in the first decision tree model, the lesion imaging features that meet the branch conditions are input into the branch, and finally the first prediction score result of each object is obtained.
[0071] Step S203: In response to the first prediction score result including at least one negative sample, or the preset update condition being met, the first decision tree model is updated based on the sample dataset to obtain the second decision tree model.
[0072] In some embodiments, the preset update conditions may include satisfying a first period, that is, updating the first decision tree model based on the sample dataset every first period.
[0073] In some embodiments, the negative samples include data in the first predicted score results where the predicted score results do not match the lesion imaging features. The negative samples are represented by lesion imaging features, that is, lesion imaging features corresponding to the mismatched predicted score results, or CT images.
[0074] In some embodiments, if the first predicted score includes at least one negative sample, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model.
[0075] In some embodiments, the data in the sample dataset includes pre-collected sample data and negative samples generated during the period from the last update of the decision tree model to the update of the first decision tree model.
[0076] For example, if the decision tree model is updated according to preset update conditions, then the negative samples generated in the first period after the last update of the decision tree model are added to the sample dataset to update the decision tree model; or, if the decision tree model is updated based on the number of negative samples, then the newly added negative samples are added to the sample dataset to update the decision tree model.
[0077] In some embodiments, the carrier updates the lesion prediction model based on a sample dataset. The lesion prediction model is used to obtain the lesion imaging features corresponding to at least one CT image result of each object in the first dataset; update the condition values of the corresponding branches in the first decision tree model based on the values of the lesion imaging features in the sample dataset; update the weights and scores of the corresponding branches in the first decision tree model based on the lesion imaging features in the sample dataset; and add corresponding branches in the first decision tree model based on the clinical information corresponding to the object.
[0078] In practice, the carrier can update the lesion prediction model based on the sample dataset. Since the CT image to the final prediction score requires passing through two models—the lesion prediction model and the decision tree model—to ensure data accuracy and avoid resource waste (the lesion prediction model has problems, yet the decision tree model is repeatedly updated), the carrier updates the lesion prediction model based on the sample dataset. After the lesion prediction model is updated, the radiographic features of the lesions corresponding to the CT images in the first dataset are re-acquired, and these re-acquired radiographic features are input into the first decision tree model to obtain the first prediction score. That is, the problems with the lesion prediction model are first updated and eliminated, and then the decision tree model is updated.
[0079] In specific implementation, the carrier updates the condition values of the corresponding branches in the first decision tree model based on the numerical values of the lesion imaging features in the sample dataset. This includes updating the condition values of the corresponding branches in the first decision tree model based on the numerical values of lesion imaging features of the same type or severity in the sample dataset. For example, taking size as an example, the first decision tree model sets a size threshold of 6, that is, it distinguishes between positive level 1 and negative based on whether the size is greater than or equal to 6 or less than 6. However, in the sample dataset, samples with a size greater than 6 and less than 8 are negative. Therefore, the size threshold in the first decision tree model can be updated to 8, where less than 8 is negative and greater than or equal to 8 is positive.
[0080] In specific implementation, the carrier updates the weights and scores of corresponding branches in the first decision tree model based on the lesion imaging features in the sample dataset, including adjusting the weights and scores of different branches in the first decision tree model. For example, the weights of features with less impact on the result are appropriately reduced, while the weights of features with greater impact are appropriately increased, to avoid the final prediction score being affected by features with smaller impact on the result due to excessively high scores and weights.
[0081] In practice, the carrier adds corresponding branches to the first decision tree model based on the clinical information corresponding to the object; wherein, the clinical information may include age, health history, smoking history, etc. To ensure the accuracy of the results, the carrier may add branches corresponding to the clinical information to the first decision tree model.
[0082] Step S204: Based on the sample dataset, obtain the first evaluation data corresponding to the first decision tree model and the second evaluation data corresponding to the second decision tree model.
[0083] In some embodiments, the carrier inputs the sample dataset into a first decision tree model to obtain evaluation data corresponding to the first decision tree model; and inputs the sample dataset into a second decision tree model to obtain evaluation data corresponding to the second decision tree model; wherein the evaluation data includes at least one of sensitivity data, specificity data, and Youden index.
[0084] In some embodiments, the sample confirms whether the first evaluation data is better than the second evaluation data; if the first evaluation data is better than the second evaluation data, it indicates that the update is invalid and step S203 is repeated; or, if the second evaluation data is better than the first evaluation data, it indicates that the update is valid and step S205 is executed.
[0085] Step S205: Confirm that the second decision tree model is the target decision tree model.
[0086] In some embodiments, if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, it indicates that the update operation is effective. Subsequently, the imaging features of the lesions corresponding to the CT impact are processed based on the second decision tree model to obtain the prediction score result. The prediction score result is used to assist doctors in early screening and cannot be used as a direct diagnosis result or treatment suggestion.
[0087] Thus, according to the data processing method provided in this disclosure, a method for updating a decision tree model is provided. During the decision tree model update process, a first dataset (data collected online) is first input into the first decision tree model for verification to obtain a first prediction score. If negative samples exist or a first cycle is met, the first decision tree model is updated based on the sample dataset (which may include negative samples). After the update, the first and second decision tree models are evaluated to determine the optimal model for use. The final decision tree model can obtain a more accurate description of lesions based on the lesion imaging features obtained from CT images, i.e., lesion-related scoring results, which can more efficiently assist doctors in early cancer screening.
[0088] Figure 4 A schematic diagram of another optional flow of the data processing method provided in the embodiments of this disclosure is shown, and will be described in accordance with each step.
[0089] Step S401: Obtain the lesion imaging features corresponding to the CT images in the first dataset based on the lesion prediction model.
[0090] In some embodiments, the first dataset may be an online dataset, i.e., CT images acquired in real time. It should be noted that the CT images in the first dataset are all of the same type, such as all lung CT images, all brain CT images, all chest CT images, or all liver CT images, etc.
[0091] In some embodiments, the first dataset includes a first data subset and a second data subset. The first dataset includes at least two CT images of an object, and the at least two CT images are acquired at different times, i.e., there is a many-to-one relationship between CT images and objects in the first dataset (i.e., one object corresponds to multiple CT images). To facilitate subsequent operations, CT images of the same object acquired at different times are set in different data subsets, and the acquisition time of the CT images in the first data subset is set to be earlier than the acquisition time of the CT images in the second data subset. It should be noted that the acquisition interval of CT images of different objects in the data subsets can be the same or different, because the purpose of introducing two subsets in this embodiment is to compare lesion changes at different times, which is unrelated to the acquisition interval.
[0092] For example, the first data subset includes CT images of object A acquired at time t1, and the corresponding second data subset includes CT images of object A acquired at time t2, where t1 is earlier than t2.
[0093] In some alternative embodiments, the first dataset may further include at least two subsets of data, the acquisition time of the CT images being different from that of the first and second subsets of data.
[0094] In some embodiments, the carrier inputs CT images from a first data subset and a second data subset into the lesion prediction model to obtain the lesion imaging features corresponding to the CT image results in the first data subset and the lesion imaging features corresponding to the CT image results in the second data subset, respectively.
[0095] Step S402: Input the lesion imaging features corresponding to at least two CT image results for each object into the first decision tree model to obtain the first prediction score result.
[0096] In some embodiments, the carrier inputs the lesion imaging features corresponding to at least two CT image results for each object into a first decision tree model to obtain a first prediction score result.
[0097] In specific implementation, the carrier can acquire at least one variation parameter based on the lesion imaging features corresponding to the CT image results in the first data subset and the lesion imaging features corresponding to the CT image results in the second data subset. The variation parameter can be obtained based on two CT image results and is used to characterize the changing trend of the lesion imaging features. The variation parameter can include at least one of solid component ratio (CTR) and volume doubling time (VDT). Variation parameter classification refers to the type of variation parameter, such as CTR or VDT.
[0098] In some embodiments, the decision tree model (including the first decision tree model, the second decision tree model, and other updated models) includes multiple branches. Each branch includes different lesion imaging feature classifications, lesion imaging feature values, variable parameter classifications, and variable parameter values. Each branch may include a lesion imaging feature classification and the corresponding lesion imaging feature value, or a variable parameter classification and the corresponding variable parameter value. Branches can be connected. For example, taking lung CT images as an example, the lesion size branch can be followed by the nodule type branch, and the nodule type branch can be followed by the VDT branch or the CTR branch.
[0099] Figure 5 Another optional schematic diagram of the decision tree model provided in this embodiment of the disclosure is shown. It should be noted that... Figure 5The decision tree model provided is merely an example and is not intended to limit the scope of protection of this disclosure. Any decision tree model can be updated using the data processing methods provided in this disclosure.
[0100] like Figure 5 As shown, the decision tree model can be used for multiple CT images. After the lesion size branch, a nodule type branch can be added. Different nodule types correspond to different branches; specifically, solid nodules correspond to the VDT branch, and ground-glass nodules correspond to the CTR branch. Lesion classification includes positive and negative, and the predicted score includes a predicted grade and a corresponding score. For example, a negative score is 0 to 30, and a positive grade 1 has scores of 30 to 40, 30 to 50, 50 to 60, etc. Optionally, after determining the lesion classification, a lesion trait probability (model scoring result) branch can also be added.
[0101] In some embodiments, the carrier inputs the lesion imaging features and change parameters corresponding to at least two CT image results of each object in the first dataset into a first decision tree model. Based on each branch of the first decision tree model, the carrier confirms the score and weight of each feature in the lesion imaging features, as well as the score and weight of each change parameter. The carrier performs a weighted summation on each feature and change parameter included in the lesion imaging features to obtain a first predicted score result corresponding to each object in the first dataset.
[0102] Alternatively, in some embodiments, the carrier inputs the lesion imaging features and change parameters corresponding to at least two CT image results of each object in the first dataset into the first decision tree model. According to different branch conditions in the first decision tree model, the lesion imaging features or change parameters that meet the branch conditions are input into the branch, and finally the first prediction score result of each object is obtained.
[0103] Step S403: In response to the first prediction score result including at least one negative sample, or the preset update condition being met, the first decision tree model is updated based on the sample dataset to obtain the second decision tree model.
[0104] In some embodiments, in response to the first predicted score not including negative samples and not meeting the preset update conditions, the first decision tree model is confirmed as the target decision tree model.
[0105] In some embodiments, the preset update conditions may include satisfying a first period, that is, updating the first decision tree model based on the sample dataset every first period.
[0106] In some embodiments, the negative samples include data in the first predicted score results where the predicted score results do not match the lesion imaging features. The negative samples are represented by lesion imaging features, that is, lesion imaging features corresponding to the mismatched predicted score results, or CT images.
[0107] In some embodiments, if the first predicted score includes at least one negative sample, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model.
[0108] In some embodiments, the data in the sample dataset includes pre-collected sample data and negative samples generated during the period from the last update of the decision tree model to the update of the first decision tree model.
[0109] For example, if the decision tree model is updated according to preset update conditions, then the negative samples generated in the first period after the last update of the decision tree model are added to the sample dataset to update the decision tree model; or, if the decision tree model is updated based on the number of negative samples, then the newly added negative samples are added to the sample dataset to update the decision tree model.
[0110] In some embodiments, the carrier updates the lesion prediction model based on a sample dataset. The lesion prediction model is used to obtain the lesion imaging features and change parameters corresponding to at least two CT image results for each object in the first dataset; update the condition values of the corresponding branches in the first decision tree model based on the values of the lesion imaging features and change parameters in the sample dataset; update the weights and scores of the corresponding branches in the first decision tree model based on the lesion imaging features and change parameters in the sample dataset; and add corresponding branches in the first decision tree model based on the clinical information corresponding to the object.
[0111] In practice, the carrier can update the lesion prediction model based on the sample dataset. Since the CT image to the final prediction score requires passing through two models—the lesion prediction model and the decision tree model—to ensure data accuracy and avoid resource waste (the lesion prediction model has problems, yet the decision tree model is repeatedly updated), the carrier updates the lesion prediction model based on the sample dataset. After the lesion prediction model is updated, the radiographic features of the lesions corresponding to the CT images in the first dataset are re-acquired, and these re-acquired radiographic features are input into the first decision tree model to obtain the first prediction score. That is, the problems with the lesion prediction model are first updated and eliminated, and then the decision tree model is updated.
[0112] In specific implementation, the carrier updates the conditional values of the corresponding branches in the first decision tree model based on the numerical values of the lesion imaging features in the sample dataset. This includes: updating the conditional values of the corresponding branches in the first decision tree model based on the numerical values of lesion imaging features of the same type or severity in the sample dataset; or, updating the conditional values of the corresponding branches in the first decision tree model based on the changing parameters of the same type or severity in the sample dataset. For example, taking size as an example, the first decision tree model sets a size threshold of 6, that is, distinguishing between positive level 1 and negative based on whether the size is greater than or equal to 6 or less than 6. However, in the sample dataset, samples with a size greater than 6 and less than 8 are negative. Therefore, the size threshold in the first decision tree model can be updated to 8, where less than 8 is negative and greater than or equal to 8 is positive. Here, type includes the type of feature, such as size, volume, density, major axis, etc.
[0113] In specific implementation, the carrier updates the weights and scores of corresponding branches in the first decision tree model based on the lesion imaging features and changing parameters in the sample dataset. This includes adjusting the weights and scores of different branches in the first decision tree model. For example, the weights of lesion imaging features and changing parameters that have a smaller impact on the result are appropriately reduced, while the weights of lesion imaging features and changing parameters that have a larger impact are appropriately increased to avoid the final prediction score being affected by features with smaller results having excessively high scores and weights.
[0114] In practice, the carrier adds corresponding branches to the first decision tree model based on the clinical information corresponding to the object; wherein, the clinical information may include age, health history, smoking history, etc. To ensure the accuracy of the results, the carrier may add branches corresponding to the clinical information to the first decision tree model.
[0115] Step S404: Based on the sample dataset, obtain the first evaluation data corresponding to the first decision tree model and the second evaluation data corresponding to the second decision tree model.
[0116] In some embodiments, the carrier inputs the sample dataset into a first decision tree model to obtain evaluation data corresponding to the first decision tree model; and inputs the sample dataset into a second decision tree model to obtain evaluation data corresponding to the second decision tree model; wherein the evaluation data includes at least one of sensitivity data, specificity data, and Youden index.
[0117] In some embodiments, the sample confirms whether the first evaluation data is better than the second evaluation data; if the first evaluation data is better than the second evaluation data, it indicates that the update is invalid and step S403 is repeated; or, if the second evaluation data is better than the first evaluation data, it indicates that the update is valid and step S405 is executed.
[0118] Step S405: Confirm that the second decision tree model is the target decision tree model.
[0119] In some embodiments, if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, it indicates that the update operation is effective. Subsequently, the imaging features of the lesions corresponding to the CT impact are processed based on the second decision tree model to obtain the prediction score result. The prediction score result is used to assist doctors in early screening and cannot be used as a direct diagnosis result or treatment suggestion.
[0120] Thus, according to the data processing method provided in this disclosure, by integrating CT images acquired at different times, the development trend and change relationship of lesions can be established, avoiding fragmentation of multiple CT images. A decision tree model is established based on the lesion prediction model's ability to automatically detect lesion imaging features and predict benign / malignant probabilities, and clinical guidelines. Iterative updates are performed on real data (i.e., sample datasets) to improve the accuracy of the decision tree model. Combining the lesion prediction model and the decision tree model allows for rapid and accurate scoring and classification of lesions, helping doctors detect lesions early, improving early detection rates, and providing patients with better treatment opportunities. It enables the completion of a large number of lung nodule scoring and classification tasks in a short time, reducing missed diagnoses and misdiagnoses, and improving the accuracy and precision of cancer diagnosis. Combined with clinical guidelines, the decision tree model becomes interpretable. It provides doctors with more information and data support, promotes the development of precision medicine, and provides patients with more personalized treatment plans and services. The continuous fusion of information from multiple CT scans allows for insight into lesion development trends and better model prediction.
[0121] Figure 6 A schematic diagram of an optional structure of the data processing apparatus provided in an embodiment of this disclosure is shown, and will be described in accordance with each step.
[0122] In some embodiments, the data processing apparatus 500 includes an acquisition unit 501, an update unit 502, an evaluation unit 503, and an iteration unit 504.
[0123] The acquisition unit 501 is used to input the lesion imaging features corresponding to at least one CT image result of each object in the first dataset into the first decision tree model to obtain the first prediction score result;
[0124] The update unit 502 is used to update the first decision tree model based on the sample dataset to obtain a second decision tree model in response to the first prediction score result including at least one negative sample or meeting a preset update condition.
[0125] The evaluation unit 503 is used to obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model based on the sample dataset.
[0126] The iteration unit 504 is used to confirm the second decision tree model as the target decision tree model if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model; or, if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, it is used to update the first decision tree model again based on the at least one negative sample and the sample dataset, and obtain evaluation data until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model.
[0127] In some embodiments, the lesion imaging features corresponding to the at least one CT image result are obtained based on a lesion prediction model; the inputs of the first decision tree model and the second decision tree model also include the lesion trait probability output by the lesion prediction model; the lesion imaging features include at least one of the following: lesion long diameter, lesion volume, and lesion density.
[0128] The acquisition unit 501 is specifically used to input the lesion imaging features corresponding to a CT image result of each object in the first dataset into the first decision tree model, determine the score and weight of each feature in the lesion imaging features based on each branch of the first decision tree model, perform weighted summation on each feature included in the lesion imaging features, and obtain the first prediction score result corresponding to each object in the first dataset.
[0129] The acquisition unit 501 is specifically used to input the lesion imaging features corresponding to at least two CT image results of each object in the first dataset into the first decision tree model, and to confirm at least one variable parameter based on the at least two CT image results.
[0130] Based on each branch of the first decision tree model, the score and weight of each feature in the lesion imaging features, as well as the score and weight of the at least one variable parameter, are confirmed.
[0131] The first prediction score result corresponding to each object in the first dataset is obtained by weighted summation of each feature included in the lesion imaging features and the at least one variation parameter.
[0132] The update unit 502 is specifically used for at least one of the following:
[0133] Based on the sample dataset, the lesion prediction model is updated. The lesion prediction model is used to obtain the lesion imaging features corresponding to at least one CT image result of each object in the first dataset.
[0134] Based on the numerical values of the lesion imaging features in the sample dataset, update the condition values of the corresponding branches in the first decision tree model.
[0135] Based on the lesion imaging features in the sample dataset, update the weights and scores of the corresponding branches in the first decision tree model;
[0136] Based on the clinical information corresponding to the object, a corresponding branch is added to the first decision tree model.
[0137] The evaluation unit 503 is specifically used to: input the sample dataset into the first decision tree model and obtain the evaluation data corresponding to the first decision tree model;
[0138] The sample dataset is input into the second decision tree model to obtain the evaluation data corresponding to the second decision tree model;
[0139] The evaluation data includes at least one of sensitivity data, specificity data, and the Youden index.
[0140] Figure 7 A schematic diagram of another optional structure of the data processing apparatus provided in the embodiments of this disclosure is shown, and will be described in terms of each part.
[0141] In some embodiments, Figure 7 The data processing device 600 shown is implemented based on the target decision tree model obtained by the method described in steps S101 to S104, S201 to S205, and S401 to S405.
[0142] In some embodiments, the data processing apparatus 600 includes an input unit 601.
[0143] The input unit 601 is used to input the imaging features of the lesion corresponding to the image to be scored into the target decision tree model to obtain the scoring result; the scoring result is used to confirm the degree of lesion.
[0144] In some optional embodiments, the input unit 601 can input the lesion imaging features corresponding to the image to be scored into the target decision tree model based on the user interaction page to obtain the scoring result.
[0145] Figure 8 This illustration shows an optional schematic diagram of a user interaction page provided in an embodiment of the present disclosure. Figure 9 Another optional schematic diagram of the user interaction page provided in this embodiment of the disclosure is shown.
[0146] like Figure 8As shown, it can be used for a single CT image. Taking a lung CT image as an example, the parameters that need to be input include nodule type (ground-glass or solid), nodule AI malignancy probability (i.e., the probability of lesion characteristics output by the lesion prediction model), maximum diameter of nodule, nodule volume, nodule CTR, and patient age. Optionally, it can also include whether it is a major lesion.
[0147] like Figure 9 As shown, it can be used for multiple CT images. Taking lung CT images as an example, the parameters that need to be input include the nodule type (ground-glass or solid) of the two images, the nodule malignancy probability of the two images (i.e., the lesion morphology probability output by the lesion prediction model), the maximum diameter of the nodule of the two images, the nodule volume of the two images, the nodule VDT of the two images, and the patient's age. Optionally, it can also include whether it is a major lesion.
[0148] Thus, through the data processing device provided in this embodiment, based on the patient's clinical information and CT images, and with the help of AI technology, the lesions in the CT are automatically extracted to calculate the probability of benign or malignant transformation. The lesion scoring model is established by referring to clinical guidelines and real data evaluation, and the lesion scoring is further graded. Different grades guide the patient's surgery or follow-up, as well as the follow-up interval. The scoring mechanism is iteratively updated on real-world data to continuously improve the performance of the scoring and grading.
[0149] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0150] Figure 10 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0151] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0152] like Figure 10As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0153] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0154] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform data processing methods by any other suitable means (e.g., by means of firmware).
[0155] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0159] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0160] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0161] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0162] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0163] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A data processing method, characterized in that, The method includes: The lesion imaging features corresponding to a single CT image of each object in the first dataset are input into the first decision tree model. Based on each branch of the first decision tree model, the score and weight of each feature in the lesion imaging features are determined. The weighted sum of each feature included in the lesion imaging features is performed to obtain the first prediction score result corresponding to each object in the first dataset. In response to the first predicted score result including at least one negative sample, or the preset update condition being met, the first decision tree model is updated based on the sample dataset to obtain a second decision tree model. Based on the sample dataset, obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model respectively; If the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model, then the second decision tree model is confirmed as the target decision tree model. If the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, then the first decision tree model is updated again based on the at least one negative sample and the sample dataset, and evaluation data is obtained until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model. Wherein, the lesion imaging features corresponding to the at least one CT image result are obtained based on the lesion prediction model; the inputs of the first decision tree model and the second decision tree model also include the lesion morphology probability output by the lesion prediction model; the lesion imaging features include at least one of the following: lesion long diameter, lesion volume, and lesion density; the negative samples include data in the first prediction score result where the prediction score result does not match the lesion imaging features, and the negative samples are represented by lesion imaging features, that is, lesion imaging features corresponding to the mismatched prediction score result, or CT images; the preset update conditions include satisfying the first cycle.
2. The method according to claim 1, characterized in that, The step of inputting the lesion imaging features corresponding to at least one CT image result of each object in the first dataset into the first decision tree model to obtain the first prediction score result includes: The lesion imaging features corresponding to at least two CT image results for each object in the first dataset are input into the first decision tree model, and at least one variable parameter is identified based on the at least two CT image results. Based on each branch of the first decision tree model, the score and weight of each feature in the lesion imaging features, as well as the score and weight of the at least one variable parameter, are confirmed. The first prediction score result corresponding to each object in the first dataset is obtained by weighted summation of each feature included in the lesion imaging features and the at least one variation parameter.
3. The method according to claim 1, characterized in that, The step of updating the first decision tree model based on the sample dataset to obtain the second decision tree model includes at least one of the following: Based on the sample dataset, the lesion prediction model is updated. The lesion prediction model is used to obtain the lesion imaging features corresponding to at least one CT image result of each object in the first dataset. Based on the numerical values of the lesion imaging features in the sample dataset, update the condition values of the corresponding branches in the first decision tree model. Based on the lesion imaging features in the sample dataset, update the weights and scores of the corresponding branches in the first decision tree model; Based on the clinical information corresponding to the object, a corresponding branch is added to the first decision tree model.
4. The method according to claim 1, characterized in that, The step of obtaining the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model based on the sample dataset includes: Input the sample dataset into the first decision tree model to obtain the evaluation data corresponding to the first decision tree model; The sample dataset is input into the second decision tree model to obtain the evaluation data corresponding to the second decision tree model; The evaluation data includes at least one of sensitivity data, specificity data, and the Youden index.
5. A data processing apparatus, characterized in that, The device includes: The acquisition unit is used to input the lesion imaging features corresponding to a CT image result of each object in the first dataset into the first decision tree model, determine the score and weight of each feature in the lesion imaging features based on each branch of the first decision tree model, and perform a weighted summation on each feature included in the lesion imaging features to obtain the first prediction score result corresponding to each object in the first dataset. The update unit is used to update the first decision tree model based on the sample dataset to obtain a second decision tree model in response to the first prediction score result including at least one negative sample or meeting a preset update condition. The evaluation unit is used to obtain the evaluation data corresponding to the first decision tree model and the evaluation data corresponding to the second decision tree model based on the sample dataset. The iteration unit is used to confirm the second decision tree model as the target decision tree model if the evaluation data corresponding to the second decision tree model is better than the evaluation data corresponding to the first decision tree model; or, if the evaluation data corresponding to the first decision tree model is better than the evaluation data corresponding to the second decision tree model, it is used to update the first decision tree model again based on the at least one negative sample and the sample dataset, and obtain evaluation data until the latest updated evaluation data is better than the evaluation data corresponding to the first decision tree model. Wherein, the lesion imaging features corresponding to the at least one CT image result are obtained based on the lesion prediction model; the inputs of the first decision tree model and the second decision tree model also include the lesion morphology probability output by the lesion prediction model; the lesion imaging features include at least one of the following: lesion long diameter, lesion volume, and lesion density; the negative samples include data in the first prediction score result where the prediction score result does not match the lesion imaging features, and the negative samples are represented by lesion imaging features, that is, lesion imaging features corresponding to the mismatched prediction score result, or CT images; the preset update conditions include satisfying the first cycle.
6. A data processing apparatus, characterized in that, Based on the target decision tree model obtained according to any one of claims 1 to 4, the apparatus comprises: The input unit is used to input the imaging features of the lesions corresponding to the image to be scored into the target decision tree model to obtain the scoring results; The scoring results are used to confirm the severity of the lesions.
7. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.
Citation Information
Patent Citations
Auxiliary analysis method and system for nerve diagnosis
CN115691794A
Method for generating a diagnosis model capable of diagnosing multi-cancer according to stratification information by using biomarker group-related value information, method for diagnosing multi-cancer by using the diagnosis model, and device using the same
US11515042B1