An esophageal cancer lesion image assessment system

By generating initial lesion area focus data and combining it with the doctor's interactive acquisition module to learn individual preference parameters, the problem of existing systems being unable to adapt to different doctors' interpretation habits has been solved, personalized output has been achieved, and the efficiency and accuracy of esophageal cancer lesion image assessment have been improved.

CN122336478APending Publication Date: 2026-07-03ZHEJIANG HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG HOSPITAL
Filing Date
2026-04-13
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing esophageal cancer lesion image assessment systems are difficult to personalize and cannot adapt to the interpretation habits of different doctors, resulting in doctors needing to make repeated corrections, increasing their workload and time costs.

Method used

The initial lesion area data is generated by the basic image processing module, and the correction operations are recorded by the doctor interaction acquisition module. Individual preference parameters are learned, and personalized output is generated by the preference learning module. The model accuracy is improved by combining self-supervised training and pseudo-labeling enhancement training.

Benefits of technology

It reduces the amount of repetitive corrections doctors need to make in case processing, improves work efficiency, slows down the rate of increase in workload, and enhances the overall throughput efficiency and result consistency of assessment tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336478A_ABST
    Figure CN122336478A_ABST
Patent Text Reader

Abstract

This invention discloses an esophageal cancer lesion image assessment system, comprising a basic image processing module for generating initial lesion area focus data from input esophageal images; a doctor interaction acquisition module for collecting doctor's correction operations on the initial lesion area focus data; a preference learning module for learning individual preference parameters corresponding to the doctor based on the correction operations; and a personalized output module for calibrating the initial lesion area focus data based on the individual preference parameters and outputting personalized lesion area focus data. This esophageal cancer lesion image assessment system improves the assessment process for esophageal cancer lesions, reducing repetitive correction operations for doctors in handling consecutive cases, decreasing the additional workload as the number of cases increases, and making the initial output more aligned with the doctor's interpretation methods, thereby improving the overall efficiency of the assessment task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image data processing technology, specifically to a lesion image assessment system for esophageal cancer. Background Technology

[0002] Lesion identification and delineation in esophageal cancer play a crucial role in clinical diagnosis, pathological evaluation, treatment planning, and follow-up. Endoscopic images are an important source of information for identifying esophageal mucosal lesions; in practical applications, it is necessary to locate and focus on the lesion area for subsequent evaluation and treatment.

[0003] For example, there is a system and method for segmenting images using a training model, as disclosed in prior art (US11176677B2). A device can identify a training dataset. This training dataset may include a region of interest for each image. This training dataset may include first annotations. The device can use the training dataset to train an image segmentation model with parameters that generates corresponding first segmented images. The device can provide the first segmented images for display on a user interface to obtain feedback. The device can receive a feedback dataset containing second annotations, which are at least a subset of the first segmented images, through the user interface. Each second annotation can label at least a second portion of the region of interest in the corresponding image within the subset. The device can retrain the image segmentation model using the feedback dataset received through the user interface.

[0004] In existing systems, the automatically generated initial focus results are insufficient to fully meet doctors' needs for interpreting lesion extent, boundary morphology, and preservation of local details. Doctors still need to modify the initial results multiple times through the interactive interface. For example, different doctors have stable differences in the tightness of lesion boundaries, the inclusion or exclusion of superficial lesions, and the preservation of structural texture regions. However, existing systems uniformly output results in a fixed format, without considering the differences in interpretation among doctors. This forces doctors to repeatedly adjust for similar deviations when processing multiple images consecutively. Summary of the Invention

[0005] The purpose of this invention is to provide a lesion image assessment system for esophageal cancer to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a lesion image assessment system for esophageal cancer, comprising: The basic image processing module is used to generate initial lesion area focus data from the input esophageal image; The doctor interaction data acquisition module is used to collect the doctor's correction operations on the initial lesion area attention data; The preference learning module is used to learn the individual preference parameters of the corresponding doctor based on the correction operation; The personalized output module is used to calibrate the initial lesion area attention data based on the individual preference parameters and output personalized lesion area attention data.

[0007] Preferably, the doctor interaction acquisition module is used to structure the doctor's correction operation into a correction vector field representing the difference between the initial lesion area attention data and the corrected lesion area annotation data.

[0008] Preferably, the preference learning module is used to train a low-dimensional preference function model based on the modified vector field, wherein the individual preference parameters are the parameters of the preference function model, and the parameter dimension is lower than the parameter dimension of the basic image processing module.

[0009] Preferably, the individual preference parameters are used to calibrate the initial lesion area focus data in subsequent image processing to reduce the amount of correction required by the physician.

[0010] Preferably, the individual preference parameters are stored separately with doctor identifiers, and the corresponding individual preference parameters are called to generate personalized lesion area attention data when used by the doctor.

[0011] Preferably, the system also includes a data-efficient training module for training the basic image processing module based on unlabeled esophageal images and a small number of labeled esophageal images.

[0012] Preferably, the data-efficient training module further utilizes pseudo-annotated esophageal images generated by doctors for enhanced training of the basic image processing module.

[0013] Preferably, the data-efficient training module includes a self-supervised training unit, a supervised fine-tuning unit, and a pseudo-labeling enhancement unit.

[0014] The system processing method includes the following steps: Step 1: Using the basic image processing module, generate initial lesion area focus data from the input esophageal image; Step 2: Collect the doctor's correction operations on the initial lesion area data, and structure the correction operations into a correction vector field; Step 3: Train a low-dimensional preference function model based on the modified vector field to obtain the individual preference parameters of the corresponding doctor; Step 4: When the doctor uses the data subsequently, the initial lesion area focus data is calibrated based on the individual preference parameters to output personalized lesion area focus data.

[0015] Preferably, the processing method further includes: Step 5: Perform self-supervised training and supervised fine-tuning of the basic image processing module using unlabeled esophageal images and a small number of labeled esophageal images; Step Six: Enhance the training of the basic image processing module using pseudo-annotated esophageal images generated by the doctor.

[0016] Compared with the prior art, the beneficial effects of the present invention are: the esophageal cancer lesion image assessment system can improve the assessment process of esophageal cancer lesion images, reduce the need for doctors to perform repeated correction operations in the continuous case processing, reduce the extra workload generated as the number of cases increases, and make the initial output more in line with the doctor's interpretation method, thereby improving the overall work efficiency of the assessment task.

[0017] 1. Reduce repetitive correction operations, improve workflow efficiency, learn from doctors' correction behavior in previous cases and calibrate the initial results in subsequent cases, so that doctors no longer need to make multiple adjustments for the same type of deviation, thereby reducing correction time, shortening the "generation-check-correction-confirmation" processing loop, and improving work efficiency in continuous processing.

[0018] 2. This application reduces the cumulative workload caused by the increase in the number of cases, enabling the correction behavior to be reused across cases. Compared with the existing technology where the correction workload increases linearly with the number of cases, this application can slow down the growth rate of workload and show better scalability in the scenario of batch processing of cases, which meets the time organization requirements in the actual workflow.

[0019] 3. Improve the fit of initial results and reduce subsequent compensation work. By using self-supervised training and pseudo-labeling enhancement mechanisms, the initial output of the basic model can be improved when the labeled samples are limited. This reduces the amount of work doctors need to do to compensate for system biases, helps maintain the rhythm and stability of continuous case processing, and improves the overall throughput efficiency of the assessment task. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the personalized calibration process in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the training phase of Embodiment 2 of the present invention; Figure 3 This is a flowchart illustrating the deployment and data feedback process of Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the image processing method according to Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the data path of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1-5 The present invention provides the following technical solution: Example 1: In the process of esophageal cancer lesion image evaluation, the initial lesion area focus data generated by the system can usually serve as the starting reference for doctors' annotation. However, there may be differences in the tightness of the boundaries and the selection of local areas compared to the annotation habits of different doctors. Doctors need to make corrections such as dragging the boundaries and adding or removing local areas through the interactive interface. Existing processes mostly treat the correction as a correction operation for the current image, lacking a mechanism to structurally express the correction differences and form low-dimensional preference parameters bound to the doctor's identification, so as to directly calibrate the initial output in subsequent image processing. Therefore, this example aims to collect correction operations and structure them into a correction vector field, learn low-dimensional preference parameters and store them bound to the doctor's identification, and call these parameters to perform calibration on the initial output in subsequent processing, so as to reduce the amount of subsequent correction and interaction overhead.

[0023] This embodiment discloses the following: The system includes a basic image processing module for generating initial lesion region focus data from an input esophageal image; a doctor interaction acquisition module for acquiring correction operations performed by doctors on the initial lesion region focus data; a preference learning module for learning individual preference parameters corresponding to the doctors based on the correction operations; and a personalized output module for calibrating the initial lesion region focus data based on the individual preference parameters and outputting personalized lesion region focus data. The doctor interaction acquisition module structures the doctor's correction operations into a correction vector field representing the difference between the initial lesion region focus data and the corrected lesion region annotation data. The preference learning module trains a low-dimensional preference function model based on the correction vector field. The individual preference parameters are the parameters of the preference function model, and their parameter dimension is lower than that of the basic image processing module. The individual preference parameters are used to perform calibration on the initial lesion region focus data in subsequent image processing to reduce the amount of correction by the doctor. The individual preference parameters are stored separately with doctor identifiers, and the corresponding individual preference parameters are called to generate personalized lesion region focus data when used by the doctor.

[0024] The esophageal cancer lesion image assessment system provided in this embodiment is deployed on a hospital local area network server. The basic image processing module runs on the server side, and doctors access the system through workstation terminals to perform esophageal image processing.

[0025] In the specific implementation, the basic image processing module can use a convolutional neural network as the backbone network for feature extraction, normalize the input original esophageal image and extract multi-scale features, and then generate a lesion area probability map with the same resolution as the original image through upsampling and feature fusion. Based on a preset threshold, the probability map is converted into initial lesion area attention data C0. C0 is stored in the form of a binary mask with the same size as the original image, which is used to represent the range of lesion areas that the system recommends to pay attention to. Furthermore, in this embodiment, the basic image processing module employs an encoder-decoder segmentation network to extract features from the input esophageal endoscopy image. The encoder extracts multi-scale texture features, color contrast features, and mucosal structure change features, while the decoder generates a pixel-level probability map P(x,y) of the same size as the input image, where P(x,y) represents the predicted probability that the corresponding pixel belongs to the lesion region. By setting a threshold, the probability map is converted into initial lesion region focus data C0.

[0026] Furthermore, in the esophageal endoscopy image scenario, a mucosal texture response enhancement channel and a vascular structure contrast enhancement factor can be introduced in the feature extraction stage to make the model more sensitive to changes in the esophageal mucosal microstructure, thereby improving the stability of the initial region prediction and providing a more reliable basis for subsequent personalized calibration.

[0027] In this embodiment, the basic image processing module maps the input esophageal endoscopy image I(x,y) into three types of feature maps, and fuses the three types of features to obtain the region response: (1) F_tex(x,y) = Phi_tex(I(x,y); W_tex) Multiscale texture feature map (2) F_col(x,y) = Phi_col(I(x,y); W_col) Color contrast feature map (3) F_muc(x,y) = Phi_muc(I(x,y); W_muc) Mucosal structure change characteristic diagram Based on this, the three types of feature maps are fused to obtain the region response R(x,y): (4) R(x,y) = aF_tex(x,y) + bF_col(x,y) + c*F_muc(x,y), where a, b, and c are fusion weight coefficients. The regional response is then mapped to a probability map, and initial lesion region focus data C0 is generated using a threshold tau. (5) P(x,y) = 1 / (1 + exp(-R(x,y))) (6) C0(x,y) = 1, if P(x,y) >= tau; 0, otherwise Where tau is the segmentation threshold, the above formulas (1)-(6) correspond to the implementation method of "extracting multi-scale texture features, color contrast features and mucosal structure change features, and generating initial lesion area focus data C0" in the instruction manual.

[0028] The doctor-interactive acquisition module provides a lesion area editing interface on the workstation terminal, overlaying the original esophageal image with C0 and providing interactive tools such as boundary dragging, area addition and subtraction, and smearing to allow doctors to correct C0 and obtain the corrected lesion area annotation data Cu. After the doctor completes the correction and confirms it, the system compares C0 and Cu pixel by pixel. For each pixel belonging to the lesion area, it calculates the state difference and position offset between the two annotations and records the difference in vector form to obtain the correction vector field V consistent with the image resolution, where V(i,j)=Cu(i,j)-C0(i,j) is used to represent the offset direction and offset magnitude at the image coordinate (i,j).

[0029] The preference learning module selects a sample set corresponding to the current doctor identifier from multiple sets of historically accumulated vector field data. These samples are divided into training and validation sets. The preference learning module internally constructs a low-dimensional preference function model. The model input is the intermediate features generated by the basic image processing module and the corresponding position C0. The output is the calibration offset at that position. The parameter dimension of the preference function model is much smaller than that of the basic image processing module. After training, the parameters are extracted as the individual preference parameter theta_d corresponding to the doctor identifier. The preference learning module adopts a training method based on error minimization, so that the error between the model output offset and the true offset of the vector field is gradually reduced.

[0030] In this embodiment, the preference function model is a low-dimensional parameter mapping model, whose parameter dimension is significantly lower than that of the basic image processing model. This model only operates in the post-processing stage of the initial region results, without changing the main segmentation network structure, and adjusts the region boundary morphology and region selection rules through parameter mapping.

[0031] When doctors process new esophageal images, the basic image processing module first generates the corresponding C0. After recognizing the currently logged-in doctor's identifier, the personalized output module loads the corresponding theta_d from storage and inputs theta_d and intermediate features from the basic image processing module into the preference function model to calculate the calibration offset of C0. The personalized output module makes minor adjustments to the lesion region boundary of C0 based on the model output to obtain personalized lesion region focus data Cd. The system overlays Cd with the original esophageal image and displays it as the starting annotation result for doctors to edit.

[0032] When multiple doctors use the system together, each doctor is assigned a unique doctor identifier. After training individual preference parameters, the preference learning module stores the parameters theta_d of different doctors in the server parameter library according to the correspondence between doctor identifier and parameter vector. When any doctor logs into the system, the personalized output module automatically loads the corresponding theta_d according to the doctor identifier to generate the doctor's unique Cd.

[0033] In this embodiment, the relevant descriptions are as follows: Initial lesion area focus data: refers to the lesion area suggestion result given by the basic image processing module after processing the input esophageal image without considering the individual preferences of doctors. It is used to represent the range of lesion areas that the system suggests to focus on in the current model state, and can be represented by a mask map or other regional representations.

[0034] Corrected vector field: refers to a data set constructed based on the difference between the initial lesion area focus data and the lesion area annotation data after doctor correction. It is used to characterize the offset direction and offset magnitude of the lesion area at different locations, in order to depict the doctor's correction behavior to the system's suggested results.

[0035] Individual preference parameters: These are a set of parameters associated with a specific doctor, generated by the preference learning module after learning the modified vector field. They describe the doctor's individual preferences in the selection of the scope and boundaries of the lesion area of ​​focus and are used to personalize the initial lesion area focus data.

[0036] Personalized lesion area focus data: This refers to the lesion area focus result obtained by the personalized output module after calibrating the initial lesion area focus data by loading specific doctor's individual preference parameters. This data serves as the doctor's initial annotation result to reduce subsequent correction operations. Specifically, the personalized output module constructs a low-dimensional preference parameter vector θ = {α, β, γ}, where α represents the boundary expansion or contraction coefficient, β represents the minimum region retention threshold, and γ represents the boundary smoothing control factor. The personalized result Cd is generated through the calibration function: Cd = F(C0, θ), where F(·) is a combined transformation function based on morphological operations and region selection rules. By changing the parameter θ, the output result can be made to more closely resemble the doctor's current correction habits in terms of spatial morphology.

[0037] Doctor Identifier: Refers to the identity information label used to distinguish different doctors. It can be a numerical code, account name, or other identifier that can uniquely correspond to the doctor's identity. It is used to index and manage the individual preference parameters corresponding to different doctors in the system.

[0038] Suppose that two doctors in a hospital are involved in annotating lesion images: Doctor A and Doctor B, and they have different habits in annotating the boundaries of lesion areas.

[0039] Doctor A tends to make tight borders on the lesion area when making annotations, that is, to make a small outward expansion near the edge of the lesion tissue; Doctor B, on the other hand, tends to make the lesion area expand moderately outward into the surrounding tissue to ensure that the area includes the ambiguous boundary area.

[0040] When using the system provided in this embodiment for the first time, the basic image processing module generates initial lesion area focus data C0 for the esophageal image. For the same image, Doctor A and Doctor B respectively correct C0 using the boundary drag tool. Doctor A shrinks part of the boundary inward and deletes the local area, while Doctor B stretches the outer edge of the lesion area outward and adds a surrounding blurred area.

[0041] The system records the correction process of the two doctors and generates correction vector fields VA and VB respectively, where: VA(i,j)=CAu(i,j)-C0(i,j) VB(i,j)=CBu(i,j)-C0(i,j) After processing by the preference learning module, the system obtains the individual preference parameter sets theta_A and theta_B for the two doctors, and binds them to the doctor identifiers by index.

[0042] In subsequent annotation processes, when Doctor A logs into the system and processes a new esophageal image, the basic image processing module first generates C0. Then, the preference function model calibrates C0 based on theta_A and outputs personalized lesion region focus data CA. CA shows a tighter pattern at the boundary, consistent with Doctor A's original annotation habits. Correspondingly, when Doctor B logs into the system, the system automatically loads theta_B and calibrates C0 to output CB, causing CB to expand appropriately at the edge region. After multiple interactions, the number of corrections required by both doctors when annotating similar images is significantly reduced compared to the first use, thereby reducing the amount of annotation operations and interaction overhead, and improving annotation efficiency.

[0043] Example 2: In the process of esophageal cancer lesion image evaluation, the initial lesion region focus data generated by the basic image processing module depends on the model training data and training method. When the scale of labeled samples is limited or the sample distribution coverage is insufficient, there may be a long-term deviation between the initial output and the correction results of doctors in the scene, so that doctors still need to make more corrections in subsequent use. Existing training processes usually focus on using a small number of labeled samples for supervised training, lacking a feature learning path that combines a large amount of unlabeled data, and a mechanism to use the correction results formed by doctors in the interaction process as usable data for continuous enhancement training. Therefore, this example aims to improve the efficiency of training data utilization by combining self-supervised training, supervised fine-tuning and pseudo-label enhancement training based on doctor corrections, so that the initial output of the basic image processing module is closer to the labeling habits in the scene, thereby further reducing the amount of subsequent corrections.

[0044] Therefore, this embodiment discloses the following solution: it also includes a data-efficient training module for training a basic image processing module based on unlabeled esophageal images and a small number of labeled esophageal images. The data-efficient training module further utilizes pseudo-labeled esophageal images generated by doctors to enhance the training of the basic image processing module. The data-efficient training module includes a self-supervised training unit, a supervised fine-tuning unit, and a pseudo-label enhancement unit.

[0045] The data-efficient training module provided in this embodiment runs on a server and is used to train and update the basic image processing module. This embodiment adopts three training paths, including a self-supervised training path, a supervised fine-tuning path, and a pseudo-annotation enhancement path, to improve the model's efficiency in utilizing limited labeled data.

[0046] In the self-supervised training path, the system selects an unlabeled esophageal image dataset U from the hospital image database, performs random occlusion, color perturbation and geometric transformation on each image, and performs self-supervised training through an autoencoder structure or a contrastive learning structure to learn the general feature representation of esophageal tissue and lesion areas. This path does not rely on manual annotation and can effectively improve the basic feature extraction capability when the sample size is large.

[0047] In the supervised fine-tuning path, the system uses a small dataset L of esophageal images with lesion area annotations to perform supervised fine-tuning training on the basic image processing module. The training objective is to generate lesion area focus data consistent with the annotations based on the annotated samples. The supervised fine-tuning path is used to adapt the general intermediate features obtained from self-supervised training to the lesion area focus task and obtain usable initial model output when annotated samples are scarce.

[0048] In the pseudo-annotation enhancement path, after the doctor corrects the initial lesion area attention data C0 in the workflow of Example 1, the system records the corrected lesion area annotation data Cu, and adds Cu and the original image together as pseudo-annotation samples to the pseudo-annotation dataset P. The pseudo-annotation enhancement unit uses the dataset P to enhance the training of the basic image processing module, so that the subsequent model generates C0 in a way that is more consistent with the actual annotation habits in medical use scenarios, thereby reducing the amount of correction.

[0049] This embodiment adopts the training sequence U→L→P, where U is used to learn general features, L is used for playback task adaptation, and P is used to strengthen the consistency between the annotation logic in the scene and the doctor's preferences. After training is completed, the basic image processing module is updated to a new model version and outputs initial lesion area attention data that is more in line with the doctor's usage habits in subsequent use.

[0050] The definitions involved in this embodiment are as follows: Unlabeled esophageal image dataset U: refers to a collection of esophageal images without lesion area annotations, used for self-supervised training.

[0051] Annotated esophageal image dataset L: refers to a collection of esophageal images containing annotations of lesion areas, used for supervised fine-tuning training.

[0052] Pseudo-annotated esophageal image dataset P: refers to the dataset corresponding to the lesion region annotation results formed after the doctor corrects the initial lesion region attention data, which is used for augmentation training.

[0053] Self-supervised training unit: refers to a module that constructs a learning task based on unlabeled data and trains it to extract general image features.

[0054] Supervised fine-tuning unit: refers to a module that is trained under supervision on labeled data, used to adapt to the task of focusing on lesion areas.

[0055] Pseudo-labeled augmentation unit: This refers to a module that uses pseudo-labeled data to enhance the training of the basic image processing module, thereby improving the model's practicality in medical scenarios.

[0056] When a hospital was carrying out image annotation work on esophageal lesions, it initially collected and annotated only a small number of esophageal image samples, with a total number of L=52 images. Each image contained manually annotated lesion area attention mask data. This dataset was used for supervised fine-tuning training to adapt to the lesion area attention task.

[0057] Meanwhile, the hospital’s historical endoscopy database contains a large number of unlabeled esophageal images, totaling approximately U≈8400 images, which contain only raw images without any annotation information. The data efficiency training module in this embodiment first performs self-supervised training on the dataset U to obtain general feature representations of esophageal tissue structure and local regional changes.

[0058] During the supervised fine-tuning phase, since the sample size of L is insufficient to cover various esophageal structures and imaging differences, there is still a certain deviation between the initial output C0 of the model and the manual annotation. During the interaction process in Example 1, the doctor corrects C0 to form the corrected lesion area annotation sample Cu. These Cu and their corresponding original images are added to the pseudo-annotation dataset P. In the initial stage of annotation, |P|=52, and its sample size is comparable to that of L.

[0059] As the usage progressed, the system continuously accumulated data corrected by doctors. After one month, doctors had corrected 217 annotation tasks, constructing a pseudo-annotated dataset P=217. At this stage, the data efficiency training module initiated the pseudo-annotation enhancement path, using P for enhancement training, thereby expanding the effective sample size of L+P to 269.

[0060] After the pseudo-annotation enhancement training is completed, the basic image processing module is updated to obtain a new model version M1. In subsequent work, the deviation between the generated C0 and the doctor's annotation tendency is significantly reduced. When the system re-annotates some samples, the number of corrections made by the doctor based on C0 is significantly reduced. The system directly generates lesion area attention data that is close to the doctor's expectations in some samples, thereby reducing the amount of interactive correction.

[0061] Example 3: In the actual use of lesion image annotation and evaluation, there are significant differences in the habits and operating methods of different doctors. Existing systems usually output lesion area focus data in a uniform style, which cannot automatically adapt to the individual usage habits of doctors. This causes doctors to repeatedly correct the uniform output during the annotation process, thereby increasing the annotation cost and work delay.

[0062] On the other hand, due to factors such as the high cost of manual annotation of lesion areas, the scale of data available for training basic image processing models is limited. Training with only a small number of labeled samples will affect the availability of data of interest in the initial lesion areas. It is necessary to improve the utilization rate of unlabeled and pseudo-labeled data during model training to improve the consistency between the initial output and the doctor's annotation preferences.

[0063] Therefore, it is necessary to provide a lesion image evaluation method that enables the system to perform personalized calibration of lesion area focus data according to the doctor's usage habits, and to achieve efficient training of the basic image processing module through a combination of unlabeled data, labeled data and pseudo-labeled data, so as to reduce the doctor's correction cost and improve the usability of the model output.

[0064] This embodiment discloses the following method: Step 1: Generate initial lesion area focus data from the input esophageal image using a basic image processing module; Step 2: Collect the doctor's correction operations on the initial lesion area focus data and structure the correction operations into a correction vector field; Step 3: Train a low-dimensional preference function model based on the correction vector field to obtain the corresponding individual preference parameters of the doctor; Step 4: When the doctor uses the data subsequently, calibrate the initial lesion area focus data based on the individual preference parameters to output personalized lesion area focus data, and further includes: Step 5: Perform self-supervised training and supervised fine-tuning of the basic image processing module using unlabeled esophageal images and a small number of labeled esophageal images; Step 6: Perform enhanced training of the basic image processing module using pseudo-labeled esophageal images generated by the doctor's corrections.

[0065] The lesion image assessment method provided in this embodiment runs on a hospital local area network server, and doctors perform image loading, correction, and confirmation operations through workstation terminals.

[0066] In step one, the doctor loads the esophageal image I to be processed on the terminal. The basic image processing module performs feature extraction and region attention inference on image I to generate initial lesion region attention data C0. C0 uses a binary mask of the same size as the original image to represent the suggested range of the lesion region.

[0067] In step two, the doctor performs a correction operation on C0 in the graphical interface to obtain the corrected lesion area annotation data Cu. The system compares C0 and Cu pixel by pixel to obtain positional offset and regional difference information, and structures this information into a correction vector field V to record the correction behavior. In step three, the system selects correction vector field samples corresponding to the current doctor's identifier from multiple historical correction records, trains a low-dimensional preference function model based on these samples, and extracts the model parameters as the individual preference parameter theta_d for the doctor.

[0068] In step four, when the doctor processes new esophageal images, the system generates a new C0 and calls theta_d corresponding to the doctor's identifier. The system then calibrates C0 using a preference function model and outputs personalized lesion area focus data Cd to reduce the amount of correction required by the doctor.

[0069] In step five, the system utilizes unlabeled esophageal image data and a small amount of labeled esophageal image data in the data efficiency training module, and performs self-supervised training and supervised fine-tuning respectively, thereby improving the availability of C0 generated by the basic image processing module.

[0070] In step six, the corrected annotation sample Cu generated by the doctor in step two is added to the pseudo-annotated dataset to enhance the training of the basic image processing module, so that the subsequently generated C0 is closer to the annotation habits in the scene and further reduces the amount of correction.

[0071] The terminology mentioned in this embodiment is as follows: Initial lesion area focus data C0: The lesion area suggestion result generated by the basic image processing module based on the input image.

[0072] Correction vector field V: Represents the two-dimensional offset field of the doctor's correction behavior, used to characterize the offset difference of Cu relative to C0.

[0073] Individual preference parameter theta_d: Individual doctor parameters used to calibrate C0, obtained from training the preference function.

[0074] Personalized lesion area focus data Cd: The lesion area focus result obtained by the system after loading theta_d and calibrating C0.

[0075] Unlabeled esophageal images: Images without labeled lesion areas, used for self-supervised training.

[0076] Annotated esophageal images: Images with annotated lesion areas are used to supervise fine-tuning training.

[0077] Pseudo-annotated esophageal images: The annotation results generated after correction by the doctor are used to enhance training.

[0078] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A lesion image assessment system for esophageal cancer, characterized in that: include: The basic image processing module is used to generate initial lesion area focus data from the input esophageal image; The doctor interaction data acquisition module is used to collect the doctor's correction operations on the initial lesion area attention data; The preference learning module is used to learn the individual preference parameters of the corresponding doctor based on the correction operation; The personalized output module is used to calibrate the initial lesion area attention data based on the individual preference parameters and output personalized lesion area attention data.

2. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: The doctor interaction acquisition module is used to structure the doctor's correction operation into a correction vector field representing the difference between the initial lesion area attention data and the corrected lesion area annotation data.

3. The esophageal cancer lesion image assessment system according to claim 2, characterized in that: The preference learning module is used to train a low-dimensional preference function model based on the modified vector field. The individual preference parameters are the parameters of the preference function model, and their parameter dimension is lower than that of the basic image processing module.

4. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: The individual preference parameters are used to calibrate the initial lesion area focus data in subsequent image processing to reduce the amount of correction required by the physician.

5. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: The individual preference parameters are stored separately with doctor identifiers, and when doctors use them, they call the corresponding individual preference parameters to generate personalized lesion area attention data.

6. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: It also includes a data-efficient training module for training the basic image processing module based on unlabeled esophageal images and a small number of labeled esophageal images.

7. The esophageal cancer lesion image assessment system according to claim 6, characterized in that: The data-efficient training module further utilizes pseudo-annotated esophageal images generated by doctors for enhanced training of the basic image processing module.

8. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: The efficient data training module includes a self-supervised training unit, a supervised fine-tuning unit, and a pseudo-labeling enhancement unit.

9. The esophageal cancer lesion image assessment system according to claim 1, characterized in that: The system processing method includes the following steps: Step 1: Using the basic image processing module, generate initial lesion area focus data from the input esophageal image; Step 2: Collect the doctor's correction operations on the initial lesion area data, and structure the correction operations into a correction vector field; Step 3: Train a low-dimensional preference function model based on the modified vector field to obtain the individual preference parameters of the corresponding doctor; Step 4: When the doctor uses the data subsequently, the initial lesion area focus data is calibrated based on the individual preference parameters to output personalized lesion area focus data.

10. The esophageal cancer lesion image assessment system according to claim 9, characterized in that: The processing method also includes: Step 5: Perform self-supervised training and supervised fine-tuning of the basic image processing module using unlabeled esophageal images and a small number of labeled esophageal images; Step Six: Enhance the training of the basic image processing module using pseudo-annotated esophageal images generated by the doctor.