Automatic segmentation quality control method and related devices

By combining visual language models with expert consensus, efficient, accurate, and interpretable quality control of automatic segmentation results in radiotherapy was achieved, solving the accuracy and interpretability issues of automatic segmentation models and improving the standardization and efficiency of quality control.

CN121505397BActive Publication Date: 2026-07-21MEVION MEDICAL EQUIPMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MEVION MEDICAL EQUIPMENT CO LTD
Filing Date
2025-11-17
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

The accuracy of existing automatic segmentation models in radiotherapy is insufficient, especially when faced with anatomical variations and imaging artifacts, which can easily lead to target omissions or oversegmentation. Furthermore, they lack interpretability, making clinical application difficult. Manual quality control is inefficient and difficult to standardize, and expert consensus has not been effectively integrated.

Method used

A pre-trained visual language model is used for cross-modal analysis. By integrating expert consensus knowledge, a medical image and text dataset is constructed to achieve efficient, accurate, and interpretable quality control of the automatic segmentation results and generate a quality control report.

Benefits of technology

It improves the accuracy and interpretability of automatic segmentation results, reduces inter-individual variability, achieves second-level quality control, optimizes the clinical decision-making process, and enhances doctors' trust and acceptance of AI results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505397B_ABST
    Figure CN121505397B_ABST
Patent Text Reader

Abstract

The application discloses an automatic segmentation quality control method and related equipment, and relates to the technical field of radiotherapy. The method comprises the following steps: acquiring a target area automatic segmentation result of a medical image; performing automatic quality evaluation on the target area segmentation result based on a pre-trained visual language model; wherein the visual language model is trained by a medical image-text data set constructed by fusing expert consensus knowledge, can perform cross-modal correlation analysis on the segmentation result and the quality control standard described by text to identify potential bias; generating a quality control report containing a quality control score and modification suggestions based on the automatic quality evaluation result; and converting the radiotherapy expert consensus into machine-understandable quality control standards and performing cross-modal analysis by using the visual language model, so that efficient, accurate and interpretable automatic quality control of the automatic segmentation result is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radiotherapy technology, and in particular to an automated segmentation quality control method and related equipment. Background Technology

[0002] Radiation therapy is one of the core methods of cancer treatment, and its success depends heavily on the precise delineation of the target area (tumor) and surrounding organs at risk. This step directly determines whether the radiation dose can accurately cover the tumor tissue while maximizing the protection of normal tissue, and is crucial to the treatment effect and patient safety.

[0003] However, current clinical practice faces multiple challenges in the delineation and quality control stages: While deep learning-based automatic segmentation models can improve delineation efficiency, their application in real-world clinical scenarios still faces significant challenges. First, the models lack accuracy, especially when dealing with anatomical variations, imaging artifacts, or rare cases, easily leading to biases such as undersegmentation (missed target areas) or oversegmentation (inclusion of too much normal tissue). Second, the model output lacks interpretability; its decision-making process is like a black box, making it difficult for clinicians to understand and trust the results. Furthermore, the models have poor adaptability to complex structures, failing to guarantee high reliability in all situations, thus limiting their widespread clinical application.

[0004] Traditional manual segmentation and quality control processes have inherent bottlenecks. On the one hand, this process heavily relies on the personal experience of the reviewers, leading to significant individual and institutional differences in the segmentation results among different doctors and medical institutions, making standardization difficult. On the other hand, manual quality control is extremely inefficient, requiring doctors to meticulously review the automated segmentation results layer by layer and describe each deviation in detail—an extremely time-consuming and labor-intensive process. Under heavy clinical workloads, this model is unsustainable, making it difficult to adequately guarantee the accuracy and quality control of the segmentation.

[0005] While expert consensus in the field of radiotherapy provides valuable reference for standardized delineation by integrating historical delineation data from multiple experts, current technologies have failed to effectively integrate this intellectual wealth with automated segmentation quality control. Most methods still rely on consensus as a simple reference, with the final judgment of automated delineation results primarily depending on manual intervention. This fails to create an efficient, interpretable, and standardized automated quality control process, thus failing to fully realize the value of expert consensus. Summary of the Invention

[0006] Based on the above problems, the purpose of this invention is to provide an automatic segmentation quality control method and related equipment, which transforms the consensus of radiotherapy experts into machine-understandable quality control standards, and uses a visual language model to perform cross-modal analysis, thereby achieving efficient, accurate and interpretable automated quality control of automatic segmentation results.

[0007] In a first aspect, the present invention provides an automatic segmentation quality control method, the method comprising: Obtain the automatic segmentation results of the target area in medical images; Based on a pre-trained visual language model, the target region segmentation results are automatically evaluated for quality. The visual language model is trained using a medical image and text dataset that integrates expert consensus knowledge. It can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases. Based on the results of automated quality assessment, a quality control report is generated that includes quality control scores and modification suggestions.

[0008] Preferably, the pre-trained visual language model-based automated quality assessment of the target region segmentation results includes: The visual features of the automatic segmentation result are extracted using the image encoder of the visual language model. Construct a text hint library based on expert consensus, including positive text hints describing standard delineation features and negative text hints describing delineation error patterns; The extracted visual features are compared with the text features of the text prompts to perform similarity calculation and matching analysis, and the quality control points and potential deviation types most relevant to the current segmentation results are identified.

[0009] Preferably, the method for constructing the text prompt library includes: Based on expert consensus on the delineation of radiotherapy target areas and organs at risk, clinical guidelines, and delineation data that has been reviewed and confirmed in the past, a structured quality control knowledge base is constructed. The generated descriptive criteria should include positive textual prompts for anatomical features, including: shape regularity, boundary clarity, spatial relationship with adjacent key anatomical structures, and size conformity. Generate negative text hints describing error patterns, including: oversegmentation, undersegmentation, jagged edges, inclusion of sensitive tissues that should not be included, and omission of critical areas.

[0010] Preferably, the step of performing similarity calculation and matching analysis between the extracted visual features and the text features of the text prompt includes: Calculate the similarity between the visual features and the text features of each text prompt; The calculated similarity is normalized to obtain the probability distribution of the visual features matching each text prompt; Based on the probability distribution, the positive text prompts with the highest matching probability to the current segmentation result are identified as the most relevant quality control points, and the negative text prompts with the highest matching probability are identified as the most significant potential deviation types.

[0011] Preferably, a multimodal attention mechanism is used to calculate the similarity between the visual features and the text features of each text prompt, including: Cross-modal attention weighting is applied to the visual and text features to generate enhanced visual-text joint features; Based on the joint features, a fine-grained similarity score between the visual features and each text prompt is calculated; The calculation process of cross-modal attention weighting synchronously outputs an attention weight map, which is used to identify the pixel region in the medical image that contributes the most to generating the current similarity score.

[0012] Preferably, the training process of the visual language model includes: Collect medical image segmentation results and their corresponding expert-annotated text descriptions to construct a medical image and text dataset; wherein, the expert-annotated text descriptions are generated based on clinical guidelines, expert consensus documents and historical quality control records. The visual language model is trained using a contrastive learning framework based on the medical image and text dataset, so that the model learns to align the matched segmentation result images with the quality control text descriptions in the feature space.

[0013] Preferably, the medical image and text dataset contains multi-level supervision signals, including: delineated images annotated by experienced physicians, delineated images annotated by domain experts, and text descriptions provided by domain experts after evaluating the physicians' delineations; The training process employs a hybrid loss function that combines contrastive learning loss and regression prediction loss.

[0014] Optionally, the hybrid loss function is constructed through the following steps: Using the image encoder of the visual language model, the outlined images annotated by the experienced doctor and the outlined images annotated by the domain expert are mapped into image embedding vectors, respectively. The text encoder of the visual language model is used to map the text description provided by the domain expert into a text embedding vector. The image embedding vectors and text embedding vectors are projected into a shared contrastive learning space; In the contrastive learning space, the similarity between all image embedding vectors and text embedding vectors is calculated and combined to obtain a similarity matrix; The similarity matrix is ​​calculated using the InfoNCE loss function to obtain the contrastive learning loss between the image drawn by the experienced doctor and the text description by the domain expert. The similarity matrix is ​​processed using convolutional regression to predict quality control scores, and the mean squared error loss between the predicted scores and the expert's actual scores is calculated. The contrastive learning loss and the mean squared error loss are weighted and combined to obtain the hybrid loss function.

[0015] Optionally, the hybrid loss function is constructed through the following steps: In the contrastive learning space, the image-image similarity matrix between the delineated images annotated by experienced doctors and those annotated by domain experts is calculated simultaneously. Based on the image-image similarity matrix, an additional image contrast learning loss term is used to constrain the consistency of expert annotations at different levels in the feature space, forming a triple contrast learning structure of image-text-image.

[0016] Preferably, the expert-annotated text description covers quality control information in at least one of the following dimensions: regularity of anatomical structure shape, clarity of boundaries, rationality of location, adequacy of target area coverage, and compliance of organ avoidance.

[0017] Preferably, the method further includes: Through the interactive interface, the user can receive modifications to the segmentation results from at least one doctor. Automated quality assessment can be performed based on the modified segmentation results, or the modified segmentation results can be sent to experts for approval; the contour data before and after modification, the identity of the modifier, and the modification timestamp can be recorded.

[0018] Preferably, the method further includes: If the quality control score is greater than or equal to the first score threshold, the current segmentation result will be sent directly to the expert for approval. If the quality control score is less than the first score threshold, the quality control report will be provided to the doctor, who will be guided to make targeted corrections based on the modification suggestions in the report. After the corrections are completed, the results will be sent to the experts for approval. If the score is still lower than the first score threshold after the same segmentation result has been corrected more than a preset number of times, the expert upgrade mechanism will be automatically triggered.

[0019] Preferably, the method further includes: When multiple experts have differing opinions on the same segmentation result, collect the text descriptions and modification records of each expert. By using clustering or voting mechanisms to integrate the opinions of multiple experts, more robust quality control standards are generated, and the text prompt library is updated. Inconsistent cases are recorded for model retraining.

[0020] Secondly, the present invention provides an automatic segmentation quality control device, the device comprising: The first acquisition module is used to acquire the automatic target area segmentation results of medical images; The quality assessment module is used to automatically assess the quality of target area segmentation results based on a pre-trained visual language model. The visual language model is trained using a medical image and text dataset constructed by integrating expert consensus knowledge. It can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases. The report generation module is used to generate quality control reports that include quality control scores and modification suggestions based on the results of automated quality assessments.

[0021] Thirdly, the present invention provides a radiotherapy system, comprising: The automatic segmentation quality control device described in this embodiment of the invention.

[0022] Fourthly, the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the methods of the present invention or the functions of the apparatus of the present invention.

[0023] Fifthly, the present invention provides a computer-readable storage medium storing computer instructions, wherein when a computer reads the computer instructions, the computer executes the steps of any of the methods described in the present invention.

[0024] Compared with existing technologies, the beneficial effects of this invention include at least the following: Through cross-modal correlation analysis of visual language models, it can deeply understand the anatomical semantics of segmentation results and identify complex biases that are difficult to detect using traditional methods; the output is no longer a simple pass / fail or a single score, but a detailed report containing specific bias types, severity, correction suggestions, and decision-making basis, greatly enhancing doctors' trust and acceptance of AI results. Automated evaluation achieves second-level quality control, freeing doctors from the tedious work of layer-by-layer review. The embedded expert consensus knowledge base ensures the objectivity and uniformity of quality control standards, effectively eliminating subjective judgment differences between different doctors and institutions, and promoting the standardization of quality control. Through quality control scoring thresholds and expert upgrade mechanisms, the system can automatically triage cases, guide workflows, and focus expert resources on the most complex and difficult cases, optimizing the overall clinical decision-making process. The interactive interface seamlessly integrates doctors' experiential modifications and AI-guided corrections into the workflow, fully leveraging the respective advantages of humans and AI to ensure the highest quality of the final delineation results. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the automatic segmentation quality control method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the automatic segmentation quality control process according to an embodiment of the present invention. Detailed Implementation

[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided to make the invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and therefore repeated descriptions of them will be omitted.

[0027] The terms used to express position and direction in this invention are illustrated with reference to the accompanying drawings, but changes can be made as needed, and all such changes are included within the scope of protection of this invention.

[0028] Example 1: This example provides an automatic segmentation quality control method, including the following steps: S1. Obtain the automatic segmentation results of the target area in medical images; S2. Based on a pre-trained visual language model (e.g., CLIP), perform automated quality assessment on the target region segmentation results; wherein, the visual language model is trained by a medical image and text dataset constructed by integrating expert consensus knowledge, and can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases; S3. Based on the analysis results of the aforementioned evaluation steps, generate a quality control report containing modification suggestions; In one possible implementation, the automatic segmentation result is achieved in the following way: Receives the patient's medical imaging data; the patient's medical imaging data includes images such as CT and MRI scans. The medical image data is processed using a deep learning segmentation model to generate automatic target region segmentation results.

[0029] The working principle of the above technical solution is as follows: A cross-modal association between medical image segmentation results and textual quality control standards is established using a visual language model pre-trained with expert consensus knowledge. Specifically, the acquired automatic segmentation results are input into the model, which, through its inherent deep learning architecture, performs semantic-level association analysis between the image-based segmentation contours and textual quality control standards (such as clear boundaries and oversegmentation). This process does not rely on complex rule writing; instead, the model directly understands and judges the degree of conformity between the segmentation results and the standards, thereby automatically identifying potential delineation deviations. Finally, based on this analysis result, a structured quality control report is automatically generated.

[0030] The beneficial effects of the above technical solution are as follows: Overcoming the bottleneck of low efficiency in traditional manual quality control, this system enables rapid, batch quality assessment of automatically segmented results, significantly liberating doctors' productivity. By embedding expert consensus knowledge into the pre-trained model, the quality control process is freed from excessive reliance on the personal experience of reviewers, effectively reducing discrepancies and misjudgments caused by human factors, and improving the reliability and consistency of results. The generated quality control report not only includes scores but also provides specific modification suggestions, making the quality control conclusions no longer black-box judgments but intelligent diagnoses with semantic interpretation, significantly enhancing the system's usability and doctors' trust.

[0031] In one possible implementation, the pre-trained visual language model in step S2 performs automated quality assessment of the target region segmentation results, including: S21. Using the image encoder of the visual language model, extract the visual features of the automatic segmentation result; S22. Construct a text prompt library based on expert consensus, including positive text prompts describing standard drawing features and negative text prompts describing typical drawing error patterns. S23. Perform similarity calculation and matching analysis between the extracted visual features and the text features of the text prompt to identify the quality control points and potential deviation types most relevant to the current segmentation result.

[0032] The working principle of the above technical solution is as follows: First, an image encoder based on a visual language model is used to extract depth visual features that characterize the shape, boundaries, and spatial relationships of the automatic segmentation results. Simultaneously, a text cue library built based on expert consensus includes standardized text describing high-quality and low-quality delineation features.

[0033] Subsequently, the similarity between the computational visual features and all text quality control standard features is essentially used to locate the current segmentation result in a quality control standard semantic space composed of expert knowledge. By finding the most matching positive and negative text, it is possible to accurately identify which positive standards the result best meets and which error type it is most likely to belong to, thereby achieving automated and semantic quality assessment.

[0034] The effects of the above technical solution are as follows: Instead of pre-setting complex geometric or intensity rules, it identifies deviations by comparing the deep semantic relationships between images and text. Its evaluation logic is closer to the thinking of human experts, enabling it to detect more complex and abstract drawing errors. The text prompt library can be easily added, deleted, and modified. When new expert consensus is formed or new quality control dimensions need to be added, only the corresponding text descriptions need to be added or modified in the prompt library, without redesigning or training complex image processing algorithms. The system is highly adaptable and maintainable. After a one-time visual feature extraction, rapid similarity matching calculations can be performed with a massive amount of text prompts. This architecture ensures comprehensive evaluation without significantly increasing computational overhead.

[0035] In one possible implementation, the method for constructing the text prompt library includes: S221. Based on expert consensus on the delineation of radiotherapy target areas and organs at risk, clinical guidelines, and delineation data verified in the past, construct a structured quality control knowledge base. S222. The generated descriptive standard outline should have positive textual prompts for anatomical features, including: shape regularity, boundary clarity, spatial relationship with adjacent key anatomical structures, and size conformity. S223. Generate negative text prompts describing typical drawing error patterns, including: oversegmentation, undersegmentation, jagged boundaries, inclusion of sensitive tissues that should not be included, and omission of key areas.

[0036] The working principle of the above technical solution is as follows: Unstructured, experience-dependent expert knowledge is systematically transformed into structured, machine-understandable semantic quality control standards. A structured quality control knowledge base is built upon this foundation by integrating radiotherapy expert consensus, clinical guidelines, and high-quality historical data. The knowledge base is then activated through two key steps: Define excellence criteria: Generate positive textual cues describing the anatomical features (such as shape, boundaries, and spatial relationships) that the model should possess. This provides a clear quality benchmark for the model.

[0037] Define an error graph: Generate negative textual hints describing typical delineation error patterns (such as oversegmentation, undersegmentation, etc.). This provides the model with an error dictionary to identify biases.

[0038] This construction process essentially creates a complete quality control semantic query dictionary covering both positive and negative aspects for subsequent visual language models, enabling AI to conduct targeted reviews of the segmentation results based on these specific text descriptions.

[0039] The effects of the above technical solution are as follows: The text prompt library is directly rooted in authoritative expert consensus and clinical guidelines, ensuring that the entire automated quality control process strictly adheres to clinical standards, fundamentally guaranteeing the clinical validity and reliability of the assessment results. By systematically generating positive and negative prompts, it covers multiple key quality control dimensions from macroscopic (size compliance) to microscopic (jawy boundaries), avoiding potential oversights or biases by individual physicians during review, making quality control checks more comprehensive and systematic. Since each suggestion in the quality control report is directly linked to a specific text prompt defined by expert consensus (e.g., detecting undersegmentation bias), the conclusions are no longer abstract numerical values ​​but diagnoses with clear clinical significance, providing physicians with clear and direct directions for correction. Transforming complex clinical knowledge engineering problems into relatively standardized text generation and management tasks allows domain experts to directly participate in the construction and optimization of quality control standards without requiring in-depth AI algorithm knowledge, greatly promoting the deep integration of artificial intelligence and clinical experience.

[0040] In one possible implementation, the step of performing similarity calculation and matching analysis between the extracted visual features and the text features of the text prompt includes: S231. Calculate the cosine similarity between the visual features and the text features of each text prompt; S232. Normalize the calculated cosine similarity to obtain the probability distribution of the visual feature matching each text prompt; S233. Based on the probability distribution, identify the positive text prompts with the highest matching probability to the current segmentation result as the most relevant quality control points, and the negative text prompts with the highest matching probability as the most significant potential deviation types.

[0041] The working principle of the above technical solution is as follows: First, cosine similarity is calculated to numerically measure the proximity of the visual features of the automatic segmentation result to each text quality control standard. Then, these similarity values ​​are normalized and transformed into a probability distribution. This probability distribution clearly quantifies the likelihood of the current segmentation result matching each quality control standard (whether positive or negative). From this probability distribution, the positive and negative prompts with the highest matching probabilities are selected. This is equivalent to performing an automated diagnosis, clearly indicating which quality standard the current result best matches and which error risk requires the most vigilance, thus condensing complex multi-dimensional feature matching into clear, actionable quality control points and deviation alerts.

[0042] The effects of the above technical solution are as follows: Transforming quality control judgments from qualitative yes / no or similar / unsimilar statements into continuous probability values ​​makes the assessment results more refined and quantifiable, facilitating physicians' understanding of the confidence level of the quality control conclusions. By focusing on the option with the highest probability, a large amount of interfering information can be eliminated, directly identifying the core quality control advantages and the most serious potential problems. This provides physicians with extremely clear and unambiguous guidance on correction priorities, avoiding information overload. The output is no longer a general evaluation of poor quality, but a specific diagnosis most likely to have oversegmentation bias. This ability to accurately pinpoint the problem type makes the generated modification suggestions incisive, greatly enhancing the clinical guidance value of the report. The entire similarity calculation and probability normalization process has low computational overhead and is fast, allowing this quality control method to be seamlessly integrated into clinical workflows, providing physicians with near real-time feedback and supporting efficient interactive correction.

[0043] In one possible implementation, a multimodal attention mechanism is used to calculate the similarity between the visual features and the text features of each text prompt, including: Cross-modal attention weighting is applied to the visual and text features to generate enhanced visual-text joint features; Based on the joint features, a fine-grained similarity score between the visual features and each text prompt is calculated; The calculation process of cross-modal attention weighting synchronously outputs an attention weight map, which is used to identify the pixel region in the medical image that contributes the most to generating the current similarity score.

[0044] The working principle of the above technical solution is as follows: By leveraging cross-modal attention mechanisms, we can achieve dynamic interaction and alignment of visual and textual features, thereby enabling accurate and interpretable similarity calculation.

[0045] Specifically, a cross-modal attention weighting process is introduced. This mechanism dynamically focuses on and weights the most relevant spatial regions in visual features based on the semantics of the current text prompt; conversely, it can also focus on the most critical words in the text prompt based on the visual content. Through this interaction, an enhanced visual-text joint feature is generated, which integrates key information from both sides.

[0046] Based on this deeply integrated joint feature, further similarity calculation yields a fine-grained similarity score that more accurately reflects the degree of matching between the image and text at a specific semantic level. More importantly, the attention weight map naturally generated during the attention weighting process intuitively reveals which specific regions in the image play a decisive role in the final quality control judgment.

[0047] The effects of the above technical solution are as follows: Through dynamic feature interaction, the model can focus on the local regions most relevant to specific quality control standards (e.g., focusing only on boundaries to assess boundary clarity), avoiding information interference from global features, making similarity calculation and deviation identification more accurate and detailed. The attention weight map, as a visualization tool, can directly overlay the model's decision-making basis onto the original image in the form of a heatmap, clearly indicating which pixel regions the model judged as undersegmented. This white-box output greatly enhances doctors' understanding and trust in the AI's judgment. For complex situations such as blurred boundaries between tumors and normal tissue, the attention mechanism helps the model learn to focus on the most discriminative radiographic features, thus making more robust and reliable judgments, overcoming the poor performance of traditional methods in marginal cases. When receiving quality control reports, doctors can quickly verify the reasonableness of the AI's judgment by combining the attention map, thus focusing their review on the most questionable areas, significantly improving the efficiency and depth of human-machine collaborative review.

[0048] In one possible implementation, after step S233, the method further includes: S234. Based on the matching probability value with the negative text prompt, classify the severity of the identified potential bias; wherein, the severity is classified based on the matching probability value with the negative text prompt corresponding to the most significant identified potential bias; the higher the matching probability value, the higher the severity level is determined. S235. Associate the identified deviations with the specific quality control points in the text prompts to determine the root cause category of the deviations; S236. Based on the highest similarity score between the visual features and the matching text prompt features, output the confidence level of the quality control evaluation result; display the confidence level in the quality control report along with the corresponding quality control score and potential deviations to indicate the reliability of the evaluation result. The confidence level of the evaluation result is calculated based on the highest matching probability value in the probability distribution, or the difference between the highest matching probability and the second highest matching probability; wherein, the higher the highest matching probability value or the larger the probability difference, the higher the output confidence level; wherein, the confidence level is used to indicate the reliability of the quality control score and potential deviation information.

[0049] The working principle of the above technical solution is as follows: The preliminary matching analysis results are subjected to multi-dimensional post-processing and meta-evaluation, which elevates the single quality control judgment into a structured and quantifiable in-depth diagnostic report.

[0050] After identifying the most significant potential biases, the process does not stop there, but rather proceeds with post-processing steps: The severity of the error is graded based on the matching probability value of the deviation. The higher the matching probability, the more closely the current segmentation result matches the semantic features of the error pattern, and therefore the higher the severity level is determined. This not only allows for the qualitative identification of the problem but also the quantitative assessment of its severity.

[0051] The identified deviations (such as jagged edges) are mapped to predefined root cause categories in the text hint library (such as insufficient image resolution or poor model smoothing). This step connects surface symptoms to underlying causes.

[0052] A confidence level is calculated by analyzing the quality of the probability distribution (such as the magnitude of the highest probability value or the degree of concentration of the probability distribution). If the model is very confident in its judgment (the probability distribution is concentrated), it outputs a high confidence level; conversely, if the model itself is hesitant (the probability distribution is flat), it outputs a low confidence level, reminding the user to consider the information carefully.

[0053] The effects of the above technical solution are as follows: By grading severity, doctors can be guided to prioritize the most serious and significant problems, optimizing the processing order of clinical workflows, improving correction efficiency, and avoiding a haphazard approach. By associating deviations with root cause categories, quality control reports are upgraded from describing problems to diagnosing them, making modification suggestions more targeted (e.g., suggestions for insufficient image resolution are completely different from suggestions for model errors), greatly enhancing the clinical utility of the reports. Introducing confidence levels proactively informs users of the reliability of their judgments, issuing warnings when uncertain, preventing users from blindly trusting potentially erroneous AI suggestions, significantly enhancing the system's robustness and doctors' trust, and providing crucial reliable evidence for human-machine collaborative decision-making.

[0054] Metadata such as severity, root cause, and confidence level provides invaluable information for subsequent system optimization. For example, high-severity but low-confidence cases can be reviewed more thoroughly, or this data can be used to retrain the model, thereby driving the continuous iterative evolution of the entire system.

[0055] In one possible implementation, the matching analysis also employs an uncertainty-aware mechanism, including: Based on the probability distribution obtained from the automatic segmentation results, calculate the information entropy of the distribution; When the information entropy exceeds a first preset threshold, it is determined that the automatic segmentation result has overall semantic ambiguity. In response to the determination, an unreliable quality indicator is added to the generated quality control report; and doctors are advised to conduct a focused manual review.

[0056] The working principle of the above technical solution is as follows: Information entropy is a quantitative indicator in information theory that measures the uncertainty of a system. A concentrated and well-defined probability distribution (where the model is highly confident) has low information entropy, while a flat and dispersed probability distribution (where the model is hesitant and considers multiple quality control standards to be possible) has high information entropy. When the information entropy exceeds a preset threshold, it is determined that the semantic features of the current automatic segmentation result are too ambiguous, causing the model to be unable to provide a confident quality assessment. In this case, instead of forcibly outputting a potentially unreliable quality control conclusion, a clear unreliable quality indicator is added to the quality control report, and human review is recommended, proactively returning the decision-making power to human experts. This is essentially a proactive risk management based on the system's self-awareness.

[0057] The effects of the above technical solution are as follows: By proactively deflecting criticism when encountering complex, ambiguous, or cases not covered by training data, AI effectively prevents the potential risks of applying unreliable AI recommendations in clinical practice.

[0058] By clearly identifying cases requiring priority review, the system provides doctors with a clear guide to work priorities. Doctors can focus their valuable time on the most challenging cases that the system is unsure about, while quickly processing cases that the system is confident in, that are of high quality, or that have clearly defined issues. This achieves optimal allocation of human resources while ensuring quality, thus improving the efficiency of the overall workflow.

[0059] All cases marked as high uncertainty serve as valuable negative examples and samples for optimization. These cases can be systematically collected for incremental training of subsequent models or algorithm iterations, thereby specifically improving the system's performance when handling similar difficult cases in the future, forming a closed-loop optimization.

[0060] In another possible implementation, the matching analysis also employs an uncertainty-aware mechanism, including: Obtain the probability distribution of visual features matching multiple text prompts; Select the K text prompts with the highest probabilities from the probability distribution, where K ≥ 2; If the K text prompts contain both positive and negative text prompts, and the probability difference between the highest probability positive text prompt and the highest probability negative text prompt is less than a preset difference threshold, then semantic ambiguity is determined and marked, and the doctor is advised to conduct a focused manual review.

[0061] The working principle of the above technical solution is as follows: After obtaining the matching probability distribution of all text prompts, the focus is on the K candidate prompts with the highest probabilities. A thorough analysis is then conducted to determine the logical consistency among these most likely options. When it is found that top-ranked candidate prompts simultaneously contain highly credible positive evaluations (e.g., regular shape) and highly credible negative evaluations (e.g., undersegmentation), and their probability values ​​are very close and difficult to distinguish, semantic ambiguity is identified. This situation indicates that the automatic segmentation results themselves may have contradictory characteristics, causing the model to fall into a decision-making dilemma and unable to provide a single, clear quality control conclusion. In this case, proactively identifying this ambiguity and suggesting manual review is essentially a more advanced form of uncertainty perception based on the detection of logical conflict in decision-making.

[0062] The effects of the above technical solution are as follows: Compared to simply calculating information entropy, this method's uncertainty report is more interpretable. It not only tells doctors that the results are unreliable, but also reveals that the unreliability is due to decision conflicts within the model, and clearly identifies which specific quality control standards (such as a positive cue and a negative cue) are contradictory, providing doctors with a clear entry point and direction for review.

[0063] In one possible implementation, the training process of the visual language model includes: Collect medical image segmentation results and their corresponding expert-annotated text descriptions to construct a medical image and text dataset; wherein, the expert-annotated text descriptions are generated based on clinical guidelines, expert consensus documents and historical quality control records. The visual language model is trained using a contrastive learning framework based on the medical image and text dataset, so that the model learns to align the matched segmentation result images with the quality control text descriptions in the feature space.

[0064] The working principle of the above technical solution is as follows: The core principle of this implementation lies in transforming human experts' quality control knowledge into the model's inherent cross-modal understanding capabilities through contrastive learning. The training process consists of two key stages: First, the system constructs a high-quality medical image-text dataset. The core value of this dataset lies in accurately pairing medical image segmentation results with text descriptions generated from authoritative knowledge such as clinical guidelines and expert consensus. This essentially transforms abstract expert experience and quality control standards into machine-readable semantic labels corresponding to specific images. Subsequently, within the contrastive learning framework, the model is forced to learn a core capability: within a shared feature space, it narrows the vector distance between matching image-text pairs (e.g., a segmentation result with blurred boundaries and its corresponding blurred text description), while simultaneously distancing itself from mismatched pairs (e.g., an image with blurred boundaries and text with regular shapes). Through training on a large amount of data, the model gradually internalizes expert consensus and ultimately learns to accurately associate visual segmentation features with semantic quality control standards, laying the foundation for subsequent automated quality control assessment.

[0065] The effects of the above technical solution are as follows: Because the model learns directly from authoritative clinical guidelines and expert consensus, the image-text association it establishes is essentially an encoding of expert knowledge. This ensures that the judgments made by the model in subsequent quality control assessments naturally conform to clinical standards, guaranteeing the reliability of the results from the source.

[0066] Contrastive learning training enables the model to understand the deeper meaning of quality control texts. It learns not only image features but also the semantics corresponding to these features. Therefore, it can understand the various visual manifestations of the concept of blurred boundaries in different images and different organs, achieving the ability to generalize from one example to another, and can also perform effective semantic assessments of unknown cases.

[0067] In one possible implementation, the medical image and text dataset is constructed for a specific anatomical location or disease type, including at least one of the following: head and neck tumors, chest tumors, abdominal tumors, and pelvic tumors; This is a medical image and text dataset for pelvic tumors, including text descriptions of quality control standards for delineating organs at risk such as the prostate, bladder, and rectum.

[0068] In one possible implementation, the medical image and text dataset contains multi-level supervision signals, including: delineated images annotated by experienced physicians, delineated images annotated by domain experts, and text descriptions provided by domain experts after evaluating the physicians' delineations. The training process employs a hybrid loss function that combines contrastive learning loss and regression prediction loss.

[0069] In one possible implementation, the hybrid loss function is constructed through the following steps: Using the image encoder of the visual language model, the outlined images annotated by the experienced doctor and the outlined images annotated by the domain expert are mapped into image embedding vectors, respectively. The text encoder of the visual language model is used to map the text description provided by the domain expert into a text embedding vector. The image embedding vectors and text embedding vectors are projected into a shared contrastive learning space; In the contrastive learning space, the cosine similarity between all image embedding vectors and text embedding vectors is calculated and combined to obtain a similarity matrix. The similarity matrix is ​​calculated using the InfoNCE loss function to obtain the contrastive learning loss between the image drawn by the experienced doctor and the text description by the domain expert. The similarity matrix is ​​processed using convolutional regression to predict quality control scores, and the mean squared error loss between the predicted scores and the expert's actual scores is calculated. The contrastive learning loss and the mean squared error loss are weighted and combined to obtain the hybrid loss function.

[0070] In practical applications: Using the visual language model image encoder, the drawn images in the dataset with domain expert annotations are processed. This is mapped to a fixed-dimensional embedding vector. ; Using the image encoder of the visual language model, the delineated images annotated by experienced doctors in the dataset are processed. Mapped to a fixed-dimensional embedding vector ; The text encoder of the visual language model is used to process the annotated text description drawn by the doctor after evaluation by domain experts. Mapped to another fixed-dimensional embedding vector ; Using the projection layer of the visual language model, the embedding vectors derived from images annotated and drawn by experienced doctors are... and embedding vectors derived from expert evaluation text descriptions Simultaneously projected onto the comparative learning space Similarly, embedding vectors and Projected into the comparative learning space The encoder generated from the delineated image annotated by experienced doctors Image embedding and text encoder generated Text embedding Perform separately Normalize, then calculate the cosine similarity between all image embeddings and text embeddings. And combine them to obtain the similarity matrix Similarly, the similarity matrix between the images drawn by domain experts and those drawn by doctors is obtained. ; The similarity matrix is ​​calculated using the loss function InfoNCE. Obtain the image drawn by the doctor and expert evaluation text Similarity score ; and using convolutional regression to process the similarity matrix. Process the predicted quality control scores and compare them with the MSE loss corresponding to the expert scores. Then the total loss function is obtained. .

[0071] The effects of the above technical solution are as follows: By organically combining contrastive learning loss and regression loss within a unified mathematical framework, the model is forced to simultaneously optimize semantic alignment and quantization scoring capabilities. This multi-task collaborative training mechanism produces a "1+1>2" effect, enabling the final visual language model to not only deeply understand the complex relationship between the drawing results and quality control text, but also output highly accurate quality control scores. Its overall performance far surpasses that of models trained using only a single loss function.

[0072] Thanks to the constraint of mean squared error (MSE) regression loss, the model learns to map a continuous similarity matrix to a scalar quality control score, enabling the model to accurately quantify the quality differences between different delineation results, achieve fine-grained quality ranking of large batches of automatic segmentation results, and provide a reliable basis for priority processing in clinical workflows.

[0073] The core function of contrastive learning loss is to widen the distance between samples of different categories in the feature space, ensuring that the image and text embeddings learned by the model have high discriminativeness. This enables the model to effectively analyze and judge complex or rare cases not covered by the training data, thus possessing excellent generalization ability.

[0074] This formalized description explicitly defines the entire training process as a series of computable mathematical operations (such as embedding, projection, normalization, similarity calculation, and loss weighting). This standardized and quantitative description greatly facilitates the accurate understanding, reproduction, and iterative optimization of the scheme by those skilled in the art, laying a solid foundation for the promotion and application of the technology.

[0075] In another possible implementation, the hybrid loss function is constructed through the following steps: In the contrastive learning space, the image-image similarity matrix between the delineated images annotated by experienced doctors and those annotated by domain experts is calculated simultaneously. Based on the image-image similarity matrix, an additional image contrast learning loss term is used to constrain the consistency of expert annotations at different levels in the feature space, forming a triple contrast learning structure of image-text-image.

[0076] The working principle of the above technical solution is as follows: The core principle of this implementation lies in constructing a triple contrastive learning structure, which provides the model with denser and more robust cross-modal supervision signals by introducing image-image consistency constraints.

[0077] This mechanism adds a parallel image-image contrast learning channel to the existing image-text contrast learning. Specifically, in the shared contrast learning space, it not only calculates the similarity matrix between the doctor's drawn image and the expert's commentary text, but also simultaneously calculates the similarity matrix between the drawn image annotated by experienced doctors and the standard drawn image annotated by domain experts. This image-image similarity matrix is ​​used to calculate an additional image contrast learning loss term. The core objective of this loss term is to narrow the distance between the doctor's drawing and the corresponding expert's drawing in the feature space, enabling the model to learn to map the drawing results of doctors at different levels, but describing the same anatomical structure, to similar positions in the feature space. This forms a more robust triple contrast learning structure composed of image-text contrast and image-image contrast, constraining and guiding the model's learning direction from multiple dimensions.

[0078] The effects of the above technical solution are as follows: Image-image contrast loss acts as a strong constraint, forcing the model to ignore the differences in the personal style of the sketcher and focus on learning the most essential anatomical features in the sketching results that conform to clinical standards. This makes the visual feature representation learned by the model less sensitive to the fluctuations of the individual sketcher, more stable and reliable, and significantly improves the robustness of the model.

[0079] By simultaneously utilizing both image-text and image-image supervision signals, the model gains richer gradient information with each parameter update. This multi-teacher supervision mechanism can more effectively guide model optimization, typically accelerating convergence during training and helping the model overcome performance bottlenecks to achieve superior final performance. For complex cases with ambiguous boundaries and intricate structures, simple image-text association may not be clear enough. In such cases, direct visual comparison of images provides the model with a crucial additional frame of reference. By allowing the model to intuitively learn the differences between physician sketches and expert gold standards in the feature space, the model can more sensitively capture subtle morphological deviations that are difficult to describe in detail with text. High-quality expert sketches cluster in the central region, while physician sketches of different qualities are distributed in different positions corresponding to the semantics of the text description based on their similarity to the expert standard, providing an ideal feature basis for achieving more refined and automated quality grading and deviation diagnosis.

[0080] In one possible implementation, the expert-annotated text description covers quality control information in at least one of the following dimensions: regularity of anatomical structure shape, clarity of boundaries, rationality of location, adequacy of target area coverage, and compliance of organ avoidance.

[0081] By covering five core clinical dimensions—shape, boundary, location, target coverage, and organ avoidance—this system can provide a comprehensive, multi-faceted evaluation of the automatic segmentation results. This avoids the blind spots that may result from traditional methods due to their singular focus, ensuring the comprehensiveness and clinical credibility of the quality control conclusions.

[0082] See attached document Figure 2 In one possible implementation, the method further includes: Through the interactive interface, the user can receive modifications to the segmentation results from at least one doctor. Automated quality assessment can be performed based on the modified segmentation results, or the modified segmentation results can be sent to experts for approval; the contour data before and after modification, the identity of the modifier, and the modification timestamp can be recorded.

[0083] In practical applications: After step S1 and before step S2, step S1-1 is also included: receiving at least one doctor's modification of the segmentation result through an interactive interface, the modification being based on the doctor's clinical experience. In step S2, an automated quality assessment is performed on the target segmentation results modified by the doctor, based on a pre-trained visual language model. After step S3, the method further includes step S4, which involves receiving modifications to the segmentation results from at least one doctor through an interactive interface; these modifications are based on a quality control report; and the modified results are sent to experts for approval.

[0084] It leverages the objectivity and efficiency of AI while retaining the domain knowledge and flexibility of human experts, resulting in a final draft quality that far surpasses purely automated or purely manual results.

[0085] The system fully records the contour data before and after modification, along with the modifiers and their reasons (based on experience or reports). This high-quality data, with clearly defined intention annotations, is the gold standard dataset for training the next generation of more accurate visual language models that better meet clinical needs, driving the entire system to form a positive feedback loop where it becomes smarter with use.

[0086] This process clearly defines the responsibilities of junior physicians and senior specialists; junior physicians are responsible for initial revisions and refinements with AI assistance, handling most routine cases; while domain experts focus on the final approval of the most complex and critical results. This division of labor greatly improves overall work efficiency and frees experts from the burden of tedious initial reviews.

[0087] In one possible implementation, the method further includes: If the quality control score is greater than or equal to the first score threshold, the current segmentation result will be sent directly to the expert for approval. If the quality control score is less than the first score threshold, the quality control report will be provided to the doctor, who will be guided to make targeted corrections based on the modification suggestions in the report. After the corrections are completed, the results will be sent to the experts for approval. If the score of the same segmentation result is still lower than the first score threshold after being corrected more than a preset number of times, the expert upgrade mechanism will be automatically triggered. This mechanism will perform the following operations: Package the case's segmentation results, quality control reports, and correction records, and attach the system's preliminary analysis of the reasons for correction failures; The aggregated report will be directly sent to a senior or high-level radiotherapy expert, who will personally review and revise it and complete the final approval. Record such cases of repeated failed corrections and use them as data to optimize the first scoring threshold or train the visual language model.

[0088] The effects of the above technical solution are as follows: The system automatically assigns processing paths based on quality control scores, ensuring that expert resources are focused on the most complex and critical cases, while routine and high-scoring cases are processed quickly. This intelligent triage significantly optimizes human resource allocation and improves the efficiency and throughput of the overall clinical workflow.

[0089] The triggering of the expert upgrade mechanism means that the system automatically identifies cognitive blind spots or technical difficulties that neither it nor junior doctors can resolve. Recording and analyzing these cases provides invaluable data for optimizing the rationality of scoring thresholds and specifically enhancing the performance of the model in its weak areas, driving the system to achieve closed-loop optimization.

[0090] This mechanism forms a three-tiered quality control barrier: initial screening by AI, collaborative correction by doctors using AI, and final decision-making by experts. In particular, the expert upgrade mechanism, as the last line of defense, ensures that any difficult cases receive the highest level of professional handling, fundamentally preventing substandard sketches from entering the clinical treatment stage and significantly improving the safety and reliability of the entire quality control system.

[0091] In one possible implementation, the method further includes: When multiple experts have differing opinions on the same segmentation result, collect the text descriptions and modification records of each expert. By using clustering or voting mechanisms to integrate the opinions of multiple experts, more robust quality control standards are generated, and the text prompt library is updated. Inconsistent cases are recorded for model retraining.

[0092] When significant differences are detected in the approval opinions of multiple experts on the same segmentation result, this is considered a valuable boundary case. At this point, the textual evaluations and specific modification operations of each expert are actively collected to form a multi-expert decision-making dataset. Subsequently, cluster analysis is used to identify common patterns in expert opinions, or a voting mechanism is used to determine the mainstream opinion, thereby integrating the perspectives of multiple experts to generate a more universal and authoritative super consensus. This optimized new quality control standard is fed back into a text prompt library, enabling dynamic updates to quality control knowledge. Simultaneously, these inconsistent cases, recording the complete decision-making process, provide highly valuable challenging samples for the iterative training of the model.

[0093] Example 2: This embodiment provides an automatic segmentation quality control device, the device comprising: The first acquisition module is used to acquire the automatic target area segmentation results of medical images; The quality assessment module is used to automatically assess the quality of target area segmentation results based on a pre-trained visual language model. The visual language model is trained using a medical image and text dataset constructed by integrating expert consensus knowledge. It can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases. The report generation module is used to generate quality control reports that include quality control scores and modification suggestions based on the results of automated quality assessments.

[0094] In one possible implementation, the automatic segmentation result is achieved in the following way: Receive medical imaging data; The medical image data is processed using a deep learning segmentation model to generate automatic target region segmentation results.

[0095] In one possible implementation, the quality assessment module includes: A visual feature extraction unit is used to extract visual features of the automatic segmentation result using the image encoder of the visual language model; The prompt library building unit is used to build a text prompt library based on expert consensus, including positive text prompts describing standard delineation features and negative text prompts describing typical delineation error patterns; The analysis unit is used to perform similarity calculation and matching analysis between the extracted visual features and the text features of the text prompt, and to identify the quality control points and potential deviation types most relevant to the current segmentation result.

[0096] The method for constructing the text prompt library includes: Based on expert consensus on the delineation of radiotherapy target areas and organs at risk, clinical guidelines, and delineation data that has been reviewed and confirmed in the past, a structured quality control knowledge base is constructed. The generated descriptive criteria should include positive textual prompts for anatomical features, including: shape regularity, boundary clarity, spatial relationship with adjacent key anatomical structures, and size conformity. Generate negative text prompts describing typical drawing error patterns, including: oversegmentation, undersegmentation, jagged edges, inclusion of sensitive tissues that should not be included, and omission of key areas.

[0097] In one possible implementation, the analysis unit performs the following steps: Calculate the cosine similarity between the visual features and the text features of each text prompt; The calculated cosine similarity is normalized to obtain the probability distribution of the visual feature matching each text prompt; Based on the probability distribution, the positive text prompts with the highest matching probability to the current segmentation result are identified as the most relevant quality control points, and the negative text prompts with the highest matching probability are identified as the most significant potential deviation types.

[0098] In one possible implementation, a multimodal attention mechanism is used to calculate the similarity between the visual features and the text features of each text prompt, including: Cross-modal attention weighting is applied to the visual and text features to generate enhanced visual-text joint features; Based on the joint features, a fine-grained similarity score between the visual features and each text prompt is calculated; The calculation process of cross-modal attention weighting synchronously outputs an attention weight map, which is used to identify the pixel region in the medical image that contributes the most to generating the current similarity score.

[0099] In one possible implementation, the analysis unit further performs the following steps: The severity of identified potential biases is graded based on the matching probability value with the negative text prompts; the severity is determined based on the matching probability value with the negative text prompts corresponding to the most significant identified potential biases; the higher the matching probability value, the higher the severity level is determined. The identified deviations are associated with the specific quality control points in the text prompts to determine the root cause category of the deviations. Based on the highest similarity score between the visual features and the matching text prompt features, the confidence level of the quality control evaluation result is output. This confidence level is displayed in the quality control report along with the corresponding quality control score and potential deviations to indicate the reliability of the evaluation result. The confidence level of the evaluation result is calculated based on the highest matching probability value in the probability distribution, or the difference between the highest and second-highest matching probabilities; wherein, the higher the highest matching probability value or the larger the probability difference, the higher the output confidence level; wherein, the confidence level is used to indicate the reliability of the quality control score and potential deviation information.

[0100] In one possible implementation, the matching analysis also employs an uncertainty-aware mechanism, including: Based on the probability distribution obtained from the automatic segmentation results, calculate the information entropy of the distribution; When the information entropy exceeds a first preset threshold, it is determined that the automatic segmentation result has overall semantic ambiguity. In response to the determination, an unreliable quality indicator is added to the generated quality control report; and doctors are advised to conduct a focused manual review.

[0101] In another possible implementation, the matching analysis also employs an uncertainty-aware mechanism, including: Obtain the probability distribution of visual features matching multiple text prompts; Select the K text prompts with the highest probabilities from the probability distribution, where K ≥ 2; If the K text prompts contain both positive and negative text prompts, and the probability difference between the highest probability positive text prompt and the highest probability negative text prompt is less than a preset difference threshold, then semantic ambiguity is determined and marked, and the doctor is advised to conduct a focused manual review.

[0102] In one possible implementation, the training process of the visual language model includes: Collect medical image segmentation results and their corresponding expert-annotated text descriptions to construct a medical image and text dataset; wherein, the expert-annotated text descriptions are generated based on clinical guidelines, expert consensus documents and historical quality control records. The visual language model is trained using a contrastive learning framework based on the medical image and text dataset, so that the model learns to align the matched segmentation result images with the quality control text descriptions in the feature space.

[0103] In one possible implementation, the medical image and text dataset is constructed for a specific anatomical location or disease type, including at least one of the following: head and neck tumors, chest tumors, abdominal tumors, and pelvic tumors; This is a medical image and text dataset for pelvic tumors, including text descriptions of quality control standards for delineating organs at risk such as the prostate, bladder, and rectum.

[0104] In one possible implementation, the medical image and text dataset contains multi-level supervision signals, including: delineated images annotated by experienced physicians, delineated images annotated by domain experts, and text descriptions provided by domain experts after evaluating the physicians' delineations. The training process employs a hybrid loss function that combines contrastive learning loss and regression prediction loss.

[0105] In one possible implementation, the hybrid loss function is constructed through the following steps: Using the image encoder of the visual language model, the outlined images annotated by the experienced doctor and the outlined images annotated by the domain expert are mapped into image embedding vectors, respectively. The text encoder of the visual language model is used to map the text description provided by the domain expert into a text embedding vector. The image embedding vectors and text embedding vectors are projected into a shared contrastive learning space; In the contrastive learning space, the cosine similarity between all image embedding vectors and text embedding vectors is calculated and combined to obtain a similarity matrix. The similarity matrix is ​​calculated using the InfoNCE loss function to obtain the contrastive learning loss between the image drawn by the experienced doctor and the text description by the domain expert. The similarity matrix is ​​processed using convolutional regression to predict quality control scores, and the mean squared error loss between the predicted scores and the expert's actual scores is calculated. The contrastive learning loss and the mean squared error loss are weighted and combined to obtain the hybrid loss function.

[0106] In another possible implementation, the hybrid loss function is constructed through the following steps: In the contrastive learning space, the image-image similarity matrix between the delineated images annotated by experienced doctors and those annotated by domain experts is calculated simultaneously. Based on the image-image similarity matrix, an additional image contrast learning loss term is used to constrain the consistency of expert annotations at different levels in the feature space, forming a triple contrast learning structure of image-text-image.

[0107] In one possible implementation, the expert-annotated text description covers quality control information in at least one of the following dimensions: regularity of anatomical structure shape, clarity of boundaries, rationality of location, adequacy of target area coverage, and compliance of organ avoidance.

[0108] In one possible implementation, the device further includes: The interaction module is used to receive modifications to the segmentation results from at least one doctor through an interactive interface; to perform automated quality assessment based on the modified segmentation results, or to send the modified segmentation results to experts for approval; and to record the contour data before and after modification, the identity of the modifier, and the modification timestamp.

[0109] The recording module is used to record the outline data before and after modification, the identity of the modifier, and the modification timestamp.

[0110] In one possible implementation, the device further includes a processing module for performing the following steps: If the quality control score is greater than or equal to the first score threshold, the current segmentation result will be sent directly to the expert for approval. If the quality control score is less than the first score threshold, the quality control report will be provided to the doctor, who will be guided to make targeted corrections based on the modification suggestions in the report. After the corrections are completed, the results will be sent to the experts for approval. If the score of the same segmentation result is still lower than the first score threshold after being corrected more than a preset number of times, the expert upgrade mechanism will be automatically triggered. This mechanism will perform the following operations: Package the case's segmentation results, quality control reports, and correction records, and attach the system's preliminary analysis of the reasons for correction failures; The aggregated report will be directly sent to a senior or high-level radiotherapy expert, who will personally review and revise it and complete the final approval. Record such cases of repeated failed corrections and use them as data to optimize the first scoring threshold or train the visual language model.

[0111] In one possible implementation, the apparatus further includes a consensus optimization module for performing the following steps: When multiple experts have differing opinions on the same segmentation result, collect the text descriptions and modification records of each expert. By using clustering or voting mechanisms to integrate the opinions of multiple experts, more robust quality control standards are generated, and the text prompt library is updated. Inconsistent cases are recorded for model retraining.

[0112] The working principle and effect of the above technical solution are the same as those of the method in Example 1, and will not be repeated here.

[0113] This invention also provides a radiotherapy system, comprising: The automatic segmentation quality control device described in Example 2.

[0114] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in Embodiment 1 of this invention or the function of the device described in Embodiment 2 of this invention.

[0115] This invention also provides a computer-readable storage medium for storing a computer program. When the computer program is executed, it implements the steps of the method in Embodiment 1 of this invention. The specific implementation method is consistent with the implementation method and the technical effects achieved in the above method embodiments, and some contents will not be repeated.

[0116] In this invention, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The program product can take the form of any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0117] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination thereof. Program code for performing operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on a user computing device, partially on an associated device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to user computing devices via any type of network, including local area networks (LANs) or wide area networks (WANs), or they can be connected to external computing devices (e.g., via the Internet using an Internet service provider).

[0118] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the invention without departing from the principles and spirit of the invention, and all such changes should fall within the protection scope of the claims of the present invention.

Claims

1. An automatic segmentation quality control method, characterized in that, The method includes: Obtain the automatic segmentation results of the target area in medical images; Based on a pre-trained visual language model, the target region segmentation results are automatically evaluated for quality. The visual language model is trained using a medical image and text dataset that integrates expert consensus knowledge. It can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases. Based on the results of automated quality assessment, a quality control report is generated that includes quality control scores and modification suggestions; The visual language model is trained using a hybrid loss function. The hybrid loss function includes contrastive learning loss and regression loss; or The hybrid loss function includes contrastive learning loss, regression loss, and image-image consistency constraint; The regression loss is obtained by performing convolutional regression on the image-text similarity matrix to predict the quality control score, and calculating the mean square error between the predicted score and the expert's actual score. The image-image consistency constraint is used to constrain the feature consistency between delineated images annotated by experts at different levels in a shared contrastive learning space, forming a triple contrastive learning structure of image-text-image.

2. The automatic segmentation quality control method according to claim 1, characterized in that, The pre-trained visual language model performs automated quality assessment of the target region segmentation results, including: The visual features of the automatic segmentation result are extracted using the image encoder of the visual language model. Construct a text hint library based on expert consensus, including positive text hints describing standard delineation features and negative text hints describing delineation error patterns; The extracted visual features are compared with the text features of the text prompts to perform similarity calculation and matching analysis, and the quality control points and potential deviation types most relevant to the current segmentation results are identified.

3. The automatic segmentation quality control method according to claim 2, characterized in that, The method for constructing the text prompt library includes: Based on expert consensus on the delineation of radiotherapy target areas and organs at risk, clinical guidelines, and delineation data that has been reviewed and confirmed in the past, a structured quality control knowledge base is constructed. The generated descriptive criteria should include positive textual prompts for anatomical features, including: shape regularity, boundary clarity, spatial relationship with adjacent key anatomical structures, and size conformity. Generate negative text hints describing error patterns, including: oversegmentation, undersegmentation, jagged edges, inclusion of sensitive tissues that should not be included, and omission of critical areas.

4. The automatic segmentation quality control method according to claim 2, characterized in that, The step of performing similarity calculation and matching analysis between the extracted visual features and the text features of the text prompt includes: Calculate the similarity between the visual features and the text features of each text prompt; The calculated similarity is normalized to obtain the probability distribution of the visual features matching each text prompt; Based on the probability distribution, the positive text prompts with the highest matching probability to the current segmentation result are identified as the most relevant quality control points, and the negative text prompts with the highest matching probability are identified as the most significant potential deviation types.

5. The automatic segmentation quality control method according to claim 4, characterized in that, A multimodal attention mechanism is used to calculate the similarity between the visual features and the text features of each text prompt, including: Cross-modal attention weighting is applied to the visual and text features to generate enhanced visual-text joint features; Based on the joint features, a fine-grained similarity score between the visual features and each text prompt is calculated; The calculation process of cross-modal attention weighting synchronously outputs an attention weight map, which is used to identify the pixel region in the medical image that contributes the most to generating the current similarity score.

6. The automatic segmentation quality control method according to claim 1, characterized in that, The training process of the visual language model includes: Collect medical image segmentation results and their corresponding expert-annotated text descriptions to construct a medical image and text dataset; wherein, the expert-annotated text descriptions are generated based on clinical guidelines, expert consensus documents and historical quality control records. The visual language model is trained using a contrastive learning framework based on the medical image and text dataset, so that the model learns to align the matched segmentation result images with the quality control text descriptions in the feature space.

7. The automatic segmentation quality control method according to claim 6, characterized in that, The medical image and text dataset contains multi-level supervision signals, including: delineated images annotated by experienced physicians, delineated images annotated by domain experts, and text descriptions provided by domain experts after evaluating the physicians' delineations.

8. The automatic segmentation quality control method according to claim 7, when the hybrid loss function includes contrastive learning loss and regression loss, the hybrid loss function is constructed through the following steps: Using the image encoder of the visual language model, the outlined images annotated by the experienced doctor and the outlined images annotated by the domain expert are mapped into image embedding vectors, respectively. The text encoder of the visual language model is used to map the text description provided by the domain expert into a text embedding vector. The image embedding vectors and text embedding vectors are projected into a shared contrastive learning space; In the contrastive learning space, the similarity between all image embedding vectors and text embedding vectors is calculated and combined to obtain a similarity matrix; The similarity matrix is ​​calculated using the InfoNCE loss function to obtain the contrastive learning loss between the image drawn by the experienced doctor and the text description by the domain expert. The similarity matrix is ​​processed using convolutional regression to predict quality control scores, and the mean squared error loss between the predicted scores and the expert's actual scores is calculated. The contrastive learning loss and the mean squared error loss are weighted and combined to obtain the hybrid loss function.

9. The automatic segmentation quality control method according to claim 7, when the hybrid loss function includes contrastive learning loss, regression loss, and image-image consistency constraint, the hybrid loss function is constructed through the following steps: In the contrastive learning space, the image-image similarity matrix between the delineated images annotated by experienced doctors and those annotated by domain experts is calculated simultaneously. Based on the image-image similarity matrix, an additional image contrast learning loss term is used to constrain the consistency of expert annotations at different levels in the feature space, forming a triple contrast learning structure of image-text-image.

10. The automatic segmentation quality control method according to claim 6, characterized in that, The expert-annotated text descriptions cover quality control information in at least one of the following dimensions: regularity of anatomical structure shape, clarity of boundaries, rationality of location, adequacy of target area coverage, and compliance of organ avoidance.

11. The automatic segmentation quality control method according to claim 1, characterized in that, The method further includes: Through the interactive interface, the user can receive modifications to the segmentation results from at least one doctor. Automated quality assessment can be performed based on the modified segmentation results, or the modified segmentation results can be sent to experts for approval; the contour data before and after modification, the identity of the modifier, and the modification timestamp can be recorded.

12. The automatic segmentation quality control method according to claim 1, characterized in that, The method further includes: If the quality control score is greater than or equal to the first score threshold, the current segmentation result will be sent directly to the expert for approval. If the quality control score is less than the first score threshold, the quality control report will be provided to the doctor, who will be guided to make targeted corrections based on the modification suggestions in the report. After the corrections are completed, the results will be sent to the experts for approval. If the score is still lower than the first score threshold after the same segmentation result has been corrected more than a preset number of times, the expert upgrade mechanism will be automatically triggered.

13. The automatic segmentation quality control method according to claim 12, characterized in that, The method further includes: When multiple experts have differing opinions on the same segmentation result, collect the text descriptions and modification records of each expert. By using clustering or voting mechanisms to integrate the opinions of multiple experts, more robust quality control standards are generated, and the text prompt library is updated. Inconsistent cases are recorded for model retraining.

14. An automatic segmentation quality control device, characterized in that, The device includes: The first acquisition module is used to acquire the automatic target area segmentation results of medical images; The quality assessment module is used to automatically assess the quality of target area segmentation results based on a pre-trained visual language model. The visual language model is trained using a medical image and text dataset constructed by integrating expert consensus knowledge. It can perform cross-modal correlation analysis between the segmentation results and the quality control standards of the text description to identify potential biases. The report generation module is used to generate quality control reports containing quality control scores and modification suggestions based on the results of automated quality assessment. The visual language model is trained using a hybrid loss function. The hybrid loss function includes contrastive learning loss and regression loss; or The hybrid loss function includes contrastive learning loss, regression loss, and image-image consistency constraint; The regression loss is obtained by performing convolutional regression on the image-text similarity matrix to predict the quality control score, and calculating the mean square error between the predicted score and the expert's actual score. The image-image consistency constraint is used to constrain the feature consistency between delineated images annotated by experts at different levels in a shared contrastive learning space, forming a triple contrastive learning structure of image-text-image.

15. A radiotherapy system, characterized in that, include: The automatic segmentation quality control device according to claim 14.

16. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method according to any one of claims 1-13.

17. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, and when the computer reads the computer instructions, the computer implements the steps of the method according to any one of claims 1-13.