Quality control method and system based on multi-scene multi-modal large model
By applying multi-scene multi-modal large models in medical imaging quality control, combined with multi-modal learning and deep learning technology, the problems of low efficiency and insufficient adaptability of traditional Chinese medicine imaging quality control in the existing technology are solved, and efficient and accurate image quality evaluation and segmentation effects are achieved.
Patent Information
- Application Number
- CN202510036212.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing medical imaging quality control methods are inefficient, highly subjective and insufficiently adaptable, making it difficult to effectively evaluate the quality of medical imaging in multimodal and multi-scenarios.
The quality control method based on multi-scene multi-modal large model is adopted, and through multi-modal learning, multi-scene adaptability and deep learning technology, the CLIP-Med pre-trained model and MM-SAM segmentation basic large model are constructed, combined with interactive prompt information, and efficient and accurate evaluation of image quality is achieved.
It realizes efficient and accurate evaluation of the medical image quality of different modes and scenarios, significantly improves segmentation accuracy and quality evaluation efficiency, and supports unified evaluation of multimodal image quality.
Smart Images

Figure CN119942314A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a quality control method and system based on a multi-scene multi-modal large model. Background Art
[0002] With the rapid development of medical imaging technology, multimodal imaging technologies such as CT, MRI, ultrasound and X-ray have become important tools for clinical diagnosis. However, in the process of acquiring, processing and using medical images of different modalities, quality issues are always important factors affecting the accuracy of diagnosis, such as image noise, artifacts, insufficient resolution and non-standard operation procedures. These problems not only increase the difficulty of diagnosis, but may also lead to misdiagnosis or missed diagnosis.
[0003] Existing medical image quality control methods mainly rely on manual inspection and subjective judgment, which are inefficient and limited by the experience level of different operators, and have significant subjectivity and differences. In addition, with the rapid growth of medical data, manual quality control can no longer meet actual needs, and automated and intelligent image quality control methods have become an urgent need for the development of the industry.
[0004] Moreover, in multi-modal and multi-scenario medical images, image quality issues are even more complex. For example, the modality differences between CT and MRI images may lead to significant differences in the manifestation of the same lesion in different images; in ultrasound images, due to differences in acquisition equipment and operating techniques, there are also large variations in image clarity and the discernibility of lesion features. How to achieve automated quality assessment of images of different modalities and provide accurate and reliable quality control results is one of the key issues in current clinical imaging applications.
[0005] In summary, the existing technology has significant shortcomings in the quality control and evaluation of medical images. There is an urgent need for a medical image quality control method and system that can be oriented to multi-modality and multi-scenario, and use deep learning and multi-modal large model technology to achieve efficient and accurate image quality control and evaluation. Summary of the invention
[0006] In order to solve the problems of low efficiency, strong subjectivity and insufficient adaptability of medical image quality control methods in the prior art, the present invention provides a quality control method and system based on a multi-scene multi-modal large model. By combining multi-modal learning, multi-scene adaptability and deep learning technology, the present invention can achieve efficient and accurate evaluation of the quality of medical images of different modalities and scenes, providing strong support for the intelligent analysis of medical images.
[0007] To achieve the above objectives, the present invention provides a quality control method and system based on a multi-scenario multi-modal large model. Among them, a quality control method based on a multi-scenario multi-modal large model includes:
[0008] Collect and obtain multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, and perform data preprocessing to obtain preprocessed data;
[0009] Construct a multimodal medical image pre-training model CLIP-Med, and use a contrastive learning mechanism to perform multimodal alignment of image features and text features to obtain pre-trained image encoders and text encoders;
[0010] The features extracted by the fine-tuned SAM image encoder are fused with the features generated by the text encoder to obtain fused features;
[0011] Based on the pre-trained model CLIP-Med and combined with the SAM encoder, we built the multi-modal medical image segmentation basic model MM-SAM.
[0012] According to the fusion features, the segmented medical image is obtained by using the multimodal medical image segmentation basic model MM-SAM in combination with interactive prompt information;
[0013] Construct a basic large model for automatic evaluation of quality control item levels in multiple parts, perform multi-level scoring and classification evaluation on the quality of segmented medical images; integrate the quality scoring results to generate a multimodal medical image quality control evaluation report.
[0014] Preferably, the medical imaging data includes CT, MRI, ultrasound and X-ray images;
[0015] The pathology and quality control reports are text descriptions related to the images;
[0016] The data preprocessing process includes:
[0017] The medical image data is standardized, including size normalization, noise filtering, and grayscale adjustment; at the same time, the text data is subjected to natural language processing operations such as word segmentation and stop word removal, and then converted into a vector representation that can be used for model training.
[0018] Preferably, the formula for the standardization process is:
[0019]
[0020] Among them, I is the original image pixel value, μ and σ are the pixel mean and standard deviation respectively, I norm is the normalized image;
[0021] The vector representation of the model training is:
[0022] T emb =Embedding(T)
[0023] Where T is the input pathology or quality control text, Embedding(·) is the text embedding operation, and T emb is the generated text vector.
[0024] Preferably, the process of multimodally aligning image features with text features through a contrastive learning mechanism includes:
[0025] The image encoder extracts image features, while the text encoder extracts features of pathology or quality control descriptions;
[0026] The contrast loss function is used to maximize the similarity between the image and the corresponding text, while minimizing the similarity of non-corresponding relationships, and multimodally align the image features with the text features to achieve the semantic association and unified expression of the image and text.
[0027] Preferably, the formula for extracting image features through the image encoder is:
[0028] f I =Encoder I (I norm )
[0029] Among them, Encoder I (·) is the image encoder;
[0030] The formula for extracting features of pathology or quality control descriptions through a text encoder is:
[0031] f T =Encoder T (T emb )
[0032] Among them, Encoder T (·) is the text encoder;
[0033] The contrast loss function formula is:
[0034]
[0035] Where N is the number of samples, τ is the temperature parameter, and f I i and are the image features and text features of the i-th sample respectively.
[0036] Preferably, the formula for fusing the features extracted by the fine-tuned SAM image encoder with the features generated by the text encoder is:
[0037]
[0038] Among them, ⊕ represents the feature concatenation operation, fSAM Image features generated by the SAM image encoder.
[0039] Preferably, the process of obtaining the segmented medical image by using the multimodal medical image segmentation basic large model MM-SAM in combination with interactive prompt information includes:
[0040] The medical images are optimized using interactive prompt information of points, boxes, and texts, and the model is optimized by combining the Dice loss function and the cross entropy loss function;
[0041] The formula expression of the model optimization is:
[0042] L seg =λ1L Dice +λ2L CE
[0043] Among them, L Dice and L CE They are Dice loss and cross entropy loss respectively, λ1 and λ2 are weight factors;
[0044] The formula expression of the Dice loss function is:
[0045]
[0046] Among them, p i is the model prediction value, g i is the true value;
[0047] The formula expression of the cross entropy loss function is:
[0048]
[0049] Preferably, the process of performing multi-level scoring and classification evaluation on the quality of the segmented medical image includes:
[0050] The segmented medical image is divided into small blocks, and the Transformer architecture is used to extract and linearly project the features of the small blocks to evaluate the image quality at a fine-grained level;
[0051] Then, the quality score is calculated based on the small block features to generate a quality assessment report:
[0052] The small block features are extracted through the Transformer model, and the formula is:
[0053] f block = Transformer(f seg )
[0054] Among them, f seg is the segmentation result feature, fblock It is a fine-grained feature of small blocks;
[0055] The formula expression of the quality score is:
[0056]
[0057] Among them, Q is the total score, w i and i are the weight and score of the i-th quality indicator respectively.
[0058] Preferably, the process of integrating the quality scoring results and generating a multimodal medical imaging quality control assessment report includes:
[0059] The quality scoring and classification results are optimized at multiple levels through a multi-layer perceptron to generate a quality control evaluation report for multimodal medical images.
[0060] The content of the quality control evaluation report includes a visual display of the image segmentation results and a quality score.
[0061] The present invention also provides a quality control system based on a multi-scenario multi-modal large model, comprising:
[0062] The data acquisition module is used to acquire multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, and perform data preprocessing to obtain preprocessed data;
[0063] A pre-training module is used to construct a multimodal medical image pre-training model CLIP-Med, and to perform multimodal alignment of image features and text features through a contrastive learning mechanism to obtain a pre-trained image encoder and text encoder; and to fuse the features extracted by the fine-tuned SAM image encoder with the features generated by the text encoder to obtain fused features;
[0064] The model building and segmentation module is used to build a multimodal medical image segmentation basic large model MM-SAM based on the pre-trained model CLIP-Med and combined with the SAM encoder; according to the fusion features, the multimodal medical image segmentation basic large model MM-SAM is used to obtain the segmented medical image in combination with the interactive prompt information;
[0065] The evaluation module is used to build a basic large model for automatic evaluation of quality control items at multiple parts, and to perform multi-level scoring and classification evaluation on the quality of segmented medical images;
[0066] The report generation module is used to integrate the quality scoring results and generate a multimodal medical image quality control assessment report.
[0067] Compared with the prior art, the present invention has the following advantages and technical effects:
[0068] The present invention supports multiple medical imaging modalities such as CT, MRI, ultrasound and X-ray, and is applicable to a variety of medical scenarios. It solves the limitation that the single modality method in the prior art cannot be promoted across modalities, and realizes the unified evaluation of multimodal image quality. Through the multimodal contrast learning mechanism, combined with medical imaging data and pathology and quality control report texts, the constructed CLIP-Med model realizes the semantic alignment of image and text features, and at the same time introduces a large amount of static image data to enhance the generalization ability of the model, overcoming the problem of insufficient training data. The fine-tuned SAM image encoder is used, combined with interactive prompt information (such as points, boxes or text), to achieve accurate segmentation of target areas in medical images, significantly improve the segmentation accuracy, and solve the problem of poor segmentation effect due to complex background or low resolution.
[0069] In addition, the present invention divides the segmentation results into small blocks through the multi-site quality control item level automatic evaluation model, and uses the Transformer architecture to perform multi-dimensional quality scoring and classification evaluation. The entire evaluation process is fully automated, efficient and stable. The final generated medical image quality control report includes segmentation result visualization and quality scoring, providing comprehensive auxiliary decision support for clinicians, especially in the case of poor image quality or abnormalities, it can quickly identify and evaluate problem images.
[0070] The present invention is significantly superior to existing technologies in terms of adaptability, multimodal feature fusion, segmentation accuracy and quality assessment efficiency. It can effectively promote the intelligence and efficiency of medical image quality control and provide strong support for assisting clinical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0072] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;
[0073] Figure 2 A schematic diagram of specific technical details of the method flow of an embodiment of the present invention;
[0074] Figure 3 Schematic diagram of the system structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0075] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0076] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0077] Embodiment 1
[0078] like Figure 1-2 As shown, this embodiment provides a quality control method and system based on a multi-scenario multi-modal large model, including the following steps:
[0079] Collect and obtain multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, and perform data preprocessing to obtain preprocessed data;
[0080] Construct a multimodal medical image pre-training model CLIP-Med, and use a contrastive learning mechanism to perform multimodal alignment of image features and text features to obtain pre-trained image encoders and text encoders;
[0081] The features extracted by the fine-tuned SAM image encoder are fused with the features generated by the text encoder to obtain fused features;
[0082] Based on the pre-trained model CLIP-Med and combined with the SAM encoder, we built the multi-modal medical image segmentation basic model MM-SAM.
[0083] According to the fusion features, the segmented medical image is obtained through the multimodal medical image segmentation basic model MM-SAM combined with interactive prompt information;
[0084] Construct a basic large model for automatic evaluation of quality control item levels in multiple parts, perform multi-level scoring and classification evaluation on the quality of segmented medical images; integrate the quality scoring results to generate a multimodal medical image quality control evaluation report.
[0085] Furthermore, data collection and preprocessing include:
[0086] Obtain multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, where the medical imaging data includes CT, MRI, ultrasound, and X-ray images, and the pathology and quality control reports are text descriptions related to the images.
[0087] Specifically, the collected image data is standardized, including size normalization, noise filtering, grayscale adjustment and other operations. The formula is as follows:
[0088]
[0089] Among them, I is the original image pixel value, μ and σ are the pixel mean and standard deviation respectively, Inorm is the normalized image.
[0090] At the same time, the text data is subjected to natural language processing operations such as word segmentation and stop word removal, and converted into a vector representation that can be used for model training. The formula is:
[0091] T emb =Embedding(T)
[0092] Where T is the input pathology or quality control text, Embedding(·) is the text embedding operation, and T emb is the generated text vector.
[0093] Furthermore, the construction of the multimodal medical imaging pre-training model CLIP-Med includes:
[0094] A multimodal pre-training model CLIP-Med is constructed using the contrastive learning mechanism. The model contains an image encoder and a text encoder, which are used to extract image features and text features respectively, and achieve semantic alignment.
[0095] Specifically, the image feature f is extracted by the image encoder I :
[0096] f I =Encoder I (I norm )
[0097] Among them, Encoder I (·) is the image encoder.
[0098] At the same time, the features of pathology or quality control description are extracted through the text encoder T :
[0099] f T =Encoder T (T emb )
[0100] Among them, Encoder T (·) is a text encoder.
[0101] The contrast loss function is used to maximize the similarity between the image and the corresponding text, while minimizing the similarity of non-corresponding relationships. The contrast loss function L contrast The formula expression is:
[0102]
[0103] Where N is the number of samples, τ is the temperature parameter, and f I i and are the image features and text features of the i-th sample respectively.
[0104] This embodiment builds a CLIP-Med model based on contrastive learning mechanism to perform multimodal alignment of image features and text features. The image features are extracted by pre-trained image encoder, and the report text features are extracted by text encoder, so as to achieve efficient association and unified expression of image and text semantics.
[0105] Furthermore, the construction of the multimodal medical image segmentation basic model MM-SAM includes:
[0106] Based on the CLIP-Med model and combined with the SAM encoder, a large multimodal medical image segmentation basic model MM-SAM is constructed.
[0107] Specifically, the features extracted by the fine-tuned SAM image encoder are fused with the features generated by the CLIP-Med text encoder:
[0108]
[0109] Among them, ⊕ represents the feature concatenation operation, f SAM Image features generated by the SAM image encoder.
[0110] In addition, the interactive prompt information of points, boxes, and texts is used to optimize the segmentation results, and the Dice loss function and cross entropy loss function are combined to optimize the model. The formula for model optimization is:
[0111] L seg =λ1L Dice +λ2L CE
[0112] Among them, L Dice and L CE are Dice loss and cross entropy loss respectively, and λ1 and λ2 are weight factors.
[0113] The formula of Dice loss function is:
[0114]
[0115] Among them, p i is the model prediction value, g i is the true value.
[0116] The formula of the cross entropy loss function is:
[0117]
[0118] This embodiment uses the fine-tuned SAM image encoder in combination with the CLIP-Med text encoder to achieve medical image feature extraction and fusion. Interactive prompt information (including points, boxes or text) is introduced to optimize the segmentation results and improve segmentation accuracy. The model can generate targeted segmentation results based on user prompts, such as organ or lesion area segmentation.
[0119] Furthermore, quality control assessment and report generation include:
[0120] Based on the segmented medical images, a basic large model for automatic evaluation of quality control item levels in multiple parts is constructed.
[0121] The segmented image is divided into small blocks, and the Transformer architecture is used to extract and linearly project the small block features to evaluate the image quality at a fine-grained level.
[0122] More specifically, the features of each small block are extracted through the Transformer model, and the formula is:
[0123] f block = Transformer(f seg )
[0124] Among them, f seg is the segmentation result feature, f block It is a small-scale fine-grained feature.
[0125] Then, the quality score is calculated based on the extracted features and a quality assessment report is generated; the formula for the quality score is:
[0126]
[0127] Among them, Q is the total score, w i and i are the weight and score of the i-th quality indicator respectively.
[0128] Furthermore, the assessment results are integrated and a report is generated including:
[0129] The quality score and classification results are optimized at multiple levels through a multi-layer perceptron (MLP) to generate a quality control evaluation report for multimodal medical images. The report includes a visual display of image segmentation results, quality scores, and a text generation model to comprehensively optimize the quality evaluation results, which can be directly used for clinical decision-making assistance.
[0130] Further applications and extensions include:
[0131] The segmentation results and quality control scores are integrated to generate a medical image quality control assessment report and provide a visual display. The segmentation results are superimposed on the original image, distinguishing tissue areas in a transparent overlay, and displaying relevant text labels (such as organ name, area, etc.). This method is applicable to a variety of medical scenarios, including lesion detection, surgical planning, and image quality control.
[0132] This embodiment significantly improves the quality assessment efficiency and accuracy of medical images by introducing contrastive learning, multimodal feature fusion and fine-grained quality control, and provides important support for intelligent analysis of medical images.
[0133] This embodiment solves the current problem of difficulty in uniformly evaluating image quality and performing efficient segmentation in multimodal medical image quality control. By constructing the CLIP-Med pre-trained model and using a large amount of static medical image data for multimodal feature alignment and model training, the model's generalization ability for multi-scene and multimodal data is enhanced; due to the difficulty of labeling dynamic image data and the limited number of samples, this embodiment uses dynamic images as external verification, fully verifying the applicability and segmentation accuracy of the segmentation model in dynamic images. Finally, this embodiment uses real-time visualization technology to superimpose the segmentation results on the original medical image, which can quickly and intuitively identify lesion areas or key structures, providing accurate support for clinical diagnosis.
[0134] The technical solution of this embodiment is relatively complete in engineering design and has high feasibility. Through the transfer learning method, the weights pre-trained in ImageNet are used for feature extraction, and combined with lightweight model design, the training speed is significantly accelerated while ensuring that the model parameters are small. The lower computational complexity not only reduces the hardware requirements, but also improves the flexibility of system deployment. In addition, a combination of Dice loss function and cross entropy loss function is used in the model optimization process to balance the model's accurate segmentation of the target area and overall segmentation performance.
[0135] In addition, the real-time visualization function provided by this embodiment can superimpose the segmentation results on the original medical image in the form of a transparent overlay, and provide intuitive text labels and regional information to provide doctors with fast image navigation. This real-time feedback mechanism helps doctors accurately locate the lesion area and improves the efficiency and safety of surgical planning. In summary, this embodiment can significantly improve the automation level and clinical applicability of medical image quality control, and provide important technical support for intelligent analysis of medical images and precision medicine.
[0136] Embodiment 2
[0137] Based on the same inventive concept, this embodiment also provides a quality control system based on a multi-scenario multi-modal large model, which implements the full process functions from data collection, model construction to segmentation and quality control report generation, and specifically includes the following modules:
[0138] Data acquisition module: used to acquire multimodal medical imaging data and its corresponding pathology or quality control reports, and supports multi-scene imaging data input; and performs data preprocessing to obtain preprocessed data;
[0139] Imaging data includes CT, MRI, ultrasound and X-ray images, covering a variety of medical scenarios and lesion types; pathology or quality control reports provide text description information related to the images to assist model understanding.
[0140] To ensure the diversity and quality of the data, this module performs standardized processing on the collected data, such as image resolution unification, noise filtering, and format conversion.
[0141] Pre-training module: used to build a multimodal medical image pre-training model CLIP-Med, perform multimodal alignment of image features and text features through a contrastive learning mechanism to obtain a pre-trained image encoder and text encoder; and fuse the features extracted by the fine-tuned SAM image encoder with the features generated by the text encoder to obtain fused features;
[0142] Specifically, the pre-training module pre-trains the CLIP-Med model and the SAM model, and achieves feature alignment between medical images and pathology and quality control reports through a comparative learning mechanism. And SAM is migrated to the field of medical image segmentation through fine-tuning. This module pre-trains the model with a large amount of collected multimodal data, so that it has relevant knowledge in the field of medical image processing and can accurately extract semantically related deep features in medical image data of different modalities.
[0143] During model training, image data is subjected to feature extraction through an image encoder, while text data is converted into comparable vector representations through a text encoder, thereby achieving semantic alignment between images and texts.
[0144] Model building and segmentation module: used to build and apply the MM-SAM model and generate segmentation results based on medical image data and prompt information;
[0145] Specifically, the segmentation module of this embodiment is based on the pre-trained model CLIP-Med, combined with the SAM encoder, to construct a multimodal medical image segmentation basic large model MM-SAM; according to the fusion features, through the multimodal medical image segmentation basic large model MM-SAM, combined with interactive prompt information, the segmented medical image is obtained;
[0146] The segmentation module of this embodiment realizes accurate segmentation of the target area of medical images through the MM-SAM model. The MM-SAM model combines the image feature extraction capability of the SAM encoder and the text semantic understanding capability of the CLIP-Med model, and introduces interactive prompt information (such as points, boxes or text) to further optimize the segmentation effect. For images of different modalities and scenes, the model can automatically adapt the segmentation parameters to generate high-precision segmentation results in images such as CT, MRI or ultrasound. At the same time, the segmentation module supports batch processing and real-time segmentation modes to meet diverse clinical needs.
[0147] Evaluation module: It is used to perform multi-level quality scoring and classification evaluation on the segmentation results based on the basic large model of automatic evaluation of multi-site quality control item levels, so as to realize multi-dimensional quality control of the segmentation results.
[0148] Specifically, the segmented image is divided into multiple small blocks, and fine-grained feature extraction is performed on each small block to evaluate multiple indicators such as image clarity, contrast, and artifact level.
[0149] This module adopts a multi-level scoring mechanism, combining segmentation quality with overall image quality to provide comprehensive data support for the final quality control report.
[0150] The module design takes into account both versatility and flexibility, and the evaluation dimensions and weights can be adjusted according to user needs.
[0151] Report generation module: used to integrate quality control evaluation results, generate visual quality control reports and provide auxiliary information.
[0152] Specifically, the report generation module is used to integrate the segmentation results and quality assessment information to generate a visual quality control report.
[0153] The report includes an intuitive display of the image segmentation results, such as marking organs or lesion areas with different colors and providing text descriptions to facilitate doctors' quick understanding.
[0154] The quality assessment section presents multiple scoring indicators of the image in the form of charts and provides diagnostic suggestions based on the assessment results, such as suggesting that the image may need to be re-acquired or the acquisition parameters may need to be corrected.
[0155] The report can be exported to multiple formats (such as PDF or electronic medical record system compatible files) to facilitate clinical application and archiving.
[0156] Furthermore, the system also supports real-time visualization and interaction functions. By directly connecting with medical imaging equipment (such as ultrasound scanners or CT machines), it processes the transmitted image data in real time and superimposes the segmentation results on the original images.
[0157] The segmentation results are displayed in a transparent overlay format, and auxiliary information such as the area and tissue name are provided for doctors to quickly refer to during surgery or diagnosis.
[0158] The system provides simple interactive functions, and doctors can optimize the segmentation results by adjusting prompt information (such as redefining the region box), thereby improving the actual application effect.
[0159] Furthermore, application scenarios and expansion capabilities include:
[0160] The system is suitable for a variety of clinical scenarios, including but not limited to:
[0161] Lesion detection and diagnosis: Through precise segmentation and quality control evaluation, it assists doctors in discovering potential lesions and provides a basis for diagnosis.
[0162] Surgical planning and guidance: Real-time visualization provides doctors with important image navigation during surgery, helping to determine resection boundaries or injection locations.
[0163] Image quality control: The automated assessment module can help imaging departments optimize the acquisition process, reduce the generation of low-quality images, and improve overall diagnostic efficiency.
[0164] The system of this embodiment supports data connection with a medical information system or an image archiving and communication system, and can update the quality control evaluation model and evaluation results in real time.
[0165] This embodiment not only provides a full-process solution from image acquisition to segmentation, evaluation and report generation, but also has high scalability and can be expanded according to the actual needs of different hospitals or departments, such as adding specialty-specific quality control indicators or introducing new image data types (such as PET images).
[0166] In summary, this embodiment realizes efficient and accurate medical image quality control and segmentation functions through the combination of modular design and multi-modal large model, providing comprehensive technical support for clinical diagnosis and treatment.
[0167] The quality control system based on a multi-scenario multi-modal large model provided in this embodiment has all the advantages of the quality control method based on a multi-scenario multi-modal large model provided in the first embodiment.
[0168] Embodiment 3
[0169] This embodiment further discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first embodiment.
[0170] Embodiment 4
[0171] This embodiment further discloses a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first embodiment are implemented.
[0172] Embodiment 5
[0173] This embodiment also discloses a computer program product, including a computer program, which implements the steps of the method described in the first embodiment when executed by a processor.
[0174] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A quality control method based on a multi-scenario multi-modal large model, characterized in that: include: Collect and obtain multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, and perform data preprocessing to obtain preprocessed data; Construct a multimodal medical image pre-training model CLIP-Med, and use a contrastive learning mechanism to perform multimodal alignment of image features and text features to obtain pre-trained image encoders and text encoders; The features extracted by the fine-tuned SAM image encoder are fused with the features generated by the text encoder to obtain fused features; Based on the pre-trained model CLIP-Med and combined with the SAM encoder, we built the multi-modal medical image segmentation basic model MM-SAM. According to the fusion features, the segmented medical image is obtained by using the multimodal medical image segmentation basic model MM-SAM in combination with interactive prompt information; Construct a basic large model for automatic evaluation of quality control item levels in multiple parts, perform multi-level scoring and classification evaluation on the quality of segmented medical images; integrate the quality scoring results to generate a multimodal medical image quality control evaluation report.
2. The method according to claim 1, characterized in that The medical imaging data include CT, MRI, ultrasound and X-ray images; The pathology and quality control reports are text descriptions related to the images; The data preprocessing process includes: The medical image data is standardized, including size normalization, noise filtering, and grayscale adjustment; at the same time, the text data is subjected to natural language processing operations such as word segmentation and stop word removal, and then converted into a vector representation that can be used for model training.
3. The method according to claim 2, characterized in that The formula expression of the standardization process is: Among them, I is the original image pixel value, μ and σ are the pixel mean and standard deviation respectively, I norm is the normalized image; The vector representation of the model training is: T emb =Embedding(T) Where T is the input pathology or quality control text, Embedding(·) is the text embedding operation, and T emb is the generated text vector.
4. The method according to claim 1, characterized in that: The process of multimodal alignment of image features and text features through contrastive learning mechanism includes: The image encoder extracts image features, while the text encoder extracts features of pathology or quality control descriptions; The contrast loss function is used to maximize the similarity between the image and the corresponding text, while minimizing the similarity of non-corresponding relationships, and multimodally align the image features with the text features to achieve the semantic association and unified expression of the image and text.
5. The method according to claim 4, characterized in that The formula for extracting image features through the image encoder is: f I =Encoder I (I norm ) Among them, Encoder I (·) is the image encoder; The formula for extracting features of pathology or quality control descriptions through a text encoder is: f T =Encoder T (T emb ) Among them, Encoder T (·) is the text encoder; The contrast loss function formula is: Where N is the number of samples, τ is the temperature parameter, and f I i and are the image features and text features of the i-th sample respectively.
6. The method according to claim 1, characterized in that The formula for fusing the features extracted by the fine-tuned SAM image encoder with the features generated by the text encoder is: Among them, ⊕ represents the feature concatenation operation, f SAM Image features generated by the SAM image encoder.
7. The method according to claim 1, characterized in that The process of obtaining the segmented medical image by using the multimodal medical image segmentation basic large model MM-SAM and combining the interactive prompt information includes: The medical images are optimized using interactive prompt information of points, boxes, and texts, and the model is optimized by combining the Dice loss function and the cross entropy loss function; The formula expression of the model optimization is: L seg =λ1L Dice +λ2L CE Among them, L Dice and L CE They are Dice loss and cross entropy loss respectively, λ1 and λ2 are weight factors; The formula expression of the Dice loss function is: Among them, p i is the model prediction value, g i is the true value; The formula expression of the cross entropy loss function is:
8. The method according to claim 1, characterized in that The process of multi-level scoring and classification evaluation of the quality of segmented medical images includes: The segmented medical image is divided into small blocks, and the Transformer architecture is used to extract and linearly project the features of the small blocks to evaluate the image quality at a fine-grained level; Then, the quality score is calculated based on the small block features to generate a quality assessment report: The small block features are extracted through the Transformer model, and the formula is: Among them, f seg is the segmentation result feature, f block It is a fine-grained feature of small blocks; The formula expression of the quality score is: Among them, Q is the total score, w i and i are the weight and score of the i-th quality indicator respectively.
9. The method according to claim 1, characterized in that: The process of integrating quality scoring results and generating a multimodal medical imaging quality control assessment report includes: The quality scoring and classification results are optimized at multiple levels through a multi-layer perceptron to generate a quality control evaluation report for multimodal medical images. The content of the quality control evaluation report includes a visual display of the image segmentation results and a quality score.
10. A quality control system based on a multi-scenario multi-modal large model, characterized in that: include: The data acquisition module is used to acquire multi-scene and multi-modal medical imaging data and their corresponding pathology and quality control reports, and perform data preprocessing to obtain preprocessed data; A pre-training module is used to construct a multimodal medical image pre-training model CLIP-Med, and to perform multimodal alignment of image features and text features through a contrastive learning mechanism to obtain a pre-trained image encoder and text encoder; and to fuse the features extracted by the fine-tuned SAM image encoder with the features generated by the text encoder to obtain fused features; The model building and segmentation module is used to build a multimodal medical image segmentation basic large model MM-SAM based on the pre-trained model CLIP-Med and combined with the SAM encoder; according to the fusion features, the multimodal medical image segmentation basic large model MM-SAM is used to obtain the segmented medical image in combination with the interactive prompt information; The evaluation module is used to build a basic large model for automatic evaluation of quality control items at multiple parts, and to perform multi-level scoring and classification evaluation on the quality of segmented medical images; The report generation module is used to integrate the quality scoring results and generate a multimodal medical image quality control assessment report.
Citation Information
Cited By
Real-time endoscope image analysis method and system based on multi-modal large model
CN120495779A
Image report quality control method and device based on large model, and storage medium
CN121439071A
Medical image analysis method based on deep learning
CN121964075A