Augmented reality-based computer-aided diagnosis system and method, medium, program product and terminal

Through an augmented reality-based computer-aided diagnosis system, wearable devices and deep learning technology are used to capture and analyze medical images in real time and generate augmented reality diagnosis reports, which solves the CAD system integration problem in existing technologies and realizes efficient and convenient medical image analysis.

CN120656690APending Publication Date: 2025-09-16SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510753549.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing computer-aided diagnosis systems face challenges such as system heterogeneity, difficulty in customization, complex compatibility testing, and difficulty in multi-algorithm collaboration when integrated into hospital information systems, which seriously restricts their promotion and updating in clinical environments.

Method used

An augmented reality-based computer-aided diagnosis system is used, which utilizes wearable devices to capture visual field changes in real time, combines deep learning and pre-trained models for medical image recognition, quality restoration, modality classification, and diagnostic analysis, and generates medical diagnostic reports displayed in augmented reality, which are integrated independently from the hospital information system.

Benefits of technology

It achieves seamless deployment of multiple AI models, overcomes the lack of system integration and flexibility, improves the development of medical imaging technology and medical service efficiency, and avoids the high cost and complexity of traditional integration methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656690A_ABST
    Figure CN120656690A_ABST
Patent Text Reader

Abstract

According to the computer-aided diagnosis system and method based on augmented reality, the medium, the program product and the terminal provided by the invention, a CAD application solution which does not depend on traditional hospital IT infrastructure and can be seamlessly deployed and efficiently integrated with various AI models is provided through the wearable device and the augmented reality technology; the defects of system integration, flexibility, expandability, clinical workflow adaptability and the like in the prior art can be overcome, and therefore development of the medical imaging technology and improvement of medical service efficiency are promoted. According to the application, the medical image is sensed in real time by using the camera integrated on the wearable device, and deep analysis and accurate interpretation are performed by using an advanced image analysis algorithm. The system is independent of a hospital information system, avoids technical obstacles during complex data interaction and integration, and provides a brand-new, efficient and convenient way for medical image analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical information technology, and in particular to a computer-aided diagnosis system, method, medium, program product, and terminal based on augmented reality. Background Art

[0002] Computer-aided diagnosis (CAD) tools have developed rapidly in recent years, with breakthroughs in deep learning significantly improving diagnostic accuracy. However, integrating these CAD tools into existing hospital information systems (HISs) remains a significant challenge. The widespread heterogeneity of HISs is a key factor contributing to the difficulty in integrating CAD models. This heterogeneity manifests itself in multiple aspects: differences in operating systems, varying database types, a lack of standardized communication protocols, a diverse software vendor base, and significant differences in functionality and modules. For example, different clinical subsystems may run on different operating systems, employ different database types, and utilize different communication protocols. These differences require extensive customization and compatibility testing when integrating CAD tools with specialized clinical subsystems, such as radiology information systems (RISs) and picture archiving and communication systems (PACSs). This challenge is particularly acute when integrating multiple CAD algorithms from different vendors with varying dependencies within the same clinical environment, often requiring large-scale modifications and upgrades to the hospital's IT infrastructure, which is costly, time-consuming, and labor-intensive. The above situation has seriously hindered the widespread application of CAD systems in clinical environments, making it difficult for many hospitals to adopt and update advanced diagnostic technologies in a timely manner, affecting the quality and efficiency of medical services.

[0003] Specifically, although current CAD technology has made significant progress in diagnostic accuracy, its integration with the hospital's existing RIS, PACS and other information systems faces many challenges, such as system heterogeneity, difficulty in customization, complex compatibility testing, and difficulty in multi-algorithm collaboration, which has severely restricted the promotion and updating of CAD technology in clinical environments. Summary of the Invention

[0004] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a computer-aided diagnosis system, method, medium, program product and terminal based on augmented reality, which are used to solve the challenges of integrating CAD systems into hospital information systems in the existing technology, such as system heterogeneity, difficulty in customization, complex compatibility testing and difficulty in multi-algorithm collaboration, which have led to serious restrictions on the promotion and updating of CAD technology in clinical environments.

[0005] To achieve the above-mentioned objectives and other related objectives, the first aspect of the present application provides a computer-aided diagnosis system based on augmented reality, comprising: a visual capture module, which is used to adaptively change the capture field of view of the camera of the wearable device as the posture of the wearer changes, and to capture the visual information of the wearer in real time; a detection and recognition module, which is used to perform image recognition using a target detection algorithm based on deep learning to identify and extract medical image information from the visual information; a quality enhancement module, which is used to perform image quality restoration processing on the medical image information based on a pre-trained image restoration network model and a prompt mechanism to obtain medical image information after quality restoration; a modality classification module, which is used to perform image modality classification on the medical image information after quality restoration based on a pre-trained modality classification model to obtain a medical image modality classification result; a diagnosis engine module, which is used to obtain a corresponding diagnosis analysis model according to the medical image modality classification result, and perform diagnostic analysis on the medical image information after quality restoration to generate a diagnosis analysis result; a report generation module, which is used to generate a medical diagnosis report using a large visual language model and combined with the diagnosis analysis result of the diagnosis engine module, and perform an augmented reality display of the medical diagnosis report on the wearable device.

[0006] In some embodiments of the first aspect of the present application, the process of performing image quality restoration processing on the medical image information based on a pre-trained image restoration network model and combined with a prompt mechanism includes: inputting a preset prompt template and medical image information into a multimodal basic model, and outputting repair prompt information of the medical image information; inputting the repair prompt information and medical image information into a pre-trained image restoration network model, and performing image quality restoration processing on the medical image information.

[0007] In some embodiments of the first aspect of the present application, the process of performing image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result includes: inputting the quality-restored medical image information and a preset medical terminology database into the pre-trained modality classification model; wherein the pre-trained modality classification model includes an image encoder and a text encoder; performing feature extraction according to the image encoder and the text encoder to obtain image features corresponding to the quality-restored medical image information and text features corresponding to the preset medical terminology database; calculating the similarity between the image features and the text features, and making a judgment based on the similarity calculation result to obtain a medical image modality classification result.

[0008] In some embodiments of the first aspect of the present application, the diagnosis engine module further includes: constructing a plurality of diagnosis analysis models corresponding to different medical image modalities to form a diagnosis analysis model library.

[0009] In some embodiments of the first aspect of the present application, a corresponding diagnostic analysis model is obtained based on the medical image modality classification result, and a diagnostic analysis is performed on the medical image information after quality restoration to generate a diagnostic analysis result. The process includes: finding the corresponding diagnostic analysis model in the diagnostic analysis model library based on the medical image modality classification result; inputting the medical image information after quality restoration into the corresponding diagnostic analysis model, and outputting the diagnostic analysis result.

[0010] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a computer-aided diagnosis method based on augmented reality, which is applied to the computer-aided diagnosis system based on augmented reality as described above, and the method includes: the capture field of view of the camera of the wearable device adaptively changes with the change of the posture of the wearing user, and the visual information of the wearing user is captured in real time; image recognition is performed using a target detection algorithm based on deep learning to identify and extract medical image information from the visual information; based on a pre-trained image restoration network model, combined with a prompt mechanism, the medical image information is subjected to image quality restoration processing to obtain medical image information after quality restoration; based on a pre-trained modality classification model, the medical image information after quality restoration is subjected to image modality classification to obtain a medical image modality classification result; according to the medical image modality classification result, a corresponding diagnostic analysis model is obtained, and a diagnostic analysis is performed on the medical image information after quality restoration to generate a diagnostic analysis result; a large visual language model is used, combined with the diagnostic analysis result of the diagnostic engine module, to generate a medical diagnostic report, and an augmented reality display of the medical diagnostic report is performed on the wearable device.

[0011] In some embodiments of the second aspect of the present application, the process of performing image quality restoration processing on the medical image information based on a pre-trained image restoration network model and combined with a prompt mechanism includes: inputting a preset prompt template and medical image information into a multimodal basic model, and outputting repair prompt information of the medical image information; inputting the repair prompt information and medical image information into a pre-trained image restoration network model, and performing image quality restoration processing on the medical image information.

[0012] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the augmented reality-based computer-aided diagnosis method when executed by a processor.

[0013] To achieve the above-mentioned objectives and other related objectives, the fourth aspect of the present application provides a computer program product, which includes a computer program code. When the computer program code is run on a computer, the computer implements the computer-assisted diagnosis method based on augmented reality.

[0014] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides an electronic terminal, including a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the augmented reality-based computer-aided diagnosis method.

[0015] As described above, the computer-aided diagnosis system, method, medium, program product, and terminal based on augmented reality provided by this application have the following beneficial effects:

[0016] This application proposes an augmented reality computer-aided diagnosis system with a wearable device as the interface through augmented reality technology. This system is a CAD application solution that does not rely on traditional hospital IT infrastructure and can seamlessly deploy and efficiently integrate multiple AI models. It can overcome the shortcomings of existing technologies in terms of system integration, flexibility, scalability, and adaptability to clinical workflows, thereby promoting the development of medical imaging technology and improving the efficiency of medical services. The present invention uses a camera integrated into a wearable device to perceive medical images in real time, and uses advanced image analysis algorithms to perform in-depth analysis and accurate interpretation of the captured images. The present invention is completely independent of the hospital information system and does not require complex data interaction with it, thereby cleverly bypassing a series of technical obstacles faced when integrating the CAD system into the hospital information system, and provides a new, efficient and convenient solution for AI analysis of medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1Shown is a structural diagram of a computer-aided diagnosis system based on augmented reality in one embodiment of the present application.

[0018] Figure 2 Shown is a diagram of a specific embodiment of a wearable device capturing visual information of the wearing user in real time in one embodiment of the present application.

[0019] Figure 3 Shown is a specific embodiment diagram of obtaining repair prompt information in one embodiment of the present application.

[0020] Figure 4 Shown is a flowchart of a computer-aided diagnosis method based on augmented reality in one embodiment of the present application.

[0021] Figure 5 Shown is a structural schematic diagram of an electronic terminal in one embodiment of the present application. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0023] Before further explaining the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations:

[0024] <1> Computer-Aided Diagnosis (CAD) refers to automated or semi-automated analysis methods based on medical images, signals, physiological parameters, and other data. It can quantify, compare, judge, and identify normal conditions, abnormalities, and diseases, as well as their severity, for prediction and diagnosis. With the advancement of medical technology, CAD has been widely used in medical imaging for initial screening, auxiliary diagnosis, and disease monitoring.

[0025] <2> Deep learning: Deep learning is a machine learning method in which algorithms automatically learn useful features from data and use these features to perform tasks such as classification and recognition. Deep learning is widely used in fields such as natural language processing, speech recognition, and computer vision. In natural language processing, deep learning can be applied to tasks such as text classification, sentiment analysis, and machine translation. In speech recognition, deep learning can automatically learn and identify speech features, improving speech recognition accuracy. In computer vision, deep learning can build complex neural networks to recognize and classify images, and even achieve advanced functions such as image generation.

[0026] <3> Large language model (LLM): Generally refers to a language model with a large number of parameters and capabilities. It learns the statistical laws and semantic relationships of language by pre-training on large-scale text data. These models usually use unsupervised learning methods to predict the next word or fill in missing words to capture the context and semantic information of the language. Large language models can generate coherent sentences, answer questions, complete translation tasks, etc. LLMs are characterized by their large scale, containing billions or even more parameters, which help them learn complex patterns in language data. The emerging capabilities of large language models include contextual learning, instruction following and step-by-step reasoning capabilities, and the ability to process multimodal data such as images.

[0027] <4> Vision-Language Models (VLMs): A VLM is a multimodal neural network model capable of processing both image and text data. Its core task is to understand and generate language information related to visual content through cross-modal representation learning, or to generate corresponding visual content from language information. Compared to traditional unimodal models, VLMs have the advantage of combining visual and language information for multifaceted reasoning and generation. VLMs not only have high theoretical research value but also demonstrate great potential in practical applications, particularly in fields such as autonomous driving, intelligent robotics, and virtual assistants.

[0028] <5> Hospital Information System: HIS is the core system of the hospital, integrating patient information, medical resources, finance and other aspects to provide comprehensive data support for hospital operations.

[0029] <6> Transformer Network: Transformer is a deep learning architecture based on the self-attention mechanism, originally used for natural language processing tasks such as machine translation and text generation. In recent years, it has also been widely used in the field of computer vision, such as image classification, object detection, and image generation.

[0030] <7> The CLIP (Contrastive Language-Image Pretraining) model achieves effective retrieval from images to text and generation from text to images by mapping images and text into the same feature space.

[0031] <8> LoRA (Low-Rank Adaptation of Large Language Models): LoRA fine-tunes the model by training only the low-rank matrix and then injecting these parameters into the original model. This approach not only reduces computational requirements but also makes the training resources much smaller than directly training the original model, making it very suitable for use in resource-limited environments.

[0032] <9> MeLo (Medical Low-Rank Adaptation): Medical image low-rank adaptation.

[0033] To facilitate understanding of the embodiments of this application, first Figure 1 Detailed description. Figure 1 The following is a schematic diagram of the structure of a computer-aided diagnosis system based on augmented reality in accordance with an embodiment of the present invention. The system in this embodiment includes a visual capture module 110, a detection and recognition module 120, a quality enhancement module 130, a modality classification module 140, a diagnosis engine module 150, and a report generation module 160. The visual capture module 110 is connected to the detection and recognition module 120, the detection and recognition module 120 is connected to the quality enhancement module 130, the quality enhancement module 130 is respectively connected to the modality classification module 140 and the diagnosis engine module 150, and the diagnosis engine module 150 is respectively connected to the modality classification module 140 and the report generation module 160.

[0034] The visual capture module 110 is used to capture the visual information of the wearer in real time based on the adaptive change of the capture field of view of the camera of the wearable device as the posture of the wearer changes.

[0035] The visual capture module, as the system's front-end component, is used to capture the wearer's visual information in real time using the AR device. Specifically, when a user wears the AR device, such as a radiologist viewing a medical display, the AR device accurately follows the radiologist's head movements, tracking changes in the radiologist's field of view in real time and automatically capturing the doctor's visual information during the diagnosis process. The visual capture module ensures that the system dynamically responds to the wearer's visual focus in real time, providing accurate input for subsequent image processing.

[0036] It's worth noting that augmented reality technology has developed rapidly in recent years, demonstrating enormous potential for application in the medical field. For example, wearable AR devices, due to early technical limitations, faced numerous shortcomings, such as poor display quality, a poor wearing experience, and insufficient computing power. This, to a certain extent, limited their widespread application in medical settings. However, with the rapid advancement of technology, today's AR devices have achieved substantial improvements in several key performance areas. For example, significant advances in sensor technology have enabled the new generation of AR devices to integrate a variety of high-precision sensors, including but not limited to cameras, depth sensors, and eye-tracking sensors. The integration of these advanced sensors greatly enhances AR devices' ability to perceive the wearer's environment, accurately capturing the user's intended actions and providing a more natural and fluid interactive experience. Consequently, AR devices are becoming an important tool with practical application value in the medical field.

[0037] Specifically, the wearable device adopts smart glasses, and a camera is integrated into the smart glasses. Figure 2 As shown, the camera adopts a high-resolution, wide-angle camera with a resolution of 2K, a horizontal viewing angle of 90 degrees, and a vertical viewing angle of 59 degrees, which is seamlessly integrated into the smart glasses. When the user wears the smart glasses, the horizontal distance from the display screen is 50-60 cm. The design of the wearable device in this embodiment can ensure that the system can accurately capture medical images when the doctor views the display at different viewing angles or in different positions. The camera itself has outstanding low-light performance and can achieve clear imaging even in dimly lit environments. At the same time, the camera also has a high frame rate and can smoothly capture dynamic images, so that it can adapt to diverse clinical environments and meet the needs of different diagnostic scenarios.

[0038] The wearable device in this embodiment can also be other devices such as a smart helmet, which can be selected according to actual conditions and is not limited in this embodiment.

[0039] The detection and recognition module 120 is used to perform image recognition using a deep learning-based target detection algorithm to identify and extract medical image information from the visual information.

[0040] The detection and recognition module is used to accurately identify medical image information within the visual information collected by the AR device based on real-time captured visual information using advanced image recognition algorithms. Specifically, the detection and recognition module can identify the area of ​​medical image information on the display and accurately segment it from surrounding interface elements, while also segmenting and excluding other non-medical image information displayed on the display page.

[0041] The image recognition algorithm adopts a target detection algorithm based on deep learning. The target detection algorithm based on deep learning includes R-CNN algorithm, Fast R-CNN algorithm, Faster R-CNN algorithm, YOLO series algorithm, SSD algorithm, RTMDet algorithm, etc., which are not limited in this embodiment.

[0042] Preferably, the deep learning-based object detection algorithm uses RTMDet and YOLOv5. A pre-trained RTMDet network is used to detect and identify the boundaries of the display in the visual information and segment the display interface; a pre-trained YOLOv5 model is used to detect and identify the medical image on the display interface to extract the medical image information.

[0043] Medical image information includes raw medical image data of a patient scanned by a medical image scanner and transmitted to a display screen for display. The medical image scanner may be a computed tomography (CT) scanner, a magnetic resonance (MR) scanner, an ultrasound scanner, a positron emission tomography (PET) scanner, or any other type of scanner used in accordance with a medical imaging modality. The medical image scanner is used to scan a patient or a portion of a patient to acquire raw medical image data representing the patient's anatomical structure.

[0044] Before using the target detection algorithm based on deep learning for image recognition, the target detection algorithm based on deep learning needs to be trained and optimized to improve the accuracy and robustness of recognition. Specifically, taking the YOLOv5 algorithm as an example, 50 common radiology workstation interface templates are collected, and the areas displaying medical image information are manually marked. Various types of medical images are embedded into various radiology workstation interface templates, and the bounding boxes of the medical images are marked to form a training sample set. The training sample set is subjected to data enhancement processing, which can improve the model's adaptability to different display conditions during subsequent training. The data enhancement processing method uses techniques such as random rotation, color jittering, random cropping, and random erasing to increase the diversity of training data and avoid model overfitting.

[0045] The YOLOv5 model was trained until convergence using the training sample set. The resulting YOLOv5 model can effectively distinguish medical image information from surrounding interface elements, such as toolbars, patient information panels, and measurement displays, enabling accurate medical image recognition across various radiology software systems. As a result, the detection and recognition module is able to quickly and accurately locate medical image information content across a variety of complex display interface layouts and different display types, providing a solid foundation for subsequent image processing.

[0046] The quality enhancement module 130 is configured to perform image quality restoration processing on the medical image information based on a pre-trained image restoration network model in combination with a prompt mechanism to obtain quality-restored medical image information.

[0047] It should be noted that in the field of medical imaging diagnosis, image quality directly affects the accuracy and reliability of disease detection. However, in the process of medical image acquisition, due to objective factors such as imaging equipment performance limitations, the complexity of the shooting environment, and the diversity of patient postures, there are inevitably problems such as viewing angle tilt, poor lighting conditions, and noise interference from scanning equipment. Specifically, viewing angle tilt causes distortion of the image's spatial structure, affecting the judgment of the spatial relationship of anatomical structures; poor lighting conditions result in uneven image brightness and reduced contrast, which can easily cause blurred details; and noise interference from scanning equipment produces random signal fluctuations, forming image artifacts, which seriously interfere with the identification of lesion features. The above problems will significantly reduce the quality of medical images, bringing difficulties and risks of misdiagnosis to subsequent diagnosis.

[0048] The quality enhancement module in this embodiment uses deep learning-based image restoration technology to construct an image restoration network model, which is then used to optimize the image quality of the identified medical image information. The image restoration network model can restore degraded input medical images to clean, clear, high-quality images. It can effectively correct image distortion caused by factors such as tilted viewing angles, poor lighting conditions, and noise interference from scanning equipment, significantly improving the clarity and contrast of medical images, making lesion edges sharper and tissue textures more delicate. This provides doctors with more accurate and reliable medical image information, helps improve the accuracy of disease diagnosis, and reduces the occurrence of missed and misdiagnoses, with significant clinical application value and economic benefits.

[0049] In one embodiment, the process of performing image quality restoration processing on the medical image information based on a pre-trained image restoration network model in combination with a prompt mechanism includes:

[0050] Inputting a preset prompt template and medical image information into a multimodal basic model, and outputting repair prompt information of the medical image information;

[0051] The repair prompt information and medical image information are input into a pre-trained image restoration network model, and image quality restoration processing is performed on the medical image information.

[0052] It should be explained that the purpose of incorporating the hint mechanism in this embodiment is to improve the quality of image restoration. By providing additional guidance to the image restoration network model, it helps the model more accurately understand and perform image quality restoration tasks, thereby improving the quality and accuracy of the restored image, bringing it closer to the true state of the original image. The restoration hint information generated by the hint mechanism is a natural language description generated by a large language model and designed specifically for image restoration tasks, such as "restore clarity to a blurred image" or "remove image noise and enhance details."

[0053] Specifically, the preset prompt template for medical image restoration processing in this embodiment includes clear instructions and context information for the image quality restoration task. The content of the preset prompt template can be "Please generate a simple prompt word", "Please generate a simple prompt paragraph", etc. Figure 3 As shown in the figure, the content of the preset prompt template is "You are an expert in medical image analysis. Please analyze the medical image I provide you. This is a degraded image. Please generate a simple prompt to help the model restore the image. Please answer strictly according to the following format:

[0054] Modal: {modal type}

[0055] Location: {anatomical part}

[0056] Perspective: {perspective}

[0057] Image quality issues:

[0058] Artifact: {specific artifact type}

[0059] Exposure: {Exposure problem description}

[0060] Noise: {noise level}

[0061] Resolution: {resolution quality}

[0062] Contrast: {Contrast problem}

[0063] Missing or truncated regions: {Are there any missing or truncated regions?}

[0064] Affected structures: {Affected anatomical structures}

[0065] Improvements needed: {Specific areas that need improvement}."

[0066] like Figure 3 The repair tips shown can be used to fix problems such as blurred images and text wrapping.

[0067] The content of the preset prompt template in this embodiment is not limited and can be adjusted according to actual needs.

[0068] The multimodal basic model adopts a large language model. The large language model learns the ability to serve human language understanding and generation by training on a large amount of text data, and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. The core idea of ​​the large language model is to learn the patterns and language structures of natural language through large-scale unsupervised training, which can better understand and generate natural text, while also being able to demonstrate certain logical thinking and reasoning abilities.

[0069] In this embodiment, a preset prompt template and medical image information are input into a large language model. The large language model generates repair prompt information corresponding to the current medical image information based on the input. The repair prompt information can be used to provide clear guidance and contextual information for image restoration processing.

[0070] The large language model adopts a GPT model, a BERT model, a Qwen model, a GLM model, an LLAMA model, etc., which is not limited in this embodiment.

[0071] Furthermore, the repair prompt information corresponding to the current medical image information generated by the large language model and the current medical image information are input into the pre-trained image restoration network model. In the pre-trained image restoration network model, the medical image information is restored according to the guidance of the repair prompt information to obtain the restored medical image information.

[0072] It's important to note that restoration hints provide the image restoration network with a clear task description and context. For example, if the hint is "enhance image details," the image restoration network will focus on the image's detailed features and enhance them during the restoration process. Restoration hints help the network better understand the task objective, thereby improving the quality and accuracy of the restored image.

[0073] Specifically, taking medical image restoration as an example, the repair prompt information includes reducing overexposure in the lower area, enhancing the contrast of the lung and mediastinum areas, suppressing artifacts, and improving the clarity and edge sharpness of the anatomical structure. During the image restoration process, problems such as overexposure in the lower area, insufficient contrast in the lung and mediastinum areas, the presence of artifacts, and insufficient clarity and edge sharpness of the anatomical structure will be specifically addressed.

[0074] The image restoration network model adopts a multi-layer Transformer network to pre-train the Transformer network. During the training process, a large number of image data pairs consisting of degraded images and their corresponding real images are used as training sample sets. Degraded images are blurred, noisy, low-resolution or other damaged images, while real images are undamaged original images. The Transformer network learns the distortion characteristics of degraded images, including but not limited to noise distribution, blur degree, color shift and other distortion characteristics, and generates targeted image restoration strategies, and then constructs a mapping relationship between degraded images and real images. The image restoration network model finally obtained can achieve effective restoration of various types of degraded images, thereby improving the accuracy and efficiency of image restoration.

[0075] The modality classification module 140 is configured to perform image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result.

[0076] The modality classification module in this embodiment uses a visual language model to construct a modality classification model. The powerful representation and reasoning capabilities of the visual language model enable in-depth analysis of restored medical image information. The modality classification model accurately extracts visual features from the restored medical image information, including image texture, grayscale distribution, and anatomical structure. It then integrates and matches these features with language information such as medical terminology and conceptual descriptions, enabling accurate identification and classification of key information such as medical image type, imaging location, imaging parameters, and lesion characteristics.

[0077] In one embodiment, the process of performing image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result includes:

[0078] Inputting the quality-restored medical image information and a preset medical terminology database into a pre-trained modality classification model; wherein the pre-trained modality classification model includes an image encoder and a text encoder;

[0079] performing feature extraction according to the image encoder and the text encoder to obtain image features corresponding to the quality-restored medical image information and text features corresponding to the preset medical terminology database;

[0080] The similarity between the image features and the text features is calculated, and a judgment is made based on the similarity calculation result to obtain a medical image modality classification result.

[0081] It should be noted that the present invention utilizes the CLIP model to construct a modality classification model. The CLIP model is pre-trained on large-scale image-text pair data and possesses excellent cross-modal feature alignment capabilities. Specifically, this embodiment pre-trains the CLIP model using medical image-text annotated data pairs to address the specific needs of medical image modality classification. By adjusting the CLIP model parameters to adapt it to the specialized knowledge and data distribution characteristics of the medical field, it enhances sensitivity to medical image modality features and classification accuracy, effectively meeting the practical application requirements of medical image modality classification tasks.

[0082] The preset medical terminology database is a collection of professional terms in the medical field, which is used to describe various attributes and features of medical images, such as image type, imaging site and other related medical information.

[0083] Specifically, image types include: X-ray, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, positron emission tomography (PET), single photon emission computed tomography (SPECT), etc.; imaging sites include the head, chest, abdomen, limbs, spine, and neck; and anatomical structures include the heart, lungs, liver, kidneys, brain, and bones. In medical image classification tasks, specialized terms from a pre-set medical terminology database can be used as labels to help the model learn how to classify images into the correct category, for example, classifying a medical image as a chest X-ray.

[0084] In this embodiment, the CLIP model includes an image encoder and a text encoder. The CLIP model maps images and text into a unified semantic space through a comparative learning method, thereby being able to calculate the similarity between the image and text. In the medical image classification task, this feature of the CLIP model is utilized to match medical images with corresponding medical terms. Specifically, the quality-restored medical image information is input into the CLIP model's image encoder to obtain the corresponding image features; then, a preset medical terminology database is input into the CLIP model's text encoder to obtain the corresponding text features; finally, the similarity between the image features and the text features is calculated, and the image modality and related anatomical parts are determined based on the magnitude of the similarity, thereby obtaining the image modality classification result. The result with the greatest similarity is usually selected as the image modality classification result.

[0085] The diagnosis engine module 150 is used to obtain a corresponding diagnosis analysis model according to the medical image modality classification result, and perform a diagnosis analysis on the medical image information after quality restoration to generate a diagnosis analysis result.

[0086] It should be noted that multiple diagnostic analysis models corresponding to different medical image modalities are constructed in the diagnostic engine module to form a diagnostic analysis model library. The medical image classification result is a medical image modality, that is, the medical image modalities include chest X-rays, breast X-rays, blood smears, etc. After obtaining the image modality classification result corresponding to the current medical image information, the diagnostic analysis model corresponding to the current medical image information is called in the diagnostic analysis model library. For example, if the image modality classification result of the current medical image information is a chest X-ray, the diagnostic analysis model corresponding to the chest X-ray is found. The medical image information after quality restoration is then input into the corresponding diagnostic analysis model for diagnostic analysis, and finally the diagnostic analysis result is output, such as the diagnostic analysis result of a chest X-ray. The diagnostic analysis results include detailed classification results of the disease type and its confidence score, etc.

[0087] The diagnostic engine module in this embodiment can intelligently call the corresponding diagnostic analysis model based on the medical image modality classification results output by the modality classification module. Each diagnostic analysis model in the diagnostic engine module is trained based on a large amount of medical imaging data, and can perform a comprehensive diagnostic evaluation of medical images, and generate detailed classification results including disease types and their confidence scores. In this embodiment, a modular design concept is adopted to construct corresponding diagnostic analysis models for different medical image modalities, form a diagnostic analysis model library, and integrate it into the diagnostic engine module. Through this design approach, the system can quickly and accurately call the most suitable diagnostic analysis model for diagnostic analysis based on the characteristics of the input medical image, thereby improving the accuracy and reliability of the diagnosis.

[0088] Specifically, the construction method of diagnostic analysis models corresponding to multiple different medical image modalities adopts the MeLo method. The MeLo method is a medical image diagnosis method based on low-rank adaptation (LoRA). It quickly adapts ViT (Vision Transformer) through low-rank decomposition technology. It only needs to introduce a small number of trainable low-rank matrices to efficiently switch between diagnostic tasks of different medical image modalities. Compared with traditional fine-tuning methods, the MeLo method significantly reduces training parameters and computational complexity while maintaining excellent performance. In addition, by adopting the MeLo method, it is possible to reduce resource consumption for training and inference while maintaining model performance, thereby improving the efficiency and feasibility of the system.

[0089] The report generation module 160 is used to use the visual language large model in combination with the diagnostic analysis results of the diagnosis engine module to generate a medical diagnosis report, and perform augmented reality display of the medical diagnosis report on the wearable device.

[0090] The visual language big model may include GLIP, BLIP, GPT, etc., which are not limited in this embodiment. The visual language big model is used to integrate and transform the diagnostic analysis results to generate a diagnostic report, with the aim of outputting diagnostic information in a clear and intuitive format. The report generation module can automatically generate a concise and clear medical diagnostic report based on key information in the diagnostic analysis results, such as disease type, confidence score, etc. The format and content of the medical diagnostic report are optimized in the report generation module to make it conform to the specifications and requirements of the medical diagnostic report, which is convenient for doctors to consult and use. In addition, the report generation module can also flexibly adjust the level of detail and expression of the report according to different user needs and preferences to meet the requirements of different doctors and medical institutions.

[0091] It should be noted that, based on medical expertise, various medical diagnostic report templates have been designed based on different medical imaging examination types (such as X-ray, CT, MRI, etc.) and disease types, forming a pre-set report template library to meet different medical imaging diagnostic needs. For example, for X-ray examinations, there are chest X-ray report templates and bone X-ray report templates; for CT examinations, there are abdominal CT report templates and brain CT report templates. Each medical diagnostic report template contains a fixed structure, such as the patient's basic information area, the examination site description area, the diagnosis results area, and the recommendation area. Within each area of ​​each medical diagnostic report template, the content format is further refined. Taking the chest X-ray report template as an example, the diagnosis results area includes a pre-set sentence framework describing common chest disease manifestations such as nodules, inflammation, and fractures, such as "The patient [anatomical site] shows [disease manifestation], approximately [specific value] in size, [description] in morphology, [description] in boundary, [description] in density, [description]."

[0092] After the diagnostic engine module outputs the diagnostic analysis results, it performs semantic analysis on them using natural language processing technology based on the visual language model. After obtaining key information such as the disease type and its confidence score, it selects a corresponding medical diagnostic report template from the preset report template library and fills the diagnostic analysis results into the corresponding medical diagnostic report template to generate a standardized medical diagnostic report. The medical diagnostic report content includes symptom description, examination results, diagnosis conclusion, treatment recommendations, etc.

[0093] For example, if a chest X-ray reveals a lung nodule, a chest X-ray report template is selected from a library of pre-set report templates. If the confidence level is low, a more cautious medical diagnosis report template is selected, containing more statements suggesting further examination. If the confidence level is high, a more definitive medical diagnosis report template is selected. Based on the chest X-ray image, the medical diagnosis report contains a detailed description of the diagnosis, such as "A nodule shadow is visible in the patient's right upper lung lobe, measuring approximately 1.2 cm x 1.0 cm, with irregular shape, blurred boundaries, and uneven density. Further examination is recommended."

[0094] Furthermore, medical knowledge graph technology can also be introduced in this application to optimize medical diagnosis reports based on medical knowledge graph technology. Specifically, a large amount of medical literature, clinical guidelines, medical databases and other materials are collected in advance, and medical knowledge is extracted. Doctors combine corresponding clinical diagnostic experience and professional knowledge to establish the relationship between diagnostic evidence, clinical manifestations and disease information, including the relationship between disease and symptoms, the correspondence between disease and examination methods, and the relationship between disease and treatment plan, thereby forming a preset medical knowledge graph.

[0095] After generating a medical diagnosis report, the results in the medical diagnosis report can be associated and matched with the preset medical knowledge graph. For example, for the diagnosis result of "right upper lobe nodule", the knowledge graph related to lung nodules is searched, and then the medical terms and descriptions in the medical diagnosis report are optimized based on the standardized terms and descriptions in the knowledge graph. If the knowledge graph stipulates that "nodular shadow" should be more accurately described as "nodular increased density shadow", then the "nodular shadow" in the medical diagnosis report is replaced with "nodular increased density shadow" to generate the final optimized medical diagnosis report to improve the accuracy and professionalism of the report.

[0096] It should be emphasized that the report generation module uses natural language generation technology to perform semantic analysis on the diagnostic analysis results and converts them into information that doctors can easily understand and use, thereby assisting doctors in making more accurate clinical decisions. Furthermore, in order to improve the accuracy and professionalism of medical diagnostic reports, medical knowledge graph technology can also be introduced into the report generation module, so that the report generation module can generate more accurate and standardized medical terms and descriptions, thereby improving the quality and credibility of medical diagnostic reports. The present invention has tested and verified the report generation module, proving that the reports it generates meet the requirements and standards of clinical practice and can provide reliable diagnostic support for doctors.

[0097] After the medical diagnosis report is generated, it is synchronously displayed on the user's wearable device in an augmented display format. For example, when a radiologist wears smart glasses and views a medical image, the camera of the smart glasses captures the medical image, and after a series of analyses and processing, a medical diagnosis report is generated. The smart glasses automatically display the contents of the medical diagnosis report for the current medical image, such as the location of the lesion and the diagnosis conclusion, in other words, the medical diagnosis report is displayed in the field of view of the user wearing the wearable device. The augmented reality display method can provide a more intuitive and convenient user experience, allowing doctors to more intuitively understand the patient's condition and improve the efficiency and convenience of medical diagnosis and treatment.

[0098] In order to facilitate the description of the computer-aided diagnosis system based on augmented reality of the present invention, the following specific embodiments are provided for illustration.

[0099] The disease classification results were evaluated using the diagnosis engine module on the PneumoniaMNIST (pneumonia two-classification) dataset and the OAI (knee joint five-classification) dataset, as shown in Table 1 below. The classification performance was used as an indicator to reflect the system performance.

[0100] Table 1 Comparison of classification performance using the quality enhancement module and other image restoration methods on the PneumoniaMNIST and OAI datasets

[0101]

[0102]

[0103] As can be seen in Table 1, the classification performance of degraded images (i.e., images captured by cameras) on both datasets differs significantly from that achieved using image restoration methods, resulting in significantly poorer classification performance. This indicates that image quality degradation significantly impacts the diagnostic task. Specifically, the F1 score for degraded images on the PneumoniaMNIST dataset is 0.805, while that on the OAI dataset is 0.522. Applying various existing image restoration methods to degraded images improves classification performance. For example, using the PneumoniaMNIST dataset, SwinIR and Restormer achieve F1 scores of 0.950 and 0.950, respectively, after image restoration. However, on the OAI dataset, the image restoration performance of each method is relatively limited, with F1 scores generally below 0.58, indicating that this task relies more heavily on subtle diagnostic features in the image.

[0104] In contrast, the present invention, using the quality enhancement module to restore degraded images on both datasets, achieved the best performance compared to other image restoration methods. On the PneumoniaMNIST dataset, the precision, recall, and F1 score all reached 0.969, significantly outperforming other image restoration methods. On the OAI dataset, it also significantly outperformed other methods, achieving the highest precision, recall, and F1 score. These results demonstrate that the image restoration method employed in the quality enhancement module of the present invention is more effective in preserving key diagnostic features in images, providing a reliable image foundation for subsequent automated medical analysis tasks.

[0105] It should be emphasized that the present invention adopts a unique modular workflow, with six modules working together to form an efficient and intelligent diagnostic auxiliary workflow. The six modules are respectively a visual capture module, a detection and recognition module, a quality enhancement module, a modality classification module, a diagnostic engine module, and a report generation module. These six modules can fully process the input medical image data, from the initial image acquisition to the final diagnostic report output, to achieve automation and intelligence of the entire process. Through its unique modular workflow, the present invention realizes the seamless deployment of CAD applications in clinical environments, providing radiologists with a convenient, efficient, and non-invasive diagnostic auxiliary tool.

[0106] This modular design enables the diagnostic system to flexibly adapt to different clinical needs, with excellent scalability and maintainability, opening up a new path for the field of medical imaging diagnosis. Compared with traditional CAD systems, the present invention has significant advantages. The present invention abandons the complex process of hospital information system integration and does not require large-scale transformation of existing systems such as HIS, RIS, and PACS. This greatly reduces the difficulty and cost of deployment, significantly shortens the implementation cycle, and effectively improves the accessibility and popularization speed of CAD technology in clinical applications. At the same time, the modular architecture of the present invention brings extremely high flexibility to the system. Taking the diagnostic engine module as an example, its function expansion and iteration can quickly integrate the latest diagnostic AI models. And with the continuous development of AI technology in the field of medical imaging, radiologists can obtain cutting-edge diagnostic auxiliary functions in a timely manner without waiting for lengthy system updates and integration processes, helping hospitals maintain the advanced nature of diagnostic technology. In addition, this system uses wearable devices as an interactive interface, combining features such as hands-free operation, real-time information overlay, remote expert collaboration, and spatial computing. Without interfering with the normal work process and field of view of doctors, it provides more convenient, efficient, and intuitive diagnostic support, which is of great significance for improving diagnostic accuracy and efficiency and improving patient diagnosis and treatment effects.

[0107] The present invention uses augmented reality technology to propose an augmented reality computer-aided diagnosis system with a wearable device as the interface. The system is a CAD application solution that does not rely on traditional hospital IT infrastructure and can seamlessly deploy and efficiently integrate multiple AI models. It can overcome the shortcomings of existing technologies in terms of system integration, flexibility, scalability, and adaptability to clinical workflows, thereby promoting the development of medical imaging technology and improving the efficiency of medical services. Specifically, the present invention uses a camera integrated in a wearable device to perceive medical images in real time, and uses advanced image analysis algorithms to perform in-depth analysis and accurate interpretation of the captured images. The present invention is completely independent of the hospital information system and does not require complex data interaction with it, thereby cleverly bypassing a series of technical obstacles faced when integrating the CAD system into the hospital information system, and provides a new, efficient and convenient solution for AI analysis of medical images.

[0108] In the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects, and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different.

[0109] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" represent examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0110] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.

[0111] Figure 4 This is a computer-aided diagnosis method based on augmented reality provided by the embodiment of the present application. Figure 4 As shown, the method includes:

[0112] Step S41: The capture field of view of the camera of the wearable device adaptively changes with the change of the posture of the wearer, thereby capturing the visual information of the wearer in real time;

[0113] Step S42: performing image recognition using a deep learning-based target detection algorithm to identify and extract medical image information from the visual information;

[0114] Step S43: performing image quality restoration processing on the medical image information based on the pre-trained image restoration network model and in combination with the prompt mechanism to obtain quality-restored medical image information;

[0115] Step S44: performing image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result;

[0116] Step S45: obtaining a corresponding diagnostic analysis model according to the medical image modality classification result, and performing diagnostic analysis on the quality-restored medical image information to generate a diagnostic analysis result;

[0117] Step S46: Using the visual language large model, combined with the diagnostic analysis results of the diagnostic engine module, to generate a medical diagnosis report, and performing an augmented reality display of the medical diagnosis report on the wearable device.

[0118] It should be understood that the specific process of the above corresponding method has been described in detail in the above system embodiment, and for the sake of brevity, it will not be repeated here.

[0119] It should also be understood that the division of modules in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.

[0120] Figure 5 : is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. Figure 5 As shown, the electronic terminal includes: at least one processor 501, a memory 502, at least one network interface 503 and a user interface 505. The various components in the device are coupled together via a bus system 504. It is understood that the bus system 504 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5In the text, various buses are labeled as bus systems.

[0121] The user interface 505 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.

[0122] It will be appreciated that the memory 502 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM) or a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0123] The memory 502 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 500. Examples of such data include: any executable program for operating on the electronic terminal 500, such as an operating system 5021 and an application 5022; the operating system 5021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 5022 can include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The computer-aided diagnosis method based on augmented reality provided in the embodiment of the present invention can be included in the application 5022.

[0124] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 501 or by software instructions. The above processor 501 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 501 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 501 may be a microprocessor or any conventional processor. The steps of the accessory optimization method provided in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium located in a memory. The processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0125] In an exemplary embodiment, the electronic terminal 500 may be configured to execute the aforementioned method using one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).

[0126] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, which, when running on a computer, enables the computer to execute the augmented reality-based computer-assisted diagnosis method of any of the embodiments shown.

[0127] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores program code. When the program code is run on a computer, the computer executes the augmented reality-based computer-assisted diagnosis method of any of the embodiments shown.

[0128] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0129] Those skilled in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0130] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0132] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0133] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0134] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (program) are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. Available media may be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0135] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.

[0136] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0137] In summary, the present application provides a computer-aided diagnosis system, method, medium, program product and terminal based on augmented reality, including: a visual capture module, which is used to adaptively change the capture field of view of the camera of the wearable device as the posture of the wearer changes, and capture the visual information of the wearer in real time; a detection and recognition module, which is used to use a deep learning-based target detection algorithm to perform image recognition to identify and extract medical image information from the visual information; a quality enhancement module, which is used to perform image quality restoration processing on the medical image information based on a pre-trained image restoration network model and a prompt mechanism to obtain medical image information after quality restoration; a modality classification module, which is used to perform image modality classification on the medical image information after quality restoration based on a pre-trained modality classification model to obtain a medical image modality classification result; a diagnosis engine module, which is used to obtain a corresponding diagnosis analysis model according to the medical image modality classification result, and perform diagnostic analysis on the medical image information after quality restoration to generate a diagnosis analysis result; a report generation module, which is used to use a large visual language model and combine the diagnosis analysis result of the diagnosis engine module to generate a medical diagnosis report, and perform an augmented reality display of the medical diagnosis report on the wearable device.

[0138] This application proposes an augmented reality computer-aided diagnosis system with a wearable device as the interface through augmented reality technology. This system is a CAD application solution that does not rely on traditional hospital IT infrastructure and can seamlessly deploy and efficiently integrate multiple AI models. It can overcome the shortcomings of existing technologies in terms of system integration, flexibility, scalability, and adaptability to clinical workflows, thereby promoting the development of medical imaging technology and improving the efficiency of medical services. The present invention uses a camera integrated into a wearable device to perceive medical images in real time, and uses advanced image analysis algorithms to perform in-depth analysis and accurate interpretation of the captured images. The present invention is completely independent of the hospital information system and does not require complex data interaction with it, thereby cleverly bypassing a series of technical obstacles faced when integrating the CAD system into the hospital information system, and providing a new, efficient and convenient solution for AI analysis of medical images. Therefore, this application effectively overcomes the various shortcomings of the existing technology and has a high industrial utilization value.

[0139] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.

Claims

1. A computer-aided diagnosis system based on augmented reality, characterized in that: include: A visual capture module is used to capture the visual information of the wearer in real time based on the wearer's camera's field of view adaptively changing with the wearer's posture; a detection and recognition module, configured to perform image recognition using a deep learning-based target detection algorithm to identify and extract medical image information from the visual information; a quality enhancement module, configured to perform image quality restoration processing on the medical image information based on a pre-trained image restoration network model and in combination with a prompt mechanism, so as to obtain quality-restored medical image information; a modality classification module, configured to perform image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result; a diagnosis engine module, configured to obtain a corresponding diagnosis analysis model according to the medical image modality classification result, and perform a diagnosis analysis on the medical image information after quality restoration to generate a diagnosis analysis result; A report generation module is used to use a large visual language model in combination with the diagnostic analysis results of the diagnostic engine module to generate a medical diagnosis report, and to perform an augmented reality display of the medical diagnosis report on the wearable device.

2. The computer-aided diagnosis system based on augmented reality according to claim 1, characterized in that The process of performing image quality restoration processing on the medical image information based on the pre-trained image restoration network model and the prompt mechanism includes: Inputting a preset prompt template and medical image information into a multimodal basic model, and outputting repair prompt information of the medical image information; The repair prompt information and medical image information are input into a pre-trained image restoration network model, and image quality restoration processing is performed on the medical image information.

3. The computer-aided diagnosis system based on augmented reality according to claim 1, characterized in that: The process of performing image modality classification on the quality-restored medical image information based on the pre-trained modality classification model to obtain a medical image modality classification result includes: Inputting the quality-restored medical image information and a preset medical terminology database into a pre-trained modality classification model; wherein the pre-trained modality classification model includes an image encoder and a text encoder; performing feature extraction according to the image encoder and the text encoder to obtain image features corresponding to the quality-restored medical image information and text features corresponding to the preset medical terminology database; The similarity between the image features and the text features is calculated, and a judgment is made based on the similarity calculation result to obtain a medical image modality classification result.

4. The computer-aided diagnosis system based on augmented reality according to claim 1, characterized in that: The diagnosis engine module further includes: constructing a plurality of diagnosis analysis models corresponding to different medical image modalities to form a diagnosis analysis model library.

5. The computer-aided diagnosis system based on augmented reality according to claim 4, characterized in that: The process of obtaining a corresponding diagnostic analysis model according to the medical image modality classification result and performing diagnostic analysis on the quality-restored medical image information to generate a diagnostic analysis result includes: Finding a corresponding diagnostic analysis model in the diagnostic analysis model library according to the medical image modality classification result; The medical image information after quality restoration is input into a corresponding diagnosis and analysis model, and a diagnosis and analysis result is output.

6. A computer-aided diagnosis method based on augmented reality, characterized in that: Applied to the computer-aided diagnosis system based on augmented reality according to any one of claims 1 to 5, the method comprising: The capture field of view of the wearable device's camera changes adaptively with the wearer's posture, capturing the wearer's visual information in real time. Performing image recognition using a deep learning-based object detection algorithm to identify and extract medical image information from the visual information; Based on a pre-trained image restoration network model and in combination with a prompt mechanism, image quality restoration processing is performed on the medical image information to obtain quality-restored medical image information; performing image modality classification on the quality-restored medical image information based on a pre-trained modality classification model to obtain a medical image modality classification result; Acquiring a corresponding diagnostic analysis model according to the medical image modality classification result, and performing diagnostic analysis on the quality-restored medical image information to generate a diagnostic analysis result; A large visual language model is used in combination with the diagnostic analysis results of the diagnostic engine module to generate a medical diagnosis report, and an augmented reality display of the medical diagnosis report is performed on the wearable device.

7. The computer-aided diagnosis method based on augmented reality according to claim 6, characterized in that: The process of performing image quality restoration processing on the medical image information based on the pre-trained image restoration network model and the prompt mechanism includes: Inputting a preset prompt template and medical image information into a multimodal basic model, and outputting repair prompt information of the medical image information; The repair prompt information and medical image information are input into a pre-trained image restoration network model, and image quality restoration processing is performed on the medical image information.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer-aided diagnosis method based on augmented reality according to any one of claims 7 or 8 is implemented.

9. A computer program product, characterized in that The computer program product includes computer program code, and when the computer program code is run on a computer, the computer is enabled to implement the augmented reality-based computer-aided diagnosis method according to any one of claims 7 or 8.

10. An electronic terminal comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the augmented reality-based computer-aided diagnosis method according to any one of claims 7 or 8.