A multi-modal medical image and AR glasses fusion display system
By integrating multimodal medical imaging with AR glasses, a system was developed that accurately collects data on the vision of medical staff and ambient light intensity. Combined with personalized prediction models and SLAM technology, it solves the problems of insufficient data accuracy and risk prediction in traditional medical operations, thereby improving the safety and efficiency of the operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-29
AI Technical Summary
In traditional medical procedures, medical staff struggle to achieve simultaneous perception and comprehensive analysis of multi-dimensional information. The lack of targeted data collection strategies results in insufficient data accuracy, making it impossible to proactively predict dynamic changes in target tissues and potential operational risks. Furthermore, the registration accuracy between virtual images and real-world scenes is limited, and the lack of personalized interactive optimization affects operational safety and efficiency.
By integrating multimodal medical imaging with AR glasses into a display system, the system collects data on the vision of medical staff and ambient light intensity. Combined with the patient's physiological-image correlation information, it uses an individualized prediction model to dynamically predict the target tissue and quantify the risk. With the help of AR dynamic optical adjustment and SLAM technology, it achieves accurate image overlay, supports diverse interactions, and iteratively optimizes the system.
It improves the accuracy and safety of medical procedures, adapts to various clinical scenarios, provides personalized risk warnings and operational guidance, reduces visual fatigue, and enhances the intelligence and standardization of operations.
Smart Images

Figure CN122111218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and AR technology, specifically to a fusion display system of multimodal medical images and AR glasses. Background Technology
[0002] In the modern medical field, precision and personalization have become the core development directions for improving medical quality and safety. The success of medical operations depends heavily on the precise control of the patient's physiological state, the dynamic changes of the target tissue, and the position of the operating instruments. The fusion of multimodal medical imaging data and real-time physiological parameters has gradually become an important direction for medical technology innovation. At the same time, AR technology, with its ability to overlay virtual information with real-world scenes, has shown broad application prospects in the medical field. Providing intuitive and real-time information support for medical staff has become an industry goal. With the advancement of medical informatization, hospital information systems and medical image archiving and communication systems have accumulated a large amount of patients' past physiological-image correlation data. How to make full use of this data and combine it with real-time collected information to provide personalized guidance and risk warnings for different medical operation scenarios has become a key issue that urgently needs to be solved in the current medical technology field, and has also laid the foundation for the development of fusion technology between multimodal medical imaging and AR glasses.
[0003] In traditional medical procedures, medical staff mainly rely on their own experience combined with independent imaging equipment and physiological monitoring instruments to obtain relevant information. The data sources are scattered and lack effective integration, making it difficult to achieve simultaneous perception and comprehensive analysis of multi-dimensional information. Different medical operation scenarios have significantly different data collection needs, but traditional technologies lack targeted collection strategy adaptation mechanisms, resulting in insufficient effectiveness and accuracy of data collection. In terms of risk assessment, traditional methods mostly rely on post-event judgment or static analysis, which cannot proactively predict the dynamic changes of the target tissue and potential operational risks, making it difficult to avoid safety hazards in advance. At the same time, the image display effect does not fully consider the differences in individual vision of medical staff and the influence of ambient light intensity, which can easily lead to problems such as unclear observation and visual fatigue, affecting operational judgment. In addition, the registration accuracy between virtual images and real scenes is limited, the interaction method is simple, and there is no dynamic iterative optimization mechanism, which cannot continuously improve adaptability based on individual patient differences and operational feedback, making it difficult to meet the needs of complex medical operations for intelligent and personalized guidance. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a multimodal medical imaging and AR glasses fusion display system. The system collects data on medical staff's vision and ambient light intensity, as well as patients' past and real-time physiological-image correlation information. It adapts to different medical operation scenarios to achieve accurate data acquisition, uses individualized prediction models to predict the dynamics of target tissues, quantifies operational risks and marks risk areas using multi-dimensional algorithms, optimizes image display effects through AR dynamic optical adjustment and ambient light adaptation, and achieves accurate overlay of virtual images and real scenes through SLAM technology. It supports diverse interactions and iteratively optimizes the system based on operational feedback. This system effectively improves the accuracy and safety of medical operations, adapts to various clinical scenarios, and provides reliable technical support for the intelligent and standardized development of medical care.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a multimodal medical imaging and AR glasses fusion display system, the system comprising: Personalized data acquisition and processing module: used to collect vision parameters of medical staff and ambient light intensity data of the current scene, retrieve the patient's previous physiological-image correlation data, and integrate and store it; Multimodal acquisition and adaptation module: used to determine the current medical operation scenario and match the acquisition strategy, synchronously acquire the patient's real-time physiological parameters, real-time medical image data and operation instrument pose data, and transmit the data to the individualized risk fusion module after completing the data timestamp alignment; Individualized Risk Fusion Module: Based on the data from the preceding modules, an individualized prediction model is constructed. The three-dimensional pose prediction of the target organization is performed through an individualized pose look-ahead prediction algorithm. The operational risk is quantitatively assessed and classified using a multi-dimensional risk feedforward quantification algorithm. Potential risk areas are marked. Finally, real-time images, semi-transparent prediction images, and risk marking information are integrated to generate a standardized fusion data package. AR calibration and rendering module: Receives visual acuity and ambient light intensity parameters from the individualized data acquisition module and fusion data packets output by the individualized risk fusion module. Corrects the refractive state of the image through the dynamic optical adjustment component of the AR glasses. Executes an ambient light adaptation strategy based on the ambient light intensity data to adjust the contrast, transparency, and brightness of the image. Performs standardized parsing and rendering processing on the fusion data packets. AR guidance optimization module: It is used to overlay the calibrated and rendered fused image onto the real scene in AR glasses, supports user interaction, provides operation advance guidance and multi-dimensional warnings, and collects operation feedback data to iteratively optimize the system algorithm parameters and prediction model.
[0006] Furthermore, in the personalized data acquisition module, the specific process of collecting medical staff's visual parameters and current ambient light intensity data, retrieving and storing the patient's past physiological-image correlation data is as follows: The medical staff's myopia, astigmatism, and astigmatism axis are obtained through the non-contact vision detection component built into the AR glasses frame or through a manual input interface; the ambient light intensity data of the current medical scene is collected and the light type is identified through the photosensitive sensor deployed at the front of the AR glasses frame; simultaneously, through the interface between the hospital information system and the medical image archiving and communication system, the patient's historical heart rate and vascular pulsation synchronous recording data, historical respiratory waveform and thoracic tissue movement correlation data, historical vascular elasticity detection data, and historical coagulation function detection data are retrieved; the collected and retrieved data are then integrated and stored.
[0007] Furthermore, in the multimodal acquisition and adaptation module, the specific process of determining the current medical operation scenario and matching the acquisition strategy is as follows: Using the scene recognition algorithm built into the AR glasses, feature markers are extracted from the current medical environment image. These feature markers include the edge contour of the operating table, the rod-like structure of the IV stand, and the specific shape of interventional diagnostic and treatment instruments. Based on the extracted feature markers, the type of the current medical operation scenario is determined. This type includes intravenous puncture scenarios, tubing care scenarios, and interventional diagnostic and treatment scenarios. For different determination results, corresponding acquisition strategies are matched. Specifically, the strategy for intravenous puncture scenarios focuses on acquiring real-time physiological parameters of the patient's upper limb region, real-time medical image data of the superficial veins of the upper limb, and the pose data of the puncture needle. The strategy for tubing care scenarios focuses on acquiring real-time physiological parameters around the tubing placement site, real-time medical image data of the tubing placement area, and the pose data of the tubing tip. The strategy for interventional diagnostic and treatment scenarios focuses on acquiring real-time physiological parameters of the patient's chest and abdomen region, real-time medical image data of the chest and abdomen during the procedure, and the pose data of the interventional catheter.
[0008] Furthermore, in the multimodal acquisition and adaptation module, the scene recognition algorithm adopts a convolutional neural network architecture. Its specific structure and processing are as follows: The algorithm includes an input layer, a feature extraction layer, a feature fusion layer, and a classification decision layer. The input layer receives color image data of the current medical environment acquired by the AR glasses and performs size standardization. The feature extraction layer consists of four concatenated convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function. The first two convolutional blocks use 3×3 convolutional kernels to extract basic edge and texture features of the image, while the latter two convolutional blocks use 5×5 convolutional kernels to extract feature vectors for the operating table edge contour, the infusion stand rod structure, and the shape of interventional diagnostic and treatment instruments. The feature fusion layer concatenates feature vectors of different dimensions and performs dimensional adjustment and feature filtering using a 1×1 convolutional kernel to obtain a global feature vector. The classification decision layer contains two fully connected layers. The first layer maps the global feature vector to a fixed-dimensional feature representation, and the second layer connects to a Softmax classifier, outputting the category probabilities of the intravenous puncture scene, the tubing care scene, and the interventional diagnostic and treatment scene to complete the scene determination.
[0009] Furthermore, in the individualized risk fusion module, the individualized prediction model is constructed based on the patient's past physiological-imaging correlation data and real-time dynamic data, including two parallel sub-models: a heartbeat-vascular pulsation prediction sub-model and a respiratory-thoracic motion prediction sub-model. Both sub-models adopt a long short-term memory network architecture, including an input layer, a hidden layer, and an output layer. The input layer receives standardized historical and real-time physiological signal data, the hidden layer is composed of multiple layers of long short-term memory units connected in series to mine temporal correlation features, and the output layer outputs pose parameters related to tissue motion. The model construction process is as follows: first, the patient's past physiological-imaging correlation data is used as training samples to train the two sub-models offline. During the training process, the input data is preprocessed with temporal alignment and normalization, and an adaptive momentum optimizer is used to adjust the model parameters. After the offline training is completed, real-time physiological parameters and imaging feature data are input into the model for online fine-tuning to adapt the model to the patient's current physiological state. The output results of the two sub-models are integrated through a data fusion unit to form unified basic data for tissue motion prediction.
[0010] Furthermore, in the individualized risk fusion module, the mathematical expression of the individualized pose look-ahead prediction algorithm is: ,in, yes The organization pose prediction vector at time t. For individualized weights, This serves as the historical reference vector for organizational pose. For the first Association weights of physiological signals For the first Physiological signals in The normalized deviation value at time 1. Weights are corrected for real-time images, and , for Organize pose vectors in real time. This is the noise compensation coefficient. Based on The pose prediction noise correction vector generated by Gaussian filtering. This indicates the number of categories of physiological signals.
[0011] Furthermore, in the individualized risk fusion module, the mathematical expression of the multi-dimensional risk feedforward quantification algorithm is: ,in For the future Quantitative value of operational risk at any given time As the pose deviation risk weight, For future organizational pose vectors, This is the real-time pose vector of the operating instrument. The pose threshold vector for safe operation. For vascular elasticity risk weight, The target vascular elasticity value, This is the baseline value for normal vascular elasticity. For coagulation function risk weight, This is a real-time coagulation function value. This is the baseline value for normal coagulation function.
[0012] Furthermore, in the individualized risk fusion module, a multi-dimensional risk feedforward quantification algorithm is used to quantitatively assess and classify operational risks, and the specific content of marking potential risk areas is as follows: based on Classify into levels: As the first level, It is the second level. It is the third level. The risk level is set to the fourth level. Based on real-time medical images, the spatial location of risk points is determined, and corresponding color labels are matched for different levels: green for the first level, yellow for the second level, orange for the third level, and red for the fourth level. The color labels are superimposed on the corresponding positions of the real-time image layer and the predicted image layer to form an image layer containing risk markers.
[0013] Furthermore, in the AR calibration rendering module, the specific process of correcting the refractive state of the image, performing ambient light adaptation, and parsing the rendering fusion data packet is as follows: the dynamic optical adjustment component adjusts the parameters of the built-in micro-liquid crystal lens array based on the received visual acuity parameters to match the myopia, astigmatism, and astigmatic axis of the medical staff, thereby completing refractive correction. This adjustment can adapt in real time to changes in the head posture of the medical staff. The ambient light adaptation strategy is as follows: image parameters are adjusted according to the ambient light intensity data and light type for different scenes. When the ambient light intensity is higher than a first preset threshold, the image contrast is improved. The brightness of risk markers is increased; when the ambient light intensity is lower than the second preset threshold, the image display brightness is increased and anti-glare processing is enabled; when natural light is detected, the image white balance is adjusted to an appropriate state; the standardized parsing and rendering process of the fused data package is as follows: first, the real-time image layer, semi-transparent prediction layer and risk marker layer contained in the fused data package are parsed, then the resolution of each layer is unified to the standard resolution and the frame rate is stabilized, the prediction layer is semi-transparent, and then the three types of layers are superimposed and rendered according to the preset hierarchical relationship to generate standardized image data that can be directly used for AR display.
[0014] Furthermore, in the AR guidance optimization module, the overlay display uses SLAM spatial positioning technology to achieve spatial registration between virtual images and real-world scenes, supporting voice commands and gesture operations. The gesture operations include pinching to zoom, swiping to switch layers, and clicking to confirm. The specific implementation process of the SLAM spatial positioning technology is as follows: the image sensor and inertial measurement unit built into the AR glasses synchronously collect environmental images and device motion data, extract stable feature points in the environment, and construct a feature dictionary; initial positioning and local map initialization are completed based on the initial frame image, and real-time pose tracking is subsequently achieved through inter-frame feature matching and inertial data fusion; the local dense map is continuously updated, and the coordinates of the virtual image in the fused data package are accurately mapped to the coordinates of the real environment to complete spatial registration.
[0015] Compared with existing technologies, this multimodal medical imaging and AR glasses fusion display system has the following advantages: I. This invention integrates medical staff's visual acuity parameters, ambient light intensity data, and patients' past and real-time physiological-image correlation information. Combined with scene recognition and adaptive acquisition strategies, it achieves accurate acquisition and temporal alignment of multi-dimensional data. Relying on individualized prediction models, it mines the correlation features between physiological signals and tissue movement. Through pose-prospective prediction algorithms, it anticipates dynamic changes in target tissues in advance. With the help of multi-dimensional risk quantification algorithms, it clarifies the operational risk level and marks potential risk areas, providing medical staff with comprehensive risk prediction basis. At the same time, through dynamic optical adjustment to correct refractive state and adapting image display parameters according to ambient light intensity, it ensures that the fused image is clearly distinguishable under different conditions, reducing the impact of visual acuity differences and environmental interference on the observation effect. This helps medical staff quickly and accurately grasp the core operational information, reduces operational errors caused by tissue movement and information deviation, and improves the accuracy and safety of medical operations.
[0016] II. This invention utilizes SLAM spatial positioning technology to achieve precise overlay of virtual fused images onto real-world scenes. It supports diverse interaction methods such as voice and gestures, and, combined with proactive operation guidance and multi-dimensional early warning functions, allows medical staff to receive real-time targeted guidance during operations, flexibly adjusting their operational rhythm and path. The system continuously collects operational feedback data, dynamically iterating and optimizing algorithm parameters and prediction models to adapt to individual patient physiological differences and dynamic changes in medical operations, enhancing the system's personalized adaptability. It matches specific data collection strategies for different medical scenarios, ensuring the targeted and effective collection of data. This allows fused images to accurately meet various operational needs, simplifying the information acquisition process for medical staff, reducing their workload, and achieving intelligent guidance and risk control for medical operations through the collaborative work of multiple modules. This promotes the development of medical operations towards greater efficiency and standardization, providing reliable technical support for various medical scenarios.
[0017] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0019] Figure 1 A flowchart illustrating the entire operation process of a multimodal medical imaging and AR glasses fusion display system; Figure 2 This is a diagram showing the overall framework connection of a multimodal medical imaging and AR glasses fusion display system. Figure 3 This is a flowchart of the data processing for the individualized risk fusion module in a multimodal medical imaging and AR glasses fusion display system. Detailed Implementation
[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below. Example
[0021] Example of a venipuncture scenario.
[0022] After medical staff put on the AR glasses that come with the system, the system activates the individualized data acquisition and processing module. The non-contact vision detection component built into the AR glasses automatically acquires the medical staff's myopia, astigmatism, and astigmatic axis, ensuring that subsequent refractive correction of images accurately matches the medical staff's vision. If the detection component malfunctions, the above vision parameters can be manually entered through an input interface to ensure the integrity of the basic data. Simultaneously, the photosensor at the front of the AR glasses frame collects real-time ambient light intensity data in the current ward, accurately identifying the light type as indoor artificial lighting, providing a basis for the execution of subsequent ambient light adaptation strategies. Subsequently, the system retrieves the patient's historical heart rate and vascular pulsation synchronous recording data, historical respiratory waveform and chest tissue movement correlation data, historical vascular elasticity test data, and historical coagulation function test data through the interface between the hospital information system and the medical image archiving and communication system. The system integrates and stores the collected vision parameters, ambient light intensity data, and retrieved patient data, providing comprehensive and accurate basic data support for the operation of subsequent modules, such as... Figure 1 As shown.
[0023] The system enters the multimodal acquisition and adaptation module, and performs scene determination through the scene recognition algorithm of the convolutional neural network architecture built into the AR glasses. The input layer of this algorithm receives the current medical environment color image data acquired by the AR glasses and performs size standardization to make subsequent feature extraction more efficient. The feature extraction layer consists of four concatenated convolutional blocks. Each convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU activation function. The first two convolutional blocks use 3×3 convolutional kernels to extract basic edge and texture features of the image, while the last two convolutional blocks use 5×5 convolutional kernels to extract feature vectors such as the edge contour of the operating table and the rod-like structure of the IV stand, ensuring accurate capture of specific landmark features in the medical scene. The feature fusion layer concatenates feature vectors of different dimensions, adjusts the dimensions and filters features through 1×1 convolutional kernels to reduce redundant information, and obtains a global feature vector. The classification decision layer processes the global feature vector through two fully connected layers, and then outputs the category probability of each scene through a Softmax classifier, ultimately accurately determining that the current scene is a venipuncture scene. Based on the judgment result, the system matches the corresponding acquisition strategy and simultaneously acquires the real-time physiological parameters of the patient's upper limb region, the real-time medical imaging data of the superficial veins of the upper limb, and the position data of the puncture needle. It focuses on the core operation area of venous puncture to ensure that the acquired data is highly consistent with the operation requirements. After completing the timestamp alignment of all data, the integrated data is transmitted to the individualized risk fusion module to ensure that subsequent risk assessment and prediction can be carried out based on time-consistent data.
[0024] After receiving data, the individualized risk fusion module constructs an individualized prediction model based on the patient's past physiological-imaging correlation data and real-time dynamic data. This model includes two parallel sub-models: a heartbeat-vascular pulsation prediction sub-model and a respiratory-thoracic motion prediction sub-model. Both sub-models employ a long short-term memory network architecture, comprising an input layer, a hidden layer, and an output layer. The system first uses the patient's past physiological-imaging correlation data as training samples to train both sub-models offline. During training, the input data undergoes temporal alignment and normalization preprocessing, and the model parameters are adjusted using an adaptive momentum optimizer to give the model preliminary tissue motion prediction capabilities. After offline training, real-time collected physiological parameters and imaging feature data are input into the model for online fine-tuning, enabling the model to accurately adapt to the patient's current physiological state and improve the specificity of tissue motion prediction. The tissue motion-related pose parameters output by the two sub-models are integrated by the data fusion unit to form unified basic data for tissue motion prediction, providing coherent and comprehensive data support for subsequent pose prediction. Subsequently, the system predicts the 3D pose of the target tissue using a personalized pose look-ahead prediction algorithm. The mathematical expression of the personalized pose look-ahead prediction algorithm is as follows: ,in, yes The organization pose prediction vector at time t. For individualized weights, This serves as the historical reference vector for organizational pose. For the first Association weights of physiological signals For the first Physiological signals in The normalized deviation value at time 1. Weights are corrected for real-time images, and , for Organize pose vectors in real time. This is the noise compensation coefficient. Based on The pose prediction noise correction vector generated by Gaussian filtering. This indicates the number of physiological signal categories, allowing for early prediction of tissue movement trends and providing a basis for proactive operational guidance. Furthermore, a multi-dimensional risk feedforward quantification algorithm is used to quantitatively assess operational risk. The mathematical expression of this algorithm is: ,in For the future Quantitative value of operational risk at any given time As the pose deviation risk weight, For future organizational pose vectors, This is the real-time pose vector of the operating instrument. The pose threshold vector for safe operation. For vascular elasticity risk weight, The target vascular elasticity value, This is the baseline value for normal vascular elasticity. For coagulation function risk weight, This is a real-time coagulation function value. Using the baseline value for normal coagulation function, risk levels are categorized based on quantification results. A quantification value less than 1 indicates Level 1, labeled with a green tag; a value greater than or equal to 1 and less than 3 indicates Level 2, labeled with a yellow tag; a value greater than or equal to 3 and less than 4 indicates Level 3, labeled with an orange tag; and a value greater than or equal to 4 indicates Level 4, labeled with a red tag. These color tags correspond to the risk level, allowing medical staff to intuitively identify the degree of risk. The system locates the spatial position of risk points based on real-time medical images, overlaying the corresponding color tags onto the real-time image layer and the predicted image layer. Finally, it integrates real-time images, semi-transparent predicted images, and risk marker information to generate a standardized fusion data package, enabling subsequent rendering to simultaneously present multiple types of key information, such as... Figure 3 As shown.
[0025] The AR calibration and rendering module receives visual acuity parameters and ambient light intensity data transmitted from the individualized data acquisition and processing module, as well as fusion data packets output by the individualized risk fusion module. This module, through the dynamic optical adjustment components of the AR glasses, adjusts the parameters of the built-in micro-liquid crystal lens array according to the medical staff's myopia, astigmatism, and astigmatic axis to complete image refractive state correction, allowing medical staff to clearly observe images without wearing additional glasses. Simultaneously, an ambient light adaptation strategy is implemented. If the current ambient light intensity is indoor artificial lighting and does not exceed a preset threshold, the system maintains the image's basic contrast, transparency, and brightness to ensure viewing comfort. If the ambient light intensity subsequently exceeds the first preset threshold, the system automatically increases image contrast and enhances the brightness of risk markers to prevent image blurring caused by strong light. If it falls below the second preset threshold, the system increases image display brightness and activates anti-glare processing to prevent weak light from affecting viewing clarity. When natural light is detected, the image white balance is adjusted to an adaptive state, making the image colors closer to the real scene. Finally, the system performs standardized parsing of the fused data package, extracting the real-time image layer, semi-transparent prediction layer, and risk marker layer. It unifies the resolution of each layer to a standard resolution and stabilizes the frame rate to reduce image flicker. After semi-transparent processing of the prediction layer, it overlays and renders the data according to a preset hierarchical relationship to generate standardized image data that can be directly used for AR display. This ensures that various types of information do not obscure each other after being overlaid, making it easier for medical staff to quickly obtain key information.
[0026] The AR guidance optimization module overlays standardized image data, after calibration and rendering, onto the real-world scene within AR glasses using SLAM spatial positioning technology. This technology synchronously acquires environmental images and device motion data through the AR glasses' built-in image sensors and inertial measurement units, extracts stable feature points from the environment, constructs a feature dictionary, and completes initial localization and local map initialization based on the initial frame image. Subsequently, real-time pose tracking is achieved through inter-frame feature matching and inertial data fusion, continuously updating the local dense map and accurately mapping virtual image coordinates to real-world environment coordinates, ensuring a misaligned overlay display and preventing misleading operations. The system supports interaction by medical staff via voice commands or gestures. Gesture operations include pinching to zoom in and out of the image, swiping to switch between different layers, and clicking to confirm operation commands, eliminating the need for manual device operation and improving ease of use. Simultaneously, based on predicted tissue pose and quantified risk levels, the system provides advance guidance for puncture procedures, clearly defining the operation path and key areas. Multi-dimensional warnings are issued for different risk levels, with an audible and visual warning for the fourth-level red risk area to prompt avoidance, reducing the probability of operational errors. During the operation, the system collects real-time feedback data from medical staff to iteratively optimize the system algorithm parameters and individualized prediction models, thereby continuously improving the accuracy and safety of subsequent operations.
[0027] In summary, in this embodiment, the multimodal medical imaging and AR glasses fusion display system operates in an orderly manner through five stages: individualized data collection and processing, multimodal acquisition and adaptation, individualized risk fusion, AR calibration and rendering, and AR-guided optimization. The system first accurately collects the vision parameters of medical staff, ambient light intensity data, and the patient's past related data. Then, it uses a scene recognition algorithm based on a convolutional neural network architecture to determine the intravenous puncture scenario and match a targeted data collection strategy. An individualized prediction model based on a long short-term memory network architecture, combined with two core algorithms, achieves tissue pose prediction and risk quantification and grading. After AR calibration and rendering, SLAM technology is used to accurately overlay the virtual and real scenes. No additional glasses adaptation is required throughout the process, supporting convenient interaction and real-time alerts. Synchronous data collection and feedback optimize the system, significantly improving the accuracy and safety of intravenous puncture.
[0028] Example 2: Example of interventional diagnosis and treatment scenario.
[0029] After medical staff put on AR glasses in the interventional operating room, the personalized data acquisition and processing module starts working. The non-contact vision detection component on the AR glasses frame quickly collects the medical staff's myopia, astigmatism, and astigmatic axis, ensuring that refractive correction can be quickly adapted to the medical staff's vision, saving preoperative preparation time. If the detection data is abnormal, it can be corrected and supplemented through a manual input interface to avoid errors in the basic data affecting subsequent procedures. The photosensor at the front of the AR glasses frame collects ambient light intensity data in real time in the interventional operating room, accurately identifying the light type as surgical shadowless lamp illumination, providing precise basis for the ambient light adaptation strategy. Next, the system retrieves the patient's historical heart rate and vascular pulsation synchronous recording data, historical respiratory waveform and chest tissue movement correlation data, historical vascular elasticity test data, and historical coagulation function test data through the interface between the hospital information system and the medical image archiving and communication system. The collected vision parameters, ambient light intensity data, and retrieved patient's past data are integrated and stored, providing comprehensive and targeted basic data for subsequent personalized prediction model construction and risk assessment, ensuring the accuracy of system operation. Figure 2 As shown.
[0030] After the multimodal acquisition and adaptation module is activated, scene determination is performed using the scene recognition algorithm based on the convolutional neural network architecture built into the AR glasses. The algorithm input layer receives color image data of the interventional operating room environment acquired by the AR glasses and performs size standardization, laying a unified foundation for feature extraction. The four cascaded convolutional blocks of the feature extraction layer work according to a predetermined structure. The first two convolutional blocks use 3×3 convolutional kernels to extract basic edge and texture features, while the last two convolutional blocks use 5×5 convolutional kernels to extract feature vectors such as the edge contour of the operating table and the specific shape of interventional diagnostic and treatment instruments, ensuring that key landmarks of the interventional diagnostic and treatment scene can be captured. The feature fusion layer concatenates feature vectors of different dimensions and processes them with a 1×1 convolutional kernel to obtain a global feature vector, simplifying effective information. The classification decision layer outputs the category probability through a fully connected layer and a Softmax classifier, ultimately accurately determining that the current scene is an interventional diagnostic and treatment scene. Based on this result, the system matches the corresponding acquisition strategy and simultaneously acquires real-time physiological parameters of the patient's chest and abdomen, real-time medical imaging data of the chest and abdomen during the operation, and the position data of the interventional catheter. It focuses on the core operation area of interventional diagnosis and treatment to ensure that the acquired data can accurately reflect the patient's key physiological state and the status of the operating instruments. After completing the timestamp alignment of all data, the data is transmitted to the individualized risk fusion module to ensure the temporal consistency of subsequent data processing.
[0031] After receiving data, the individualized risk fusion module constructs an individualized prediction model comprising a heartbeat-vascular pulsation prediction sub-model and a respiratory-thoracic motion prediction sub-model. Both sub-models employ a long short-term memory network architecture. The system first uses the patient's past physiological-imaging correlation data as training samples to train the two sub-models offline. During training, data temporal alignment and normalization preprocessing are performed, and parameters are adjusted using an adaptive momentum optimizer to allow the model to grasp the patient's historical tissue movement patterns. After offline training, real-time physiological parameters and imaging feature data are input for online fine-tuning, enabling the model to quickly adapt to the patient's current physiological state and improve prediction accuracy. The outputs of the two sub-models are integrated by the data fusion unit to form the basic data for tissue movement prediction, ensuring the integrity of the prediction data. Subsequently, an individualized pose prospective prediction algorithm predicts the three-dimensional pose of the target tissue, allowing for advance understanding of the tissue movement trajectory and providing predictive support for interventional procedures. The mathematical expression of the individualized pose prospective prediction algorithm is: ,in, yes The organization pose prediction vector at time t. For individualized weights, This serves as the historical reference vector for organizational pose. For the first Association weights of physiological signals For the first Physiological signals in The normalized deviation value at time 1. Weights are corrected for real-time images, and , for Organize pose vectors in real time. This is the noise compensation coefficient. Based on The pose prediction noise correction vector generated by Gaussian filtering. This represents the number of categories of physiological signals; operational risk is quantified using a multi-dimensional risk feedforward quantization algorithm, the mathematical expression of which is: ,in For the future Quantitative value of operational risk at any given time As the pose deviation risk weight, For future organizational pose vectors, This is the real-time pose vector of the operating instrument. The pose threshold vector for safe operation. For vascular elasticity risk weight, The target vascular elasticity value, This is the baseline value for normal vascular elasticity. For coagulation function risk weight, This is a real-time coagulation function value. Using normal coagulation function as a baseline, risk levels are categorized according to standards: a value less than 1 corresponds to Level 1 (green label), 1 ≤ value < 3 corresponds to Level 2 (yellow label), 3 ≤ value < 4 corresponds to Level 3 (orange label), and a value ≥ 4 corresponds to Level 4 (red label), visually presenting the degree of risk. The system locates risk points spatially based on real-time medical images, overlays corresponding color labels, and integrates real-time images, semi-transparent predictive images, and risk marker information to generate a standardized fusion data package. This ensures the orderly integration of various key information, providing a clear data foundation for subsequent rendering.
[0032] After receiving visual acuity parameters, ambient light intensity data, and fusion data packets, the AR calibration rendering module adjusts the parameters of the micro-liquid crystal lens array through a dynamic optical adjustment component. Based on the medical staff's visual acuity parameters, it corrects the image refractive state, allowing them to clearly observe the image during surgery without needing to focus on visual adaptation. For the ambient light intensity of the current surgical shadowless lamp, the system determines it exceeds a first preset threshold and executes an ambient light adaptation strategy, enhancing image contrast and increasing the brightness of risk markers to prevent blurring of image details under strong light and ensure clear visibility of risk markers. Subsequently, the system analyzes the real-time image layer, semi-transparent prediction layer, and risk marker layer in the fusion data packet, unifying the resolution of each layer and stabilizing the frame rate to reduce image jitter. The prediction layer is semi-transparent, and the layers are overlaid and rendered according to a preset hierarchical relationship to generate standardized AR display image data. This ensures clear distinction between real-time images, predicted images, and risk markers, facilitating quick differentiation of various information by medical staff in complex surgical scenarios.
[0033] The AR-guided optimization module employs SLAM spatial positioning technology to precisely overlay calibrated and rendered images onto the real-world scene of the interventional operating room. This technology uses the image sensors and inertial measurement units of the AR glasses to collect environmental images and equipment motion data, extracting stable feature points to construct a feature dictionary. After initial localization and local map initialization, real-time pose tracking is achieved through inter-frame feature matching and inertial data fusion, continuously updating the local dense map to achieve precise mapping between virtual image coordinates and real-world environment coordinates. This ensures a perfect fit between the overlaid image and the surgical area, avoiding operational deviations. Medical staff can interact with the system via voice commands or gestures. Pinch to zoom for better detail observation, swipe to switch layers to focus on different information, and click to confirm for quick execution of commands, improving the continuity of surgical procedures. Based on tissue pose predictions and risk levels, the system provides forward-looking guidance for interventional catheter procedures, clarifying the catheter insertion path and precautions, issuing warnings for high-risk areas, reminding medical staff to avoid risks, and reducing the probability of surgical complications. During the operation, the system collects real-time feedback data from medical staff to iteratively optimize the system algorithm parameters and individualized prediction models, making the system more aligned with actual operational needs in subsequent interventional diagnosis and treatment scenarios, and continuously improving the safety and accuracy of the surgery.
[0034] In summary, this embodiment demonstrates that the system is fully adapted for interventional diagnosis and treatment scenarios. From the individualized data acquisition stage, which accurately obtains visual acuity, ambient light, and patient history data; to the multimodal acquisition and adaptation stage, which uses specific algorithms to accurately identify the scene and collect core chest and abdominal data; to the individualized risk fusion stage, which relies on dual parallel sub-models and core algorithms to achieve pose prediction and risk grading; subsequently, AR calibration and rendering are applied to adapt to the strong light environment of the operating room; and finally, SLAM technology is used to achieve accurate overlay of images onto reality. The system supports voice and gesture interaction, provides pre-operation guidance and high-risk warnings, and synchronously collects feedback data to iteratively optimize algorithms and models, fully meeting the needs of interventional diagnosis and treatment operations, effectively improving surgical continuity and accuracy, and reducing the probability of complications.
[0035] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A fusion display system for multimodal medical imaging and AR glasses, characterized in that, The system includes: Personalized data acquisition and processing module: used to collect medical staff's vision parameters and current ambient light intensity data, retrieve patients' past physiological-image correlation data, and integrate and store them; Multimodal acquisition and adaptation module: used to determine the current medical operation scenario and match the acquisition strategy, synchronously acquire the patient's real-time physiological parameters, real-time medical image data and operation instrument pose data, and transmit the data to the individualized risk fusion module after completing the data timestamp alignment; Individualized Risk Fusion Module: Based on the data from the preceding modules, an individualized prediction model is constructed. The three-dimensional pose prediction of the target organization is performed through an individualized pose look-ahead prediction algorithm. The operational risk is quantitatively assessed and classified using a multi-dimensional risk feedforward quantification algorithm. Potential risk areas are marked. Finally, real-time images, semi-transparent prediction images, and risk marking information are integrated to generate a standardized fusion data package. AR calibration and rendering module: Receives visual acuity and ambient light intensity parameters from the individualized data acquisition module and fusion data packets output by the individualized risk fusion module. Corrects the refractive state of the image through the dynamic optical adjustment component of the AR glasses. Executes an ambient light adaptation strategy based on the ambient light intensity data to adjust the contrast, transparency, and brightness of the image. Performs standardized parsing and rendering processing on the fusion data packets. AR guidance optimization module: It is used to overlay the calibrated and rendered fused image onto the real scene in AR glasses, supports user interaction, provides operation advance guidance and multi-dimensional warnings, and collects operation feedback data to iteratively optimize the system algorithm parameters and prediction model.
2. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, The specific process of collecting medical staff's vision parameters and current ambient light intensity data, retrieving and storing the patient's past physiological-image correlation data in the personalized data acquisition module is as follows: The medical staff's myopia, astigmatism, and astigmatism axis are obtained through the non-contact vision detection component built into the AR glasses frame or through a manual input interface; the ambient light intensity data of the current medical scene is collected and the light type is identified through the photosensitive sensor deployed at the front of the AR glasses frame; simultaneously, through the interface between the hospital information system and the medical image archiving and communication system, the patient's historical heart rate and vascular pulsation synchronous recording data, historical respiratory waveform and thoracic tissue movement correlation data, historical vascular elasticity detection data, and historical coagulation function detection data are retrieved; the collected and retrieved data are then integrated and stored.
3. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the multimodal acquisition and adaptation module, the specific process of determining the current medical operation scenario and matching the acquisition strategy is as follows: Using the scene recognition algorithm built into the AR glasses, feature markers are extracted from the current medical environment image. These feature markers include the edge contour of the operating table, the rod-like structure of the IV stand, and the specific shape of interventional diagnostic and treatment instruments. Based on the extracted feature markers, the type of the current medical operation scenario is determined. This type includes intravenous puncture scenarios, tubing care scenarios, and interventional diagnostic and treatment scenarios. For different determination results, corresponding acquisition strategies are matched. Specifically, the strategy for intravenous puncture scenarios focuses on acquiring real-time physiological parameters of the patient's upper limb region, real-time medical image data of the superficial veins of the upper limb, and the pose data of the puncture needle. The strategy for tubing care scenarios focuses on acquiring real-time physiological parameters around the tubing placement site, real-time medical image data of the tubing placement area, and the pose data of the tubing tip. The strategy for interventional diagnostic and treatment scenarios focuses on acquiring real-time physiological parameters of the patient's chest and abdomen region, real-time medical image data of the chest and abdomen during the procedure, and the pose data of the interventional catheter.
4. The multimodal medical imaging and AR glasses fusion display system according to claim 3, characterized in that, In the multimodal acquisition and adaptation module, the scene recognition algorithm adopts a convolutional neural network architecture. Its specific structure and processing are as follows: The algorithm includes an input layer, a feature extraction layer, a feature fusion layer, and a classification decision layer. The input layer receives color image data of the current medical environment acquired by the AR glasses and performs size standardization. The feature extraction layer consists of four concatenated convolutional blocks. Each convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU activation function. The first two convolutional blocks use 3×3 convolutional kernels to extract basic edge and texture features of the image, while the last two convolutional blocks use 5×5 convolutional kernels to extract feature vectors for the operating table edge contour, the infusion stand rod structure, and the shape of interventional diagnostic and treatment instruments. The feature fusion layer concatenates feature vectors of different dimensions and performs dimensional adjustment and feature filtering using a 1×1 convolutional kernel to obtain a global feature vector. The classification decision layer contains two fully connected layers. The first layer maps the global feature vector to a fixed-dimensional feature representation, and the second layer connects to a Softmax classifier, outputting the category probabilities of intravenous puncture scenes, tubing care scenes, and interventional diagnostic and treatment scenes to complete scene determination.
5. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the individualized risk fusion module, the individualized prediction model is constructed based on the patient's past physiological-imaging correlation data and real-time dynamic data. It includes two parallel sub-models: a heartbeat-vascular pulsation prediction sub-model and a respiratory-thoracic movement prediction sub-model. Both sub-models adopt a long short-term memory network architecture, which includes an input layer, a hidden layer, and an output layer. The input layer receives standardized historical and real-time physiological signal data. The hidden layer is composed of multiple layers of long short-term memory units connected in series to mine temporal correlation features. The output layer outputs the pose parameters related to tissue movement. The model building process is as follows: First, the patient's previous physiological-image correlation data is used as training samples to train the two sub-models offline. During the training process, the input data is preprocessed by temporal alignment and normalization, and the model parameters are adjusted by an adaptive momentum optimizer. After the offline training is completed, real-time physiological parameters and image feature data are input into the model for online fine-tuning to make the model adapt to the patient's current physiological state. The outputs of the two sub-models are integrated through a data fusion unit to form unified basic data for predicting tissue movement.
6. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the individualized risk fusion module, the mathematical expression of the individualized pose look-ahead prediction algorithm is: ,in, yes The organization pose prediction vector at time t. For individualized weights, This serves as the historical reference vector for organizational pose. For the first Association weights of physiological signals For the first Physiological signals in The normalized deviation value at time 1. Weights are corrected for real-time images, and , for Organize pose vectors in real time. This is the noise compensation coefficient. Based on The pose prediction noise correction vector generated by Gaussian filtering. This indicates the number of categories of physiological signals.
7. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the individualized risk fusion module, the mathematical expression of the multi-dimensional risk feedforward quantification algorithm is: ,in For the future Quantitative value of operational risk at any given time As the pose deviation risk weight, For future organizational pose vectors, This is the real-time pose vector of the operating instrument. The pose threshold vector for safe operation. For vascular elasticity risk weight, The target vascular elasticity value, This is the baseline value for normal vascular elasticity. For coagulation function risk weight, This is a real-time coagulation function value. This is the baseline value for normal coagulation function.
8. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the individualized risk fusion module, a multi-dimensional risk feedforward quantification algorithm is used to quantitatively assess and classify operational risks. The specific content of marking potential risk areas is as follows: based on... Classify into levels: It is the first level. It is the second level. It is the third level. The risk level is set to the fourth level. Based on real-time medical images, the spatial location of risk points is determined, and corresponding color labels are matched for different levels: green for the first level, yellow for the second level, orange for the third level, and red for the fourth level. The color labels are superimposed on the corresponding positions of the real-time image layer and the predicted image layer to form an image layer containing risk markers.
9. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the AR calibration rendering module, the specific process of correcting the refractive state of the image, performing ambient light adaptation, and parsing the rendering fusion data package is as follows: The dynamic optical adjustment component adjusts the parameters of the built-in micro-liquid crystal lens array based on the received visual acuity parameters to match the myopia degree, astigmatism degree, and astigmatic axis of the medical staff, thereby completing refractive correction. This adjustment can adapt in real time to changes in the head posture of the medical staff. The ambient light adaptation strategy is as follows: image parameters are adjusted according to the ambient light intensity data and light type for different scenes. When the ambient light intensity is higher than a first preset threshold, the image contrast is increased and enhanced. High risk marker brightness; when the ambient light intensity is lower than the second preset threshold, increase the image display brightness and enable anti-glare processing; when it is identified as natural light, adjust the image white balance to the appropriate state; the standardized parsing and rendering process of the fused data package is as follows: first, parse the real-time image layer, semi-transparent prediction layer and risk marker layer contained in the fused data package, then unify the resolution of each layer to the standard resolution and stabilize the frame rate, after semi-transparent processing of the prediction layer, the three types of layers are superimposed and rendered according to the preset hierarchical relationship to generate standardized image data directly used for AR display.
10. The multimodal medical imaging and AR glasses fusion display system according to claim 1, characterized in that, In the AR guidance optimization module, the overlay display uses SLAM spatial positioning technology to achieve spatial registration between virtual images and real-world scenes, supporting voice commands and gesture operations. The gesture operations include pinching to zoom, swiping to switch layers, and clicking to confirm. The specific implementation process of the SLAM spatial positioning technology is as follows: the image sensor and inertial measurement unit built into the AR glasses synchronously collect environmental images and device motion data, extract stable feature points in the environment, and construct a feature dictionary; initial positioning and local map initialization are completed based on the initial frame image, and real-time pose tracking is subsequently achieved through inter-frame feature matching and inertial data fusion; the local dense map is continuously updated, and the coordinates of the virtual image in the fused data package are accurately mapped to the coordinates of the real environment to complete spatial registration.