An interactive system based on augmented reality to improve the doctor-patient communication experience

Through an augmented reality-based interactive system, combined with deep perception devices and lesion recognition models, the problems of information complexity and knowledge level differences in traditional doctor-patient communication are solved, and more efficient and accurate communication effects are achieved.

CN119650104BActive Publication Date: 2025-05-16重庆至道科技股份有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510186913.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-16
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

There are time limitations, information complexity and knowledge level differences in traditional doctor-patient communication, which makes it difficult for patients to understand the professional terms of medical staff, which in turn affects communication efficiency and effectiveness.

Method used

Using an augmented reality-based interaction system, the patient's image information and imaging data are collected through depth perception devices, combined with the lesion recognition model and a three-dimensional human model, the lesion marks are displayed in real time, and dynamically adjusted according to the patient's movement to achieve "looking in the mirror"-like interaction.

Benefits of technology

It improves the efficiency and effectiveness of doctor-patient communication, enables patients to understand their condition more intuitively, enhances patients' sense of participation and trust, and reduces the time for doctors to speak, providing a more accurate treatment plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119650104B_ABST
    Figure CN119650104B_ABST
Patent Text Reader

Abstract

This solution belongs to the field of medical information technology, and specifically relates to an interactive system based on augmented reality to enhance the doctor-patient communication experience. This solution includes: an acquisition module, a processing module and a delivery module; the acquisition module includes a depth perception device, the depth perception device is used to collect image information of patients in a medical environment, and the medical environment includes a number of reference objects whose postures are fixed to the medical environment; used to collect the patient's imaging data and its corresponding diagnostic data; the processing module is used to construct a three-dimensional model of the medical environment as an environmental model based on the posture of the reference object in the medical environment and the three-dimensional model of each reference object, and establish an environmental coordinate system based on the environmental model, and the processing module also includes a three-dimensional model of the human body; the processing module segments the point cloud data of the reference object and the patient in the image information based on the reference object and the three-dimensional model of the human body. This solution improves the communication efficiency and effectiveness between medical staff and patients in a medical environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This solution belongs to the field of medical information technology, and specifically involves an interactive system based on augmented reality to enhance the doctor-patient communication experience. Background Art

[0002] In modern medical services, the quality of doctor-patient communication directly affects patient satisfaction and treatment outcomes. Traditional communication methods are often constrained by time constraints, information complexity, and differences in knowledge levels between medical staff and patients. When communicating with patients, medical staff are accustomed to using professional terms, but the professional terms used by medical staff are usually obscure and difficult for patients to understand. Without the help of other forms, patients often cannot understand what medical staff want to express in most cases. Therefore, in clinical practice, in order to communicate with patients and enable them to understand the ideas of medical staff, medical staff usually try to fully express their ideas by making gestures on paper, etc., but even so, patients cannot intuitively understand what medical staff mean.

[0003] At present, medical staff usually use specific popular science animations or virtual human models to explain pathology to patients. Popular science animations use vivid images and dynamic effects to intuitively show patients complex pathological processes. For example, when explaining heart valve lesions, animations can clearly show the opening and closing process of the valve and the changes in blood flow, so that patients can intuitively understand the function of the heart valve and the impact of lesions; virtual human models are personalized human models built from the patient's personal imaging data. Medical staff show the virtual human model displayed on the computer to the patient, and use the knob function of the model software to show the lesion model from different angles, so that patients can fully understand the location, shape and size of the lesion. For example, when explaining liver tumors, the relationship between the tumor and the liver can be observed from multiple perspectives such as front and back, left and right, and up and down. However, there are differences in the cognitive levels and needs of different patients. Some patients find it difficult to combine their specific conditions with animations or virtual human models to understand their own conditions and pathology, resulting in low communication efficiency between medical staff and patients. Summary of the invention

[0004] The purpose of this solution is to provide an interactive system based on augmented reality to enhance the doctor-patient communication experience, so as to improve the efficiency and effectiveness of communication between medical staff and patients in the medical environment.

[0005] In order to achieve the above objectives, this solution provides an interactive system based on augmented reality to enhance the doctor-patient communication experience, including:

[0006] The acquisition module includes a depth perception device, wherein the depth perception device is used to acquire image information of the patient in the treatment environment, wherein the treatment environment includes a plurality of reference objects whose positions are fixed to the treatment environment; and is used to acquire the patient's imaging data and the corresponding diagnostic data;

[0007] A processing module, used to construct a three-dimensional model of the medical environment as an environmental model based on the posture of the reference object in the medical environment and the three-dimensional models of each reference object, and establish an environmental coordinate system based on the environmental model, wherein the processing module also includes a three-dimensional model of a human body; the processing module segments the point cloud data of the reference object and the patient in the image information based on the reference object and the three-dimensional model of the human body, and then extracts key feature points from the point cloud data of the patient as first feature points, matches the three-dimensional model of each reference object with the point cloud data of the reference object, and obtains the key feature points of each reference object as second feature points; the processing module establishes a shooting coordinate system based on the image information and a preset fixed position, obtains the positions of the first feature point and the second feature point in the shooting coordinate system, and then obtains the position of the second feature point in the environmental coordinate system based on the environmental model, matches the environmental coordinate system and the shooting coordinate system based on the positions of the second feature point in the environmental coordinate system and the shooting coordinate system, and calculates the position of the first feature point in the environmental coordinate system;

[0008] A delivery module includes a lesion recognition model constructed by a deep learning model, wherein the delivery module puts the imaging data of the patient into the lesion recognition model for processing to obtain a lesion mark, then compares the lesion mark with the diagnostic data corresponding to the imaging data, modifies the lesion mark with inconsistent comparison results according to the diagnostic data, and uses the modified lesion mark and the imaging data to train the lesion recognition model; corresponds the position of the imaging data in the patient's three-dimensional model of the human body to the position of the lesion mark in the imaging data, determines the position of the lesion mark in the patient's three-dimensional model of the human body, and then determines the relative position of the lesion mark and the first feature point as the first position according to the patient's point cloud data and the patient's three-dimensional model of the human body; the delivery module is used to obtain the position of the first feature point in the image data in the environmental coordinate system, then determines the position of the lesion mark in the environmental coordinate system according to the first position, and determines the position of the lesion mark in the shooting coordinate system as the second position according to the matching result of the environmental coordinate system and the shooting coordinate system; the delivery module is used to display image information, and display the lesion mark in the image information according to the second position.

[0009] The principle and effect of this solution are as follows: First, this solution can display the patient's imaging data and lesion marks in real time, and when displaying, the lesion marks will move with the patient's movement (rotation, movement, etc.). Since the lesion marks are bound to the patient's three-dimensional human body model (i.e., the first position), the patient can observe and move at the same time, realizing the interaction like "looking in the mirror". This interactive method is more in line with the daily living habits of most patients, making the lesion marks displayed synchronously and in the same proportion in the image information (i.e., the "mirror") easier for patients to quickly understand and accept, thereby improving the doctor-patient communication experience. At the same time, this method reduces the pathological content that doctors need to describe, shortens the time for doctors to communicate pathology, and improves the efficiency and effectiveness of communication between medical staff and patients in the medical environment.

[0010] Secondly, this solution generates a personalized 3D model based on the patient's imaging data, allowing the patient to see the location and shape of the lesion that fully matches their body, enhancing the patient's sense of participation and trust. Through intuitive display, patients have a clearer expectation of the treatment process and reduce anxiety and fear caused by the unknown.

[0011] Furthermore, this solution uses the doctor's diagnostic data to modify the lesion mark, and the modified content is used to train the lesion recognition model. This reliable method can accurately locate the lesion and mark it on the patient's three-dimensional model, providing more accurate reference information for medical staff and helping to formulate more accurate treatment plans. At the same time, this method of integrating the patient's imaging data, diagnostic data and lesion marks makes it more convenient for medical staff to view patient information and reduces the time spent on reading medical records. Doctors can watch the same display content at the same time as patients, and quickly provide patients with more comprehensive and accurate supplementary explanations (no need to check the imaging data one by one before explaining to the patient, and no need to switch back and forth between popular science animations, virtual human models, actual imaging data and diagnostic data), further improving the efficiency and effectiveness of communication between medical staff and patients in the medical environment.

[0012] In summary, the interactive system based on augmented reality to enhance the doctor-patient communication experience provided by this solution improves the efficiency and effectiveness of communication between medical staff and patients in the medical environment.

[0013] Furthermore, the processing module is also used to obtain the patient's physiological characteristic information, combine the physiological characteristic information with the imaging data to predict the patient's body shape, combine the patient's body shape with the patient's point cloud data, divide the patient's point cloud data into clothing point cloud and torso point cloud, and combine the torso point cloud and the human body three-dimensional model to divide the first feature point according to the human torso structure, divide the patient's three-dimensional model into different torso areas with the first feature point as a node, and then determine the torso area as the target area according to the first position of the lesion mark; the processing module uses the torso area corresponding to the patient's hand as the operation area, and when the delivery module displays the lesion mark, the pixel points of the lesion mark are used to replace the pixel points corresponding to the target area in the image information; when the target area overlaps with other torso areas in the shooting coordinate system, the transparency of the torso area is increased to highlight the display effect of the target area, and then the effect of superimposing the torso area and the target area is displayed in the image information; when the target area overlaps with the operation area in the shooting coordinate system, the edge contour of the operation area is highlighted.

[0014] First, the patient's body shape is predicted by combining the patient's physiological characteristics information and imaging data, and the body point cloud and clothing point cloud are divided accordingly, so that the lesion mark can be accurately superimposed on the patient's body model to provide a personalized display. In addition, by dividing the patient's point cloud data and the feature point division of the human body three-dimensional model, the system can accurately locate the lesion and provide more accurate reference information for medical staff. Secondly, by combining the lesion mark with the patient's three-dimensional model and replacing the pixel points of the target area in the image information, the patient can intuitively see the specific location and range of the lesion, so as to better understand the condition. When the target area overlaps with other torso areas, the transparency of the torso area is increased to highlight the display effect of the target area, so that the patient can see the lesion location more clearly. When the target area overlaps with the operation area, the edge contour of the operation area is highlighted to guide the patient to focus on the key area and enhance the patient's sense of participation. This method of displaying different effects based on the overlap of different areas enables patients to focus on lesion marks during exercise. As the patient's posture changes during exercise, the display effects of different torso areas of the patient in the image information also change accordingly, increasing the interactivity and intuitiveness of the program with the patient. At the same time, this intuitive interaction can give patients movement feedback that fits themselves and the lesion marks, increasing their experience of using the program and thus enhancing the physical experience of doctor-patient communication.

[0015] Furthermore, the delivery module is also used to enlarge and display the image information, and when enlarging and displaying the image information, enlarge the lesion model according to the enlargement ratio of the image information, and then determine the enlarged second position according to the enlargement ratio and the first position, and display the lesion mark in the image information according to the enlarged second position.

[0016] By magnifying the image information, patients can see the details of the lesion more clearly, especially for some smaller lesions or complex lesions. The magnification function can help patients better understand their condition. While magnifying the display, the lesion model is magnified according to the magnification ratio, and the second position after magnification is determined according to the magnification ratio and the first position, ensuring that the lesion mark is always accurately superimposed on the image information, providing more accurate reference information for medical staff and patients. This method of adjusting the display effect of lesion marks in real time allows medical staff to instantly adjust the explanation content based on patient feedback and improve communication efficiency.

[0017] Furthermore, the shooting posture of the depth perception device is fixed to the medical environment, and the delivery module includes a delivery device for displaying image information, and the delivery device is a first terminal carried by the patient. The delivery module sends the image information, the second position and the lesion mark to the first terminal, and the first terminal receives the image information, the second position and the lesion mark and displays the lesion mark in the image information according to the second position. When displaying the image information, the first terminal collects sound information in the environment and performs screen recording to obtain a first video stream. The sound information and the first video stream are combined to obtain a medical consultation playback video, and the medical consultation playback video is stored.

[0018] Furthermore, the delivery module is also used to send the patient's three-dimensional body model, the first position and the lesion mark to the first terminal. After receiving the patient's three-dimensional body model, the first position and the lesion mark, the first terminal superimposes the lesion mark on the patient's three-dimensional body model according to the first position to form a first model, and stores the first model.

[0019] The first terminal collects sound information in the environment, and performs screen recording to obtain the first video stream. The sound information and the first video stream are combined to obtain a medical consultation playback video, and the medical consultation playback video is stored, so that the patient can review the entire communication process after the consultation and better understand and remember the doctor's advice. After the first terminal receives the patient's three-dimensional human body model, the first position and the lesion mark, it superimposes the lesion mark on the patient's three-dimensional human body model according to the first position to form a first model, and stores the first model, so that the patient can view the lesion mark and the three-dimensional human body model at any time after the consultation, thereby enhancing the patient's self-management ability. Separating the depth perception device from the delivery device (the patient's first terminal) not only reduces the hospital's delivery cost, but also increases the patient's sense of participation. It also helps patients retain medical consultation data, making it convenient for patients and their families to understand and review the patient's condition more accurately, and further improves the medical services provided to patients by this solution during the doctor-patient communication process.

[0020] Furthermore, the delivery module also includes a second delivery device, which is fixed to the treatment environment. The delivery module is used to display the lesion mark in the image information according to the second position, and then vertically flip the image information and display it on the second device.

[0021] The image information after vertical flipping is more in line with the patient's "mirror" perspective, making it easier for patients to immerse themselves in understanding their condition; at the same time, vertical flipping provides doctors with another perspective to observe lesion markings, helping doctors to understand the condition more comprehensively.

[0022] Furthermore, the acquisition module is preset with a number of key frames. The acquisition module selects the image with the latest acquisition time from the continuously acquired image information as the key image according to the number of key frames, obtains the movement trajectory of the first feature point in the key image in the environmental coordinate system, obtains the patient's movement speed according to the length of the movement trajectory and the acquisition frequency of the image information, and synchronously regulates the acquisition frequency according to the change of the movement speed; the acquisition module is provided with a movement speed threshold. When the movement speed exceeds the movement speed threshold, FAST or ORB is used to extract the patient's key feature points, and Tiny-YOLO or MobileNet is used to obtain various torso areas of the patient.

[0023] By presetting the number of key frames, the acquisition module can filter out the image with the latest acquisition time from the continuously acquired image information as the key image, ensuring that the solution can capture the patient's latest dynamics in a timely manner, especially when the patient moves quickly, reducing information lag caused by image delays. The acquisition frequency is dynamically adjusted according to the patient's movement speed to ensure that the acquisition frequency of key frames is increased when the patient moves quickly, so as to capture the patient's dynamic changes more timely. At the same time, such a method can accurately evaluate the patient's movement speed and dynamically adjust the acquisition frequency accordingly, reducing the processing of redundant data, avoiding system freezes caused by processing too much unnecessary data, and enabling the solution to process image information more efficiently, thereby improving the smoothness of image display and reducing image tearing or freezes caused by rapid patient movement.

[0024] The FAST algorithm is a fast corner detection algorithm that can detect key points in an image in a short time; the ORB algorithm combines the FAST key point detector and the BRIEF descriptor, which not only has a fast detection speed, but also has rotation invariance, making it suitable for real-time applications. When the patient moves quickly, these efficient algorithms can quickly extract key feature points and reduce the loss or misdetection of feature points caused by the patient's rapid movement. Tiny-YOLO is a lightweight target detection algorithm that can significantly reduce the amount of calculation while maintaining high detection accuracy; MobileNet further optimizes the computational efficiency of the model through technologies such as deep separable convolution. When the patient moves quickly, these lightweight algorithms can quickly and accurately identify the torso area, ensuring the real-time and accuracy of the system in high-dynamic scenarios. This method not only improves the efficiency and accuracy of feature point extraction, but also enhances the robustness of the system, improves user experience, optimizes system performance, and improves the efficiency of doctor-patient communication. These effects work together to make this solution more efficient and reliable in doctor-patient communication.

[0025] Furthermore, when extracting key feature points, the processing module performs confidence evaluation on the key feature points, and screens and divides the key feature points according to the human body joint areas; classifies and sorts the divided key feature points according to the confidence, and selects the key feature point with the highest confidence in each category as the first feature point.

[0026] By evaluating the confidence of each key feature point, feature points with higher confidence can be screened out, thereby improving the accuracy of feature point extraction. This solution can effectively filter out erroneous feature points caused by noise, occlusion or misdetection, and further improve the reliability of feature points and reduce the possibility of misjudgment. By screening and dividing key feature points, this solution can reduce the number of feature points to be processed, thereby reducing the computational burden and improving the operating efficiency of the system. This approach can also reduce the response time of the system and improve overall performance.

[0027] Furthermore, when the processing module matches the environment coordinate system and the shooting coordinate system, the position of the first feature point is represented by homogeneous coordinates, and a 4×4 transformation matrix is ​​used to realize the transformation from the environment coordinate system to the shooting coordinate system, as shown in formula (1). Formula (1) is as follows:

[0028] (1),

[0029] Where R is the rotation matrix, t is the translation vector, The matrix consists of a 3x3 rotation matrix R and a 3x1 translation vector t, which together form a homogeneous coordinate transformation matrix for coordinate transformation in three-dimensional space. The last row [0, 0, 0, 1] is to maintain the consistency of homogeneous coordinates so that the matrix can handle rotation and translation operations at the same time;

[0030] Alternatively, the optimal rotation and translation parameters can be solved by minimizing the position difference of the first feature point in the two coordinate systems. The function is shown in formula (2). Formula (2) is as follows:

[0031] (2),

[0032] in, is the feature point in the environment coordinate system, is the corresponding first feature point in the shooting coordinate system;

[0033] The processing module enhances the discrimination of the first feature point by calculating the color histogram or color gradient of the pixels around the first feature point during the matching process. The discrimination is calculated as shown in formula (3):

[0034] (3),

[0035] in, and They are the color information around the two first feature points.

[0036] Using homogeneous coordinates and a 4×4 transformation matrix can unify translation and rotation operations into one matrix, simplifying the calculation process of coordinate transformation and improving calculation efficiency. Moreover, through homogeneous coordinate transformation, the feature points in the shooting coordinate system can be accurately converted to the environment coordinate system to ensure that the positional relationship of the feature points in the two coordinate systems is accurate, which is crucial for subsequent feature point matching and model superposition. Furthermore, the homogeneous transformation matrix not only supports translation and rotation operations, but can also be extended to other geometric transformations such as scaling and projection, providing greater flexibility for this solution and being able to cope with different scenarios and needs.

[0037] By minimizing the position difference of the feature points in the two coordinate systems, the optimal rotation and translation parameters can be solved. This method can ensure that the first feature point is as close to its true position as possible in the converted coordinate system, thereby improving the accuracy of coordinate matching. The least squares optimization method can effectively reduce the matching error caused by noise or measurement error. Through iterative optimization, this scheme can gradually adjust the transformation parameters to minimize the position difference of the first feature point in the two coordinate systems, thereby improving the accuracy and reliability of the overall matching. In practical applications, the position of the first feature point may be affected by many factors, such as illumination changes, occlusion, etc. The method of minimizing the position difference can adapt to these complex scenarios and compensate for these influences by optimizing the transformation parameters, thereby improving the robustness of the system.

[0038] By calculating the color histogram or color gradient of the pixels around the feature points, the distinguishability of the feature points can be enhanced. This method can make the feature points more prominent in the image, thereby improving the recognition ability and matching accuracy of the feature points. At the same time, the introduction of color information can effectively reduce the mismatch of feature points caused by lighting changes or background interference. By comparing color information, this solution can more accurately identify and match the first feature point, thereby improving the anti-interference ability of this solution. As part of the visual feature, color information can be fused with data from other modalities (such as depth information, texture information, etc.). This multimodal data fusion method can further improve the distinguishability of feature points and provide richer information for subsequent image processing and analysis.

[0039] Further, the delivery module determines the display position of the lesion mark in the image information according to the second position calculated by the processing module, and then according to the first position of the lesion mark, the delivery module obtains the color information of the pixel points adjacent to the lesion mark in the image information as , the delivery module displays the difference according to the preset color ,Will The value of is used as the color of the lesion mark. The value of is displayed in the image information; the delivery module captures the patient's facial expression image in real time through a depth perception device or camera, and uses a pre-trained deep learning-based expression recognition model to identify whether the patient has squinting, frowning, or leaning in. The expression recognition model outputs the expression recognition result, including the expression type and its confidence; the delivery module also has a difference adjustment unit When the recognition model detects that the patient is squinting, frowning, or leaning in, The value of is used as the color of the lesion mark, and the new lesion mark color is applied to the image information, and then the patient's facial expression is continuously monitored to capture the expression changes in real time; if the patient squints, frowns, or moves closer again, the color of the lesion mark is adjusted again according to The color of the lesion mark is calculated and the display of the lesion mark is updated until the patient no longer appears to be squinting, frowning, or leaning in.

[0040] By obtaining the color information of the adjacent pixels of the lesion mark and setting the color of the lesion mark according to the preset color display difference, it is ensured that the lesion mark has sufficient contrast with the background color. This color adjustment method based on actual image information can effectively avoid confusion between the lesion mark and the patient's body color or clothing color, making the lesion mark more clearly visible. At the same time, the color histogram or color gradient of the pixels around the first feature point is calculated for local features, and there is no need to perform large-scale complex operations on the entire image. This localized calculation strategy helps to reduce the resource usage of this solution in the image processing process, improve the system operation efficiency, and save system resources. This method of highlighting lesion marks can also distinguish the lesion features in the image information from other irrelevant information, making it easier for patients to identify the key parts and features related to their condition, and will not be misled by other interfering information, which helps them understand their condition and pathology more clearly, thereby improving the effectiveness and efficiency of doctor-patient communication. At the same time, the mechanism of dynamically adjusting the color of lesion markings enables patients to feel that this program pays attention to their visual needs and enhances their sense of participation. Medical staff can also promptly discover where patients need assistance from changes in contrast and provide patients with more timely help or services, further improving patient satisfaction and making patients more willing to accept explanations and suggestions from medical staff, thereby improving the communication efficiency between medical staff and patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a schematic diagram of the structure of an interactive system based on augmented reality to enhance the doctor-patient communication experience.

[0042] Figure 2 The flowchart for matching the environment coordinate system and the shooting coordinate system for the processing module.

[0043] Figure 3 This is a flowchart for the delivery module to adjust the display color of the lesion mark according to the changes in the patient's expression. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the concept and technical effects of the present invention in combination with the embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative work are all within the scope of protection of the present invention:

[0045] Example 1

[0046] like Figure 1 As shown, the interactive system based on augmented reality to improve the doctor-patient communication experience includes:

[0047] The acquisition module includes a depth sensing device, which is used to acquire image information of the patient in the treatment environment, including a number of reference objects whose positions are fixed to the treatment environment; and is used to acquire the patient's imaging data and the corresponding diagnostic data;

[0048] A processing module is used to construct a three-dimensional model of the medical environment as an environmental model based on the position of the reference object in the medical environment and the three-dimensional models of each reference object, and to establish an environmental coordinate system based on the environmental model. The processing module also includes a three-dimensional model of the human body; Figure 2 As shown, the processing module segments the point cloud data of the reference object and the patient in the image information according to the reference object and the three-dimensional model of the human body, then extracts the key feature points from the point cloud data of the patient as the first feature points, matches the three-dimensional model of each reference object with the point cloud data of the reference object, and obtains the key feature points of each reference object as the second feature points; the processing module establishes a shooting coordinate system according to the image information and a preset fixed position, obtains the positions of the first feature point and the second feature point in the shooting coordinate system, then obtains the position of the second feature point in the environment coordinate system according to the environment model, matches the environment coordinate system and the shooting coordinate system according to the positions of the second feature point in the environment coordinate system and the shooting coordinate system, and calculates the position of the first feature point in the environment coordinate system;

[0049] The delivery module includes a lesion recognition model constructed by a deep learning model. The delivery module puts the patient's imaging data into the lesion recognition model for processing to obtain a lesion mark, then compares the lesion mark with the diagnostic data corresponding to the imaging data, modifies the lesion mark with inconsistent comparison results according to the diagnostic data (diagnosis text input by the doctor), and uses the modified lesion mark and imaging data to train the lesion recognition model; corresponds the position of the imaging data in the patient's three-dimensional model of the human body to the position of the lesion mark in the imaging data, determines the position of the lesion mark in the patient's three-dimensional model of the human body, and then determines the relative position of the lesion mark and the first feature point as the first position according to the patient's point cloud data and the patient's three-dimensional model of the human body; the delivery module is used to obtain the position of the first feature point in the image data in the environmental coordinate system, then determines the position of the lesion mark in the environmental coordinate system according to the first position, and determines the position of the lesion mark in the shooting coordinate system as the second position according to the matching result of the environmental coordinate system and the shooting coordinate system; the delivery module is used to display image information, and display the lesion mark in the image information according to the second position.

[0050] Among them, the processing module is also used to obtain the patient's physiological characteristic information, combine the physiological characteristic information with the imaging data to predict the patient's body shape, combine the patient's body shape with the patient's point cloud data, divide the patient's point cloud data into clothing point cloud and body point cloud, and combine the body point cloud and the human body three-dimensional model to divide the first feature point according to the human trunk structure, divide the patient's three-dimensional model into different trunk areas with the first feature point as the node, and then determine the trunk area as the target area according to the first position of the lesion mark.

[0051] The processing module uses the torso area corresponding to the patient's hand as the operation area. When the delivery module displays the lesion mark, the pixel points of the lesion mark are used to replace the pixel points corresponding to the target area in the image information; when the target area overlaps with other torso areas in the shooting coordinate system, the transparency of the torso area is increased to highlight the display effect of the target area, and then the effect of superimposing the torso area and the target area is displayed in the image information; when the target area overlaps with the operation area in the shooting coordinate system, the edge contour of the operation area is highlighted.

[0052] Among them, the processing module extracts features of image information through a pre-trained CNN (convolutional neural network) model (such as ResNet, InceptionV3 or Vision Transformer, where: ResNet is a residual network; InceptionV3 is a convolutional neural network architecture that can more efficiently extract image features by using parallel calculations of convolution kernels of different sizes; Vision Transformer's Chinese name is "visual converter", which is a deep learning model based on the Transformer architecture) and converts the image information into a high-dimensional feature vector, which is then input into a language generation model (such as RNN, LSTM, Transformer or an encoder-decoder model based on Transformer, where: RNN is a recurrent neural network; LSTM is a long short-term memory network; Transformer is a Transformer architecture) to generate a detailed text description of the image information, and then uses a multimodal model (such as CLIP, MiniGPT-4, where: CLIP is a comparative language-image pre-training model; MiniGPT-4 is a multimodal model) to directly convert the image into a text description. Through comparative learning, these models can understand the semantic relationship between image information and diagnostic data and generate more accurate descriptions. They then perform word segmentation, part-of-speech tagging, and syntactic analysis on the generated text descriptions to extract key information, such as object names, actions, scenes, etc. NLP (natural language processing) models (such as BERT and GPT, where BERT is a bidirectional encoder representation model based on Transformer; GPT is a generative pre-trained Transformer, a series of large-scale language models based on the Transformer architecture) are used to perform semantic analysis on the text descriptions, understand the contextual relationships in the descriptions, judge the accuracy and completeness of the descriptions, and finally add corresponding tags to the images based on the key information in the text descriptions.

[0053] Among them, the delivery module is also used to enlarge and display the image information, and when enlarging and displaying the image information, the lesion model is enlarged according to the enlargement ratio of the image information, and then the enlarged second position is determined according to the enlargement ratio and the first position, and the lesion mark is displayed in the image information according to the enlarged second position.

[0054] The acquisition module has a preset number of key frames (the default is 10 frames, which can be set by the administrator). The acquisition module selects the image with the latest acquisition time from the continuously acquired image information as the key image based on the number of key frames, obtains the movement trajectory of the first feature point in the key image in the environmental coordinate system, and processes the patient's movement speed according to the length of the movement trajectory and the acquisition frequency of the image information (the default is 10FPS, which can be set by the administrator, where FPS is the abbreviation of "Frames Per Second", which means "frames per second" in Chinese) and adjusts the acquisition frequency synchronously according to the change of the movement speed. In general, the acquisition frequency increases by 5FPS for every 0.3 m / s increase in the movement speed, and decreases by 5FPS for every 0.3 m / s decrease in the movement speed. When the movement speed is 0 m / s, the acquisition frequency is fixed at 10FPS. The administrator or medical staff can also set the way to adjust the acquisition frequency: when the movement speed increases or decreases, the acquisition frequency increases or decreases year-on-year according to the proportion of the increase or decrease in the movement speed.

[0055] The acquisition module is equipped with a movement speed threshold (generally 0.5 m / s. Medical staff can reduce the movement speed threshold to 0.3 m / s or increase it to 0.8 m / s according to actual usage; administrators can change the movement speed threshold according to computing resources). When the movement speed exceeds the movement speed threshold, FAST or ORB is used to extract the patient's key feature points, and Tiny-YOLO or MobileNet is used to obtain the patient's various torso areas.

[0056] Among them, when extracting key feature points, the processing module evaluates the confidence of the key feature points, and screens and divides the key feature points according to the human body joint areas; the divided key feature points are classified and sorted according to the confidence, and the key feature point with the highest confidence in each category is selected as the first feature point.

[0057] When implementing:

[0058] In a hospital clinic, a patient with lung disease comes to see a doctor. The administrator arranges several reference objects with fixed positions in the clinic, such as signboards at specific locations and fixtures at corners.

[0059] The depth perception device is turned on, and the image information of the patient in the treatment environment is continuously collected at a frequency of 10FPS (default acquisition frequency). These images cover the patient's activities in the clinic and the surrounding scenes with reference objects. At the same time, the patient's lung imaging data (such as X-rays, CT scan data, etc.) and its corresponding diagnostic data (diagnosis text entered by the doctor, such as "shadows found in the left lobe of the lungs, initially diagnosed as pneumonia") are collected.

[0060] The number of key frames preset in the acquisition module is 10 frames. The acquisition module selects the image with the latest acquisition time from the continuously acquired image information as the key image. Assume that after processing, the length of the moving trajectory of the first feature point in the key image in the environmental coordinate system is 3 meters. Since the acquisition frequency defaults to 10FPS, the patient's movement speed is calculated according to the formula to be 3 meters / second (here it is assumed that the key image interval is 10 frames, that is, 1 second). Because the acquisition frequency increases by 5FPS for every 0.3 m / s increase in movement speed, the acquisition frequency is adjusted to 60FPS after 1 second.

[0061] The processing module constructs a three-dimensional model of the medical environment as an environmental model according to the position of the reference object in the medical environment and the three-dimensional models of each reference object, and establishes an environmental coordinate system. The processing module already includes a three-dimensional model of the human body.

[0062] The processing module segments the point cloud data of the reference object and the patient in the image information according to the reference object and the three-dimensional model of the human body. Key feature points are extracted from the point cloud data of the patient as the first feature points, such as points of joints such as the head, shoulders, and elbows. The three-dimensional model of each reference object is matched with the point cloud data of the reference object to obtain the key feature points of each reference object as the second feature points, such as the corner points of the signboard, the specific end points of the corner device, etc.

[0063] The processing module establishes a shooting coordinate system according to the image information and the preset fixed position, obtains the positions of the first feature point and the second feature point in the shooting coordinate system, obtains the position of the second feature point in the environment coordinate system according to the environment model, matches the environment coordinate system and the shooting coordinate system through these position information, and calculates the position of the first feature point in the environment coordinate system.

[0064] The processing module obtains the patient's physiological characteristics (such as height, weight, etc.), and combines it with the imaging data to predict the patient's body shape. The patient's body shape is combined with the patient's point cloud data, and the patient's point cloud data is divided into clothing point cloud and body point cloud. Combined with the body point cloud and the human body 3D model, the first feature point is divided according to the human body trunk structure (such as chest cavity, abdominal cavity, etc.), and the patient's human body 3D model is divided into different trunk areas, such as chest area, abdominal area, etc., with the first feature point as the node. Then, the trunk area is determined as the target area based on the first position of the lesion mark. Assuming that the lesion mark is in the lung position, the chest area is determined as the target area.

[0065] The processing module extracts features from image information through a pre-trained ResNet model and converts the image information into a high-dimensional feature vector. The high-dimensional feature vector is then input into the Transformer-based encoder-decoder model to generate a detailed text description of the image information, such as "the patient stands in the middle of the consulting room, facing the doctor, with his left hand hanging naturally and his right hand slightly raised". The CLIP model is then used to directly associate and match the image information with the diagnostic data, and the semantic relationship between the two is understood through comparative learning to generate a more accurate description. The generated text description is processed by word segmentation, part-of-speech tagging, and syntactic analysis to extract key information, such as "patient", "station", "consulting room", etc. The BERT model is used to perform semantic analysis on the text description, understand the contextual relationship, and judge the accuracy and completeness of the description. Finally, according to the key information in the text description, corresponding tags are added to the image, such as adding an indicator tag to the position where the patient is standing.

[0066] The delivery module puts the patient's imaging data into the lesion recognition model built by the deep learning model to obtain lesion marks, such as marking the shadow position on the lung imaging data. The lesion marks are then compared with the diagnostic data corresponding to the imaging data. The lesion marks with inconsistent comparison results are modified according to the diagnostic data ("A shadow was found in the left lobe of the lung, and it is initially diagnosed as pneumonia"), and the modified lesion marks and imaging data are used to train the lesion recognition model to improve the accuracy of the model.

[0067] The delivery module matches the position of the imaging data in the patient's three-dimensional model with the position of the lesion mark in the imaging data, determines the position of the lesion mark in the patient's three-dimensional model, and then determines the relative position of the lesion mark and the first feature point as the first position based on the patient's point cloud data and the patient's three-dimensional model. The position of the first feature point in the image data in the environmental coordinate system is obtained, and then the position of the lesion mark in the environmental coordinate system is determined based on the first position, and the position of the lesion mark in the shooting coordinate system is determined as the second position based on the matching result of the environmental coordinate system and the shooting coordinate system. The image information is displayed, and the lesion mark is displayed in the image information according to the second position, such as accurately displaying the mark of the lung shadow on the image of the patient's chest.

[0068] When the target area (chest area) in the shooting coordinate system overlaps with other torso areas, the delivery module increases the transparency of other torso areas to highlight the display effect of the target area, and then displays the superimposed effect of the torso area and the target area in the image information, allowing doctors and patients to see the lesion marks in the target area more clearly.

[0069] The processing module uses the torso area corresponding to the patient's hand as the operation area. When the target area (chest area) in the shooting coordinate system overlaps with the operation area, the delivery module highlights the edge contour of the operation area to facilitate the doctor to understand the relationship between the operation area and the target area.

[0070] The delivery module also supports magnified display of image information. Assuming the magnification ratio is 2, when the image information is magnified, the delivery module magnifies the lesion model according to the magnification ratio of the image information, and then determines the second position after magnification according to the magnification ratio and the first position, and displays the lesion mark in the magnified image information according to the second position after magnification, ensuring that the lesion mark is accurately positioned and clearly visible in the magnified image.

[0071] The acquisition module has a moving speed threshold, which is currently set to 0.5 m / s. Since the patient's moving speed is calculated to be 3 m / s, it exceeds the moving speed threshold. At this time, the FAST algorithm is used to extract the patient's key feature points, and Tiny-YOLO is used to obtain the patient's various torso areas to track the patient's movements and regional changes more quickly and accurately to meet the real-time requirements of the system.

[0072] Example 2

[0073] The only difference between this embodiment and Embodiment 1 is that the shooting posture of the depth perception device is fixed to the medical environment, the delivery module includes a delivery device for displaying image information, the delivery device is a first terminal carried by the patient, the delivery module sends the image information, the second position and the lesion mark to the first terminal, the first terminal receives the image information, the second position and the lesion mark and then displays the lesion mark in the image information according to the second position, the first terminal collects sound information in the environment when displaying the image information, and simultaneously performs screen recording to obtain a first video stream, combines the sound information and the first video stream to obtain a medical consultation playback video, and stores the medical consultation playback video.

[0074] Among them, the delivery module also includes a second delivery device, the second delivery device is fixed to the treatment environment, and the delivery module is used to display the lesion mark in the image information according to the second position, and then vertically flip the image information and display it on the second device.

[0075] The delivery module is also used to send the patient's three-dimensional body model, the first position and the lesion mark to the first terminal. After receiving the patient's three-dimensional body model, the first position and the lesion mark, the first terminal superimposes the lesion mark on the patient's three-dimensional body model according to the first position to form a first model, and stores the first model.

[0076] When implementing:

[0077] A 55-year-old male patient, Mr. Zhang, came to the hospital for treatment due to long-term cough and chest pain. The hospital adopted an interactive system based on augmented reality to enhance the doctor-patient communication experience, aiming to allow patients to understand their own conditions more intuitively and facilitate communication between doctors and patients.

[0078] The depth sensing device is installed at a fixed position in the consulting room, and its shooting posture is fixed to the consulting environment. During Mr. Zhang's consultation, the depth sensing device continuously collects image information of Mr. Zhang in the consulting environment, and at the same time collects Mr. Zhang's lung imaging data (such as high-resolution CT images) and its corresponding diagnostic data (the doctor initially diagnosed it as a lung nodule, the nature of which is yet to be determined). The processing module constructs a three-dimensional model and environmental coordinate system of the consulting environment based on the reference objects in fixed positions in the consulting room, processes the collected image information, segments the point cloud data of Mr. Zhang and the reference objects, extracts key feature points, and completes a series of operations such as matching the environmental coordinate system with the shooting coordinate system, and determines the second position of the lesion mark in the shooting coordinate system.

[0079] The delivery module sends the image information, the second position and the lesion mark to the first terminal (such as smart glasses, mobile phone or tablet) carried by Lao Zhang. After Lao Zhang puts on the smart glasses, he can see the image information of his medical environment in real time, and the lesion mark of the lungs is accurately displayed in the image according to the second position, allowing him to intuitively see the parts of his lungs where there may be problems. While the doctor is explaining the condition to Lao Zhang, the first terminal displays the image information while collecting the sound information in the environment, and simultaneously performs screen recording to obtain the first video stream. The first terminal combines the sound information and the first video stream to obtain a medical consultation playback video. After Lao Zhang finishes the medical consultation, the medical consultation playback video is stored in the first terminal, which is convenient for him to review the doctor's explanation later and deepen his understanding of the condition.

[0080] In addition, the delivery module also sends the patient's three-dimensional human body model, the first position and the lesion mark to the first terminal. After receiving it, Mr. Zhang's first terminal superimposes the lesion mark on the patient's three-dimensional human body model according to the first position to form a first model. Mr. Zhang can view this three-dimensional human body model with lesion marks from different angles by operating the first terminal, and have a more comprehensive understanding of the location and morphology of the lung nodules in his body. This first model is also stored in the first terminal for Mr. Zhang to view at any time.

[0081] The delivery module also includes a second delivery device fixed to the treatment environment (such as a large display screen on the wall of the clinic). The delivery module displays the lesion mark in the image information according to the second position, and then vertically flips the image information with the lesion mark and displays it on the second delivery device. In this way, when the doctor communicates with Mr. Zhang, he can view Mr. Zhang's condition information from different perspectives through the second delivery device, which is convenient for the doctor to compare and analyze when explaining the condition, and also convenient for other medical staff to intuitively understand the patient's condition when necessary.

[0082] Example 3

[0083] The difference between this embodiment and Embodiment 1 and Embodiment 2 is that when the processing module matches the environment coordinate system and the shooting coordinate system, the position of the first feature point is represented by homogeneous coordinates, and a 4×4 transformation matrix is ​​used to realize the conversion from the environment coordinate system to the shooting coordinate system, as shown in Formula (1): The matrix consists of a 3x3 rotation matrix R and a 3x1 translation vector t, which together form a homogeneous coordinate transformation matrix for coordinate transformation in three-dimensional space. The last row [0, 0, 0, 1] is to maintain the consistency of homogeneous coordinates so that the matrix can handle rotation and translation operations at the same time. Formula (1) converts the environment coordinate system into Convert to a point in the shooting coordinate system ; Formula (1) is as follows:

[0084] (1),

[0085] Where R is the rotation matrix and t is the translation vector;

[0086] Alternatively, the optimal rotation and translation parameters can be solved by minimizing the position difference of the first feature point in the two coordinate systems. The function is shown in formula (2). Formula (2) is as follows:

[0087] (2),

[0088] in, is the feature point in the environment coordinate system, is the corresponding first feature point in the shooting coordinate system;

[0089] The processing module enhances the discrimination of the first feature point by calculating the color histogram or color gradient of the pixels around the first feature point during the matching process. The discrimination is calculated as shown in formula (3):

[0090] (3),

[0091] in, and They are the color information around the two first feature points.

[0092] The delivery module determines the display position of the lesion mark in the image information according to the second position calculated by the processing module, and then determines the display position of the lesion mark according to the first position of the lesion mark. Figure 3 As shown, the delivery module obtains the color information of the pixel points adjacent to the lesion mark in the image information as , the delivery module displays the difference according to the preset color ,Will The value of is used as the color of the lesion mark. The value of is displayed in the image information; the delivery module captures the patient's facial expression image in real time through a depth perception device or camera, and uses a pre-trained deep learning-based expression recognition model to identify whether the patient has squinting, frowning, or leaning in. The expression recognition model outputs the expression recognition result, including the expression type and its confidence; the delivery module also has a difference adjustment unit When the recognition model detects that the patient is squinting, frowning, or leaning in, the color of the current lesion mark is increased. (The current lesion mark color is The new lesion marking color is applied to the image information, and the patient's facial expression is continuously monitored to capture expression changes in real time; if the patient squints, frowns, or moves closer again, the new lesion marking color is applied to the image information. The color of the lesion mark is calculated and the display of the lesion mark is updated until the patient no longer appears to be squinting, frowning, or leaning in.

[0093] When implementing:

[0094] A 60-year-old patient came to see a doctor for suspected cerebral vascular disease.

[0095] After constructing the three-dimensional model of the treatment environment and establishing the environmental coordinate system, the processing module begins to match the environmental coordinate system with the shooting coordinate system. For the first feature point extracted from the patient's point cloud data, its position is represented by homogeneous coordinates, and the transformation from the shooting coordinate system to the environmental coordinate system is achieved with the help of a 4×4 transformation matrix (Formula (1)). During the calculation process, the specific parameters of the rotation matrix R and the translation vector t are determined by analyzing the correspondence between the reference object and the patient's point cloud data. For example, according to the position difference of the sign on the wall in the two coordinate systems, R and t are accurately adjusted to make the coordinate system conversion more accurate. In addition, the optimal rotation and translation parameters can be solved by minimizing the position difference of the first feature point in the two coordinate systems (Formula (2)). The processing module calculates and optimizes multiple first feature points to improve the accuracy of the overall conversion.

[0096] In the matching process, in order to enhance the distinguishability of the first feature point, the processing module calculates the color histogram and color gradient of the pixels around the first feature point. The difference value of the color information around different first feature points is calculated by formula (3): For example, the color differences around different feature points on the patient's head contour are calculated to more clearly distinguish each feature point, providing a guarantee for subsequent accurate processing.

[0097] The delivery module accurately determines the display position of the lesion mark in the image information based on the second position calculated by the processing module. Then, according to the first position of the lesion mark, the color information of the pixel points adjacent to the lesion mark in the image information is obtained as . Assume the preset color shows the difference is 8 (can be adjusted according to the actual display effect), then The value of is used as the color of the lesion mark, and the lesion mark is displayed in the image information with this color, so that the patient can intuitively see the location of the stenosis of the brain artery.

[0098] The delivery module uses a depth sensing device to capture the patient's facial expression image in real time. When the doctor shows the patient an image with lesion marks and explains the condition, if the expression recognition model recognizes that the patient has squinting, frowning, or leaning closer (for example, the patient leans closer to the screen because he cannot see the lesion mark clearly), and the confidence of the output expression recognition result is high (for example, more than 0.7), the difference adjustment unit Start working. The value of is used as the new lesion marking color and applied to the image information. After that, the patient's expression is continuously monitored. If the patient shows such an expression again, the color of the lesion is marked again according to the Calculate and update the color of the lesion mark until the patient no longer has these expressions, ensuring that the patient can clearly observe the lesion mark and improving communication effectiveness.

[0099] The above is only an embodiment of the present invention, and the common knowledge such as the known specific structure and characteristics in the scheme is not described in detail here. It should be pointed out that for those skilled in the art, several deformations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. An interactive system based on augmented reality to enhance the doctor-patient communication experience, characterized by: include: The acquisition module includes a depth perception device, wherein the depth perception device is used to acquire image information of the patient in the treatment environment, wherein the treatment environment includes a plurality of reference objects whose positions are fixed to the treatment environment; and is used to acquire the patient's imaging data and the corresponding diagnostic data; A processing module, used to construct a three-dimensional model of the medical environment as an environmental model based on the posture of the reference object in the medical environment and the three-dimensional models of each reference object, and establish an environmental coordinate system based on the environmental model, wherein the processing module also includes a three-dimensional model of a human body; the processing module segments the point cloud data of the reference object and the patient in the image information based on the reference object and the three-dimensional model of the human body, and then extracts key feature points from the point cloud data of the patient as first feature points, matches the three-dimensional model of each reference object with the point cloud data of the reference object, and obtains the key feature points of each reference object as second feature points; the processing module establishes a shooting coordinate system based on the image information and a preset fixed position, obtains the positions of the first feature point and the second feature point in the shooting coordinate system, and then obtains the position of the second feature point in the environmental coordinate system based on the environmental model, matches the environmental coordinate system and the shooting coordinate system based on the positions of the second feature point in the environmental coordinate system and the shooting coordinate system, and calculates the position of the first feature point in the environmental coordinate system; A delivery module includes a lesion recognition model constructed by a deep learning model, wherein the delivery module puts the patient's imaging data into the lesion recognition model for processing to obtain a lesion mark, then compares the lesion mark with the diagnostic data corresponding to the imaging data, modifies the lesion mark with inconsistent comparison results according to the diagnostic data, and uses the modified lesion mark and the imaging data to train the lesion recognition model; corresponds the position of the imaging data in the patient's three-dimensional model to the position of the lesion mark in the imaging data, determines the position of the lesion mark in the patient's three-dimensional model, and then determines the relative position of the lesion mark and the first feature point as the first position according to the patient's point cloud data and the patient's three-dimensional model; The delivery module is used to obtain the position of the first feature point in the image data in the environmental coordinate system, determine the position of the lesion mark in the environmental coordinate system according to the first position, and determine the position of the lesion mark in the shooting coordinate system as the second position according to the matching result of the environmental coordinate system and the shooting coordinate system; The delivery module is used to display the image information and display the lesion mark in the image information according to the second position.

2. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 1, characterized in that: The processing module is also used to obtain physiological characteristic information of the patient, combine the physiological characteristic information with the imaging data to predict the patient's body shape, combine the patient's body shape with the patient's point cloud data, divide the patient's point cloud data into clothing point cloud and body point cloud, and combine the body point cloud and the human body three-dimensional model to divide the first feature point according to the human body trunk structure, divide the patient's three-dimensional human body model into different trunk areas with the first feature point as a node, and then determine the trunk area as the target area according to the first position of the lesion mark; The processing module uses the trunk area corresponding to the patient's hand as the operation area. When the delivery module displays the lesion mark, the pixel points corresponding to the target area are replaced with the pixel points of the lesion mark in the image information; when the target area overlaps with other trunk areas in the shooting coordinate system, the transparency of the trunk area is increased to highlight the display effect of the target area, and then the effect of superimposing the trunk area and the target area is displayed in the image information; when the target area overlaps with the operation area in the shooting coordinate system, the edge contour of the operation area is highlighted.

3. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 2, characterized in that: The delivery module is also used to amplify and display the image information, and when amplifying and displaying the image information, amplify the lesion model according to the amplification ratio of the image information, and then determine the enlarged second position according to the amplification ratio and the first position, and display the lesion mark in the image information according to the enlarged second position.

4. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 1, characterized in that: The shooting posture of the depth perception device is fixed to the medical environment. The delivery module includes a delivery device for displaying image information. The delivery device is a first terminal carried by the patient. The delivery module sends the image information, the second position and the lesion mark to the first terminal. The first terminal receives the image information, the second position and the lesion mark and then displays the lesion mark in the image information according to the second position. When displaying the image information, the first terminal collects sound information in the environment and performs screen recording to obtain a first video stream. The sound information and the first video stream are combined to obtain a medical consultation playback video, and the medical consultation playback video is stored.

5. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 4, characterized in that: The delivery module is also used to send the patient's three-dimensional body model, the first position and the lesion mark to the first terminal. After receiving the patient's three-dimensional body model, the first position and the lesion mark, the first terminal superimposes the lesion mark on the patient's three-dimensional body model according to the first position to form a first model, and stores the first model.

6. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 5, characterized in that: The delivery module also includes a second delivery device, which is fixed to the treatment environment. The delivery module is used to display the lesion mark in the image information according to the second position, and then vertically flip the image information and display it on the second device.

7. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 1, characterized in that: The acquisition module is preset with a number of key frames. The acquisition module selects the image with the latest acquisition time from the continuously acquired image information as the key image according to the number of key frames, obtains the moving trajectory of the first feature point in the key image in the environmental coordinate system, obtains the moving speed of the patient according to the length of the moving trajectory and the acquisition frequency of the image information, and synchronously adjusts the acquisition frequency according to the change of the moving speed; A movement speed threshold is set in the acquisition module. When the movement speed exceeds the movement speed threshold, FAST or ORB is used to extract the key feature points of the patient, and Tiny-YOLO or MobileNet is used to obtain various torso areas of the patient.

8. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 1, characterized in that: When extracting key feature points, the processing module evaluates the confidence of the key feature points, and screens and divides the key feature points according to the human body joint areas; classifies and sorts the divided key feature points according to the confidence, and selects the key feature point with the highest confidence in each category as the first feature point.

9. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 1, characterized in that: When the processing module matches the environment coordinate system and the shooting coordinate system, the position of the first feature point is represented by homogeneous coordinates, and a 4×4 transformation matrix is ​​used to realize the transformation from the environment coordinate system to the shooting coordinate system, as shown in formula (1). Formula (1) is as follows: (1), Where R is the rotation matrix, t is the translation vector, The matrix consists of a 3x3 rotation matrix R and a 3x1 translation vector t; Alternatively, the optimal rotation and translation parameters can be solved by minimizing the position difference of the first feature point in the two coordinate systems. The function is shown in formula (2). Formula (2) is as follows: (2), in, is the feature point in the environment coordinate system, is the corresponding first feature point in the shooting coordinate system; The processing module enhances the discrimination of the first feature point by calculating the color histogram or color gradient of the pixels around the first feature point during the matching process. The discrimination is calculated as shown in formula (3): (3), in, and They are the color information around the two first feature points.

10. The interactive system for improving doctor-patient communication experience based on augmented reality according to claim 9, characterized in that: The delivery module determines the display position of the lesion mark in the image information according to the second position calculated by the processing module, and then obtains the color information of the pixel points adjacent to the lesion mark in the image information according to the first position of the lesion mark as , the delivery module displays the difference according to the preset color ,Will The value of is used as the color of the lesion mark. The value of is shown in the image information; The delivery module uses a depth sensing device or camera to capture the patient's facial expression image in real time, and uses a pre-trained deep learning-based expression recognition model to identify whether the patient has squinting, frowning, or leaning in. The expression recognition model outputs the expression recognition result, including the expression type and its confidence level. The delivery module also has a difference adjustment unit. When the recognition model detects that the patient is squinting, frowning, or leaning in, The value of is used as the color of the lesion mark, and the new lesion mark color is applied to the image information, and then the patient's facial expression is continuously monitored to capture the expression changes in real time; if the patient squints, frowns, or moves closer again, the color of the lesion mark is adjusted again according to The color of the lesion mark is calculated and the display of the lesion mark is updated until the patient no longer appears to be squinting, frowning, or leaning in.

Citation Information

Patent Citations

  • Method and system for realizing surgical navigation by using real-time structured light technology

    CN113229937A