Interaction system of endoscope and virtual reality equipment
Through a system combining endoscopy and virtual reality equipment, three-dimensional images are reconstructed using binocular cameras and deep learning algorithms, solving the misjudgment of traditional endoscopy and patient discomfort problems, and achieving efficient and accurate surgical operations and training.
Patent Information
- Application Number
- CN202510659277.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional 2D endoscopic images are difficult to meet the needs of modern precision medicine, especially in the high rate of misjudgment in the depth of the lesion, the lack of real-time three-dimensional navigation increases the difficulty of surgery, and the patient feels strong discomfort, which affects diagnosis and treatment cooperation.
Endoscope equipped with binocular cameras is used, a three-dimensional image model of the digestive tract is reconstructed by combining deep learning algorithms and optical flow method, and the image data is displayed and marked through virtual reality devices, supporting doctors to operate input and feedback, and using physiological parameters to adjust patient images.
Improve surgical accuracy and efficiency, reduce misjudgment, improve patient experience, improve training results, shorten surgical time and reduce anxiety.
Smart Images

Figure CN120570543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical devices, and in particular to an interaction system between an endoscope and a virtual reality device. Background Art
[0002] With the advancement of medical technology, endoscopy has become increasingly important as a crucial tool for diagnosing and treating digestive system diseases. However, traditional 2D endoscopic imaging has limitations and cannot meet the demands of modern precision medicine. When operating traditional endoscopes, physicians primarily rely on two-dimensional images, making it difficult to assess lesion depth and spatial relationships. This is particularly true for critical information such as the depth of invasion in early-stage cancers, with an error rate as high as 15%-20%. Furthermore, the lack of real-time 3D navigation during complex procedures such as ESD (endoscopic submucosal dissection) increases both operative time and operational difficulty. Furthermore, patients often experience discomfort or even pain during endoscopic examinations, particularly children and those with sensitive skin conditions. This leads to approximately 10%-15% of patients refusing repeat examinations. Furthermore, patients' limited understanding of the diagnostic and treatment process also hinders their cooperation with their physicians.
[0003] To address these challenges, existing industry solutions primarily focus on enhancing endoscopist training through VR technology. For example, the LAP Mentor simulator launched by 3D Systems has improved the training environment to some extent, but its application in actual surgeries remains very limited. A small number of existing studies have attempted to use VR in endoscopic images, but due to issues such as high latency and insufficient accuracy, it has not yet been widely adopted in clinical practice. Furthermore, while many companies have already made investments in endoscopic AI diagnostics, they have yet to fully integrate VR interaction technology. Summary of the Invention
[0004] An embodiment of the present invention provides an endoscope and virtual reality device interaction system, which can solve at least one of the technical problems mentioned above.
[0005] An embodiment of the present invention provides an endoscope and virtual reality device interaction system, including: An endoscope is configured with a binocular camera for collecting two-dimensional original image data in the digestive tract based on the binocular camera; a data processing device is used to process the two-dimensional original image data based on a deep learning algorithm and an optical flow method to determine a three-dimensional image model of the digestive tract, wherein the three-dimensional image model is marked with different tissue layers including a mucosal layer, a vascular network, and a muscle layer, and is marked with a lesion infiltration range; and is used to perform two-dimensional annotation of the two-dimensional original image data with respect to the tissue layers and the lesion infiltration range based on the three-dimensional image model to obtain two-dimensional annotated image data; a virtual reality device includes a first display end worn by an operator, which is used to display the two-dimensional annotated image data.
[0006] In one embodiment, the processing of the two-dimensional raw image data based on a deep learning algorithm and an optical flow method to determine a three-dimensional image model of the digestive tract includes: performing classification coloring on the two-dimensional raw image data based on tissue layers and pathological layers based on a deep learning algorithm to obtain coloring information for the two-dimensional raw image data; wherein the deep learning algorithm is pre-trained based on labeled data and includes an image data input head, an output head for tissue layer coloring, and an output head for pathological layer coloring; performing pixel displacement analysis on continuous frames of two-dimensional raw image data captured by a monocular camera in the binocular camera based on the optical flow method to obtain displacement information of the endoscope; performing difference value analysis on the two-dimensional raw image data simultaneously captured by the binocular camera to obtain stereoscopic vision information of the binocular camera; determining a depth estimate based on the displacement information of the endoscope and the stereoscopic vision information of the digestive tract, and compensating the depth estimate based on an adaptive mechanism to obtain compensated depth information; wherein the adaptive mechanism is pre-configured based on the impact of mucosal surface reflection on the depth estimate during endoscopic operation; and constructing a three-dimensional image model of the digestive tract based on the coloring information and the compensated depth information.
[0007] In one embodiment, the virtual reality device further includes an operating terminal worn by the operator, which is used to receive the operator's operation input and forward it to the data processing device.
[0008] In one embodiment, the data processing device is also used to: control the endoscope to perform a specified action matching the operation input, and update the two-dimensional annotated image data based on the specified action; wherein the specified action includes at least one of moving, taking pictures, grabbing and cutting; the first display end is also used to: display the updated two-dimensional annotated image data.
[0009] In one embodiment, the data processing device is also used to: determine the specified action that needs to be completed by the endoscope based on the operation input, and perform image deduction on the two-dimensional annotated image data according to the specified action to obtain the deduced two-dimensional annotated image data; the first display end is also used to: display the deduced two-dimensional annotated image data.
[0010] In one embodiment, the data processing device is also used to: in response to the result of the image deduction being that the endoscope has passed, control the endoscope to perform a specified action matching the operation input, and update the two-dimensional annotated image data based on the specified action; wherein the specified action includes at least one of moving, taking pictures, grabbing and cutting; the first display end is also used to: display the updated two-dimensional annotated image data.
[0011] In one embodiment, the data processing device is further configured to perform force feedback simulation according to the specified action; and the operating end is further configured to perform resistance simulation based on the force feedback simulation.
[0012] In one embodiment, the virtual reality device further includes a second display terminal; the second display terminal is worn on the subject and is used to display customized images; wherein the customized images are adjusted in real time based on the subject's physiological parameters, and the physiological parameters include heart rate and / or electromyographic signals.
[0013] In one embodiment, the second display end is further configured to replay the two-dimensional annotated image data previously displayed by the first display end.
[0014] In one embodiment, the data processing device is configured in the endoscope or the virtual reality device, or is configured independently. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0016] Figure 1 A structural block diagram of an interaction system between an endoscope and a virtual reality device is shown; Figure 2 A schematic diagram of a wearing method of a virtual reality device is shown. DETAILED DESCRIPTION
[0017] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0018] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "an," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.
[0019] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0020] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0021] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or device. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or device comprising the element.
[0022] It should be noted in particular that any symbols and / or numbers in the specification that are not marked in the accompanying drawings are not drawing marks.
[0023] The optional embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0024] The embodiment provided in this application is an embodiment of an interaction system between an endoscope and a virtual reality device.
[0025] The following combination Figure 1 The embodiments of the present application are described in detail.
[0026] Figure 1 A structural block diagram of an endoscope and virtual reality device interaction system 1 is shown. Figure 1 As shown, it includes an endoscope 11, a data processing device 12 and a virtual reality device 13.
[0027] The endoscope 11 is equipped with a binocular camera for collecting two-dimensional original image data in the digestive tract based on the binocular camera.
[0028] The data processing device 12 is configured to process the 2D raw image data based on a deep learning algorithm and optical flow methods to determine a 3D image model of the digestive tract. The 3D image model is labeled with different tissue layers, including the mucosa, vascular network, and muscularis, and the extent of lesion invasion. Furthermore, based on the 3D image model, the 2D raw image data is 2D labeled with respect to the tissue layers and the extent of lesion invasion, thereby obtaining 2D labeled image data.
[0029] The virtual reality device 13 includes a first display terminal 131 worn by the operator, and is used to display two-dimensional annotated image data.
[0030] The interactive system between an endoscope 11 and a virtual reality device 13 provided by the present invention comprises an endoscope 11, a data processing device 12, and a virtual reality device 13. The system uses an endoscope 11 equipped with a binocular camera to collect two-dimensional raw image data from within the digestive tract. The data processing device 12 converts this data into a high-precision three-dimensional image model based on a deep learning algorithm and optical flow methods, while also marking tissue layers and the extent of lesion infiltration. Ultimately, this information is displayed in the form of two-dimensional annotations on the first display terminal 131 of the virtual reality device 13 worn by the operator. This design greatly enhances the doctor's understanding of the spatial relationship of the surgical area, reduces the possibility of misjudgment, and improves surgical efficiency.
[0031] In an embodiment of the present invention, the following method can be used to process the two-dimensional original image data based on a deep learning algorithm and an optical flow method to determine a three-dimensional image model of the digestive tract.
[0032] For example, a deep learning algorithm is used to perform classification and colorization based on tissue and pathology levels on the two-dimensional raw image data, obtaining colorization information for the two-dimensional raw image data. The deep learning algorithm is pre-trained based on labeled data and includes an image data input header, an output header for tissue colorization, and an output header for pathology colorization. Pixel displacement analysis is performed on consecutive frames of two-dimensional raw image data captured by the monocular camera in the binocular camera system using an optical flow method to obtain displacement information of the endoscope 11. Difference analysis is performed on the two-dimensional raw image data simultaneously captured by the binocular cameras to obtain stereoscopic vision information from the binocular cameras. A depth estimate is determined based on the displacement information of the endoscope 11 and the stereoscopic vision information of the digestive tract. This depth estimate is then compensated using an adaptive mechanism to obtain compensated depth information. The adaptive mechanism is pre-configured based on the impact of mucosal surface reflections on the depth estimate during operation of the endoscope 11. A three-dimensional image model of the digestive tract is constructed based on the colorization information and the compensated depth information. The algorithm separately configures output heads for tissue-level coloring and pathology-level coloring in order to output coloring results obtained in different classification methods through different output heads. This can, on the one hand, avoid classification interference caused by coloring with the same color, and on the other hand, facilitate direct observation of coloring information using a single classification method, which facilitates subsequent further processing of the coloring information.
[0033] In the system provided by the present invention, the process of converting two-dimensional images to three-dimensional models is further refined by using deep learning algorithms and optical flow methods. First, the two-dimensional raw image data is classified and colored by a deep learning algorithm to obtain information on different tissue layers and pathological layers. Next, the optical flow method is used to analyze the pixel displacement in the continuous frame image data of the monocular camera to obtain the displacement information of the endoscope 11. In addition, stereoscopic vision information is obtained by performing difference value analysis on the data collected simultaneously by the binocular cameras. Finally, the above information is combined and an adaptive mechanism is used to compensate for the errors caused by the reflection of the mucosal surface to construct an accurate three-dimensional image model. This meticulous and efficient method significantly improves the accuracy and reliability of the model, which helps to improve the success rate of surgery and the safety of patients.
[0034] For example, the virtual reality device 13 also includes an operating terminal 132 worn by the operator, which is used to receive the operator's operational input and forward it to the data processing device 12. The purpose of providing the operating terminal 132 is to enhance the interactivity of the system, allowing the operator to input instructions to the data processing device 12. This approach not only allows the doctor to more directly control the actions of the endoscope 11, such as movement, imaging, grasping, or cutting, but also allows for real-time updates of the displayed content, enhancing flexibility and immediate responsiveness during the procedure. This significantly improves the doctor's operating experience while also enhancing the accuracy and safety of the procedure.
[0035] In the embodiment of the present invention, the system supports actual operation or virtual simulation based on the virtual reality device 13.
[0036] In one embodiment, direct control of the endoscope 11 by the data processing device 12 can enable actual operation of the endoscope 11 via the virtual reality device 13. For example, the data processing device 12 is further configured to control the endoscope 11 to perform a designated action corresponding to the operational input, and to update the two-dimensional annotated image data based on the designated action. The designated action includes at least one of moving, capturing, grabbing, and cutting. The first display terminal 131 is further configured to display the updated two-dimensional annotated image data.
[0037] In this embodiment, when the system simulates and passes the endoscope 11 operation, the data processing device 12 controls the endoscope 11 to perform the corresponding specified actions according to the operation input received during the simulation, and updates the two-dimensional annotated image data in real time. This simulation function is particularly valuable for novice surgeons because it provides a safe learning environment for practicing various complex surgical techniques without increasing risks, thereby improving their practical skills.
[0038] In another embodiment, the data processing device 12 can control and simulate the virtual reality device 13 to achieve a virtual simulation of the operation of the endoscope 11. For example, the data processing device 12 is further configured to determine a specific action to be performed by the endoscope 11 based on the operation input, and to simulate the two-dimensional annotated image data according to the specific action to obtain simulated two-dimensional annotated image data. The first display terminal 131 is further configured to display the simulated two-dimensional annotated image data.
[0039] Furthermore, the data processing device 12 can determine the specific action that the endoscope 11 needs to perform based on the operator's input, and then perform a visual interpretation of the two-dimensional annotated image data accordingly, and then display the interpretation result on the first display terminal 131. This method allows doctors to preview possible results before actual operation, allowing them to make more informed decisions. It not only improves the quality of surgical planning but also provides strong support for the training of new doctors, helping them to master complex skills more quickly.
[0040] On this basis, we can go a step further and directly replicate the simulated actions in practice if the simulation passes.
[0041] For example, the data processing device 12 is also used to control the endoscope 11 to perform a designated action that matches the operation input in response to the endoscope 11 operation simulation, and to update the two-dimensional annotated image data based on the designated action. The designated action includes at least one of moving, taking a picture, grabbing, and cutting. Accordingly, the first display terminal 131 is also used to display the updated two-dimensional annotated image data. When the system simulates the operation of the endoscope 11 and passes, the data processing device 12 will control the endoscope 11 to perform the corresponding designated action in accordance with the operation input received during the simulation process, and update the two-dimensional annotated image data in real time to achieve smooth and proactive endoscope 11 operation indication operation.
[0042] Furthermore, in some embodiments, the data processing device 12 is further configured to perform force feedback simulation based on a specified action. Accordingly, the operating terminal 132 is further configured to perform resistance simulation based on the force feedback simulation. Force feedback simulation is another important feature of this system. It enables the operating terminal 132 to generate a corresponding resistance simulation effect based on a specified action, enhancing the realism of the operation and improving the surgeon's feel when performing delicate movements, helping to reduce errors, which is particularly important during high-precision surgery.
[0043] In some embodiments, the virtual reality device 13 further includes a second display terminal 133 , which is worn by the subject.
[0044] On this basis, second display terminal 133 is used, on the one hand, to display customized images. These customized images are adjusted in real time based on the patient's physiological parameters, including heart rate and / or electromyographic signals. Furthermore, second display terminal 133 is also used to replay the two-dimensional annotated image data previously displayed on first display terminal 131.
[0045] In this embodiment of the present invention, taking into account the patient's comfort, the second display terminal 133 is specifically used to display customized images to the patient. These images are adjusted in real time based on the patient's physiological parameters (such as heart rate and electromyographic signals) to reduce the patient's anxiety and pain. This humanized design not only enhances the patient experience, but also helps reduce unexpected situations caused by tension, ensuring the smooth progress of the treatment process. In addition, the second display terminal 133 can also replay the two-dimensional annotated image data previously displayed on the first display terminal 131. This is of great significance for postoperative review and education. It allows doctors to analyze the surgical process in detail and summarize experiences and lessons learned. It can also serve as teaching materials to help medical students better understand the key points and technical details of actual operations.
[0046] In some embodiments, the data processing device 12 is integrated into the endoscope 11 or the virtual reality device 13, or can be configured independently. This design takes into account the needs of various application scenarios, ensuring the system's portability and adaptability while also leaving room for future technological upgrades and expansion. Overall, this multifunctional integrated design concept demonstrates a high degree of technological innovation and practicality.
[0047] Figure 2 A schematic diagram of a wearing method of a virtual reality device is shown.
[0048] In this embodiment, Figure 2 As shown, the operator can observe the real-time conditions within the patient's digestive tract through the first display terminal 131. By manipulating the operating terminal 132, the endoscope 11 can be simulated within the first display terminal 131, thereby pre-simulating the operation of the endoscope 11. Alternatively, the system can directly receive operational inputs from the operating terminal 132, or replicate previously received inputs if the simulation succeeds, thereby physically controlling the endoscope 11, causing it to perform thrust or surgical maneuvers according to the operator's instructions. The operating terminal 132 can be a tactile glove or operating handle worn by the operator, or other devices capable of operational input, such as a recognition and response device based on eye movement or voice data. Furthermore, a second display terminal 133, worn by the patient, is specifically designed to display customized images. These images are adjusted in real time based on the patient's physiological parameters (such as heart rate and electromyographic signals) to alleviate the patient's anxiety and pain. In addition, the second display terminal 133 can also replay the two-dimensional annotated image data previously displayed on the first display terminal 131, which is of great significance for postoperative review and education.
[0049] The present invention provides a comprehensive system that integrates real-time 3D endoscopic image reconstruction, tactile feedback surgical navigation, immersive patient analgesia and education, and a multi-expert remote collaboration platform. In the real-time 3D endoscopic image reconstruction module, a dual-camera endoscope is used in conjunction with a deep learning algorithm, specifically an optical flow field dynamic depth estimation model, to convert 2D images into 3D models with millimeter-level accuracy, ensuring latency of less than 50 milliseconds. This allows for a clear display of layered anatomical structures, including the mucosal layer, vascular network, and muscular layer, in a VR headset, and accurately marks the extent of lesion infiltration. For the tactile feedback surgical navigation module, when a physician wears tactile gloves to operate a virtual endoscope, the system simulates tissue resistance (e.g., tumor hardness) through force feedback and automatically plans a safe resection path. Furthermore, patients can view customized scenarios, such as deep-sea exploration, while wearing VR devices. The system dynamically adjusts content based on physiological parameters such as heart rate and respiration to reduce pain perception. Postoperatively, VR animations can be used to replay the diagnosis and treatment process, improving follow-up compliance.
[0050] In terms of multi-expert remote collaboration, the system supports multiple doctors using VR technology to enter the same virtual operating room, enabling real-time lesion annotation, endoscope manipulation, and other functions, with the help of 5G networks to achieve low latency of less than 10 milliseconds. Research results show that on the doctor side, novice doctors improved their operating accuracy by 40% after receiving virtual ESD surgery training, with the error controlled within 0.5 mm, while the error under traditional training methods was 1.2 mm. Real-time 3D navigation also shortened the time for complex polyp removal by 25%, reducing the average time from 24 minutes to 18 minutes. On the patient side, after using VR for analgesia, the patient's pain score (VAS) dropped to 2.1, a significant decrease from 6.5 in the control group, and the incidence of anxiety was also reduced by 70%.
[0051] In terms of key technologies, the present invention innovatively applies a dynamic light field reconstruction algorithm, which can extract depth information from monocular endoscopic videos in real time, solve the problem of mucosal reflection interference through adaptive optical flow compensation, and achieve a high precision of 0.2 mm. The bio-signal linkage VR engine can dynamically adjust the VR scene according to the patient's heart rate and electromyography signals, with a response time of less than 100 milliseconds. Cross-platform tactile synchronization technology achieves a force feedback delay of less than 5 milliseconds between tactile gloves and virtual endoscopic operations, and supports multi-level force simulation. In the face of practical difficulties, such as the balance between computing power and power consumption, the present invention ensures that power consumption is less than 5 watts by deploying lightweight models on edge devices; in response to medical data compliance issues, patient images and physiological data are anonymized and encrypted for transmission; in terms of hardware adaptability, the interface standardization between VR devices and existing endoscopic systems is achieved, such as DICOM protocol extension.
[0052] Although operations are described in a particular order in the drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in serial order, or that all shown operations be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.
[0053] The methods, apparatus, devices, and storage media of the present invention can be implemented using standard programming techniques, utilizing rule-based logic or other logic to implement the various method steps. It should also be noted that the terms "apparatus" and "module" as used herein and in the claims are intended to include implementations using one or more lines of software code and / or hardware implementations and / or devices for receiving input.
[0054] Any steps, operations or procedures described herein may be performed or implemented using one or more hardware or software modules, either alone or in combination with other devices. In one embodiment, the software modules are implemented using a computer program product comprising a computer-readable medium containing computer program code, which can be executed by a computer processor to perform any or all of the steps, operations or procedures described.
[0055] The foregoing description of the implementation of the present invention has been given for the purpose of illustration and description. The foregoing description is not intended to be exhaustive nor to limit the invention to the exact form disclosed, and various variations and modifications may exist in accordance with the above teachings, or may be obtained from the practice of the invention. These embodiments are selected and described in order to illustrate the principles of the present invention and its practical application, so that those skilled in the art can utilize the present invention in various embodiments and various modifications suitable for the specific purpose conceived. With respect to the apparatus in the above-mentioned embodiments, the specific manner in which each module performs the operation has been described in detail in the embodiments related to the method and will not be elaborated here.
[0056] It is further understood that, unless otherwise specified, “connection” includes a direct connection where there are no other components between the two elements, and also includes an indirect connection where there are other elements between the two elements.
[0057] It should be further understood that, although operations are described in a particular order in the accompanying drawings in the embodiments of the present invention, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.
[0058] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the field of the invention not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0059] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present invention is limited only by the scope of the appended rights. The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An interactive system between an endoscope and a virtual reality device, characterized in that: include: an endoscope, configured with a binocular camera, for collecting two-dimensional raw image data in the digestive tract based on the binocular camera; a data processing device for processing the two-dimensional raw image data based on a deep learning algorithm and an optical flow method to determine a three-dimensional image model of the digestive tract, wherein the three-dimensional image model is labeled with at least one tissue layer including a mucosal layer, a vascular network, and a muscular layer, and is labeled with a lesion infiltration range; and for performing two-dimensional annotation on the two-dimensional raw image data based on the three-dimensional image model with respect to the tissue layer and the lesion infiltration range to obtain two-dimensional annotated image data; The virtual reality device includes a first display end worn by an operator and used for displaying the two-dimensional annotated image data.
2. The endoscope and virtual reality device interaction system according to claim 1, characterized in that: The processing of the two-dimensional original image data based on the deep learning algorithm and the optical flow method to determine the three-dimensional image model of the digestive tract includes: Based on a deep learning algorithm, the two-dimensional original image data is subjected to classification coloring based on tissue layers and pathological layers, respectively, to obtain coloring information of the two-dimensional original image data; wherein the deep learning algorithm is obtained based on pre-training of labeled data and includes an image data input head, an output head for tissue layer coloring, and an output head for pathological layer coloring; Performing pixel displacement analysis on the continuous frames of two-dimensional raw image data collected by the monocular camera in the binocular camera based on the optical flow method to obtain displacement information of the endoscope; Performing difference value analysis on the two-dimensional original image data simultaneously collected by the binocular cameras to obtain stereoscopic vision information of the binocular cameras; determining a depth estimate based on displacement information of the endoscope and stereoscopic visual information of the digestive tract, and compensating the depth estimate based on an adaptive mechanism to obtain compensated depth information; wherein the adaptive mechanism is preconfigured based on the effect of mucosal surface reflection on the depth estimate during endoscopic operation; A three-dimensional image model of the digestive tract is constructed based on the coloring information and the compensated depth information.
3. The endoscope and virtual reality device interaction system according to claim 1, characterized in that: The virtual reality device further includes an operating terminal worn by the operator, which is used to receive the operator's operation input and forward it to the data processing device.
4. The endoscope and virtual reality device interaction system according to claim 3, characterized in that: The data processing device is further configured to: control the endoscope to perform a designated action matching the operation input, and update the two-dimensional annotated image data based on the designated action; wherein the designated action includes at least one of moving, photographing, grabbing, and cutting; The first display end is further used to display updated two-dimensional annotated image data.
5. The endoscope and virtual reality device interaction system according to claim 3, characterized in that: The data processing device is further configured to: determine a designated action to be performed by the endoscope based on the operation input, and perform image deduction on the two-dimensional annotated image data according to the designated action to obtain deduced two-dimensional annotated image data; The first display end is further used to display the deduced two-dimensional annotated image data.
6. The endoscope and virtual reality device interaction system according to claim 5, characterized in that: The data processing device is further configured to: In response to the result of the image deduction being that the endoscope has passed, controlling the endoscope to perform a specified action matching the operation input, and updating the two-dimensional annotated image data based on the specified action; wherein the specified action includes at least one of moving, photographing, grabbing, and cutting; The first display end is further used to display updated two-dimensional annotated image data.
7. The endoscope and virtual reality device interaction system according to claim 4 or 5, characterized in that: The data processing device is further configured to: perform force feedback simulation according to the specified action; The operating end is further used to perform resistance simulation based on the force feedback simulation.
8. The endoscope and virtual reality device interaction system according to claim 1, characterized in that: The virtual reality device further includes a second display terminal; the second display terminal is worn by the subject and is used to display customized images; The customized image is adjusted in real time based on the physiological parameters of the patient, and the physiological parameters include heart rate and / or electromyographic signals.
9. The endoscope and virtual reality device interaction system according to claim 1, characterized in that: The second display end is further used to replay the two-dimensional annotated image data previously displayed by the first display end.
10. The endoscope and virtual reality device interaction system according to claim 1, characterized in that: The data processing device is configured in the endoscope or the virtual reality device, or is configured independently.