A gastric tube insertion guidance system with AI real-time navigation function

CN122557918APending Publication Date: 2026-08-14SHANXI BETHUNE HOSPITAL (SHANXI ACAD OF MEDICAL SCI SHANXI HOSPITAL OF TONGJI HOSPITAL AFFILIATED TO TONGJI MEDICAL COLLEGE OF HUAZHONG UNIV OF SCI & TECH SHANXI MEDICAL UNIV THIRD HOSPITAL SHANXI MEDICAL UNIV THIRD CLINICAL COLLEGE OF MEDICINE)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-29
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

本发明提供的专用于胃管置入操作的智能导引系统及方法,以解决传统“盲插法”风险高、损伤大、成功率低,以及现有可视化方案智能化程度低、操作依赖经验、成本高昂的问题

Benefits of technology

1.本发明通过专用导航算法模块,在胃管置入过程中实现了对气道(声门、气管环)的AI实时自动识别与预警。系统采用灵敏度优先的安全策略,确保对“偏离路径”的检出率不低于99%,并结合双重确认告警机制。从根本上解决了盲插法误入气管这一致命风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122557918A_ABST
    Figure CN122557918A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of computer-aided medical device technology and discloses a gastric tube insertion guidance system with AI real-time navigation function. Traditional blind insertion methods are prone to tracheal intubation, causing significant damage and having a low success rate. Existing visualization solutions only provide images and cannot intelligently identify risks and guide the procedure. This invention includes an intelligent guidewire, a controller, and a computing module. The computing module runs a dedicated navigation algorithm module. Its target recognition unit identifies multiple anatomical landmarks in the intraluminal video stream through a neural network model. Then, the path decision unit uses a temporal decision model to determine whether the guidewire tip is on the correct path or off-path, and then generates AI real-time navigation instructions. Finally, it provides intelligent guidance for the gastric tube insertion procedure through navigation voice prompts. This invention fundamentally eliminates the fatal risk of tracheal intubation through AI real-time navigation risk warning, and greatly reduces the operational threshold with navigation voice prompts, achieving standardized nursing care.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer-aided medical device technology, specifically to a gastric tube insertion guidance system with AI real-time navigation function. Background Technology

[0002] With the accelerating pace of my country's deep aging population, the demand for long-term nutritional support among disabled and semi-disabled elderly people has surged, making nasogastric tube placement a crucial clinical nursing procedure. Meanwhile, the promotion of "Internet + nursing services" has extended this procedure from hospitals to the home, placing unprecedented demands on its safety and convenience.

[0003] Currently, the mainstream method for gastric tube insertion in clinical practice is still the traditional "blind insertion method," which relies on the operator's feel and the patient's swallowing movements. This method has significant clinical drawbacks: High risk: Traditional methods are prone to causing the tube to be accidentally inserted into the airway, while traditional confirmation methods such as auscultation and inhalation auscultation have limited reliability and pose a risk of fatal complications such as aspiration pneumonia or even suffocation.

[0004] High risk of damage: Repeated trial and error under "blind insertion" can easily damage the fragile mucous membranes of the nasal cavity, pharynx, esophagus, etc., leading to bleeding, pain and increasing the patient's suffering.

[0005] Low success rate: For patients with coma, weakened swallowing reflex, limited neck movement, or anatomical variations, the success rate of first-time catheterization is low. The success of the procedure depends excessively on the operator's personal experience and feel, making it difficult to standardize.

[0006] To address the drawbacks of the "blind insertion method," visualization gastric tubes or guidewires have emerged in existing technologies. While these techniques provide direct visualization, they also have significant limitations: The existing solutions suffer from weak interactivity and low intelligence: they only provide image display and cannot automatically identify anatomical structures, determine the position of the catheter tip, or warn of the risk of accidental airway entry. Operators still need to rely on their own experience to interpret complex cavity images, which does not fundamentally reduce the technical threshold and cognitive load of the operation. For inexperienced community nurses or home-visit nurses, the challenge remains significant.

[0007] High cost and lack of flexibility: Many visualization solutions integrate the camera with the gastric tube or guidewire in a single setup, resulting in the need for various models of "visual gastric tubes," leading to high inventory and management costs. The entire system is discarded after each procedure, wasting medical resources and creating environmental pressure. Furthermore, the complex structure and high cost of these devices make them completely unsuitable for standardized, rapid gastric tube insertion procedures.

[0008] The path for gastric tube insertion is relatively fixed (nasal cavity → pharynx → esophagus → stomach). The core clinical requirement is to quickly, accurately, and safely confirm that the tube is in the correct path and avoid accidental entry into the airway.

[0009] Therefore, there is an urgent need in clinical applications for a lightweight, cost-effective solution specifically designed for gastric tube placement that can provide intelligent guidance, in order to fundamentally improve the safety, success rate and accessibility of the procedure. Summary of the Invention

[0010] The purpose of this invention is to provide a gastric tube insertion guidance system with AI real-time navigation function to solve the problems mentioned in the background art.

[0011] The main design concept of this invention is as follows: This invention provides an intelligent guidance system and method specifically for gastric tube insertion, addressing the problems of high risk, significant damage, and low success rate associated with traditional blind insertion methods, as well as the low level of intelligence, reliance on experience, and high costs of existing visualization solutions. Specifically, it offers a solution that seamlessly integrates into existing nursing procedures, significantly reduces reliance on operator experience through AI-driven real-time voice and image navigation, and allows for the reusability of core guidance components to substantially reduce costs.

[0012] Specifically, this invention aims to achieve the following core objectives: Improve the safety and accuracy of intubation: Through real-time high-definition visualization and AI real-time navigation, the traditional "blind intubation" is transformed into "direct visualization", providing real-time warning of the risk of airway misentry, eliminating fatal complications, improving the success rate of single intubation, and avoiding mechanical damage caused by repeated operations.

[0013] Lowering the operational threshold and standardizing procedures: By using AI to help identify key anatomical structures in real time and convert them into operational instructions, the operation process is transformed from an "art" that relies on "feel" and rich clinical experience into a "standardized technology" with clear navigation guidance. This significantly reduces the learning curve for non-specialist or young nurses and promotes the standardization and homogenization of nursing operations.

[0014] Optimize patient experience and expand application scenarios: Reduce patient pain and discomfort during catheterization, especially suitable for difficult catheterization cases that are hard to handle with traditional methods (such as coma, loss of swallowing reflex, cervical spine injury, esophageal stricture, etc.). Its portability, ease of use, and intelligent features make it perfectly applicable to diverse scenarios such as ICU, emergency room, and home care.

[0015] Achieving both economic and environmental benefits: The core guiding component of the system is designed as a reusable instrument that can withstand high-temperature and high-pressure sterilization. Compared with disposable visual gastric tubes, it can significantly reduce long-term medical costs and the pressure of medical waste disposal, which is in line with the concept of sustainable medical development.

[0016] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a gastric tube insertion guidance system with AI real-time navigation function, comprising: The image acquisition module is used to acquire real-time image sequences along the gastric tube insertion path; The computing and processing module runs a dedicated navigation algorithm module, which generates AI real-time navigation instructions for gastric tube insertion. The dedicated navigation algorithm module includes: The target recognition unit deploys a trained neural network model to identify preset anatomical landmarks in each frame of a real-time image sequence, and obtains a sequence of anatomical landmark recognition results for each frame. The path decision unit, based on the sequence of anatomical marker recognition results from multiple consecutive frames of images, obtains the path status judgment result of gastric tube insertion by deploying a pre-trained temporal decision model; The instruction output unit generates and outputs AI real-time navigation instructions based on the path status judgment results.

[0017] More preferably, the image acquisition module includes: The intelligent guidewire has an image acquisition unit and an illumination unit at its distal end and a connector at its proximal end. The image acquisition unit and the illumination unit are respectively connected to the connector for communication. The intelligent guidewire is wrapped with a tube, and the material stiffness of the tube gradually decreases from the proximal end to the distal end. A controller, connected to the connector, is used to control the movement of the intelligent guidewire. The controller is communicatively connected to the computing module.

[0018] More preferably, the preset anatomical landmarks include: oral cavity, epiglottis, piriform recess, esophageal inlet, esophageal body, cardia, gastric cavity, glottis, and tracheal rings.

[0019] More preferably, the path status determination result of the gastric tube insertion includes the target path status and the deviation from the path status; The target path is as follows: the gastric tube is inserted sequentially through the oral cavity, epiglottis, pyriform recess, esophageal inlet, esophageal body, cardia, and gastric cavity. The deviation from the path refers to the gastric tube being inserted into the glottis or tracheal rings after passing through the epiglottis.

[0020] More preferably, the instruction output unit further includes a warning function, including: When the sequence of anatomical marker recognition results for a single frame image is determined to be off-path, a Level 1 alarm is triggered. When the sequence of anatomical marker recognition results for multiple frames image is determined to be off-path, a Level 2 alarm is triggered.

[0021] More preferably, the temporal decision model is any one of a lightweight temporal convolutional network, a long short-term memory network, or a rule-based finite state machine; The state transition rule of the rule-based finite state machine is as follows: when the anatomical marker recognition result sequence of a single frame image is the glottis or tracheal ring, the state machine is driven to transition to the off-path state.

[0022] More preferably, the calculation processing module further includes a display unit, which is used to display visual prompts for the path status judgment results; The visual cues include bounding boxes for anatomical landmarks and forward arrows indicating the target path status.

[0023] More preferably, the computing processing module further includes a voice unit, which is used to synthesize AI real-time navigation instructions into navigation voice prompts for broadcast.

[0024] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention utilizes a dedicated navigation algorithm module to achieve real-time AI-powered automatic identification and early warning of the airway (glottis, tracheal rings) during gastric tube insertion. The system employs a sensitivity-priority safety strategy, ensuring a detection rate of at least 99% for "deviations from the path," and incorporates a dual-confirmation alarm mechanism. This fundamentally solves the fatal risk of accidental tracheal intubation during blind insertion.

[0025] 2. The core interactive innovation of this invention lies in AI real-time voice navigation. The system converts complex intracavitary image recognition results into clear action commands in real time, guiding the operation through voice prompts. This allows operators to complete catheter placement without long-term training, simply by following voice prompts. This greatly promotes the homogenization of nursing quality in medical institutions at all levels and even in home settings.

[0026] 3. This invention employs a reusable intelligent guidewire for initial probing, followed by the insertion of a disposable gastric tube. This procedure is entirely consistent with long-established clinical practices for tube placement. Medical staff do not need to change the core operational steps; they only need to add the simple action of "following voice guidance." The learning cost is extremely low, making it easy to be widely accepted and rapidly promoted.

[0027] 4. The intelligent guidewire of this invention adopts a gradient stiffness tube body and a fully sealed sterilization design, achieving safe and reliable reusability. Compared with disposable visual gastric tubes, the cost per use is significantly reduced, alleviating the economic burden on medical institutions and patients. It also reduces the generation of disposable medical waste, aligning with the concept of green and low-carbon medical development.

[0028] 5. The path decision unit of this invention not only provides neural network time-series models (such as TCN and LSTM) solutions, but also proposes a hybrid decision architecture of "finite state machine (FSM) + neural network". It uses transparent and deterministic rules (state transition diagrams) for final path determination, ensuring that each alarm can be traced back to a specific image frame and recognition result, making the decision-making process fully auditable and traceable.

[0029] 6. This invention is not only applicable to traditional scenarios such as ICUs, emergency rooms, and wards, but can also be extended to off-site scenarios such as community health service centers, elderly care institutions, and rehabilitation wards. This enables patients with difficult catheter placement (such as those who are comatose, have weak swallowing reflexes, or have anatomical variations) to receive high-success-rate catheter placement services regardless of their location, greatly expanding the coverage of safe and reliable nursing services. Attached Figure Description

[0030] Figure 1 This is a block diagram of the gastric tube insertion guidance system with AI real-time navigation function of the present invention; Figure 2 This is a diagram of the gastric tube placement guidance system with AI real-time navigation function according to the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] In the description of this invention, it should be noted that the terms "upper," "lower," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, the terms "installation" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0033] Example 1 like Figure 1 and Figure 2As shown, this embodiment provides a gastric tube placement guidance system with AI real-time navigation function, including the following modules: Image acquisition module 101 is used to acquire real-time image sequences along the gastric tube insertion path; The computing and processing module 102 has a dedicated navigation algorithm module running on it, which is used to generate AI real-time navigation instructions for gastric tube insertion operations; The dedicated navigation algorithm module includes: The target recognition unit deploys a trained neural network model to recognize preset anatomical landmarks in each frame of the real-time image sequence, and obtains a sequence of anatomical landmark recognition results for each frame. The path decision unit, based on the sequence of anatomical marker recognition results from multiple consecutive frames of images, deploys a trained temporal decision model to obtain the path status judgment result for gastric tube insertion; The instruction output unit generates and outputs AI real-time navigation instructions based on the path status judgment results to perform gastric tube insertion.

[0034] The technical features of each module are described in detail below.

[0035] S10, Intelligent Guidewire The intelligent guidewire is the core component of this system that contacts the patient and acquires images. Its total length is typically between 100cm and 120cm, and its diameter between 2.0mm and 3.5mm, ideally allowing it to pass smoothly through a standard-sized gastric tube. The intelligent guidewire consists of three parts: a distal structure, a tube body structure, and a sealing and sterilization structure.

[0036] 1. Remote structure The distal structure features a streamlined design with a smooth surface to reduce insertion resistance. A miniature camera and illumination source (4-6 miniature LEDs) are integrated on the end face of the distal structure. The miniature camera preferably uses a CMOS image sensor smaller than 2mm x 2mm, with a resolution of 1920x1080 pixels (1080P) or higher and a frame rate of at least 30fps to ensure image clarity and smoothness.

[0037] 2. Pipe structure The internal cables (including signal and power lines) of the guide tube are specially designed to enable the reusability of this invention. Specifically, the cables use high-strength, flexible miniature coaxial cables to transmit signals and power, and then the cables are coated with a gradient composite material co-extrusion molding process on the outside.

[0038] The proximal portion of the guidewire body is made of a relatively hard material, such as polyetheretherketone (PEEK), to provide sufficient pushing stiffness and torque transmission capacity, facilitating operator control of direction. The middle portion of the guidewire body is made of a medium-hardness material, such as polyurethane (PU). The distal portion of the guidewire body is made of the softest material, such as silicone or soft polyurethane, to ensure gentle conformability to the curves of the pharynx and esophageal inlet, avoiding damage to the mucosa.

[0039] This "gradient stiffness" design, with its continuous or stepwise decrease in stiffness from proximal to distal, was optimized through extensive mechanical simulations and in vitro experiments, maximizing patient comfort and safety while ensuring operability.

[0040] 3. Sealed sterilization structure To achieve reusability, the entire smart guidewire must be able to withstand high-pressure steam sterilization. Specific implementation solutions include: Material Selection: The outer coating material is made entirely of medical-grade polymer materials with excellent high-temperature resistance, such as fluorosilicone rubber (FVMQ), perfluoroether rubber (FFKM), or high-temperature vulcanized silicone rubber (HTV). These materials can maintain stable performance for a long time at high temperatures of 135℃-145℃.

[0041] Sealing process: The distal seal uses high-temperature cured medical-grade epoxy resin. The tube body's coating is tightly bonded to the internal cables via co-extrusion or injection molding. The proximal connector is made of medical-grade metal or high-temperature plastic, and the connection to the tube body is permanently sealed using laser welding or high-temperature adhesive.

[0042] Circuit potting: Inside the connector, all solder joints and precision components are fully potted with biocompatible, high-temperature resistant potting compound (such as silicone potting compound) to prevent high-temperature, high-pressure steam from entering and causing short circuits or corrosion.

[0043] The intelligent guidewires processed by the above process must undergo rigorous sterilization validation tests to prove that they can withstand multiple standard vacuum high-pressure steam sterilization cycles. After each sterilization, functional tests (image clarity, illumination uniformity, electrical safety, biocompatibility) and leakage tests must be performed to ensure reliable performance and the integrity of the sterile barrier.

[0044] S11, Controller The controller serves as the physical interface for human-computer interaction. Its housing is made of medical-grade ABS or PC material with a non-slip surface. Internally integrated are: Main control board: Contains a microcontroller for managing power, controlling LED brightness, processing button signals, and communicating with the computing unit via a built-in Wi-Fi / Bluetooth module or a physical interface (such as USB-C).

[0045] Power module: A rechargeable lithium battery that powers the camera, lighting source, and controller itself of the smart guide wire.

[0046] User interface includes a power switch, brightness button, and photo / video recording button. There is also a status indicator light that displays different colors depending on the system status (e.g., blue for standby, green for running, and flashing red for alarm).

[0047] S12, Calculation and Processing Unit The computing processing unit is a standalone mobile terminal, such as a tablet or smartphone, or it can be an embedded system integrated into a controller with a small display screen. Its core task is to run a dedicated navigation algorithm module and provide a user interface.

[0048] The tablet is mounted on a stand in a position easily observed by the operator. Its hardware configuration needs to meet the requirements of real-time AI inference, such as being equipped with a high-performance ARM processor or a dedicated edge AI computing chip.

[0049] Example 2 This embodiment provides a detailed description of the dedicated navigation algorithm module, which is the core software component for achieving real-time AI navigation in this invention. It is deployed on the computing processing unit and includes a complete software flow encompassing data preprocessing, AI inference, decision logic, and feedback generation.

[0050] S20: Building the dataset High-quality, targeted datasets are the cornerstone of training specialized navigation algorithm models. This includes collecting data from multiple sources, establishing data labeling systems, and partitioning the dataset.

[0051] S201: Collecting multi-source data Multi-source data is divided into three levels, specifically including: 1. Real-world clinical data (core data): With ethical review and informed consent from patients, the entire process of gastric tube insertion was recorded in the hospital's endoscopy center, ICU, and other departments using prototype equipment or equivalent miniature endoscopes.

[0052] The goal is to collect no fewer than 2,000 videos, covering adults, the elderly, different body positions (supine, semi-recumbent), and different pathological states (normal anatomy, postoperative changes, tumor lesions), to ensure the model's generalization ability. Video recording will begin from the oral cavity insertion point and continue until confirmation of entry into the gastric cavity.

[0053] 2. High-fidelity simulator data (auxiliary data): Using high-fidelity upper respiratory tract-digestive tract models such as the Laerdal laryngotrachea model, standard insertion paths and intentionally misplaced trachea paths were simulated under controlled laboratory conditions, and comparative videos were recorded. This data is crucial for obtaining sufficient and standardized negative samples of "deviation from the path".

[0054] 3. Data Augmentation and Synthesis: Offline augmentation is performed on the acquired real images to expand the dataset and improve model robustness. Offline augmentation includes: geometric transformations (random horizontal / vertical flipping, small-angle rotation, scaling, affine transformation); photometric transformations (randomly adjusting brightness, contrast, saturation, and hue); simulated noise (adding Gaussian noise and salt-and-pepper noise); and simulated occlusion (simulating the effect of the lens being partially obscured by mucus, blood, or bubbles, achieved by randomly adding black or semi-transparent patches).

[0055] The data collection specifications are shown in Table 1 below: Table 1: Data Collection Specifications S201: Data Labeling System The annotation work was completed by a professional medical data annotation team under the guidance of senior clinicians, and a "three-level annotation + arbitration" mechanism was adopted to ensure quality. Junior annotators (trained medical image annotators) complete the first round of annotation; Senior annotators (resident physicians in gastroenterology / anesthesiology) will review and correct the errors. Expert arbitration (associate chief physician or above) makes the final ruling on the disputed labeling.

[0056] Data annotation comprises two parallel tasks. Task A is used to annotate the detection boxes for anatomical landmarks. Frames are extracted from the video data at fixed intervals (e.g., 1 frame per second). For each frame, a bounding box is drawn using an annotation tool, and its category is labeled. This invention defines a total of 9 categories (C0-C8), which are described in detail in Table 2 below.

[0057] Table 2: Categories of Anatomical Markers Task B is used for path state segment annotation, which is performed on video segments (such as sliding windows, 2 seconds long, approximately 30-60 frames). Annotators observe the segments and determine whether the guidewire tip is generally on the correct path or off-path. For segments containing state transitions, segmentation and annotation are performed. The definitions of two-state path segment categories are shown in Table 3 below.

[0058] Table 3: Classification of Two-State Path Fragments To ensure the reliability and consistency of data annotation results, this invention sets strict evaluation criteria for each annotation task. For Task A, the intersection-union ratio (IoU) is used as a consistency metric; the annotation is considered consistent only when the IoU value between the annotation result and the baseline is ≥0.75.

[0059] For Task B, Cohen's Kappa coefficient is used as an indicator to evaluate consistency among annotators, requiring a Kappa coefficient ≥ 0.85. These two thresholds indicate that Task A requires a high degree of consistency in the spatial location and range of the labeled targets, while Task B requires that the classification judgments have almost completely consistent credibility, together ensuring a high-quality foundation for the labeled data used in subsequent model training and evaluation.

[0060] S202: Dataset Partitioning Images are strictly segmented according to the "patient / procedure case" dimension, rather than randomly shuffling. Ensure that video frames from the same patient appear only in one of the training, validation, or test sets to prevent data leakage that could lead to inflated model evaluation results.

[0061] This invention divides the dataset into training, validation, and test sets in a 70%:15%:15% ratio. The test set must include a subset of "hard samples" specifically designed to evaluate the model's performance in challenging scenarios such as low light, abundant mucus, and anatomical variations.

[0062] S21: Design Model Architecture and Training The dedicated navigation algorithm model of this invention adopts a two-stage cascaded architecture of "target detection + temporal decision". Among them, target detection is used to identify anatomical landmarks, and temporal decision is used to determine binary path segments.

[0063] S210: Anatomical Marker Recognition Model (Phase 1) To achieve real-time inference (>20fps) on terminal devices, the lightweight single-stage object detection network YOLOv8-nano or YOLOv8-small was selected as the backbone. The input image size was fixed at 640x480 pixels, aligned with the output of the miniature camera.

[0064] 1. Model Structure Backbone: Employs an improved CSPDarknet structure for extracting multi-scale feature maps from input images. It includes multiple CBS (Conv-BN-SiLU) modules and C2f modules, controlling computational cost while maintaining feature extraction capabilities.

[0065] Neck (Feature Fusion): Employs a Path Aggregation Network (PANet) or Bidirectional Feature Pyramid Network (BiFPN) structure. It receives feature maps from different layers of the Backbone and performs top-down and bottom-up fusion, combining deep semantic information with shallow detail information, which is beneficial for detecting anatomical landmarks of different sizes.

[0066] Head (Detection Head): A decoupled head is used to separate the classification and regression tasks. The final output consists of three parts: the coordinates of each predicted box, the confidence score of containing the target, and the probability of belonging to each of the nine categories (C0-C8).

[0067] 2. Domain Adaptation Strategy Because medical images differ greatly from general datasets (such as COCO), direct training yields poor results. Therefore, this invention employs the following domain adaptation strategy.

[0068] Pre-trained weight initialization: The model's backbone parameters are initialized using YOLOv8-nano weights pre-trained on the large general-purpose dataset COCO. This leverages the primary visual features learned by the model in general object recognition, accelerating convergence and improving final performance.

[0069] Progressive unfreezing and fine-tuning: In the first stage, all layers of the backbone are frozen, and only the neck and head parts are trained. A small learning rate is used, and the network is trained on a gastric tube placement dataset for approximately 50 epochs. This step allows the network's "neck" and "head" to adapt to the feature distribution of medical images.

[0070] The second stage involves unfreezing some or all of the backbone layers, using a slightly larger learning rate, and employing a cosine annealing scheduler for end-to-end fine-tuning, training for approximately 300 epochs. An early stopping strategy is also employed: training stops when the validation set loss no longer decreases within 50 consecutive epochs to prevent overfitting.

[0071] Class-weighted loss: Due to the significantly smaller sample size of certain key classes (such as C7 glottis and C8 tracheal rings) in the dataset compared to other classes, a class imbalance problem exists. In the loss function, higher weights are applied to the classification loss for these key classes, forcing the model to focus more on these few but crucial classes during training, thereby improving its detection accuracy.

[0072] S211: Path State Decision Model (Phase Two) Detection of a single frame image may result in false positives, false negatives, or occlusions. Therefore, it is necessary to make decisions based on temporal context, that is, to analyze the "sequence of occurrence of anatomical landmarks" formed by multiple consecutive frames of images to determine whether the current path is correct.

[0073] The correct order of anatomical landmarks for the correct pathway is as follows: Oral cavity (C0) → Epiglottis (C1) → Piriform recess (C2) → Esophageal inlet (C3) → Esophageal body (C4) → Cardia (C5) → Gastric cavity (C6); The characteristic signal of deviation from the path is: If the glottis (C7) and / or tracheal rings (C8) appear after the epiglottis (C1), it is considered a deviation from the path.

[0074] 1. Input Construction The trained anatomical landmark recognition model is used to infer each frame of a continuous video (e.g., 16 frames). For the t-th frame, a feature vector is obtained. The feature vectors of the consecutive T frames are stacked in chronological order to form a T x D feature matrix (D is the dimension of the feature vectors), which is used as the input to the path state decision model.

[0075] 2. Model Structure Selection: Lightweight Temporal Convolutional Network (TCN) is preferred. TCN uses one-dimensional causal convolution to process temporal data, which can effectively capture local patterns (such as "C1 epiglottis appears, followed by C7 glottis", which is a strong signal of deviation from the path).

[0076] Its structure consists of multiple residual blocks, each containing dilated convolution, weight normalization, ReLU activation, and Dropout. Finally, a scalar is output through a global average pooling layer and a fully connected layer. After passing through a sigmoid function, the probability P that the current segment is in the "correct path state" is obtained. A preset threshold θ (e.g., 0.5) is used; P is considered positive if it ≥ θ, and negative otherwise.

[0077] Alternative model architecture options include Long Short-Term Memory (LSTM) networks. Single-layer or two-layer LSTMs can be used to model temporal dependencies. LSTMs can learn long-range dependencies through their gating mechanism. The temporal feature vector sequence is sequentially input into the LSTM, the hidden state at the last time step is taken, and the probability is output through a fully connected layer and a sigmoid function.

[0078] Alternatively, a hybrid approach combining a finite state machine (FSM) and a neural network can be used for the model structure. Its core is a predefined deterministic state machine, whose states include: S0 - initiation (oral cavity), S1 - epiglottis, S2 - esophageal inlet, S3 - esophageal body, S4 - cardia, S5 - gastric cavity (in position), and S_ERR - deviation alarm.

[0079] The triggering conditions for state transitions are entirely driven by the neural network's recognition results of a single frame of image. The rules are deterministic and transparent, and the state transition diagram is defined as follows: S0 (oral cavity) — C1→S1 (epiglottic area) was detected; S1 (epiglottic area) — C2 / C3 detected → S2 (esophageal inlet) → PATH_CORRECT maintained; S1 (epiglottic area) — Detection of C7→S_ERR (glottis)→PATH_DEVIATED triggers an alarm; S2 (esophageal inlet) — C4 detected → S3 (esophageal body) → PATH_CORRECT maintained; S3 (body of the esophagus) — C5 detected → S4 (cardia) → PATH_CORRECT maintained; S4 (cardia) — C6 detected → S5 (gastric cavity in position) → PATH_CORRECT final confirmation.

[0080] The neural network (anatomical landmark recognition model) is responsible for providing the detection results for each frame, and the path state decision logic is clearly defined by the state machine. This makes the system's decision-making process fully auditable and traceable, and any alarm can be mapped to a specific image frame and recognition result, meeting the stringent requirements of transparency and interpretability for medical products.

[0081] 3. Security thresholds and alarm policies Given the stringent safety requirements of medical applications, this system's decision threshold setting follows the principle of "sensitivity (Recall) priority." It is better to generate false alarms (false positives) than to miss alarms (false negatives). The specific solution is as follows: Plot the receiver operating characteristic curve (ROC curve) of the two-state decision model on the validation set, and select the threshold corresponding to the point with the highest specificity, provided that the sensitivity (detection rate of deviation from the path) is not less than 99%, as the operating point.

[0082] A dual verification mechanism is implemented to balance sensitivity and specificity: Level 1 Warning: When a high-risk marker (C7 or C8) is detected in a single frame image and the confidence level is higher than a threshold (e.g., 0.85), the system immediately highlights the area on the screen with a flashing yellow box and plays a prompt voice message: "Please note, there may be an airway ahead."

[0083] Level 2 Confirmation Alarm: If high-risk markers are continuously detected within the following N consecutive frames (e.g., N=3), the system determines it as "confirmed deviation". At this time, the warning box on the screen turns red and flashes strongly, while a clear voice command is played: "Warning! You have deviated from your airway. Please back up immediately!"

[0084] S22: Model Training Process and Optimization After completing the model architecture design and dataset construction, a carefully designed training process is needed to transform the data into usable model weights. The training in this invention is divided into two stages: offline training and online optimization.

[0085] S221: Training Environment Configuration Model training is performed on a high-performance computer, and the specific configuration requirements are shown in Table 4 below. S221: Training of Anatomical Marker Recognition Model This phase trains a deep learning model (such as YOLOv8-nano) to perform object detection, with the goal of accurately identifying various anatomical landmarks in a single frame image.

[0086] 1. Set training parameters The initial learning rate is 0.01. Combined with the cosine annealing scheduler, the learning rate decreases smoothly with each training round, which helps the model converge to a better local optimum.

[0087] Batch size: 32. Scales linearly with the number of GPUs when using multi-GPU distributed training.

[0088] Optimizer: Stochastic gradient descent optimizer is used with a momentum of 0.937 and a weight decay of 5e-4, which is the default efficient configuration of the YOLOv8 framework.

[0089] Training cycles: A total of 300 training cycles, with an early stopping strategy. If the validation set loss does not decrease within 50 consecutive cycles, training is terminated early to prevent overfitting.

[0090] Input size: fixed at 640×480 pixels, aligned with the camera resolution of the smart guidewire.

[0091] Data augmentation includes Mosaic (probability 1.0), MixUp (probability 0.1), random HSV hue perturbation, and horizontal flipping to improve the model's robustness to different lighting, angles, and local features.

[0092] Loss functions include CIoU Loss for bounding box regression, binary cross-entropy loss for classification, and Distribution Focal Loss for improved classification.

[0093] 2. Training Steps Data preparation: Divide the labeled dataset into training, validation, and test sets according to the patient dimension. Convert the annotations to YOLO format and generate a data configuration file.

[0094] Pre-trained weight loading: Load the weights of the YOLOv8 model pre-trained on the COCO dataset, using its learned general image features as a starting point.

[0095] The first training phase involves freezing the first 10 layers of the backbone network and training it for approximately 50 epochs using a low learning rate (0.001). The goal of this phase is to allow the model's "neck" and "head" to adapt to the feature distribution of medical images while retaining its general feature extraction capabilities.

[0096] The second training phase involves unfreezing all network layers and performing end-to-end fine-tuning using an initial learning rate of 0.01 and a cosine annealing strategy, training for 300 epochs or triggering early stop. This phase allows for deep optimization of the entire model parameters for the gastric tube placement scenario.

[0097] Model evaluation: Evaluate the performance of the final model on the test set. Key metrics include mean accuracy.

[0098] S221: Training of Two-State Decision Model 1. Construct training data Using a pre-trained anatomical landmark recognition model, inference is performed on all videos in the training set to obtain the detection results (confidence and location for each category) for each frame. Then, the video is divided into multiple segments using a sliding window approach. Each segment contains T frames (T=16) and its corresponding detection result sequence. The label of this segment (training sample) is the "correct path" or "off-path" status marked by "Task B".

[0099] 2. Training parameters Taking the temporal convolutional network scheme as an example, the training parameters are shown in Table 5 below.

[0100] Table 5: Training Parameters of the Two-State Decision Model S222: Training Monitoring and Iterative Optimization During model training, key metrics are monitored and visualized in real time: Monitoring metrics: For anatomical biomarker models, monitor the training / validation loss curve, mAP curve, and precision and recall curves for each category. For binary decision models, monitor loss, accuracy, sensitivity, specificity, F1 score, and AUC-ROC curve.

[0101] Adversarial testing: Regularly evaluate the model on a subset of “hard samples” that include occlusion, blurring, and anomalous anatomy to ensure its robustness.

[0102] Iterative strategy: If the performance on the validation set does not meet expectations, iterative optimization should be performed according to the following priorities: 1) Increase labeled data, especially samples of underperforming categories; 2) Adjust data augmentation strategies; 3) Adjust model capacity; 4) Adjust hyperparameters such as learning rate.

[0103] S22: Model Deployment and Performance Optimization The dedicated navigation algorithm model trained through the above process needs to be optimized and transformed before it can be efficiently and stably deployed on the embedded hardware of the terminal.

[0104] S221: Model Optimization and Conversion Process To ensure cross-platform compatibility and inference efficiency, the following standardized model conversion process is adopted: Export format: Export the model weight file (.pt) obtained from PyTorch training as an open ONNX format intermediate representation (.onnx). The input size of the model is fixed at 640×480 pixels, and tools such as onnx-simplifier are used to simplify the computation graph.

[0105] Based on the computing power characteristics of the target deployment platform, perform in-depth optimization: For high-performance edge computing devices, such as the NVIDIA Jetson series, use the TensorRT SDK to convert the ONNX model into an inference engine (.engine) and apply FP16 half-precision quantization.

[0106] For general-purpose embedded CPU platforms: use the ONNX Runtime inference engine, with optional INT8 integer quantization.

[0107] Mobile alternative: Convert to TensorFlow Lite or Core ML format to adapt to mobile devices such as smartphones or tablets.

[0108] S221: Deployment Performance Requirements The system after model deployment must meet the following performance indicators to ensure clinical real-time performance and reliability.

[0109] Table 6: Deployment Performance Metrics S221: Terminal Multi-threaded Deployment Architecture To achieve low-latency, high-throughput real-time processing, a multi-threaded pipeline architecture is deployed on the terminal device, specifically including the following threads: Image acquisition thread: The miniature camera of the smart guidewire captures video frames and stores them in a circular buffer.

[0110] Preprocessing thread: retrieves frames from the buffer and performs standardization operations such as scaling, pixel value normalization, and conversion to tensors.

[0111] Anatomical marker detection inference thread: calls the optimized TensorRT engine, performs YOLO model inference, outputs the detection results of the current frame, and puts them into the result queue.

[0112] Two-state decision reasoning thread: retrieves detection results from multiple consecutive frames in the queue, inputs them into the temporal convolutional network (TCN) or finite state machine (FSM) module, and obtains the final path state judgment.

[0113] Feedback execution thread: Triggers the corresponding navigation feedback immediately based on the judgment result. If the value is PATH_CORRECT, the system indicator light will turn green and trigger a voice prompt to "continue".

[0114] If it is PATH_DEVIATED, the red indicator light will flash, a voice alarm will play saying "Please note: the catheter may have deviated from the esophageal path. It is recommended to remove the catheter for inspection," and a warning message will be displayed on the screen.

[0115] S23: System Testing and Solution Verification To ensure the system's safety and effectiveness, it must undergo multi-level and rigorous testing and verification.

[0116] S231: Laboratory bench testing In a controlled laboratory environment, the dedicated navigation algorithm model was quantitatively evaluated. The specific test plan is shown in Table 7 below.

[0117] Table 7: Deployment Performance Metrics S232: Preclinical Validation Using Human Simulators Operational validation was performed on a highly realistic upper digestive tract mannequin to evaluate the system's performance in a human-computer interaction environment: Participants: Operators with different experience levels (senior physicians, junior physicians, and nurses) were invited to complete 30 catheter placement procedures each.

[0118] Scenario design: 30% of the operation settings are designed to create obstacles, leading the operator to accidentally enter the trachea, in order to test the system's real-time warning capability.

[0119] Evaluation method: Record the system's judgment results. After the operation is completed, another expert who did not participate in the operation confirms the "gold standard" path through video playback. The two results are compared to calculate the confusion matrix, and clinical performance indicators such as sensitivity, specificity, positive predictive value, and negative predictive value are obtained.

[0120] S233: Clinical Validation Design and execute prospective, multicenter clinical trials in accordance with the requirements for medical device clinical trials: Primary endpoint: the rate of agreement between the algorithm's bistate judgment results and the clinical gold standard (confirmation of gastric tube position by X-ray imaging).

[0121] Secondary endpoints: the improvement in first-time catheterization success rate, the reduction in average operation time, and the incidence of adverse events.

[0122] Sample size: determined based on statistical power calculations, estimated to require 200-500 subjects.

[0123] S24: Continuous Learning and Version Management To address new situations encountered in clinical practice and continuously improve performance, the system possesses continuous learning and strict version management capabilities.

[0124] S241: Data Closed-Loop Mechanism Establish a closed-loop data mechanism of "use-collection-improvement-update".

[0125] Difficult sample collection: In compliant clinical use, with the patient's consent, "difficult frames" with low detection confidence (e.g., <0.5) or those judged to be in the uncertain range are automatically screened.

[0126] Manual annotation and data entry: The anonymized difficult data is uploaded to a secure server, where it is reviewed and finely annotated by an expert team before being added to the central training set.

[0127] Model retraining and testing: The model is iteratively retrained using the augmented dataset, and the new model must pass the full regression test suite.

[0128] Security Update Push: After passing the test, the new model version will be pushed to all active terminal devices via OTA through a secure wireless network.

[0129] S242: Version Control Strategy To meet the lifecycle management requirements of medical device software, the following version control strategy is implemented.

[0130] Semantic version number: Each model and its accompanying software version is bound to a unique version number (e.g., v1.2.3) to clearly identify major updates, feature improvements or patches.

[0131] The entire process is traceable: all training datasets, annotation files, training configurations, and model weight files are included in a version control system (such as Git).

[0132] Regression testing assurance: Before any model update, a regression test suite containing all historical test cases must be used to ensure that the performance of the new version is no less than that of the old version in all tested scenarios.

[0133] Compliant with SaMD requirements: The entire continuous learning and version management process follows the relevant international standards and regulatory guidelines for Medical Device Software as Medical Devices (SaMD).

[0134] Example 3 like Figure 1 and Figure 2 As shown, based on the gastric tube placement guidance system with AI real-time navigation function in Embodiment 1, this embodiment provides a complete workflow of a gastric tube placement guidance method with AI real-time navigation function, including the following steps: The smart guidewire is inserted into the gastric tube to form an assembly; The assembly was inserted into the patient's nasal cavity, and real-time image sequences were acquired along the gastric tube insertion path. AI-powered real-time navigation commands are generated based on real-time image sequences. AI-powered real-time navigation command broadcasting and navigation voice prompts; Perform gastric tube insertion based on navigation voice prompts; Once the assembly reaches the target position, the smart guidewire is withdrawn from the gastric tube.

[0135] 1. Connection and Self-Test The operator removes the sterilized smart guidewire and inserts its proximal connector into the corresponding port on the controller. The controller and computing unit (such as a tablet) are then powered on. The system automatically establishes a connection and performs a self-test: checking if the miniature camera image is normal, if the illumination source is adjustable, and if the dedicated navigation algorithm module is successfully loaded. After the self-test passes, the tablet screen displays the real-time camera feed.

[0136] 2. Assembly and Preparation The operator takes a suitable size of disposable sterile gastric tube and slowly inserts the smart guidewire through the tail end of the tube until the camera portion at the tip of the guidewire protrudes slightly from the side hole at the front of the tube. Then, a sufficient amount of water-based medical lubricant is applied to the tip of the smart guidewire / gastric tube assembly.

[0137] 3. Placement operation under AI real-time navigation The operator stands to the patient's right, holding the controller in one hand and supporting the gastric tube with the other. The assembly is then slowly inserted through the patient's nasal cavity.

[0138] Initial stage: Entering the nasal cavity and nasopharynx. The internal structure of the nasal cavity is displayed on the screen, and the operator mainly relies on touch to slowly advance the device.

[0139] Entering the oropharynx: As the distal end of the intelligent guidewire crosses the nasopharynx and enters the oropharynx, the AI ​​identifies the C0 oral / oropharyngeal cavity features. The system's voice prompt states, "Entering the pharynx, please continue advancing slowly." Simultaneously, the identified anatomical landmarks are marked with green boxes on the screen.

[0140] Key Decision Point—Epiglottis and Esophagus Inlet: When the C1 epiglottis appears in the field of view, the screen highlights this structure. This is the first key decision point, and the system will continue to analyze subsequent frames. Correct path: If the C2 piriform recesses on either side of the epiglottis or the C3 esophageal inlet below are detected, the system determines that the path is correct. The screen will point to the esophageal inlet with a green arrow and play the instruction: "Esophagus inlet found, please continue gently." Deviation from Path (Risk Warning): If the system detects an inverted V-shaped C7 glottis within the C1 glottis area, a deviation warning will be triggered immediately. The C7 glottis area on the screen will be prominently marked with a flashing red box.

[0141] Simultaneously, a warning voice will play: "Warning! Airway ahead. Please stop advancing, gently back up and adjust your direction downwards." The operator must follow the instructions. If the operator enters deeper and the C8 tracheal ring is detected, the alarm will be more intense.

[0142] Passing through the esophagus: After entering the C3 esophageal inlet, the AI ​​identifies it as the C4 esophageal body. A voice prompt states, "Entering the esophagus, proceed at a steady pace." A green forward arrow appears on the screen.

[0143] Reaching the stomach cavity: Upon entering the stomach cavity through the C5 cardia, the C6 gastric mucosal fold appears. Voice prompt: "Stomach cavity reached, preparing for final location confirmation," the screen displays "Target reached - Stomach cavity" and may include a confirmation icon.

[0144] Throughout the catheterization process, the operator's attention should be focused on the patient and the procedure being performed. They only need to listen to clear, concise voice instructions and do not need to constantly look at the screen. Visual feedback on the screen serves as an aid, used for confirmation when in doubt.

[0145] 3. Placement operation under AI real-time navigation Upon hearing the signal that the gastric cavity has been reached, the operator holds the guidewire with one hand and pushes the gastric tube along the guidewire to the predetermined mark with the other. Then, a final confirmation is made using traditional methods (such as using a syringe connected to the gastric tube to aspirate and observe gastric fluid). After confirmation, the gastric tube is secured to the patient's nose and cheek with adhesive tape.

[0146] Finally, the smart guidewire is removed from the gastric tube. After cleaning, the removed smart guidewire is sent to the sterilization supply center for standard cleaning, sterilization, and autoclaving. After drying, it is placed in sterile packaging for future use.

[0147] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0148] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A gastric tube insertion guidance system with AI real-time navigation function, characterized in that, include: The image acquisition module is used to acquire real-time image sequences along the gastric tube insertion path; The computing and processing module runs a dedicated navigation algorithm module, which generates AI real-time navigation instructions for gastric tube insertion. The dedicated navigation algorithm module includes: The target recognition unit deploys a trained neural network model to identify preset anatomical landmarks in each frame of a real-time image sequence, and obtains a sequence of anatomical landmark recognition results for each frame. The path decision unit, based on the sequence of anatomical marker recognition results from multiple consecutive frames of images, obtains the path status judgment result of gastric tube insertion by deploying a pre-trained temporal decision model; The instruction output unit generates and outputs AI real-time navigation instructions based on the path status judgment results.

2. The gastric tube placement guidance system with AI real-time navigation function according to claim 1, characterized in that, The image acquisition module includes: The intelligent guidewire has an image acquisition unit and an illumination unit at its distal end and a connector at its proximal end. The image acquisition unit and the illumination unit are respectively connected to the connector for communication. The intelligent guidewire is wrapped with a tube, and the material stiffness of the tube gradually decreases from the proximal end to the distal end. A controller, connected to the connector, is used to control the movement of the intelligent guidewire. The controller is communicatively connected to the computing module.

3. A gastric tube insertion guidance system with AI real-time navigation function according to claim 2, characterized in that, The preset anatomical landmarks include: oral cavity, epiglottis, piriform recess, esophageal inlet, esophageal body, cardia, gastric cavity, glottis, and tracheal rings.

4. A gastric tube insertion guidance system with AI real-time navigation function according to claim 3, characterized in that, The path status determination result of the gastric tube insertion includes the target path status and the deviation from the path status; The target path is as follows: the gastric tube is inserted sequentially through the oral cavity, epiglottis, pyriform recess, esophageal inlet, esophageal body, cardia, and gastric cavity. The deviation from the path refers to the gastric tube being inserted into the glottis or tracheal rings after passing through the epiglottis.

5. A gastric tube insertion guidance system with AI real-time navigation function according to claim 4, characterized in that, The instruction output unit also includes a warning function, including: When the sequence of anatomical marker recognition results for a single frame image is determined to be off-path, a Level 1 alarm is triggered. When the sequence of anatomical marker recognition results for multiple frames image is determined to be off-path, a Level 2 alarm is triggered.

6. A gastric tube insertion guidance system with AI real-time navigation function according to claim 4, characterized in that, The temporal decision model can be any one of the following: a lightweight temporal convolutional network, a long short-term memory network, or a rule-based finite state machine. The state transition rule of the rule-based finite state machine is as follows: when the anatomical marker recognition result sequence of a single frame image is the glottis or tracheal ring, the state machine is driven to transition to the off-path state.

7. A gastric tube insertion guidance system with AI real-time navigation function according to claim 4, characterized in that, The calculation and processing module also includes a display unit, which is used to display visual prompts for the path status judgment results; The visual cues include bounding boxes for anatomical landmarks and forward arrows indicating the target path status.

8. A gastric tube insertion guidance system with AI real-time navigation function according to claim 1, characterized in that, The computing processing module also includes a voice unit, which is used to synthesize AI real-time navigation instructions into navigation voice prompts for broadcast.