An interactive teaching demonstration method and system for gastrointestinal endoscopy

By acquiring video streams and keyframe data of gastrointestinal endoscopy procedures for time-series analysis, identifying action deviations, and generating augmented reality error correction guidance, the problem of lack of objective assessment and real-time feedback in gastrointestinal endoscopy training is solved, thereby improving the standardization and safety of training.

CN122090670APending Publication Date: 2026-05-26SECOND AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SECOND AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE
Filing Date
2026-02-10
Publication Date
2026-05-26

Smart Images

  • Figure CN122090670A_ABST
    Figure CN122090670A_ABST
Patent Text Reader

Abstract

This invention discloses a teaching demonstration and interactive method and system for gastrointestinal endoscopy. The method includes: acquiring continuous video streams and keyframe image data of a trainee operating a simulated gastrointestinal endoscopy procedure using an image acquisition device; performing temporal analysis and feature extraction on the continuous video streams and keyframe image data to identify action deviations and delayed steps between the trainee's operating techniques and the standard procedure; generating augmented reality error correction guidance content that integrates real-time video footage based on the action deviations and delayed steps; overlaying the augmented reality error correction guidance content onto the display interface corresponding to the trainee's operating field of view in real time; and adaptively adjusting the presentation intensity and prompting rhythm of the augmented reality error correction guidance content based on real-time video image feedback from the trainee's subsequent operations. Using this invention, objective evaluation, immediate feedback, and personalized guidance of the trainee's operation process can be achieved, improving the standardization, teaching efficiency, and safety of gastrointestinal endoscopy skills training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data interaction technology, specifically a teaching demonstration and interactive method and system for gastrointestinal endoscopy. Background Technology

[0002] Currently, training in gastrointestinal endoscopy skills relies heavily on on-site guidance from instructors and practical experience on mannequins or patients. Traditional methods have several significant shortcomings: First, the teaching process is highly subjective, making it difficult for instructors to objectively and quantitatively assess every detail of the trainee's operation (such as technique, angle, and sequence). Second, feedback is delayed, typically only providing a summary review after the procedure, preventing trainees from immediately correcting errors during the process. Third, training resources are limited; standardized three-dimensional anatomical structures and standardized operating procedures are difficult to integrate intuitively and personally into real-time training scenarios, resulting in low teaching efficiency and potential risks. Summary of the Invention

[0003] The purpose of this invention is to provide an interactive teaching demonstration method and system for gastrointestinal endoscopy to address the shortcomings of existing technologies. This system enables objective evaluation, real-time feedback, and personalized guidance of trainees' operation processes, thereby improving the standardization, teaching efficiency, and safety of gastrointestinal endoscopy skills training.

[0004] One embodiment of this application provides an interactive teaching demonstration method for gastrointestinal endoscopy, the method comprising: The continuous video stream and keyframe image data of the trainee operating the simulated gastrointestinal endoscopy are acquired through image acquisition equipment. Time-series analysis and feature extraction are performed on the continuous video stream and keyframe image data to identify action deviations and delayed steps between the trainee's operation techniques and the standard procedure. Based on the aforementioned action deviations and delayed steps, the system dynamically matches and calls upon a pre-set 3D anatomical model and a standard operation image library to generate augmented reality error correction guidance content that integrates real-time images. The augmented reality error correction guidance content is overlaid on the display interface corresponding to the student's operating field of view in real time, forming a virtual and real combined operation guidance view; Based on real-time video image feedback from the student's subsequent operations, the presentation intensity and prompting rhythm of the augmented reality error correction guidance content are adaptively adjusted until the student's operation matches the standard process to a preset teaching qualification threshold.

[0005] Another embodiment of this application provides a teaching demonstration and interactive system for gastrointestinal endoscopy, the system comprising: The acquisition module is used to acquire continuous video streams and keyframe image data of trainees operating simulated gastrointestinal endoscopy through image acquisition equipment; The recognition module is used to perform time-series analysis and feature extraction on the continuous video stream and key frame image data, and to identify the action deviations and delayed steps between the trainee's operation techniques and the standard procedures. The matching module is used to dynamically match and call a preset three-dimensional anatomical model and standard operation image library based on the action deviation and step lag segment to generate augmented reality error correction guidance content that integrates real-time images. The display module is used to overlay the augmented reality error correction guidance content onto the display interface corresponding to the student's operating field of view in real time, forming a virtual and real combined operation guidance view. The adjustment module is used to adaptively adjust the presentation intensity and prompting rhythm of the augmented reality error correction guidance content based on the real-time video image feedback of the student's subsequent operations, until the matching degree between the student's operation and the standard process reaches the preset teaching qualification threshold.

[0006] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0007] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0008] Compared with existing technologies, the present invention provides an interactive teaching demonstration method for gastrointestinal endoscopy, which enables objective evaluation, real-time feedback and personalized guidance of trainees' operation process, thereby improving the standardization, teaching efficiency and safety of gastrointestinal endoscopy skills training. Attached Figure Description

[0009] Figure 1 A hardware structure block diagram of a computer terminal for a teaching demonstration and interactive method for gastrointestinal endoscopy provided in an embodiment of the present invention; Figure 2 A flowchart illustrating an interactive teaching demonstration method for gastrointestinal endoscopy provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an interactive teaching demonstration system for gastrointestinal endoscopy provided in an embodiment of the present invention. Detailed Implementation

[0010] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0011] The present invention first provides an interactive teaching demonstration method for gastrointestinal endoscopy, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0012] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for an interactive teaching demonstration method for gastrointestinal endoscopy provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0013] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any interactive teaching demonstration method for gastrointestinal endoscopy.

[0014] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0015] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any interactive teaching demonstration method for gastrointestinal endoscopy.

[0016] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0017] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0018] See Figure 2 The present invention provides an interactive teaching demonstration method for gastrointestinal endoscopy, which may include the following steps: S201 acquires continuous video streams and keyframe image data of the trainee operating the simulated gastrointestinal endoscope through image acquisition equipment; Specifically, a multi-angle camera array can be deployed on the simulated gastrointestinal endoscopy operating platform to simultaneously acquire raw video signals from three perspectives: the operator's hand, the movement trajectory of the instrument, and the endoscope simulation display screen, generating multiple synchronous raw video streams. The core of this step is to achieve full-dimensional capture of the operation process through a multi-view camera layout, ensuring that the raw video signal covers the three core dimensions of hand movements, instrument trajectory, and simulated images, providing complete data support for subsequent analysis. The specific implementation method is as follows: The simulated gastrointestinal endoscope operating platform features a modular structure. The camera array is deployed according to functional zones, with a total of 6 cameras forming three independent viewing angle acquisition channels. Two cameras are deployed for the operator's hand view, symmetrically mounted on both sides of the operating platform. The lenses are 30cm away from the hand operating area, with a 60° field of view, a resolution of 1920×1080 pixels, a frame rate of 30fps, and a focal length of 8mm. This ensures clear capture of subtle movements such as finger joint bending and wrist rotation, avoiding obstruction of crucial operational details by the hand. Two cameras are deployed for the instrument movement trajectory view, mounted above and in front of the platform respectively. The upper camera shoots vertically downwards, while the front camera is at a 45° angle to the instrument's movement direction. Both cameras work together to capture the complete trajectory of the gastrointestinal endoscope instrument from insertion to movement. The resolution is also 1920×1080 pixels, but the frame rate is increased to 60fps to meet the capture requirements of rapid instrument movement. The focal length is 12mm, and the depth of field range is 10-50cm, ensuring the instrument tip remains in the clear imaging area at all times. Two cameras are deployed to simulate the view of the endoscope display screen, symmetrically installed on both sides of the screen. The lenses are close to the edge of the screen, parallel to the screen plane, 5cm away from the screen surface, with a field of view of 90°, a resolution of 1920×1080 pixels, a frame rate of 30fps, and a focal length of 5mm. This ensures complete acquisition of anatomical images and operation marks on the simulated display screen. During the acquisition process, autofocus is turned off and the focal length is fixed to avoid screen image distortion.

[0019] All cameras employ a global shutter mode to avoid motion blur caused by rolling shutter. They are also uniformly equipped with infrared fill lights, whose intensity automatically adjusts according to ambient light (range 100-500 lux) to ensure uniform image brightness under different lighting conditions. Acquisition synchronization is achieved through a unified trigger signal. The platform controller generates a 10MHz synchronization pulse, and all cameras simultaneously start acquisition upon receiving the pulse, ensuring that the start time of acquisition for multiple video signals is completely consistent. This generates three independent raw video streams, corresponding to the hand view, instrument trajectory view, and display view, respectively. Each frame of the video stream contains a pixel matrix, an acquisition timestamp (accurate to milliseconds), and a viewpoint identifier. Example of a raw video stream segment for the hand view: timestamp 00:01:02.345, viewpoint identifier "hand - left side," and pixel matrix corresponding to the initial posture of the fingers holding the instrument.

[0020] The multi-channel synchronous raw video streams are time-aligned and encoded for compression. A hardware encoder is used to unify the timestamps and convert them into a standard video format to generate a synchronous standardized video stream. The core of this step is to eliminate timing discrepancies and format differences between multiple video streams, generating a standardized video stream to provide normalized data for subsequent keyframe extraction and timing analysis. The specific implementation method is as follows: Timing alignment is further calibrated based on a hardware synchronization triggering mechanism. By extracting the timestamp of each video stream, linear interpolation is used to correct minute timing deviations (synchronization error controlled within ±1ms). For time differences caused by camera transmission delays, dynamic compensation is achieved through the platform controller's clock synchronization module, ensuring that the timestamps of corresponding frames for actions performed at the same moment are completely consistent across the three video streams. In the example, the timestamp of a frame from the hand perspective is 00:01:02.345, while the timestamp of the corresponding frame from the instrument trajectory perspective is 00:01:02.346. Through interpolation compensation, the timestamp of that frame from the instrument trajectory perspective is corrected to 00:01:02.345, achieving strict timing alignment.

[0021] Encoding compression is achieved using a hardware encoder, balancing compression efficiency and image quality preservation. The encoder supports the H.265 encoding standard, with a compression ratio set to 50:1. While ensuring lossless image quality (peak signal-to-noise ratio ≥40dB), the bitrate of each video stream is controlled within 4Mbps, reducing data storage and transmission pressure. During encoding, differentiated compression parameters are used for video streams from different perspectives: the quantization parameter is set to 20 to retain higher motion details from the hand and equipment trajectory perspectives; the quantization parameter is set to 25 for the display screen perspective due to its relatively stable image, balancing compression efficiency and detail preservation.

[0022] The timestamps are uniformly formatted in UTC, accurate to the microsecond. During encoding, the corrected timestamps are embedded in the frame header information of the video stream, along with metadata such as viewpoint identifier, resolution, and frame rate. The standard video format generated after format conversion is MP4, and the encapsulation method uses ISOBaseMediaFileFormat to ensure compatibility. The final generated synchronized standardized video stream contains three sub-streams, each carrying a uniform microsecond-level timestamp and complete metadata. Example standardized video stream frame header information: timestamp 2024-05-20T14:30:00.123456Z, viewpoint identifier "equipment trajectory - above", resolution 1920×1080, frame rate 60fps, encoding format H.265, ensuring that subsequent processing can accurately associate data from different viewpoints at the same moment.

[0023] Based on a synchronized standardized video stream, a motion saliency detection algorithm is applied to automatically identify landmark operation moments, including instruments entering key anatomical sites, performing biopsies, or rinsing, and generate a sequence of key operation time points. The core of this step is to accurately capture landmark operation moments through algorithms, filter out key time-series nodes with analytical value, and provide a basis for subsequent keyframe extraction. The specific implementation method is as follows: The motion saliency detection algorithm combines inter-frame differencing with a Gaussian mixture model (GMM). First, inter-frame differencing extracts the motion region, then the GMM removes background noise while preserving the true motion characteristics. Inter-frame differencing uses a three-frame differencing method, calculating the difference in grayscale values ​​between the current frame and the previous and next frames. A difference threshold T=30 is set (grayscale value range 0-255; a threshold of 30 effectively distinguishes motion regions from static backgrounds). Regions with grayscale differences greater than T are marked as candidate motion regions. The GMM uses three Gaussian components with a learning rate α=0.001 to model the background pixel distribution. By calculating the matching degree between candidate motion region pixels and the background model, noise interference (such as false motion regions caused by slight light fluctuations) is eliminated, preserving true motion regions such as machine movements and hand gestures, generating the final saliency motion mask.

[0024] The identification of key operations is achieved by matching motion region features with preset templates. There are three types of preset motion feature templates for key operations: when the instrument enters a key anatomical site, the movement trajectory of the instrument tip is linear, the movement area is concentrated in the center of the simulated display screen and the central axis area of ​​the instrument trajectory view, and the movement speed is v=5-10cm / s; during biopsy operations, a pressing action is seen in the hand view, the movement area is concentrated in the finger area, and at the same time, a slight vertical movement is seen in the instrument trajectory view; during rinsing operations, liquid flow texture is seen in the simulated display screen view, the movement area is diffused, and there is no obvious positional change in the instrument trajectory view.

[0025] The algorithm traverses a synchronized, standardized video stream using a sliding window, with the window size set to 10 frames (corresponding to 0.167 seconds, adapted to a 60fps frame rate). For each frame, it calculates the matching degree between the position, shape, and velocity of the moving area and a preset template. A matching degree ≥ 0.8 is considered a landmark operation. In the example, in the instrument trajectory view video stream, the instrument's end-effector is detected advancing linearly along the central axis at a speed of 7 cm / s, with a matching degree of 0.85. Simultaneously, a duodenal anatomical marker appears in the simulated display view, indicating "instrument entering a key anatomical site," and the timestamp 2024-05-20T14:30:15.678901Z is recorded. From the hand view, a finger pressing motion is detected with a matching degree of 0.82, indicating a "biopsy operation," and the timestamp 2024-05-20T14:30:30.123456Z is recorded. All landmark operation moments are arranged chronologically to generate a key operation time point sequence in the format "timestamp-operation type-matching degree," ensuring that the operation attributes of each key moment are clearly defined.

[0026] Based on the key operation time point sequence, the static images of the corresponding frames are extracted from the synchronous standardized video stream and bound and encapsulated with timestamps and viewpoint tags to finally generate keyframe image data packets with spatiotemporal annotations.

[0027] The core of this step is to extract multi-view static images at key moments, achieve data association through spatiotemporal annotation binding, and generate structured data packages to provide precise data units for subsequent action analysis. The specific implementation method is as follows: Keyframe extraction is based on the sequence of key operation time points. For each time point, frames corresponding to the three video streams from different perspectives are extracted. To ensure the complete capture of the operation, additional images are extracted from the frame before and after each key time point, forming a "pre-middle-after" three-frame combination. Each perspective corresponds to three frames, resulting in a set of nine frames across the three perspectives. Precise indexing using timestamps ensures strict correspondence between images from different perspectives at the same key time point, with extraction accuracy controlled within one frame to avoid analytical bias caused by frame misalignment. In the example, the key timestamp is 2024-05-20T14:30:15.678901Z. Frames 1800, 1801, and 1802 are extracted from the hand perspective, 3600, 3601, and 3602 from the instrument trajectory perspective, and 1800, 1801, and 1802 from the display screen perspective, ensuring that the images from the three perspectives correspond to the same instant of operation and the preceding and following states.

[0028] Image binding and encapsulation employs a structured data format, with each image frame associated with complete spatiotemporal annotation information, including a UTC timestamp (accurate to microseconds), viewpoint labels (e.g., "hand - left side", "instrument trajectory - top"), operation type (e.g., "instrument enters key anatomical site"), frame number, image resolution, pixel depth (24-bit RGB), and other metadata. Image data is compressed using JPEG format with a compression quality factor of 90 (ensuring clear image details and a compression ratio of approximately 10:1), and the size of a single image frame is kept below 200KB.

[0029] The keyframe image data packet with spatiotemporal annotations is encapsulated in binary format, comprising four parts: a packet header, metadata area, image data area, and checksum area. The packet header (16 bytes) contains the packet ID, data length, and generation time; the metadata area (256 bytes) stores the spatiotemporal annotation information for each frame, arranged in viewpoint order; the image data area stores the compressed data of 9 frames, encapsulated in the order of "hand-device trajectory-display screen"; the checksum area (4 bytes) uses the CRC-32 checksum algorithm to calculate the checksum values ​​of the header and metadata area, ensuring no damage during packet transmission and storage. Example packet: ID K001, generation time 2024-05-20T14:30:16.000000Z, containing 9 frames and corresponding annotations, checksum 0x12345678, total size approximately 1.8MB, can be directly used for subsequent motion feature extraction and deviation analysis.

[0030] S202, perform time-series analysis and feature extraction on the continuous video stream and key frame image data to identify the action deviation and step lag segments between the trainee's operation method and the standard procedure. Specifically, keyframe image data packages with spatiotemporal annotations can be loaded, and the coordinates of the trainee's hand joints, the pose of the device's end effector, and the direction vector of motion can be extracted through a pre-trained pose estimation network to generate a sequence of operational action features. The core of this step is to complete data packet parsing and accurate extraction of action features, transforming visualized image data into quantitative features to provide a data foundation for subsequent deviation analysis. The specific implementation method is as follows: When loading a data packet, an integrity check is performed first. This is done by parsing the CRC-32 checksum at the end of the data packet and comparing it to a locally recalculated checksum. If they match, the data packet is considered undamaged; otherwise, a reloading mechanism is triggered (up to two reloads). After successful verification, the data packet is parsed in the order of "header-metadata area-image data area," extracting spatiotemporal annotation information such as the UTC timestamp, viewpoint label, and operation type for each frame of image. Simultaneously, JPEG format image data is decoded to reconstruct a 24-bit RGB static image, maintaining a resolution of 1920×1080 pixels. In the example, loading the data packet with ID K001 yields 9 frames of images, covering three views: hand, instrument trajectory, and display screen, corresponding to the "instrument entering key anatomical site" operation. The timestamp range is from 2024-05-20T14:30:15.678901Z to 2024-05-20T14:30:15.685234Z.

[0031] The pre-trained pose estimation network uses the lightweight HRNet model, balancing extraction accuracy and speed, with inference latency controlled within 30ms, adapting to real-time teaching needs. The network is jointly pre-trained on the COCO dataset and a medical operation-specific dataset, optimized for gastrointestinal endoscopy scenarios, and can accurately identify 21 hand joints (including 5 joints of the thumb and 4 joints each of the other four fingers) and key points at the ends of gastrointestinal endoscopes. During extraction, the hand-view image is used as the core, combined with complementary calibration using the instrument trajectory-view image to avoid feature loss due to hand occlusion.

[0032] The hand joint coordinates are based on the image pixel coordinate system, with the origin set at the top left corner of the image. The x-axis points horizontally to the right, the y-axis points vertically downward, and the z-axis represents the joint depth (calculated based on binocular vision parallax, in mm). In the example, the coordinates of the student's thumb fingertip joint are extracted as (890, 540, 25), the index finger middle joint is (920, 580, 23), and the wrist joint is (780, 650, 30), with a coordinate accuracy of ±1 pixel and a depth accuracy of ±2 mm. The device end-effector pose is described using six degrees of freedom parameters, including three-dimensional position coordinates (x, y, z) and three-dimensional attitude angles (roll, pitch, yaw). The attitude angles are calculated based on the angle between the device axis and the image coordinate system, in degrees. In the example, the device end-effector position coordinates are (1050, 480, 40), and the attitude angles are roll = 5.2°, pitch = 3.1°, and yaw = 1.8°, representing the current spatial attitude of the device.

[0033] The motion direction vector is calculated from the difference between the joint coordinates and the end-effector coordinates of adjacent frames, using the formula V=(x_k-x_k-1,y_k-y_k-1,z_k-z_k-1), with units of pixels / frame (the spatial direction vector is synchronously converted to mm / frame). In the example, the x-coordinate of the end-effector changes from 1045 to 1050, the y-coordinate from 478 to 480, and the z-coordinate from 38 to 40 in two adjacent frames, resulting in a motion direction vector of (5,2,2) pixels / frame, corresponding to a spatial vector of (2.5,1.0,1.0) mm / frame. This represents the propulsive motion of the device, primarily along the positive x-axis and secondarily along the positive y and z-axis.

[0034] The sequence of operation action features is organized in time stamp order, with each time point corresponding to a set of feature vectors. The vectors are 72-dimensional (21 hand joints × 3D coordinates + 6D pose of the instrument end effector + 3D motion direction vector). They are also associated with viewpoint labels and operation types. Example feature sequence fragment: timestamp 2024-05-20T14:30:15.678901Z, operation type "instrument propulsion", feature vector [890,540,25,...,1050,480,40,5.2,3.1,1.8,5,2,2], ensuring that the features strictly correspond to the operation time sequence.

[0035] The operation action feature sequence is dynamically time-aligned with the expert operation template in the standard teaching process database, and the difference in motion trajectory, speed and angle of each time segment is calculated to generate a time difference map. The core of this step is to eliminate the time scale differences between trainees and experts through time alignment, quantify the deviations in action details, and form a visual difference map. The specific implementation method is as follows: The standard teaching process database stores multiple expert operation templates. Each template corresponds to a feature sequence of a complete gastrointestinal endoscopy procedure. The operation process is recorded by at least three associate chief physicians and generated after extraction by a posture estimation network. The template includes information such as the motion trajectory baseline, velocity threshold range, and standard posture angle values. For the "instrument entering key anatomical sites" operation, the corresponding expert template was called. The template is 8 seconds long and contains feature vectors of 240 time-series nodes. Since the template's duration differs from the trainee's feature sequence (3 seconds, 90 time-series nodes), alignment is achieved through Dynamic Time Warping (DTW).

[0036] The Dynamic Time Warping algorithm solves the problem of inconsistent time series lengths by constructing a distance matrix to find the optimal alignment path between the trainee feature sequence and the expert template. The elements of the distance matrix are the Euclidean distances between the feature vectors of the corresponding time series nodes, expressed as d(i,j)=√[Σ(x_ij-y_ij)]. 2 (where x_ij is the j-th feature of the i-th node of the trainee, and y_ij is the j-th feature of the i-th node of the expert template). The smaller the distance, the more similar the actions. The algorithm constraint is set to a local slope range of [0.5, 2] to avoid excessive distortion of the alignment path. A dynamic programming strategy is used during the alignment process to minimize the total distance cost. The total distance cost formula is D = Σd(i,j) × w_j (where w_j is the weight of the j-th feature, the joint coordinate and pose weights are set to 0.6, and the motion vector weight is set to 0.4). In the example, after aligning 90 nodes of the trainee feature sequence with 240 nodes of the expert template, 90 optimal matching pairs are obtained, with a total distance cost D = 128.5, providing a basis for subsequent difference calculation.

[0037] The difference calculation is carried out separately for three dimensions: motion trajectory, speed, and angle. The weighted sum is used to obtain the comprehensive difference score. The weights are allocated as follows: trajectory 0.4, speed 0.3, and angle 0.3. The difference score range is [0,1], where 0 indicates complete consistency and 1 indicates extreme difference. The motion trajectory difference score is calculated by the Euclidean distance between the end-effector coordinates of the corresponding nodes of the trainee and the expert, and normalized to the [0,1] interval. The formula is D_t=d_t / d_max (d_t is the distance of the current node, and d_max is the preset maximum distance threshold of 50mm). In the example, the end-effector coordinates of the trainee at a certain node are (1050,480,40), and the expert template coordinates are (1048,475,38). The distance d_t=√[(2)] 2 +(5) 2 +(2) 2 The trajectory difference is approximately 5.7 mm, and the trajectory difference is D_t = 5.7 / 50 = 0.114.

[0038] The speed difference is calculated based on the magnitude of the motion direction vector, where the student's speed is v_s = √(V_x). 2 +V_y 2 +V_z 2 The expert standard speed v_e is taken as the mean of the corresponding interval of the template. The speed difference D_v = |v_s - v_e| / v_e_max (v_e_max is the preset maximum speed threshold of 20mm / frame). In the example, the student speed v_s = √(2.5) 2 +1.0 2 +1.0 2 The speed difference is approximately 2.92 mm / frame, the expert standard speed is v_e = 2.5 mm / frame, and the speed difference is D_v = |2.92-2.5| / 20≈0.021. The angle difference is calculated by the average deviation of the attitude angle, D_a = (|roll_s-roll_e|+|pitch_s-pitch_e|+|yaw_s-yaw_e|) / (3×θ_max) (θ_max is the maximum deviation threshold of a single angle, 10°). In the example, the deviations of the student's attitude angle from the expert's are 0.2°, 0.1°, and 0.3°, respectively. The angle difference is D_a = (0.6) / (30) = 0.02, and the comprehensive difference is D = 0.4×0.114+0.3×0.021+0.3×0.02≈0.056.

[0039] The temporal difference map uses time as the horizontal axis (in seconds) and overall difference as the vertical axis (range [0,1]). It also labels the single-dimensional difference curves for trajectory, velocity, and angle, and adds a difference threshold line (preset threshold 0.3) to visually present the deviation of each time segment. In the example map, the overall difference for 0-1.5 seconds remains between 0.05 and 0.08, close to expert level; the difference for 1.5-2.0 seconds rises to 0.25, mainly due to increased velocity deviation; the difference for 2.0-3.0 seconds falls back to below 0.1. The map uses colored curves to distinguish different dimensions: red for overall difference, blue for trajectory difference, green for velocity difference, and yellow for angle difference, facilitating quick location of the source of deviation.

[0040] Based on the temporal difference map, an abnormal segment detection algorithm is used to identify continuous time periods that exceed a preset threshold. These continuous time periods correspond to operational errors or instrument position deviations, and preliminary deviation segment identifiers are generated. The core of this step is to accurately capture abnormal deviation segments using algorithms, eliminate interference from accidental fluctuations, and clarify the initial range and type of deviation. The specific implementation method is as follows: The abnormal segment detection algorithm employs a sliding window-based thresholding method, balancing detection accuracy and real-time performance. The sliding window size is set to 10 frames (corresponding to 0.33 seconds, adapted to a 30fps frame rate), and the window step size is set to 1 frame (0.033 seconds) to ensure no detection blind spots. A preset difference threshold of 0.3 is used. When the overall difference of all temporal nodes within the window exceeds the threshold, and three consecutive windows meet this condition, the segment is identified as abnormal, preventing accidental fluctuations in a single frame (such as feature extraction deviations caused by lighting interference) from being misjudged as abnormal.

[0041] During algorithm execution, the temporal difference map is first smoothed using a 5-point moving average method to filter out noise and make the difference curve smoother. The preprocessing formula is D_smooth(k)=(D(k-2)+D(k-1)+D(k)+D(k+1)+D(k+2)) / 5 (where k is the current frame number, and edge frames are padded with zeros). In the example, the original difference of a certain frame is 0.32, which is smoothed to 0.31 after preprocessing. This still exceeds the threshold, ensuring that abnormal segments are not masked by noise.

[0042] After identifying abnormal segments, the core information of the segments is extracted, including the start timestamp, end timestamp, duration, average difference, and main deviation dimension, generating a preliminary deviation segment identifier. In the example, one abnormal segment was detected, with a start timestamp of 2024-05-20T14:30:16.234567Z, an end timestamp of 2024-05-20T14:30:17.012345Z, a duration of 0.778 seconds, and an average comprehensive difference of 0.38. Through single-dimensional difference analysis, it was found that the average trajectory difference was 0.45, which is the main deviation dimension. Therefore, the segment was determined to correspond to the deviation type of "instrument position deviating from the correct anatomical path".

[0043] To further verify the authenticity of the deviation, cross-checking was performed using images from both the instrument trajectory perspective and the display screen perspective: the instrument trajectory perspective image showed that the instrument tip deviated from the central axis by approximately 8mm, exceeding the standard path range (±5mm); the display screen perspective image showed that the instrument marker in the simulated anatomical image deviated from the duodenal inlet region, consistent with the algorithm's recognition result, thus ruling out false anomalies. If multiple consecutive abnormal windows were detected but no actual deviation was found in the image verification, it was determined to be a feature extraction error. The threshold was automatically adjusted and the detection was repeated to ensure the accuracy of the deviation marker.

[0044] Preliminary deviation segment identification is recorded in a structured format, including segment ID, spatiotemporal range, deviation type, average difference, main deviation dimension, and verification result. Example identification content: ID is E001, start time is 2024-05-20T14:30:16.234567Z, end time is 2024-05-20T14:30:17.012345Z, deviation type is "instrument position deviation", average difference is 0.38, main deviation dimension is "motion trajectory", and verification result is "true deviation", providing a foundation for subsequent accurate calibration.

[0045] By combining the time requirements of the standard process at each stage, we analyze the completion time of key steps before and after the initial deviation segment identification, determine whether there are problems such as disordered execution order of steps or excessive time consumption, and finally accurately identify the action deviation and delayed steps.

[0046] The core of this step is to integrate time-dimensional constraints, correct initial deviation indicators, accurately distinguish between action deviations and step lags, and form a complete deviation calibration result. The specific implementation method is as follows: The standard procedure's time requirements for each stage are based on statistical analysis of expert operation templates. The gastrointestinal endoscopy procedure is divided into multiple key stages, each with a minimum, standard, and maximum time threshold. Exceeding the maximum time threshold is considered a delayed step, and deviations from the standard procedure's execution order are considered disordered. For the "instrument entry into key anatomical sites" stage, the standard time is 5-8 seconds, with a minimum of 4 seconds and a maximum of 10 seconds. The preceding step is "instrument insertion into the esophagus" (standard time 3-5 seconds), and the subsequent step is "instrument positioning in the pylorus" (standard time 2-3 seconds). The time nodes for each step are linked to the operation type, forming a standard timeline.

[0047] First, the completion times of key steps before and after the initial deviation segment were analyzed. The start and end timestamps of each stage in the trainee's operation were extracted and compared with the standard timeline. In the example, the trainee's "instrument insertion into the esophagus" stage took 4.5 seconds (meeting the standard). The "instrument entering the key anatomical site" stage started at 2024-05-20T14:30:12.000000Z and ended at 2024-05-20T14:30:21.500000Z, taking 9.5 seconds, which is close to the maximum time threshold of 10 seconds. The initial deviation segment (0.778 seconds) was in the middle of this stage, mainly due to the slowdown caused by the deviation of the instrument position.

[0048] When a step is judged to be lagging, the difference between the stage time and the standard time is calculated using the formula ΔT = T_actual - T_standard (where T_actual is the student's actual time, and T_standard is the average standard time of 6.5 seconds). If ΔT ≥ 2 seconds and does not exceed the maximum time threshold, it is judged as a slight lag; if ΔT ≥ 3 seconds and exceeds the maximum time threshold, it is judged as a severe lag. In the example, the student's stage time was 9.5 seconds, ΔT = 3 seconds, which does not exceed the maximum time of 10 seconds, and is judged as a slight lag. The cause of the lag is directly related to the deviation of the instrument position in the initial deviation segment. The deviation leads to a decrease in operational efficiency and an increase in time.

[0049] When determining the execution order of steps, the sequence of key operation time points is compared between the trainee's steps and the standard procedure. If the start time of a subsequent step in the trainee's operation is earlier than the end time of a previous step, or if the logical relationship between steps is inconsistent with the standard, it is determined to be an incorrect sequence. In the example, the trainee started the "instrument positioning of pylorus" step only after the "instrument enters key anatomical site" stage was completed, which is consistent with the standard and there is no issue with incorrect sequence.

[0050] Final accurate calibration requires integrating motion deviation and step lag information, correcting initial deviation markers, and supplementing information such as the degree of lag, cause of lag, and correlation between deviation and lag. Example calibration results: Motion deviation segment ID is E001, spatiotemporal range 2024-05-20T14:30:16.234567Z to 2024-05-20T14:30:17.012345Z, deviation type "instrument position deviates from correct anatomical path", average difference 0.38, main deviation dimension "motion trajectory"; step lag is slight, lag duration 3 seconds, cause of lag "instrument position deviation leads to increased adjustment time", associated with deviation segment E001; no step sequence disorder issues.

[0051] Simultaneously, the deviation segments are graded and labeled according to their overall difference degree, into three levels: slight deviation (0.3≤D<0.5), moderate deviation (0.5≤D<0.7), and severe deviation (D≥0.7). In the example, the average difference degree of the deviation segments is 0.38, and it is labeled as slight deviation. The final deviation calibration results need to be associated with the corresponding keyframe images to facilitate the subsequent generation of targeted error correction guidelines, ensuring that the calibration results are accurate and complete, and providing a clear basis for teaching guidance.

[0052] S203, based on the aforementioned action deviation and step lag segments, dynamically match and call the preset three-dimensional anatomical model and standard operation image library to generate augmented reality error correction guidance content that integrates real-time images; Specifically, it can analyze the type and location of action deviations and step delays, call the corresponding three-dimensional models of gastrointestinal segments from the three-dimensional anatomical model library, highlight the correct anatomical position of the current instrument, and generate a three-dimensional spatial error correction model. The core of this step is to accurately locate the core information of the deviation, match it with the corresponding anatomical model, and strengthen the guidance of the correct position, so as to provide a spatial benchmark for the subsequent generation of error correction content. The specific implementation method is as follows: The analysis of movement deviations and step lag segments requires combining the results of previous calibration to extract core information, including the type of deviation, the anatomical location of occurrence, the degree of deviation, and the characteristics of lag. Deviation types are categorized into three types based on the operational dimension: instrument position deviation (e.g., deviation from the anatomical path, improper depth), operational angle deviation (e.g., instrument posture angle exceeding the standard range), and manual movement deviation (e.g., incorrect hand force application). Step lag is categorized into slight lag (time exceeding the standard by 2-3 seconds), moderate lag (3-5 seconds), and severe lag (≥5 seconds), with the cause of lag correlated with the deviation type. In the example, the previously calibrated deviation was "slight instrument position deviation," occurring in the duodenal inlet region, with a deviation distance of 8mm, accompanied by slight step lag (lag of 3 seconds). The cause of the lag was the increased adjustment time due to instrument deviation.

[0053] The 3D anatomical model library is constructed according to gastrointestinal segments, covering key parts such as the esophagus, stomach, duodenum, and small intestine. Each segment model is reconstructed based on real human anatomical data with an accuracy of 0.1mm. It includes details such as mucosal texture, blood vessel distribution, and anatomical landmarks (e.g., pylorus and duodenal papilla). The models support scaling, rotation, and local highlighting. Based on the location of the deviation, "duodenal entrance," the corresponding 3D model of the duodenum and gastric antrum segment is called. The model format is a general 3D model format, including vertex, texture, and topological data. The loading time is ≤50ms, adapting to real-time teaching needs.

[0054] The highlighting of the correct anatomical location employs a multi-level rendering strategy. First, the anatomical landmark at the duodenal inlet is located (coordinates are based on the model coordinate system, with the origin at the center of the pylorus of the stomach, the x-axis pointing towards the duodenum, the y-axis perpendicular to the gastrointestinal axis, and the z-axis horizontal). In the example, the correct location coordinates are (20mm, 5mm, 3mm). The highlighting method is divided into three levels: the bottom layer is a semi-transparent red glowing area (70% transparency, 1.2 times the luminous intensity), covering a 5mm radius around the correct location to indicate the location range; the middle layer is a white outline (2mm wide, flashing 1 time / second), outlining the correct path along the gastrointestinal wall; the top layer is a yellow marker (3mm in diameter, statically displayed), precisely marking the core location that the instrument's end should reach. Simultaneously, details in the model unrelated to the current operation (such as the distal small intestine structure) are hidden to reduce visual interference.

[0055] The generated 3D spatial error correction model integrates segmented models and highlighted elements, and adds coordinate system calibration information to ensure that the model is consistent with the simulated image perspective in the trainee's field of vision. The model scaling ratio matches the screen ratio of the endoscope simulation display. In the example, the model scaling factor is 1.5, so that the duodenal inlet area fills the core area of ​​the screen, and the highlighted elements are always in the visual focus, making it easy for trainees to quickly identify the correct position.

[0056] Based on the type of deviation, a matching short video clip of expert standard operation is retrieved from the standard operation video library. This clip demonstrates the corrective action from the current error state to the correct state, and a standard corrective action video is generated. The core of this step is to accurately retrieve targeted expert operation clips to provide trainees with intuitive references for action correction. The clips must be relevant to the current deviation state, achieving a complete demonstration of "error-correction-correction". The specific implementation method is as follows: The standard operating procedure image library stores short videos of procedures recorded by at least three associate chief physicians, covering common deviation correction scenarios throughout the entire gastrointestinal endoscopy process. Each segment is 3-5 seconds long, with a resolution of 1920×1080 pixels, a frame rate of 30fps, and uses H.265 encoding format with a bitrate of 4Mbps. Segments are categorized and labeled according to "deviation type - anatomical location - correction technique," and keyword tags are added (e.g., "instrument position deviation - duodenum - fine-tuning technique," "angle deviation - pylorus - rotation adjustment"). Simultaneously, motion feature vectors (consistent with the output format of the previous pose estimation network) are extracted from the segments for similarity matching.

[0057] The retrieval employed a dual strategy of "keyword matching + feature similarity calculation." First, based on the deviation type "instrument position deviation" and the anatomical location "duodenum," 10 candidate segments were retrieved. All candidate segments contained corrective actions after the instrument deviated from the duodenal inlet. Feature similarity calculation used the cosine similarity algorithm, with the formula cosθ=(A・B) / (||A||×||B||), where A is the feature vector of the student's deviation action (instrument end pose, movement direction vector), and B is the feature vector of the starting frame of the candidate segment's corrective action. The similarity value ranged from [0,1], with values ​​closer to 1 indicating greater similarity in action states. The retrieval threshold was set to 0.85.

[0058] In the example, the feature vector A of the student's deviation action is [1050,480,40,5.2,3.1,1.8,5,2,2] (the coordinates of the instrument's end point, attitude angle, and motion vector). The feature vector B of the starting frame of the candidate segment numbered S012 is [1048,478,39,5.0,3.2,1.9,4.8,2.1,2.2]. The calculated cosine similarity cosθ = 0.92, which exceeds the retrieval threshold, thus it is determined to be the optimal matching segment. This segment demonstrates how an expert, after discovering that the instrument has deviated from the entrance of the duodenum, corrects the movement by slightly rotating the wrist to the left (2.5°), slowly retracting the instrument by 3mm, and then precisely advancing it. The correction time from the incorrect state to the correct position is 1.2 seconds, with a moderate range of motion, suitable for students to imitate.

[0059] When generating standard corrected motion footage, the retrieved segments are preprocessed, irrelevant elements (such as the background other than the expert's hands) are cropped, the core operational areas are preserved, and brightness and contrast are adjusted to make the motion details clearer (brightness increased by 10%, contrast increased by 15%). Simultaneously, timeline markers are added to indicate key nodes in the corrected motion (e.g., "0.2 seconds: retract the device," "0.8 seconds: rotate and adjust the angle," "1.2 seconds: precisely advance to the correct position"), facilitating trainees' understanding of the motion sequence. The preprocessed segment is 1.8 seconds long and approximately 1.2MB in size, and can be directly used for subsequent blending and rendering.

[0060] The three-dimensional spatial error correction model is fused and rendered with the standard corrected motion image. Augmented reality overlay technology is used to synthesize the three-dimensional arrow guidance, the correct path outline and the standard motion animation into the same visual layer to generate a basic augmented reality guidance layer. The core of this step is to organically combine virtual models, motion graphics, and guidance elements through fusion rendering, constructing a clear and well-defined basic layer to ensure that trainees can both see the spatial position and master the corrective actions. The specific implementation method is as follows: The fusion rendering employs a layer-based rendering strategy. The bottom layer is a 3D spatial error-correcting model (duodenal segment model and highlighted areas), the middle layer is standard corrected motion images (semi-transparently embedded), and the top layer is augmented reality guidance elements (3D arrows, path outlines). The transparency of each layer is adjusted hierarchically to avoid mutual occlusion (model layer transparency 80%, motion image transparency 60%, guidance element transparency 90%). A low-latency rendering engine is used, with a rendering latency of ≤20ms to ensure real-time performance. It supports synchronous viewpoint adjustment; when the student rotates the endoscopic simulator, the layer viewpoint changes synchronously, always consistent with the operating field of view.

[0061] The 3D arrows indicate the direction of movement of the instrument. Based on the standard correction motion trajectory settings, there are two arrows, corresponding to the "retract" and "advance" actions respectively. The retract arrow is blue, 15mm long, with a 30° arrowhead angle, pointing backward along the gastrointestinal axis (opposite to the current direction of instrument movement), and its animation effect is a uniform flashing (1 time / second), labeled "Retract 3mm". The advance arrow is green, 15mm long, with a 30° arrowhead angle, pointing to the correct position at the duodenal inlet, and its animation effect is a gradual extension (extending from the starting point to the correct position, lasting 0.5 seconds), labeled "Advance to Marker Point". The arrow positions are calibrated based on the 3D model coordinate system to ensure precise correspondence with the current position of the instrument's end. In the example, the retract arrow is located 3mm in front of the instrument's end, and the advance arrow is located 5mm in front of the correct position.

[0062] The correct path outline is drawn on the inner wall of the gastrointestinal tract in the 3D anatomical model. It is a solid white line, 2mm wide and 50mm long, extending from the current position of the instrument to the correct position. The outline is illuminated (1.1 times the intensity) to enhance visual visibility. The curvature of the outline matches the natural arc of the gastrointestinal tract, conforming to the anatomical structure and avoiding misleading the trainee's operating path. Simultaneously, a dynamic effect is added to the outline, gradually illuminating from the correct position to the current position of the instrument (lasting 0.8 seconds) to guide the trainee's attention to the path's direction.

[0063] When embedding standard corrected motion images, a picture-in-picture mode is used, placed in the upper right corner of the display screen (1 / 4 the size of the screen, 480×270 pixels resolution). The image view is synchronized with the 3D model's perspective. When the trainee adjusts their field of vision, the image view adjusts accordingly, ensuring that the motion demonstration corresponds to the spatial position. The video playback uses a loop mode to continuously demonstrate the corrected motion, while a semi-transparent mask is added to only display the expert's hands and the equipment area, minimizing background interference.

[0064] After the fusion rendering is completed, a basic augmented reality guidance layer is generated. The layer contains four main elements: 3D model, motion image, arrow guidance, and path outline. The elements work together in sequence (arrow flashing, outline lighting, and image looping are synchronized), with no visual flickering or misalignment. The guidance logic is clear, providing a basic framework for subsequent customized adjustments.

[0065] Depending on the severity and type of the deviation, dynamic text prompts, key step numbers, or warning labels are overlaid on the basic augmented reality guidance layer, and the transparency and flashing frequency of the visual elements are adjusted to ultimately generate customized augmented reality error correction guidance content.

[0066] The core of this step is to customize the guidance intensity based on the deviation level. Slight deviations result in simplified prompts, while moderate and severe deviations require stronger guidance. This ensures that the prompts accurately match the learner's needs, neither interfering with the operation nor hindering error correction. The specific implementation method is as follows: Different customization strategies are assigned to different levels of deviation severity. Mild deviation (overall difference score 0.3-0.5) uses "basic guidance + concise prompts," moderate deviation (0.5-0.7) uses "enhanced guidance + detailed prompts," and severe deviation (≥0.7) uses "highlighted guidance + warning prompts + voice assistance." The deviation in the example is mild, so it is customized according to the mild strategy, focusing on optimizing the parameters of the visual elements and the conciseness of the text prompts.

[0067] Dynamic text prompts are overlaid at the bottom center of the display screen, using bold white font (16pt) with a semi-transparent black overlay (50% transparency) to avoid conflict with the screen. The text content corresponds to the type of deviation and the corrective action, and is concise and clear. An example prompt is "Adjust the instrument slightly to the left by 2.5°, retract 3mm, and then advance it to the duodenal inlet." The text uses a word-by-word animation (display speed 0.1 seconds / word), and also includes a disappearing animation (the text fades out after the corrective action is completed, lasting 0.5 seconds). For moderate deviations, the text prompt will explain the principle of the action (e.g., "Fine-tuning the angle can avoid damage to the mucosa"); for severe deviations, a warning text (in red font, such as "The deviation is too large, do not force advancement!") is added.

[0068] Key steps are numbered sequentially according to the correction action, using circular markers (12mm in diameter, yellow background, black text) labeled "1", "2", and "3", corresponding to the three key nodes of the correction action: number 1 is located next to the retraction arrow, labeled "Retraction 3mm"; number 2 is located next to the end of the device, labeled "Turn 2.5° to the left"; and number 3 is located next to the correct position marker, labeled "Advance Positioning". The step numbers light up sequentially as the correction action progresses (step 2 lights up after step 1 is completed, and so on), guiding trainees to operate in sequence. Minor deviations only require marking the core steps (up to 3), while moderate and severe deviations require marking the detailed steps (up to 5).

[0069] Visual element parameter adjustments have been optimized for minor deviations, balancing guidance clarity and operational visibility: the transparency of highlighted areas in the 3D model has been reduced from 70% to 60% to avoid obscuring the simulated image; the arrow flashing frequency has been reduced from 1 time / second to 0.8 times / second to reduce visual interference; and the path outline's luminous intensity has been adjusted from 1.1 times to 1.0 times, maintaining legibility without being glaring. For moderate deviations, transparency is reduced by 10%, the flashing frequency is increased to 1.2 times / second, and the luminous intensity is increased to 1.2 times. For severe deviations, transparency is reduced by 20%, the flashing frequency is increased to 1.5 times / second, a red warning sign (triangular border with an exclamation mark inside) is added, and a voice prompt is triggered (e.g., "Please note, the instrument has deviated from the correct path; adjust immediately").

[0070] The final customized augmented reality error correction guidance content integrates all elements and performs overall calibration to ensure precise alignment between virtual elements and the anatomical images on the simulated display screen, with no misalignment (alignment error ≤ 1 pixel), rendering latency ≤ 25ms, and adaptability to real-time teaching interaction needs. In the example, the guidance content includes a duodenal segment model, blue retraction arrows, green advancement arrows, white path outlines, a looping correction motion image in the upper right corner, dynamic text prompts below, and three key step numbers. The overall visual coordination is harmonious, the guidance is clear, and it can be directly overlaid into the student's field of vision.

[0071] S204, The augmented reality error correction guidance content is superimposed on the display interface corresponding to the student's operating field of view in real time to form a virtual and real combined operation guidance view; Specifically, it can obtain the screen coordinate parameters and spatial orientation data of the endoscope simulation display screen corresponding to the student's current operating field of view, and generate display interface spatial registration information; The core of this step is to accurately collect the spatial attribute data of the display screen, establish a spatial correlation benchmark between the virtual guide and the real screen, and provide data support for subsequent viewpoint matching and overlay rendering. The specific implementation method is as follows: The screen coordinate parameters of the endoscope simulation display are obtained through a combination of the screen's built-in sensors and external calibration tools. Key parameters include physical resolution, screen size, pixel density, and coordinate origin position. The physical resolution is set to 1920×1080 pixels by default, with a pixel depth of 24-bit RGB to ensure display accuracy. The screen's physical size is 50cm×28cm (width×height), calibrated using a laser rangefinder with an error controlled within ±0.1cm. The pixel density is calculated as 1920 pixels / 50cm = 38.4 pixels / cm, representing the number of pixels per centimeter of screen length, used for subsequent size scaling calculations. The coordinate origin is defined as the top-left corner of the screen, establishing a two-dimensional screen coordinate system with the x-axis horizontally to the right and the y-axis vertically downwards, with the coordinate unit being pixels.

[0072] Orientation data is acquired via an attitude sensor mounted on the back of the display screen. The sensor supports measuring three-dimensional Euler angles (roll, pitch, yaw) with an accuracy of ±0.05° and a sampling frequency of 30Hz, consistent with the video capture frame rate to ensure timing synchronization. The roll angle represents the screen's rotation around the x-axis (horizontal tilt), the pitch angle represents the rotation around the y-axis (vertical tilt), and the yaw angle represents the rotation around the z-axis (horizontal deflection). In the example, the display screen is horizontally positioned during the student's operation, and the acquired attitude data are roll=0°, pitch=0°, and yaw=5°. yaw=5° indicates that the screen is tilted 5° to the right, and this viewing angle deviation needs to be corrected in subsequent transformations.

[0073] To eliminate inherent sensor errors, a checkerboard calibration method was used for data calibration. A 10×8 checkerboard (each square 2cm on each side) was printed and affixed to the display screen surface. Calibration images were captured using a camera, and the deviation between the sensor's measured values ​​and the actual values ​​was calculated to generate calibration coefficients. In the example, after calibration, a deviation of +0.1° in the yaw angle measurement was found. Using the calibration formula yaw_calibrated = yaw_measured - 0.1°, the actual yaw angle was corrected to 4.9°, ensuring the accuracy of the spatial orientation data.

[0074] The registration information displayed on the interface is ultimately integrated into structured data, containing three main categories: coordinate system parameters, screen physical properties, and calibrated attitude data. Example registration information: coordinate origin (0,0) pixels, x-axis horizontal to the right, y-axis vertical downwards; screen resolution 1920×1080, physical size 50×28cm, pixel density 38.4 pixels / cm; calibrated attitude angles roll=0°, pitch=0°, yaw=4.9°; registration timestamp 2024-05-20T14:30:22.123456Z. This information is updated in real-time at a frequency of 30Hz to ensure it always matches the spatial state of the student's current operating field of view.

[0075] Based on the spatial registration information of the display interface, the customized augmented reality error correction guidance content is subjected to perspective transformation and size scaling to make it completely match the perspective and proportion of the anatomical image on the display screen, and the guidance content after spatial registration is generated. The core of this step is to eliminate perspective deviation through geometric transformation, adjust the size of the guiding content, and achieve precise alignment between virtual elements and real anatomical images, avoiding problems such as perspective misalignment or proportional distortion. The specific implementation method is as follows: Perspective transformation is based on attitude angle data from spatial registration information and is achieved by constructing a perspective projection matrix. The matrix parameters include camera intrinsic and extrinsic parameters. The camera intrinsic parameters are calculated from the display screen size and pixel density, while the extrinsic parameters are determined by the attitude angles (roll, pitch, yaw) and the camera position. The formula for the perspective projection matrix is ​​P=K×[R|t], where K is the camera intrinsic parameter matrix (including focal length and principal point coordinates), R is the rotation matrix (obtained from the attitude angle transformation), and t is the translation vector (representing the relative position of the camera and the display screen, fixed at a distance of 10cm from the center of the screen to the camera).

[0076] In the example, the camera intrinsic parameter matrix K has a focal length f_x = 38.4 pixels / cm × 10cm = 384 pixels, f_y = 384 pixels (horizontal and vertical focal lengths are consistent), and principal point coordinates (u_0, v_0) = (960, 540) pixels (screen center); the rotation matrix R is obtained by converting yaw = 4.9°, roll and pitch are both 0°, and the matrix elements correspond to the sine and cosine values ​​of the angles; the translation vector t = (0, 0, 10) cm. The 3D coordinates of the customized augmented reality guidance content are substituted into the perspective projection matrix to calculate the transformed 2D screen coordinates, correcting the viewing angle deviation caused by screen deflection. For example, the original coordinates of the 3D arrow (20mm, 5mm, 3mm) are transformed to correspond to screen coordinates (1020, 480) pixels, ensuring that the arrow's direction is consistent with the anatomical location of the duodenal inlet on the display screen.

[0077] Size scaling is based on the screen pixel density and the scale of the anatomical image. First, key dimensions of the real-time anatomical image on the display (such as the display width of the duodenal inlet) are extracted and compared with the actual dimensions of the 3D anatomical model to calculate the scaling factor. In the example, the display width of the duodenal inlet in the anatomical image is 5cm (corresponding to 192 screen pixels), while the actual width of the duodenal inlet in the 3D model is 2cm. The scaling factor k = 5cm / 2cm = 2.5. All elements of the customized guidance content are enlarged by a scaling factor of 2.5 to match the size of the 3D model and arrow guidance with the scale of the anatomical image on the screen. During scaling, the element proportions are kept constant to avoid stretching and distortion. For example, the original length of the 3D arrow is 15mm, and after scaling, its length is 37.5mm (corresponding to 144 screen pixels), coordinating with the size of the anatomical image.

[0078] To ensure transformation accuracy, a feature point matching verification method was employed. Three key feature points (the center of the duodenal inlet, the display position of the instrument tip, and the pyloric landmark) were selected. Their screen coordinates in the anatomical image and their corresponding coordinates in the transformed guidance content were extracted. The coordinate deviation was calculated, with a deviation threshold set at ±1 pixel. If the deviation exceeded the threshold, the perspective matrix and scaling factor were readjusted. In the example, the feature point matching deviation was 0.8 pixels, meeting the accuracy requirements. The spatially registered guidance content was generated, which was fully adapted to the current display screen's viewpoint and scale, and could be directly used for subsequent overlay rendering.

[0079] Using a low-latency graphics rendering engine, the guidance content after spatial registration is rendered and superimposed on the endoscopic simulation video screen in real time as a semi-transparent layer, generating an original image that blends the virtual and real worlds. The core of this step is to achieve low-latency, high-quality layer overlay, ensuring that the virtual guide is synchronized with the real video footage in terms of timing and visual consistency, avoiding issues such as stuttering, flickering, or layer misalignment. The specific implementation method is as follows: The low-latency graphics rendering engine employs a hardware-accelerated rendering pipeline, leveraging the parallel computing capabilities of the graphics processing unit to keep rendering latency below 20ms, thus meeting the needs of real-time interactive teaching. The rendering engine supports an alpha blending rendering mode, enabling the overlay of semi-transparent layers. It also supports multi-layer management, allowing the registered guidance content to be rendered as an independent layer, layered with the endoscopic simulation video feed, facilitating subsequent adjustments and optimizations.

[0080] The transparency parameter of the semi-transparent layer is dynamically set based on the deviation type and the operational scenario. For mild deviations, the guide layer transparency is set to 60% (ensuring visibility without obscuring the actual anatomical view); for moderate deviations, it's set to 70%; and for severe deviations, it's set to 80%. The example shows a mild deviation with a transparency set to 60%. The Alpha blending formula C_final = C_virtual × α + C_real × (1 - α) (where C_final is the blended color, C_virtual is the virtual guide color, C_real is the actual image color, and α is the transparency coefficient) is used to ensure a natural blend between the virtual element and the actual image. For example, if the virtual arrow is green (RGB value 0, 255, 0), and the corresponding location in the actual image is mucosal pink (RGB value 255, 192, 203), the blended color is (102, 233, 121), preserving the arrow's recognizability while matching the color tone of the actual anatomical view.

[0081] Rendering timing synchronization is achieved through timestamp alignment. The rendering engine extracts the timestamp of each frame of the endoscopic simulation video stream and synchronously calls the corresponding timestamp's spatial registration guidance content, ensuring that each video frame corresponds precisely to the guidance content without timing misalignment. The rendering frame rate is kept consistent with the video capture frame rate at 30fps to avoid stuttering or screen tearing caused by frame rate mismatch. Simultaneously, the rendering engine supports frame buffer optimization, employing a double buffering mechanism: the first buffer displays the current frame, and the second buffer renders the next frame, with a switching time ≤1ms, eliminating visual flicker.

[0082] When generating the original image that blends virtual and real elements, it is necessary to ensure the positional stability of the virtual guide elements. Minor positional fluctuations are corrected using inter-frame interpolation algorithms. For example, when the screen coordinate deviation of guide elements in adjacent frames is ≤0.5 pixels, linear interpolation is used for a smooth transition to avoid element jitter. In the example, the coordinate deviation of the 3D arrows in adjacent frames is 0.3 pixels. After interpolation, the image transition is smooth with no obvious jitter, resulting in a natural blend of virtual and real elements. The original image retains the realism of the endoscopic simulation image while also providing clear augmented reality guidance.

[0083] Edge smoothing and anti-aliasing are applied to the original image that blends virtual and real visuals to ensure a natural transition between the virtual guide and the real image, without visual flickering or misalignment. The final output is a stable and clear operation guide view displayed on the student's screen.

[0084] The core of this step is to optimize the visual effects of the image, eliminate problems such as jagged edges and harsh junctions of virtual elements, improve image clarity and stability, and provide students with a good visual experience. The specific implementation method is as follows: Edge smoothing employs a bilateral filtering algorithm, which preserves edge details while smoothing noise and jagged edges, preventing image blurring. The filtering parameters are set as follows: window size 5×5 pixels (balancing smoothing effect and processing speed), color standard deviation σ_color=10 (controlling color similarity weight; the larger the value, the more difficult it is to smooth pixels with large color differences), and spatial standard deviation σ_space=5 (controlling spatial distance weight; the larger the value, the more difficult it is to smooth distant pixels). In the example, bilateral filtering is applied to the edges of the 3D arrows and path contours in the original image of the virtual-real fusion. After filtering, jagged edges are eliminated, lines are smooth, and the transition between the edges and the real anatomical image is natural, without any obvious discontinuity.

[0085] Anti-aliasing employs multi-sampling anti-aliasing technology with a sampling rate set to 4x (balancing anti-aliasing effect and processing efficiency). It eliminates edge jaggedness and color banding by averaging the colors of the four sampling points surrounding each pixel. For vector elements such as virtual text prompts and key step numbers, an additional vector anti-aliasing algorithm is used to ensure smooth text edges without jagged edges. In the example, the dynamic text prompt "Retract 3mm and then advance" is processed with anti-aliasing, resulting in clear, jagged edges, high readability, and seamless integration with the real-world image.

[0086] Visual flicker and misalignment detection employs an inter-frame comparison algorithm. It continuously extracts 10 frames of fused virtual-real images, calculating the relative positional deviation between the virtual guide element and the real anatomical feature points. If the deviation is ≤1 pixel and shows no continuous fluctuation, no misalignment is determined. If there are 3 consecutive frames with a deviation >1 pixel, or a single frame with a deviation >2 pixels, position calibration is triggered, and the perspective transformation and scaling process is re-executed to correct the misalignment. In the example, the positional deviation of all 10 frames is ≤0.8 pixels, with no significant fluctuations, indicating no misalignment and good image stability.

[0087] The final output operation guidance view needs to undergo brightness and contrast calibration to ensure clear visibility under different ambient lighting conditions. The calibration formula is: Brightness L_final = L_original × 1.1 (appropriately increase brightness to enhance guidance visibility), Contrast C_final = C_original × 1.2 (enhance the sense of depth in the image, making the distinction between virtual elements and the real image more obvious). After calibration, the output is sent to the student's display screen, with the screen resolution set to 1920×1080 pixels and a refresh rate of 60Hz. The image is flicker-free and misaligned, the boundary between the virtual guidance and the real anatomical image is natural, the lines are smooth, and the text is clear, allowing for precise guidance for students to correct operational deviations.

[0088] S205, based on the real-time video image feedback of the student's subsequent operations, adaptively adjust the presentation intensity and prompting rhythm of the augmented reality error correction guidance content until the student's operation and the standard process match the preset teaching qualification threshold.

[0089] Specifically, after presenting the operation guidance view, real-time video images of the student's subsequent operations can be continuously collected, and steps can be executed in real time to perform time-series analysis and feature extraction on the continuous video stream and key frame image data. This process identifies the action deviation between the student's operation technique and the standard procedure, as well as the action feature extraction and difference calculation process in the delayed segment of the step, and generates a real-time curve of the matching degree of subsequent operations. The core of this step is to continuously monitor the dynamics of trainees' operations. Through real-time feature extraction and difference calculation, the operation matching degree is quantified and visualized, providing data support for subsequent adaptive adjustments. The specific implementation method is as follows: After the operation guidance view is presented, the multi-angle camera array continues to acquire data with the same acquisition parameters as before: resolution 1920×1080 pixels, frame rate of 30fps for hand and display screen views, frame rate of 60fps for machine trajectory views, global shutter mode to avoid motion blur, and infrared fill light automatically adjusting intensity (100-500 lux) according to ambient light. The acquired real-time video stream is transmitted to the processing unit via an industrial fieldbus with a transmission delay of ≤10ms to ensure data real-time performance. Simultaneously, the system extracts one keyframe image every 10 frames, synchronously binding a timestamp and viewpoint label to form a real-time keyframe sequence for feature extraction and verification.

[0090] Real-time motion feature extraction reuses a pre-trained lightweight HRNet pose estimation network, with inference latency controlled within 30ms, adapting in tandem with the acquisition frame rate. The extraction process is consistent with previous steps: using the hand-view image as the core, combined with calibration of the device trajectory view image, it extracts the 3D coordinates of 21 hand joints, the six-DOF pose of the device end effector, and the 3D motion direction vector, generating a 72-dimensional feature vector for each frame. In the example, the feature vector of the 10th frame of the student's subsequent operation is [885,538,24,...,1049,476,39,5.0,3.0,1.7,4.5,1.8,1.9], representing the state of fine-tuning the device end effector position and slowing down the movement speed.

[0091] The difference calculation uses a weighted model of three dimensions: trajectory, speed, and angle. The weights are allocated as follows: trajectory 0.4, speed 0.3, and angle 0.3, with a comprehensive difference D ∈ [0,1]. The matching degree is defined as M = 1 - D, where M ∈ [0,1]. The closer M is to 1, the more consistent the operation is with the standard procedure. In the example, in frame 10, the trajectory difference is 0.09, the speed difference is 0.018, and the angle difference is 0.015. The comprehensive difference D = 0.4 × 0.09 + 0.3 × 0.018 + 0.3 × 0.015 ≈ 0.045, and the matching degree M = 0.955. In frame 15, due to excessive fine-tuning by the student, the trajectory difference increases to 0.12, and the matching degree decreases to 0.942.

[0092] The real-time matching curve for subsequent operations is plotted with time on the horizontal axis (in seconds, with a precision of 0.033 seconds, corresponding to a frame rate of 30fps) and matching degree on the vertical axis (range [0,1]), with a preset qualified threshold line (0.85) and a trend line added. The curve is drawn in real-time, updating one data point per frame, and a 5-point moving average method is used to smooth the curve to eliminate interference from accidental fluctuations. The smoothing formula is M_smooth(k)=(M(k-2)+M(k-1)+M(k)+M(k+1)+M(k+2)) / 5. In the example curve, the matching degree slowly fluctuates from 0.955 to 0.948 from 0 to 2 seconds, showing a slight downward trend overall. From 2 to 3 seconds, the matching degree rises to 0.963 due to the student's corrective actions. The curve visually presents the improvement in operation.

[0093] Based on the changing trend of the real-time curve of the matching degree in subsequent operations, an adaptive adjustment strategy is designed. When the matching degree increases, the transparency and prompt frequency of the virtual guidance are gradually reduced. When the matching degree stagnates or decreases, the visual salience of the guidance is enhanced and detailed explanations are added, generating a dynamic adjustment parameter set. The core of this step is to formulate a differentiated adjustment strategy based on the trend of matching degree changes, so as to achieve accurate adaptation of the guidance content. This avoids excessive prompts that interfere with operation, while ensuring that trainees can receive sufficient guidance when they encounter difficulties. The specific implementation method is as follows: The matching accuracy trend is determined by the curve slope and fluctuation range, with three trend thresholds set: an upward trend is defined as a matching accuracy increase of ≥0.02 for 3 consecutive frames and a slope ≥0.005 / frame; a stagnant trend is defined as a matching accuracy fluctuation of ≤0.01 for 5 consecutive frames and an absolute slope <0.003 / frame; and a downward trend is defined as a matching accuracy decrease of ≥0.02 for 3 consecutive frames and a slope ≤-0.005 / frame. In the example, the matching accuracy of the student's operation increased from 0.948 to 0.963 in 2-3 seconds, an increase of 0.015 for 3 consecutive frames, with a slope of 0.005 / frame, which is determined to be an upward trend; the matching accuracy fluctuated by 0.013 in 0-2 seconds, with a slope of -0.0007 / frame, which is determined to be a stagnant trend.

[0094] The adaptive adjustment strategy is designed according to trend categories. For an improving trend, a "weakening guidance" strategy is adopted: gradually increasing the transparency of virtual guidance (reducing visual proportion), increasing transparency by 2% per frame until it drops to the base value (40% transparency for slight deviation); simultaneously reducing the prompt frequency, increasing the text prompt interval from 2 seconds / time to 5 seconds / time, and reducing the arrow flashing frequency from 0.8 times / second to 0.3 times / second, gradually reducing interference with the trainee's operation. For a stagnant trend, a "maintain and refine" strategy is adopted: maintaining the current transparency value (60%), keeping the prompt frequency unchanged, and supplementing the text with detailed explanations of the action (such as "keep the fine adjustment within 2° to avoid over-adjustment") to help trainees break through the bottleneck. For a declining trend, a "strengthening guidance" strategy is adopted: reducing transparency to 80% (increasing visual salience), shortening the text prompt interval to 1 second / time, increasing the arrow flashing frequency to 1.2 times / second, adding voice explanations (such as "if the equipment deviates from the path, immediately fine-tune to the left and retract"), and highlighting the outline of the correct path.

[0095] The dynamic adjustment parameter set integrates three main categories of parameters: trend judgment results, visual attribute adjustment values, and prompt rules. Parameter values ​​are dynamically updated frame-by-frame, with accuracy consistent with the rendering frame rate. In the example, the parameter set corresponding to the "Improvement" trend has the following characteristics: Trend type "Improvement," current matching degree 0.963, transparency adjustment target 40% (currently 60%, 2% increase per frame), arrow blinking frequency 0.8 → 0.3 times / second, text prompt interval 2 → 5 seconds / time, text detail "Concise," and no voice narration. The parameter set corresponding to the "Stagnation" trend has the following characteristics: Trend type "Stagnation," transparency 60%, blinking frequency 0.8 times / second, prompt interval 2 seconds / time, text detail "Refined," and supplementary detailed explanations. The parameter set includes a timestamp and frame number to ensure accurate synchronization with real-time curves and guidance content.

[0096] Based on a dynamic adjustment parameter set, the visual attributes of augmented reality elements in the operation guidance view are adjusted in real time, including transparency, color saturation, blinking mode, and the level of detail of text prompts, to generate an adaptively optimized operation guidance view. The core of this step is to transform dynamically adjustable parameters into visual effects. By precisely adjusting the attributes of virtual elements, the guidance content is adapted to the student's operational status in real time, ensuring that the adjustment effect is intuitive and natural. The specific implementation method is as follows: The transparency adjustment uses a gradient transition to avoid abrupt changes that could cause visual discomfort. Each frame is adjusted in increments set by the parameter set (increasing trend +2%, decreasing trend -5%), with the transition formula being α_current = α_prev + Δα (α_current is the current transparency, α_prev is the transparency of the previous frame, and Δα is the adjustment increment). In the example, the initial transparency is 60%, increasing by 2% each frame under the increasing trend. The first frame adjusts to 62%, the second frame to 64%, until the tenth frame drops to 40% and stabilizes, with a smooth and stutter-free transition. The transparency adjustment is synchronized with the alpha blending formula to ensure a consistently natural blend between virtual elements and the real-world image.

[0097] Color saturation adjustment is linked to trends. When the trend is up, the saturation gradually decreases from 100% to 80% (weakening visual impact), the trend remains at 100%, and the trend is down, increasing to 120% (enhancing recognizability). In the example, under the uptrend, the green arrow's saturation decreases from 100% to 80% at a rate of 2% per frame, changing the color from bright green to a softer green, retaining its guiding function while reducing visual interference; under the downtrend, the arrow's saturation increases to 120%, making the color brighter and attracting students' attention.

[0098] The blinking mode adjusts according to the prompt frequency, with four modes: constant, slow blink (0.3-0.5 times / second), medium blink (0.8 times / second), and fast blink (1.2-1.5 times / second). An upward trend transitions from medium blink to slow blink, eventually remaining constant; a stagnant trend maintains medium blink; and a downward trend transitions from medium blink to fast blink. In the example, the blinking frequency of the down arrow in an upward trend gradually decreases from 0.8 times / second to 0.3 times / second, switching to constant blink mode after the 10th frame to reduce dynamic interference.

[0099] The level of detail in the text prompts is adjusted in three levels: Concise (rising trend) retains only the core action prompts, such as "advance along the path"; Standard (stagnant trend) adds the operation range, such as "turn left 2°, retract 3mm and advance"; Detailed (falling trend) adds the principle and precautions, such as "turn left 2° to avoid damaging the mucosa, retract 3mm and then accurately advance to the duodenal entrance." The text length is controlled between 10-25 characters, using word-by-word display and fade-out animation to improve readability. In the example, under the rising trend, the text is simplified from the Standard level "turn left 2°, retract 3mm and advance" to the Concise level "advance along the path," and the prompt interval increases from 2 seconds to 5 seconds.

[0100] After all attributes are adjusted, an adaptive and optimized operation guidance view is generated. This view is fully adapted to the current student's operation status. In the example, the optimized view under the upward trend is as follows: the arrow is always on, the transparency is 40%, the saturation is 80%, the concise text prompts are every 5 seconds, and the virtual elements blend naturally with the real screen. It does not interfere with the student's independent operation, but provides basic guidance.

[0101] After returning to the step of continuously capturing real-time video images of the student's subsequent operations after the operation guidance view is presented, the system continuously monitors the matching degree of the student's operations until the average matching degree of the student in a complete operation process exceeds the preset teaching qualification threshold. Then, the system automatically fades out and turns off all augmented reality error correction guidance, marking the completion of this teaching guidance.

[0102] The core of this step is to build a cyclical monitoring mechanism to ensure that students' operations meet the teaching standards, achieve orderly exit from the guidance content, and complete the teaching loop. The specific implementation method is as follows: The cyclical monitoring is achieved through process iteration. The system continuously cycles according to the logic of "collection-analysis-adjustment-presentation," updating the matching degree data once per frame and dynamically adjusting the guidance content without manual intervention until the termination conditions are met. During the cycle, it determines in real time whether the trainee's operation has entered the complete operation process. The complete operation process is defined as the entire stage from "instrument insertion into the esophagus" to "instrument positioning at the pylorus," taking 15-20 seconds and including all key operation nodes to ensure comprehensive evaluation.

[0103] The preset teaching qualification threshold is set based on expert operation data statistics, with an average matching degree threshold of 0.85. It also requires a matching degree of no less than 0.80 for three consecutive frames to avoid misjudgment based on a single frame meeting the standard. The average matching degree is calculated using a sliding window method, with the window size equal to the total number of frames in the complete operation process (30fps × 18 seconds = 540 frames). The average value of all data points within the window is calculated in real time, using the formula M_avg = ΣM(k) / N (where N is the number of frames within the window). In the example, after the student's operation enters the complete process, the average matching degree for the first 100 frames is 0.82, which is below the standard, and the system continues to adjust cyclically. At the 300th frame, the average matching degree rises to 0.86, and the matching degrees for the next three frames are 0.87, 0.86, and 0.88, respectively, all above 0.80, meeting the qualification requirements.

[0104] The automatic fade-out closing guide adopts a gradual exit strategy to avoid sudden disappearance causing discomfort to trainees. The fade-out duration is set to 0.5 seconds and is executed in three stages: the first stage (0-0.2 seconds) is when the transparency of the virtual elements quickly increases from the current value (40%) to 80%, while the color saturation decreases to 50%; the second stage (0.2-0.4 seconds) is when the transparency increases from 80% to 100%, the blinking mode is turned off, and the text prompts fade out word by word; the third stage (0.4-0.5 seconds) is when all virtual elements completely disappear, the layer resources are released, and the original image of the endoscope simulation is restored.

[0105] After the guided instruction is completed, the system generates a teaching summary, including information such as total operation time, number of deviation segments, average matching degree, and time to reach the target. Example summary: total operation time 18.5 seconds, 1 slightly deviation segment, average matching degree 0.88, and time to reach the target starting from frame 300, marking the successful completion of this gastrointestinal endoscopy operation guided instruction. The system simultaneously resets all parameters to prepare for the next teaching interaction.

[0106] Another embodiment of the present invention provides an interactive teaching demonstration system for gastrointestinal endoscopy, see [link to relevant documentation]. Figure 3 The system may include: The acquisition module 301 is used to acquire continuous video streams and key frame image data of the trainee operating the simulated gastrointestinal endoscopy process through the image acquisition device; The recognition module 302 is used to perform time-series analysis and feature extraction on the continuous video stream and key frame image data, and to identify the action deviation and step lag segments between the trainee's operation method and the standard procedure. The matching module 303 is used to dynamically match and call a preset three-dimensional anatomical model and standard operation image library based on the action deviation and step lag segment to generate augmented reality error correction guidance content that integrates real-time images. Display module 304 is used to overlay the augmented reality error correction guidance content onto the display interface corresponding to the student's operating field of view in real time, forming a virtual and real combined operation guidance view. The adjustment module 305 is used to adaptively adjust the presentation intensity and prompting rhythm of the augmented reality error correction guidance content based on the real-time video image feedback of the student's subsequent operations, until the matching degree between the student's operation and the standard process reaches the preset teaching qualification threshold.

[0107] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0108] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0109] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0110] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A teaching demonstration and interactive method for gastrointestinal endoscopy, characterized in that, The method includes: The continuous video stream and keyframe image data of the trainee operating the simulated gastrointestinal endoscopy are acquired through image acquisition equipment. Time-series analysis and feature extraction are performed on the continuous video stream and keyframe image data to identify action deviations and delayed steps between the trainee's operation techniques and the standard procedure. Based on the aforementioned action deviations and delayed steps, the system dynamically matches and calls upon a pre-set 3D anatomical model and a standard operation image library to generate augmented reality error correction guidance content that integrates real-time images. The augmented reality error correction guidance content is overlaid on the display interface corresponding to the student's operating field of view in real time, forming a virtual and real combined operation guidance view; Based on real-time video image feedback from the student's subsequent operations, the presentation intensity and prompting rhythm of the augmented reality error correction guidance content are adaptively adjusted until the student's operation matches the standard process to a preset teaching qualification threshold.

2. The method according to claim 1, characterized in that, The acquisition of continuous video streams and keyframe image data of the trainee's operation of the simulated gastrointestinal endoscopy process through image acquisition equipment includes: A multi-angle camera array is deployed on a simulated gastrointestinal endoscopy operating platform to simultaneously acquire raw video signals from three perspectives: the operator's hand, the movement trajectory of the instrument, and the endoscope simulation display screen, generating multiple synchronous raw video streams. The multi-channel synchronous raw video streams are time-aligned and encoded for compression. A hardware encoder is used to unify the timestamps and convert them into a standard video format to generate a synchronous standardized video stream. Based on a synchronized standardized video stream, a motion saliency detection algorithm is applied to automatically identify landmark operation moments, including instruments entering key anatomical sites, performing biopsies, or rinsing, and generate a sequence of key operation time points. Based on the key operation time point sequence, the static images of the corresponding frames are extracted from the synchronous standardized video stream and bound and encapsulated with timestamps and viewpoint tags to finally generate keyframe image data packets with spatiotemporal annotations.

3. The method according to claim 2, characterized in that, The step of performing time-series analysis and feature extraction on the continuous video stream and keyframe image data to identify action deviations and delayed steps between the trainee's operation techniques and the standard procedure includes: Load keyframe image data packets with spatiotemporal annotations, and extract the trainee's hand joint coordinates, device end pose, and motion direction vector through a pre-trained pose estimation network to generate a sequence of operational action features. The operation action feature sequence is dynamically time-aligned with the expert operation template in the standard teaching process database, and the difference in motion trajectory, speed and angle of each time segment is calculated to generate a time difference map. Based on the temporal difference map, an abnormal segment detection algorithm is used to identify continuous time periods that exceed a preset threshold. These continuous time periods correspond to operational errors or instrument position deviations, and preliminary deviation segment identifiers are generated. By combining the time requirements of the standard process at each stage, we analyze the completion time of key steps before and after the initial deviation segment identification, determine whether there are problems such as disordered execution order of steps or excessive time consumption, and finally accurately identify the action deviation and delayed steps.

4. The method according to claim 3, characterized in that, The process of dynamically matching and calling a pre-set 3D anatomical model and standard operation image library based on the aforementioned action deviations and step lag segments to generate augmented reality error correction guidance content that integrates real-time images includes: The system analyzes the type and location of movement deviations and delayed steps, retrieves the corresponding 3D model of the gastrointestinal segment from the 3D anatomical model library, highlights the correct anatomical position of the current instrument, and generates a 3D spatial error correction model. Based on the type of deviation, a matching short video clip of expert standard operation is retrieved from the standard operation video library. This clip demonstrates the corrective action from the current error state to the correct state, and a standard corrective action video is generated. The three-dimensional spatial error correction model is fused and rendered with the standard corrected motion image. Augmented reality overlay technology is used to synthesize the three-dimensional arrow guidance, the correct path outline and the standard motion animation into the same visual layer to generate a basic augmented reality guidance layer. Depending on the severity and type of the deviation, dynamic text prompts, key step numbers, or warning labels are overlaid on the basic augmented reality guidance layer, and the transparency and flashing frequency of the visual elements are adjusted to ultimately generate customized augmented reality error correction guidance content.

5. The method according to claim 4, characterized in that, The step of overlaying the augmented reality error correction guidance content onto the display interface corresponding to the student's field of vision in real time to form a virtual-real combined operation guidance view includes: Obtain the screen coordinate parameters and spatial orientation data of the endoscope simulation display screen corresponding to the student's current operating field of view, and generate display interface spatial registration information; Based on the spatial registration information of the display interface, the customized augmented reality error correction guidance content is subjected to perspective transformation and size scaling to make it completely match the perspective and proportion of the anatomical image on the display screen, and the guidance content after spatial registration is generated. Using a low-latency graphics rendering engine, the guidance content after spatial registration is rendered and superimposed on the endoscopic simulation video screen in real time as a semi-transparent layer, generating an original image that blends the virtual and real worlds. Edge smoothing and anti-aliasing are applied to the original image that blends virtual and real visuals to ensure a natural transition between the virtual guide and the real image, without visual flickering or misalignment. The final output is a stable and clear operation guide view displayed on the student's screen.

6. The method according to claim 5, characterized in that, The step of adaptively adjusting the presentation intensity and prompting rhythm of the augmented reality error correction guidance content based on real-time video image feedback from the student's subsequent operations until the student's operation matches the standard process to a preset teaching qualification threshold includes: After presenting the operation guidance view, the system continuously collects real-time video images of the student's subsequent operations and performs time-series analysis and feature extraction on the continuous video stream and key frame image data in real time. It identifies the action deviation between the student's operation method and the standard process, as well as the action feature extraction and difference calculation process in the step lag segment, and generates a real-time curve of the matching degree of subsequent operations. Based on the changing trend of the real-time curve of the matching degree in subsequent operations, an adaptive adjustment strategy is designed. When the matching degree increases, the transparency and prompt frequency of the virtual guidance are gradually reduced. When the matching degree stagnates or decreases, the visual salience of the guidance is enhanced and detailed explanations are added, generating a dynamic adjustment parameter set. Based on a dynamic adjustment parameter set, the visual attributes of augmented reality elements in the operation guidance view are adjusted in real time, including transparency, color saturation, blinking mode, and the level of detail of text prompts, to generate an adaptively optimized operation guidance view. After returning to the step of continuously capturing real-time video images of the student's subsequent operations after the operation guidance view is presented, the system continuously monitors the matching degree of the student's operations until the average matching degree of the student in a complete operation process exceeds the preset teaching qualification threshold. Then, the system automatically fades out and turns off all augmented reality error correction guidance, marking the completion of this teaching guidance.

7. A teaching demonstration and interactive system for gastrointestinal endoscopy, characterized in that, The system includes: The acquisition module is used to acquire continuous video streams and keyframe image data of trainees operating simulated gastrointestinal endoscopy through image acquisition equipment; The recognition module is used to perform time-series analysis and feature extraction on the continuous video stream and key frame image data, and to identify the action deviations and delayed steps between the trainee's operation techniques and the standard procedures. The matching module is used to dynamically match and call a preset three-dimensional anatomical model and standard operation image library based on the action deviation and step lag segment to generate augmented reality error correction guidance content that integrates real-time images. The display module is used to overlay the augmented reality error correction guidance content onto the display interface corresponding to the student's operating field of view in real time, forming a virtual and real combined operation guidance view. The adjustment module is used to adaptively adjust the presentation intensity and prompting rhythm of the augmented reality error correction guidance content based on the real-time video image feedback of the student's subsequent operations, until the matching degree between the student's operation and the standard process reaches the preset teaching qualification threshold.

8. The system according to claim 7, characterized in that, The acquisition module is specifically used for: A multi-angle camera array is deployed on a simulated gastrointestinal endoscopy operating platform to simultaneously acquire raw video signals from three perspectives: the operator's hand, the movement trajectory of the instrument, and the endoscope simulation display screen, generating multiple synchronous raw video streams. The multi-channel synchronous raw video streams are time-aligned and encoded for compression. A hardware encoder is used to unify the timestamps and convert them into a standard video format to generate a synchronous standardized video stream. Based on a synchronized standardized video stream, a motion saliency detection algorithm is applied to automatically identify landmark operation moments, including instruments entering key anatomical sites, performing biopsies, or rinsing, and generate a sequence of key operation time points. Based on the key operation time point sequence, the static images of the corresponding frames are extracted from the synchronous standardized video stream and bound and encapsulated with timestamps and viewpoint tags to finally generate keyframe image data packets with spatiotemporal annotations.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-6.