An endoscope-based AI-assisted diagnosis method and system

By identifying potential lesion areas and generating scoring information in the endoscopic AI-assisted diagnostic system, relevant information is immediately recorded and displayed when the scoring information exceeds a threshold, thus solving the problem of missed diagnosis of early lesions and improving the detection rate of early gastrointestinal tumors.

CN122223017APending Publication Date: 2026-06-16FUJIAN ZHIDE MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN ZHIDE MEDICAL TECH CO LTD
Filing Date
2026-05-15
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing AI-assisted endoscopic diagnostic systems may miss diagnoses in complex and dynamic environments due to time smoothing mechanisms that can filter out early, occult lesions.

Method used

By acquiring endoscopic images, potential lesion areas are identified and scoring information is generated. When the scoring information is greater than a preset score threshold, a record information containing time points, image segments, probe positions, and identification results is immediately generated and stored, and this record information is displayed at the end of the examination.

Benefits of technology

It effectively solves the problem that AI systems filter out instantaneous high-confidence recognition results due to time smoothing mechanisms, improves the detection rate of early gastrointestinal tumors, and ensures that even lesion features with extremely short durations can be captured and preserved, providing more comprehensive and reliable diagnostic evidence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223017A_ABST
    Figure CN122223017A_ABST
Patent Text Reader

Abstract

The application provides an endoscopic AI-assisted diagnosis method and system, relates to the field of medical diagnosis, and identifies potential lesion areas and generates score information by acquiring endoscopic images. When the score information is greater than a preset score threshold, the system immediately generates and stores record information containing a time point, an image segment, a probe position and an identification result. After the examination is completed, the record information is presented to the doctor, effectively solving the problem in the prior art that the AI system filters out instantaneous high-confidence identification results due to the time smoothing mechanism, resulting in missed diagnosis of early occult lesions. By recording and presenting all potential lesion information that meets the preset threshold in real time, even very short-lasting lesion characteristics can be captured and retained, thereby avoiding the loss of diagnostic clues caused by a short observation window and significantly improving the detection rate of early tumors in the digestive tract.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical diagnostic technology, and more specifically, to an endoscopic AI-assisted diagnostic method and system. Background Technology

[0002] In the field of medical diagnosis, especially in the screening of early gastrointestinal tumors, endoscopy is a crucial tool. To improve the accuracy and efficiency of diagnosis, AI-assisted diagnostic systems have been introduced to analyze endoscopic images in real time and provide alerts to doctors.

[0003] Specifically, the inherent instability of real-time endoscopic video streams can lead to inconsistencies in AI system judgments. Existing systems typically generate a stable, flicker-free prompt on the screen. This effectively filters out isolated, erroneous judgments caused by momentary poor image quality.

[0004] However, this smoothing mechanism, introduced to enhance user experience, has created hidden dangers in certain clinical situations. In the digestive tract, there exists a special type of early lesion, such as flat or depressed adenomas, whose characteristics are very subtle, showing minimal difference from the surrounding normal mucosa. When a doctor performs a rapid scan, the endoscopic probe quickly sweeps across a section of intestinal mucosa. In a very brief moment, the camera angle and lighting may perfectly capture the characteristics of a flat lesion. The AI ​​system's core recognition method accurately identifies this fleeting lesion feature and generates a high-confidence judgment internally. However, because the doctor's scanning action is very fast, this observation window is extremely short, possibly covering only 3 to 5 frames of video, far below the trigger threshold set by the time smoothing mechanism. The doctor sees no indication on the screen and naturally assumes the area is normal, continuing to move the probe for further examination, ultimately leading to the omission of this early lesion. Summary of the Invention

[0005] This application discloses an AI-assisted diagnostic method and system based on endoscopy, aiming to solve the technical problem that existing AI-assisted diagnostic systems for endoscopy may filter out early occult lesions due to time smoothing mechanisms in complex dynamic environments, thus causing missed diagnoses.

[0006] The technical solution of this application is as follows:

[0007] In a first aspect, this application discloses an endoscopic AI-assisted diagnostic method, including:

[0008] Acquire endoscopic images;

[0009] Identify potential lesion areas in the endoscopic images and generate scoring information corresponding to the potential lesion areas;

[0010] Determine whether the scoring information is greater than a preset score threshold;

[0011] When the scoring information is greater than a preset score threshold, a record information containing the time point, image segment, probe position, and recognition result of the corresponding potential lesion area is generated and stored.

[0012] When the inspection is complete, the recorded information is displayed.

[0013] Secondly, this application also discloses an endoscopic AI-assisted diagnostic system, which includes:

[0014] Image acquisition module, used to acquire endoscopic images;

[0015] The region recognition module is used to identify potential lesion regions in the endoscopic image and generate scoring information corresponding to the potential lesion regions;

[0016] The threshold determination module is used to determine whether the scoring information is greater than a preset score threshold;

[0017] The information recording module is used to generate and store recording information containing the time point, image segment, probe position, and recognition result of the corresponding potential lesion area when the scoring information is greater than a preset score threshold.

[0018] The information display module is used to display the recorded information when the inspection is completed.

[0019] Beneficial effects:

[0020] This application discloses an AI-assisted diagnostic method based on endoscopy. By acquiring endoscopic images, it identifies potential lesion areas and generates scoring information. When the score exceeds a preset threshold, the system immediately generates and stores a record containing time points, image segments, probe positions, and identification results. This record is then presented to the physician after the examination. This technical solution effectively solves the problem in existing technologies where AI systems filter out instantaneous high-confidence identification results due to time smoothing mechanisms, leading to missed diagnoses of early, occult lesions. By recording and displaying all potential lesion information reaching the preset threshold in real time, even lesion features with extremely short durations can be captured and preserved, thus avoiding the loss of diagnostic clues due to the brevity of the "perfect observation window." This significantly improves the detection rate of early gastrointestinal tumors, overcomes the contradiction between the reliability of information and the early lesion detection rate in complex and dynamic examination environments of existing systems, and provides physicians with more comprehensive and reliable diagnostic evidence. Attached Figure Description

[0021] Figure 1This is a schematic diagram of an AI-assisted diagnostic method based on endoscopy, as provided in this application.

[0022] Figure 2 This is a schematic diagram of an AI-assisted diagnostic system based on endoscopy, provided for the purposes of this application. Detailed Implementation

[0023] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0024] Reference Figure 1 The diagram illustrates an embodiment of an AI-assisted diagnostic method under endoscopy according to the present invention, which may specifically include the following steps:

[0025] S101, acquire endoscopic images;

[0026] S102, identify potential lesion areas in the endoscopic image and generate scoring information corresponding to the potential lesion areas;

[0027] S103, determine whether the scoring information is greater than the preset score threshold;

[0028] S104, When the scoring information is greater than the preset score threshold, generate and store the recording information containing the time point, image segment, probe position and recognition result of the corresponding potential lesion area;

[0029] S105, When the inspection is completed, the recorded information is displayed.

[0030] In the field of medical diagnosis, especially in the screening of early gastrointestinal tumors, endoscopy is a crucial tool. To improve diagnostic accuracy and efficiency, AI-assisted diagnostic systems have been introduced to analyze endoscopic images in real time and provide alerts to doctors. However, the complexity and dynamism of actual examination environments, along with doctors' need for rapid and comprehensive scans, pose significant challenges to the stability and reliability of AI system alerts. Traditional AI-assisted diagnostic systems, when processing real-time endoscopic video streams, may exhibit inconsistencies in their judgments due to factors such as natural organ peristalsis and fluid or food residue within the digestive tract, resulting in flickering alert boxes. To address this issue, existing systems typically employ a temporal result smoothing mechanism, where a stable, flicker-free alert box is only generated on the screen when the AI ​​method consecutively identifies the same area as a suspicious lesion in a specific number of video frames. However, in certain clinical scenarios, such as for occult lesions like flat or depressed adenomas, this smoothing mechanism may cause the AI ​​system to identify the lesion but filter it out because the duration is shorter than a threshold, ultimately leading to the omission of early lesions.

[0031] To address this issue, this application proposes an AI-assisted diagnostic method based on endoscopy. This method acquires endoscopic images, identifies potential lesion areas within the images, and generates corresponding scoring information. Subsequently, it determines whether this scoring information exceeds a preset score threshold. When the scoring information exceeds the preset threshold, the system generates and stores a record containing the time point, image segment, probe position, and identification result of the corresponding potential lesion area. Finally, when the examination is completed, the system displays this record information. This application aims to solve the potential omission problem in existing AI-assisted diagnostic systems when processing instantaneous high-confidence identification results, ensuring that even transient lesion features can be effectively recorded and displayed, thereby improving the detection rate of early lesions.

[0032] To make the technical solution of this application easier and clearer to understand, the key terms and implementation environment involved will be explained in detail below.

[0033] "Endoscopic images" refer to visual data of the inside of the digestive tract that are acquired or recorded in real time using endoscopic equipment. These images can be static images or video frames. They form the basis for AI-assisted diagnosis and are used for subsequent lesion identification and analysis.

[0034] "Potential lesion areas" refer to areas in endoscopic images that, after preliminary analysis by the AI ​​model, are considered to potentially contain abnormalities or lesions. These areas require further evaluation and confirmation.

[0035] "Scoring information" refers to a quantitative indicator generated by the AI ​​model after assessing the risk of identified potential lesion areas. This scoring information can reflect the probability, severity, or confidence level of the lesion and is used for subsequent judgment.

[0036] "Preset score threshold" refers to a score limit set in advance before the system runs. When the score information of a potential lesion area exceeds this threshold, it indicates that the area has a high risk of lesion and needs to be recorded.

[0037] "Recorded information" refers to a series of relevant data automatically generated and stored by the system after identifying high-risk potential lesion areas. This includes the time point of lesion appearance, corresponding image segments, spatial location of the endoscopic probe, and the recognition results of the AI ​​model. This information is of great significance for subsequent review, diagnosis, and teaching.

[0038] The implementation environment of this application typically includes an endoscope, an image processing unit, an AI diagnostic module, a data storage module, and a display interface. The endoscope is responsible for acquiring endoscopic images, the image processing unit preprocesses the images, the AI ​​diagnostic module performs the core recognition and scoring functions, the data storage module stores and records information, and the display interface presents the diagnostic results and recorded information to the doctor.

[0039] The core of the AI-assisted diagnostic method based on endoscopy in this application lies in the intelligent analysis of endoscopic images and the effective recording and display of key information.

[0040] Firstly, various methods can be used to acquire endoscopic images. For example, the endoscopic device can directly transmit real-time video streams to the image processing unit, which then extracts the endoscopic images frame by frame or at a fixed frame rate. Alternatively, the endoscopic device can store the acquired video data in a local storage medium, which is then read and processed by the image processing unit. Furthermore, image data from remote endoscopic devices can be received via a network interface.

[0041] Secondly, deep learning models can be used to identify potential lesion areas in endoscopic images and generate corresponding scoring information. For example, a convolutional neural network (CNN) model can be used to analyze endoscopic images. This model, trained on a large amount of labeled data, can identify visual features related to lesions in the image. After identifying a potential lesion area, the model outputs a confidence score, which is the scoring information, indicating the probability that the area is a lesion. For example, if the probability of an area being identified as a polyp is 0.8, then 0.8 is the scoring information.

[0042] Next, to determine whether the score exceeds a preset threshold, the system compares the score generated by the AI ​​model with the preset threshold. For example, if the preset threshold is set to 0.7, a score of 0.8 for a potential lesion area is considered to exceed the preset threshold. This threshold can be adjusted based on clinical experience and diagnostic needs to balance the detection rate and false alarm rate.

[0043] Then, when the score exceeds a preset threshold, the system generates and stores a record containing the time point, image segment, probe position, and recognition result for the corresponding potential lesion area. For example, when the AI ​​model identifies a region whose score exceeds the threshold, the system immediately records the current timestamp, extracts a few seconds of image segment containing that region, records the spatial coordinates of the endoscopic probe within the digestive tract, and stores the AI ​​model's recognition result (e.g., "suspected adenoma") in the database. This information can be packaged into a structured data object for storage.

[0044] Finally, when the examination is complete, the system displays the recorded information. For example, after an endoscopy, the doctor can access the stored records through the user interface. The system can present all recorded potential lesion areas in a list format, with each entry including a time point, a thumbnail of the image segment, a diagram of the probe position, and the identification result. The doctor can select any entry to view detailed image segments and identification results for retrospective analysis or further diagnostic decisions.

[0045] This application presents an AI-assisted diagnostic method based on endoscopy. It acquires endoscopic images and utilizes an AI model to identify potential lesion areas and generate scoring information. When the scoring information exceeds a preset threshold, the system immediately generates and stores a record containing time points, image segments, probe positions, and identification results. This record information is then presented to the physician after the examination.

[0046] Specifically, during an endoscopic examination, the endoscopic device continuously acquires real-time video streams of the digestive tract, which are then transmitted in real-time to the AI ​​diagnostic module. The AI ​​diagnostic module analyzes each frame or set of frames to identify potential lesions. Once a potential lesion is identified, the AI ​​model evaluates it and generates a quantitative score reflecting the probability or confidence level of the lesion. The system then compares this score with a preset score threshold.

[0047] The key to this application is that once the score exceeds a preset threshold, the system immediately triggers a recording mechanism even if the high-confidence identification result lasts only a very short time (e.g., less than the number of frames required by the smoothing mechanism in existing technologies). This means that transiently high-confidence lesion identifications that might be filtered out by the time smoothing mechanism in existing technologies can be captured and recorded in a timely manner in this application. The system accurately records the time point of lesion appearance, the corresponding image segment, the specific spatial location of the endoscopic probe in the digestive tract, and the identification result of the AI ​​model. This recorded information is securely stored, forming a complete diagnostic clue.

[0048] After the entire endoscopic examination is completed, doctors can easily access and review all recorded information on potential lesions through the system's user interface. The system presents this information clearly and intuitively, for example, by sorting it chronologically or by lesion severity, and provides functions such as image playback and probe position visualization. Doctors can use this recorded information to focus on re-examining any transient lesions that might have been missed during the examination, thus avoiding the misdiagnosis of early lesions.

[0049] This application presents an AI-assisted diagnostic method based on endoscopy, aiming to address the potential omission issues in existing AI-assisted diagnostic systems when processing instantaneous high-confidence identification results. Traditional AI-assisted diagnostic systems typically employ a time-dimension smoothing mechanism to avoid interference from flickering alert boxes. This means that a stable alert box is only generated on the screen when the AI ​​method consecutively identifies the same area as a suspicious lesion in a specific number of video frames. However, this mechanism can lead to the omission of early lesions, even if the AI ​​system accurately identifies the lesion within a very short time, because the duration of the error does not reach the threshold of the smoothing mechanism.

[0050] This application effectively overcomes the limitations of existing technologies by introducing a core innovation: "when the scoring information exceeds a preset score threshold, the system generates and stores recorded information containing the time point, image segment, probe position, and recognition result of the corresponding potential lesion area." Unlike the strategy in existing technologies where "the system only generates a stable, non-flickering prompt box on the screen when the AI ​​method continuously identifies the same area as a suspicious lesion in a specific number of video frames," this application no longer relies on the duration of the lesion recognition result but focuses on the confidence level of the recognition result (i.e., the scoring information). As long as the AI ​​model's scoring information for a potential lesion area reaches or exceeds the preset score threshold, regardless of the duration, the system will immediately generate and store detailed recorded information.

[0051] The advantages of this improvement are twofold: First, it can capture fleeting but high-confidence early lesion clues, avoiding missed diagnoses due to time smoothing mechanisms. For example, when an endoscope probe quickly passes over a occult lesion, the AI ​​system may only identify high-confidence lesion features within 3-5 frames, which would be filtered out in existing systems but recorded in this application. Second, by storing recorded information including time points, image segments, probe positions, and recognition results, it provides doctors with comprehensive retrospective diagnostic evidence. Even if lesions are not detected in real time during the examination, doctors can review these recorded information after the examination to focus on suspicious areas, thereby significantly improving the detection rate of early lesions. This application's solution, while ensuring the stability of the AI-assisted diagnostic system, greatly enhances the ability to identify and record occult and transient lesions, providing more reliable and comprehensive support for clinical diagnosis.

[0052] This application further proposes an endoscopic AI-assisted diagnostic method, which includes:

[0053] The flatness value of the feature extraction layer in the potential lesion area is calculated to obtain the evaluation result;

[0054] Under different endoscopic imaging modes, lesion features of potential lesion areas are identified, the consistency between multiple lesion features is verified, and the verification results are obtained.

[0055] Based on the evaluation results and the verification results, the uncertainty level of the identification result is determined;

[0056] When the uncertainty level of the identification result is higher than a preset level threshold, feature extraction is performed on the image segment corresponding to the potential lesion area to obtain the feature extraction result;

[0057] Based on the feature extraction results, various diagnostic information and corresponding probability scores are generated;

[0058] The system displays various diagnostic information and their corresponding probability scores, and based on the selected events for the diagnostic information, it displays the evidence information corresponding to the diagnostic information, including image features, identification criteria, and patient history information.

[0059] Specifically, the flatness value of the feature extraction layer in the potential lesion region is calculated to assess the stability or ambiguity of the feature representation when the AI ​​model processes this region. The flatness value can be understood as the density of the data points in the feature space; a lower flatness value may indicate a more dispersed feature distribution, suggesting lower confidence in the AI ​​model's recognition of this region. This evaluation result is used to quantify the inherent uncertainty of the AI ​​model in recognizing potential lesion regions.

[0060] This process involves identifying lesion features under different endoscopic imaging modes, such as white light mode, narrow-band imaging (NBI) mode, or fluorescence mode, extracting features such as color, vascular morphology, and surface structure of the lesion area. Verifying the consistency between multiple lesion features involves comparing whether the features extracted in different modes corroborate each other. For example, if white light mode shows mucosal protrusion and NBI mode shows an abnormal vascular network, and the two highly match in spatial location and morphology, then the consistency is considered high. This cross-validation of multimodal features helps improve the robustness and accuracy of lesion feature identification, thus yielding the validation results.

[0061] In practical applications, the uncertainty level of the recognition result can be comprehensively judged based on the evaluation and verification results. For example, if the flatness value is low and the consistency of multimodal features is poor, the uncertainty level may be high. The uncertainty level can be divided into multiple levels, such as low, medium, and high, or represented by continuous numerical values. When the uncertainty level of the recognition result is higher than the preset level threshold, it indicates that the AI ​​model lacks confidence in the current recognition result and requires further analysis and auxiliary information.

[0062] Furthermore, feature extraction can be performed on image segments corresponding to potential lesion areas using deeper, more refined deep learning models or specific feature extraction algorithms to obtain richer features that aid in differential diagnosis. These feature extraction results form the basis for generating various diagnostic information.

[0063] Therefore, based on the feature extraction results, various possible diagnostic information can be generated, such as "early-stage cancer," "adenoma," and "inflammation," and a probability score can be provided for each diagnostic information, indicating the likelihood of that diagnosis. This provides doctors with a more comprehensive diagnostic perspective, rather than a simple "yes / no" judgment.

[0064] Finally, the system displays various diagnostic information and their corresponding probability scores. Based on the doctor's selection of events for specific diagnostic information, it further reveals the supporting evidence. This evidence includes imaging features (such as specific textures, colors, and vascular patterns of the lesion area), the basis for identification (key features or decision paths used by the AI ​​model to make the judgment), and patient history information (such as medical history, family history, and laboratory test results). This approach aims to provide doctors with transparent and interpretable diagnostic support, helping them understand the AI's judgment logic and make final decisions based on their own experience.

[0065] This application's solution effectively addresses the lack of confidence assessment in AI diagnostic results found in the aforementioned basic solutions by introducing an evaluation mechanism for the uncertainty of AI recognition results. Specifically, by calculating the flatness value of the feature extraction layer and verifying the consistency of lesion features across different imaging modalities, the uncertainty of AI recognition results can be comprehensively quantified from two dimensions: internal model feature representation and multimodal information fusion. When this uncertainty is high, the system no longer simply provides a simple recognition result but actively triggers deeper feature extraction and generates multiple possible diagnostic information and their probability scores. This multidimensional, probabilistic diagnostic information, combined with traceable evidence (including imaging features, identification criteria, and patient history information), enables doctors to gain a deeper understanding of the AI's judgment logic and make comprehensive judgments based on their own professional knowledge. This not only improves the transparency and interpretability of AI-assisted diagnosis but also provides doctors with richer and more reliable decision support, thereby compensating for potential diagnostic blind spots in complex or atypical cases in basic solutions.

[0066] In some preferred embodiments, a specific example is given below. Suppose that during an endoscopic examination, the AI ​​system initially identifies a region of gastric mucosa as a potential lesion and assigns a high score. However, in further analysis, the system finds that the flatness value of the feature extraction layer for this potential lesion region is low, indicating that the AI ​​model's representation of the region's features is not stable enough. Simultaneously, there is some inconsistency between the mucosal bulges observed in white light mode and the vascular patterns observed in NBI mode, resulting in low consistency in the verification results. Based on these evaluation and verification results, the system determines the uncertainty level of this identification result to be "moderately high," exceeding the preset level threshold.

[0067] At this point, the system no longer simply displays a single identification result of "potential lesion," but instead performs deeper feature extraction on the image segment. Based on the extracted refined features, the system generates multiple diagnostic information and their probability scores, such as: "gastric adenoma (probability: 70%)", "early gastric cancer (probability: 20%)", and "chronic gastritis (probability: 10%)". When a doctor selects the diagnosis of "gastric adenoma", the system immediately displays corresponding evidence, including: specific imaging features of the area (such as regular glandular openings and mild redness), key evidence used by the AI ​​model to identify it as an adenoma (such as the matching degree of specific texture patterns with the known adenoma feature library), and the patient's medical history (such as no family history and a history of gastric polyp removal). Based on this detailed and interpretable evidence, combined with their own clinical experience, doctors can make a final judgment on the diagnosis of "gastric adenoma" or further explore other diagnostic possibilities. This approach provides doctors with more comprehensive and reliable diagnostic support.

[0068] Specifically, this application further proposes that the above method also includes:

[0069] The evidence information is weighted and sorted to obtain the weighted and sorted results;

[0070] Based on the weighted and sorted results, an evidence presentation path is generated; the evidence presentation path represents the display order or display process of the evidence.

[0071] Based on the evidence presentation path, the image feature regions in the image segments are displayed in layers, and an interactive evidence navigation bar is generated;

[0072] When a selection event is received for the evidence navigation bar, the image feature area is magnified locally, and the identification basis and patient history information corresponding to the image feature area are displayed; as well as diagnostic information, probability score, and overview information; wherein, the overview information refers to an overview display of the evidence information.

[0073] Specifically, weighting and ranking evidence information refers to assigning different weights to different types of evidence information (such as imaging features, identification criteria, patient history information, etc.) according to preset rules or models, and prioritizing the evidence information based on these weights. For example, weighting can be based on factors such as the relevance of the evidence information to the diagnostic result, the reliability of the evidence, and the clinical importance of the evidence. The purpose is to highlight key evidence and help operators quickly focus on the most important diagnostic evidence.

[0074] The evidence presentation path generated based on weighted and ranked results can be understood as the system automatically planning the display order and layout of evidence according to its priority and relevance. For example, high-weight evidence can be displayed first, or related evidence can be grouped together. The evidence presentation path represents the display order or process of evidence, aiming to guide the operator to view and understand the evidence step by step in a logically clear manner, avoiding information overload.

[0075] In practical applications, layered display of image feature regions in image segments refers to overlaying different image features (such as color changes, vascular patterns, surface structures, etc.) related to potential lesions on endoscopic images at different visual levels or in different ways. For example, different feature regions can be distinguished by different colors, transparency, or outlines. Simultaneously, generating an interactive evidence navigation bar aims to provide an intuitive tool that allows operators to select, switch, or adjust the display of different evidence layers as needed, thereby enabling refined exploration of evidence.

[0076] When a selection event is received for the evidence navigation bar, the system will magnify the specific image feature area selected by the operator to facilitate observation of details. Simultaneously, it will display the corresponding identification criteria and patient history information for that image feature area, as well as diagnostic information, probability scores, and overview information. The overview information provides a summary of the evidence information, aiming to offer background information closely related to the feature area and an overall diagnostic overview while magnifying the image. This helps the operator switch between and connect details and the overall picture, thereby forming a comprehensive diagnostic judgment.

[0077] This application's solution, through weighted and sorted processing of evidence information, effectively filters and highlights evidence crucial to diagnosis, preventing operators from getting lost in a sea of ​​information. Based on this weighting and sorting result, an evidence presentation path is generated, ensuring that the presentation of evidence is not simply a collection of facts, but rather organized according to logic and importance, guiding operators to browse and understand efficiently. Furthermore, by displaying image feature regions in image segments in layers, coupled with an interactive evidence navigation bar, operators can flexibly explore different levels of image features according to their needs, achieving refined observation from macro to micro. When an operator selects a specific image feature region for local magnification, the system not only provides detailed visual information of that region but also simultaneously displays its corresponding identification criteria, patient history information, diagnostic information, probability score, and overview information, thus closely integrating local details with the overall diagnostic background, greatly enhancing the relevance and comprehensibility of the evidence. It is precisely this structured and interactive evidence presentation method that enables operators to analyze potential lesions more comprehensively and deeply, effectively solving the problems of insufficient evidence presentation or difficulty in effectively utilizing evidence information in traditional methods.

[0078] In some preferred embodiments, it is assumed that during an endoscopy, the AI ​​system identifies a potential lesion area in the digestive tract and generates various diagnostic information and corresponding probability scores, as well as rich evidence information, including various imaging features of the area (such as abnormal mucosal color, disordered vascular texture, surface bulges or depressions, etc.), the identification basis of the AI ​​model (such as deep learning model activation maps, feature vectors, etc.), and the patient's past medical history and family history.

[0079] To efficiently present this evidence to physicians, the system in this application first weights and ranks this evidentiary information. For example, the system can assign different weights based on the typicality of imaging features, the confidence level of AI identification, and the relevance of patient history information (such as whether there is a family history of cancer). Evidence with higher weights (such as imaging features highly suggestive of malignancy and AI identification evidence with high confidence) will be prioritized.

[0080] Based on this weighted and ranked result, the system generates an evidence presentation path. This path might be set as follows: first, displaying the most critical image features; then, the AI's recognition criteria; and finally, the patient's medical history. On the display interface, the system will display image feature regions in the image segment layer by layer according to this path. For example, on the original endoscopic image, the system can overlay highlighted areas of different colors to represent features such as color abnormalities, vascular irregularities, and surface bulges, and these layers can be independently turned on or off. Simultaneously, an interactive evidence navigation bar will be generated on the side of the screen, containing options such as "Color Abnormalities," "Vascular Texture," "Surface Structure," "AI Criteria," and "Patient History."

[0081] When a doctor clicks the "Vascular Texture" option on the evidence navigation bar, the system zooms in on the area in the image related to vascular texture irregularities and immediately displays the identification criteria for that vascular texture (e.g., an explanation of the abnormal vascular pattern identified by the AI ​​model), the patient's history of vascular disease, current diagnostic information (e.g., "high probability of early-stage cancer"), the corresponding probability score (e.g., 85%), and an overview (briefly summarizing all the evidence). Doctors can quickly switch to other feature areas for viewing using other options on the navigation bar, thus comprehensively and systematically assessing potential lesions.

[0082] In some embodiments described above, a preset threshold was used to determine the uncertainty level of the identification result. However, in actual endoscopic examinations, the physiological environment inside the digestive tract is dynamically changing, and the digestive tract mucosa may undergo transient deformation, such as deformation caused by peristalsis, respiration, or probe movement. This transient deformation may lead to a decrease in endoscopic image quality, affecting the visual characteristics of potential lesion areas. Consequently, determining the uncertainty level of the identification result based on a fixed preset threshold may be inaccurate, and could even lead to misdiagnosis or missed diagnosis of lesion areas. If the above problems are not addressed, the reliability and accuracy of AI-assisted diagnosis may be reduced, especially when lesion features are not obvious or the image is interfered with.

[0083] In response, this application further proposes an AI-assisted diagnostic method based on endoscopy, the method further comprising:

[0084] Deformation assessment was performed on the digestive tract mucosa region in the endoscopic images to obtain deformation assessment results;

[0085] When the deformation assessment results show that there is transient deformation of the digestive tract mucosa, the preset level threshold is adjusted according to the deformation assessment results.

[0086] Specifically, deformation assessment of the digestive tract mucosa in endoscopic images refers to detecting and quantifying changes in the position, shape, or texture of the digestive tract mucosa over a short period of time using image processing and analysis techniques. This can be achieved using various techniques, such as detecting the velocity and direction of mucosal movement by analyzing the optical flow field between consecutive frames, or monitoring the displacement of specific mucosal regions using feature point tracking algorithms. Alternatively, deep learning models can be employed, trained to identify and quantify different types of mucosal deformation, such as peristaltic waves, contraction, or expansion. The deformation assessment result can be a numerical value representing the degree, speed, or type of deformation.

[0087] When deformation assessment results indicate transient deformation of the digestive tract mucosa, it means the detected deformation is temporary and non-persistent, typically related to physiological activities (such as peristalsis) or external disturbances (such as slight probe contact). In this case, adjusting the preset level threshold based on the deformation assessment results aims to enable the system to adapt to changes in image quality caused by deformation. For example, if significant transient deformation is detected, it may mean that the features of potential lesion areas in the image are blurred or distorted, leading to an overestimation of uncertainty in the AI ​​model's recognition results. To avoid missed diagnoses in this situation, the preset level threshold can be appropriately lowered to improve the system's sensitivity to potential lesion areas, triggering further detailed analysis even with poor image quality. Conversely, if the deformation assessment results indicate that the mucosa is in a relatively stable state, the threshold can be maintained or appropriately increased to reduce unnecessary detailed analysis.

[0088] This application's solution introduces an assessment of deformation in the digestive tract mucosa and dynamically adjusts preset threshold levels based on the assessment results, enabling the AI-assisted diagnostic system to better adapt to the complex physiological environment during endoscopic examinations. When there is transient deformation of the digestive tract mucosa, image quality may be affected, potentially leading to deviations in the uncertainty level of the AI ​​model's identification of potential lesion areas. By dynamically adjusting the threshold, the system can compensate for this image instability caused by deformation, ensuring that the judgment of uncertainty in the identification results remains accurate and effective even when image quality fluctuates. This avoids misjudgments or missed judgments caused by fixed thresholds, especially when lesion features are temporarily obscured or distorted by transient deformation.

[0089] As a specific implementation, suppose that during an endoscopic examination, the AI ​​system identifies a potential lesion area and calculates the uncertainty level of its identification result. Without the aforementioned adjustment mechanism, if this uncertainty level is below a fixed preset threshold, the system may not trigger further detailed diagnostic procedures. However, if significant transient peristaltic deformation is occurring in the digestive tract mucosa at this time, this deformation may cause the image features of the lesion area to temporarily become blurred or atypical, resulting in the AI ​​model's assessment of its uncertainty level being lower than it should be. With the solution of this application, the system first assesses the deformation of the digestive tract mucosa area in the endoscopic image. If the assessment result shows significant transient deformation, the system dynamically lowers the preset threshold according to the degree of deformation. At this point, even if the uncertainty level of the potential lesion area identification result might be ignored at the original fixed threshold, it may exceed the adjusted lower threshold. As a result, the system will trigger subsequent detailed feature extraction, generation and display of various diagnostic information and probability scores, as well as display of evidence information, thereby ensuring that potential lesion areas can be fully focused on and analyzed under dynamic physiological conditions, avoiding diagnostic omissions due to fluctuations in image quality.

[0090] This application further proposes an optimization scheme, wherein the aforementioned image feature region includes a highlighted display region; the aforementioned evidence navigation bar includes a feature selection tool; and the aforementioned method of hierarchically displaying image feature regions in image segments according to the evidence presentation path and generating an interactive evidence navigation bar includes:

[0091] Identify lesion features in highlighted areas of an image segment; wherein the lesion features are correlated with different diagnostic information;

[0092] Generate visual identifiers for each lesion feature;

[0093] The lesion features are displayed in layers and overlaid according to the visual identifiers;

[0094] When a selection event for the lesion feature is received, the display effect of other lesion features is adjusted through the feature selection tool, and the identification basis and patient history information corresponding to the selected lesion feature are displayed at the same time.

[0095] Specifically, image feature areas are defined as highlighted areas. This means that after the system identifies potential lesion areas, it will visually highlight these areas (e.g., through color, borders, flashing, etc.) so that doctors can quickly locate them. The evidence navigation bar has been further enhanced, integrating a feature selection tool designed to provide refined interactive capabilities for specific lesion features within the highlighted areas.

[0096] In the implementation process, the system first performs in-depth analysis on the highlighted areas of the image segment to identify various lesion features. These lesion features can be polyps, ulcers, erosions, vascular abnormalities, etc., and each lesion feature has a clear association with one or more diagnostic information; for example, a specific lesion feature may strongly suggest a certain type of tumor. To facilitate differentiation and understanding by doctors, the system generates unique visual identifiers for each identified lesion feature, such as different colors, shapes, icons, or text labels. Subsequently, these visually identifiable lesion features are displayed in a layered, overlaid manner on the highlighted area according to a preset hierarchy or importance, allowing doctors to clearly see the distribution and interrelationships of different lesion features.

[0097] When an operator selects a specific lesion feature using the feature selection tool, the system responds to this selection event. At this time, the feature selection tool not only highlights the selected lesion feature but also intelligently adjusts the display of other unselected lesion features, such as reducing their transparency, lightening their color, or temporarily hiding them, to reduce visual interference and help doctors focus their attention on the currently selected lesion feature. Simultaneously, the system immediately displays the identification criteria directly corresponding to the selected lesion feature (such as the reasoning behind the AI ​​model's identification of the feature, key visual cues, etc.) and the patient's historical information (such as medical history, family history, medication history, etc.), thus providing doctors with comprehensive diagnostic support.

[0098] This application's solution addresses the lack of interaction regarding specific lesion features within image feature regions in basic solutions by refining image feature regions into highlighted display areas and introducing a feature selection tool. Specifically, the system first identifies lesion features associated with different diagnostic information within the highlighted display areas and generates a unique visual identifier for each feature. The layered overlay display of these visual identifiers allows doctors to intuitively understand the complexity of the lesion region. When a doctor selects a specific lesion feature, the feature selection tool intelligently adjusts the display of other features, effectively guiding the doctor's gaze and attention, and avoiding information overload. Simultaneously, the system can instantly retrieve and display the identification criteria and patient history information related to that specific lesion feature. This allows doctors to obtain all necessary contextual information in a centralized view, enabling more accurate and efficient diagnostic decisions. This refined interaction method significantly improves doctors' depth of understanding of lesion features and diagnostic efficiency.

[0099] In some preferred embodiments, assuming that during an endoscopy, the AI ​​system identifies multiple potential lesion features within a highlighted area, such as a suspected polyp, an abnormal vascular area, and a minor erosion. The system generates different visual labels for these three lesion features; for example, a polyp is marked with a red circle, an abnormal vascular area with a blue box, and an erosion with a yellow dashed line, displayed in a layered overlay on the endoscopic image. When the physician is interested in the suspected polyp marked with a red circle, they can click on the red circle using the feature selection tool. At this time, the system immediately dims or makes transparent the blue box and yellow dashed line, making the polyp marked with the red circle the visual focus. Simultaneously, next to or below the image, the system displays detailed identification criteria for the polyp (e.g., the probability of the AI ​​model determining it to be a polyp, key morphological features, etc.) and relevant patient history information (e.g., whether the patient has a history of polyps, family history, etc.). If the physician subsequently wants to view the abnormal vascular area, they can again click on the blue box using the feature selection tool, and the system will adjust the display accordingly and provide detailed information about the vascular abnormality. This interactive method allows doctors to switch flexibly and efficiently between different lesion characteristics and obtain targeted auxiliary diagnostic information.

[0100] In some embodiments described above, this application proposes displaying corresponding identification criteria, patient history information, diagnostic information, probability scores, and overview information when zooming in on a localized image feature area. However, in its implementation, if the overview information and the zoomed-in area are displayed simultaneously in a fixed manner, it may cause visual interference. Alternatively, when the user needs to focus on local details, the overview information may obscure some key areas, thus affecting diagnostic efficiency and user experience. Furthermore, switching between local detail observation and global overview may require additional manual operation, reducing the ease of use.

[0101] In response, this application proposes an optimized method for, upon receiving a selection event for the aforementioned evidence navigation bar, to locally magnify the aforementioned image feature region and display the corresponding identification criteria and patient history information for the aforementioned image feature region; as well as diagnostic information, probability scores, and overview information, specifically including:

[0102] Set the overview information display area;

[0103] The overview information display area displays diagnostic information, probability scores, and overview information;

[0104] When an interactive operation is received targeting the feature region of the magnified image, the transparency of the overview information display area is adjusted.

[0105] When the gaze duration on the overview information display area exceeds the gaze dwell time threshold, the overall overview view and the identification results corresponding to the potential lesion areas are displayed.

[0106] Specifically, setting up an overview information display area refers to reserving or dynamically creating a dedicated area in the user interface for displaying overview information. This area can be a floating window, a sidebar, or a specific panel on the screen, and its position and size can be preset or adjusted according to the actual application scenario. The overview information refers to a summary display of evidence information, aiming to provide a quick overview of the current diagnostic situation, such as key information like lesion type, severity, and confidence level.

[0107] Furthermore, displaying diagnostic information, probability scores, and overview information in the aforementioned overview information display area means that the diagnostic results generated by the AI ​​model, the corresponding confidence or probability scores, and a brief summary of all evidence information are presented in this specific area so that the operator can easily obtain the core diagnostic data.

[0108] In a preferred embodiment, when an interactive operation is received targeting the magnified image feature area, the transparency of the overview information display area can be adjusted. For example, when the operator drags, zooms, or clicks on the magnified area, the system can automatically reduce the transparency of the overview information display area, making it semi-transparent or even completely hidden, to ensure that the operator can clearly observe the details of the magnified area and avoid information obstruction. Conversely, when the operator stops interacting with the magnified area, the transparency of the overview information display area can be restored to its initial state so that the operator can regain access to the overview information.

[0109] Furthermore, this application proposes that when the gaze duration on the overview information display area exceeds a gaze dwell time threshold, an overall overview view and the corresponding identification results of potential lesion areas are displayed. The gaze dwell time threshold is a preset time length used to determine the operator's level of attention to the overview information display area. When the operator's gaze lingers on this area for longer than this threshold, the system determines that the operator needs a more macroscopic perspective, and automatically switches the display mode to show an overall overview view of the entire digestive tract or endoscopy examination, marking all identified potential lesion areas and their identification results, helping the operator understand the lesion distribution and diagnostic status from a global perspective.

[0110] This application's solution effectively addresses the problem of overview information potentially interfering with the observation of local details when magnifying image feature areas by introducing a dynamic management mechanism for the overview information display area. Specifically, by setting up a dedicated overview information display area, diagnostic information, probability scores, and overview information can be centrally displayed, resulting in a more rational information layout. When the operator needs to focus on the details of a magnified area, the system can intelligently adjust the transparency of the overview information display area based on the operator's interaction, such as making it semi-transparent or temporarily hiding it, thereby preventing the overview information from obscuring key image features and ensuring the clear visibility of local details. Simultaneously, by detecting the operator's gaze duration on the overview information display area and comparing it with a preset gaze duration threshold, the system can intelligently determine the operator's need for global information. Once the gaze duration exceeds the threshold, the system automatically switches to the overall overview view and displays the identification results of all potential lesion areas. This achieves seamless switching and efficient integration of local details and global overview information without increasing the operator's workload, greatly improving the smoothness and convenience of the diagnostic process.

[0111] In some preferred embodiments, a specific example is given below. Suppose a doctor is examining a patient using an AI-assisted endoscopic diagnostic system. The system first identifies a potential lesion area in the endoscopic image and magnifies it for the doctor to observe in detail. At this time, the system sets up an overview information display area in the upper right corner of the screen, which displays the diagnostic information of the lesion area (e.g., "high probability of early cancer"), probability score (e.g., "92%)", and overview information (e.g., "lesion size: 5mm, shape: irregular, color: reddish").

[0112] When a doctor uses the mouse to drag or zoom in on a magnified area of ​​the image to examine the edges or internal structures of a lesion more closely, the system immediately detects these interactions. In response, the transparency of the overview information display area automatically decreases, making it semi-transparent, allowing the doctor to see the image details below and preventing information obstruction. When the doctor stops interacting with the magnified area and shifts their focus to the overview information display area in the upper right corner, the system starts a timer. If the doctor's gaze time on this overview information display area exceeds a preset gaze duration threshold (e.g., 2 seconds), the system determines that the doctor needs to understand the overall situation. At this point, the system automatically switches the display mode to show a three-dimensional spatial model of the entire digestive tract (if supported by the system), clearly marking all identified potential lesion areas and their corresponding identification results in this overall overview view. For example, multiple lesion points may be displayed in the antrum or body of the stomach, along with brief identification results, helping the doctor assess the distribution and overall condition of the lesions from a macroscopic perspective. This seamless switching allows doctors to efficiently integrate information between detailed observation and a global overview, leading to a more comprehensive and accurate diagnosis.

[0113] This application further proposes the following steps for displaying the overall overview view and the identification results corresponding to the potential lesion area when the gaze duration for the overview information display area exceeds the gaze dwell time threshold:

[0114] Detect the operator's visual focus and fixation time;

[0115] Obtain the type and duration of interactive behaviors in the zoomed-in area;

[0116] Adjust the gaze dwell time threshold according to the interaction type and interaction time;

[0117] Based on the gaze focus, the gaze duration, and the adjusted gaze dwell time threshold, determine whether the gaze duration for the overview information display area exceeds the adjusted gaze dwell time threshold.

[0118] When the gaze duration for the overview information display area exceeds the adjusted gaze dwell time threshold, an overall overview view and the identification results corresponding to the potential lesion areas are displayed.

[0119] Specifically, detecting the operator's gaze focus and fixation time refers to using eye-tracking devices or related technologies to monitor the operator's gaze point and duration on the screen in real time. The gaze focus can be understood as the area of ​​the screen the operator is currently focusing on, while fixation time is the duration the gaze remains on that area; the purpose is to obtain information about the operator's visual attention. Acquiring the types and duration of interactive behaviors in magnified areas refers to the system recording the operator's actions when magnifying feature areas of the image, such as mouse clicks, dragging, scrolling, and keyboard input, and recording the time and duration of these actions. Interactive behavior types can include, but are not limited to, zooming, panning, selection, and annotation; the purpose is to capture the operator's specific operational intentions during detailed examination.

[0120] The adjustment of the gaze dwell time threshold based on the interaction behavior type and interaction time can be understood as the system dynamically adjusting the gaze dwell time threshold that triggers the display of the overall overview view based on the operator's current interaction state. For example, when the system detects that the operator is performing frequent and detailed zooming or panning operations, it indicates that the operator may be observing details in depth. In this case, the gaze dwell time threshold can be appropriately extended to avoid the overall overview view popping up prematurely and interfering with the operation. Conversely, when the system detects that the operator's interaction behavior in the local zoom-in area is less frequent or that the operator is in a static state, and the gaze lingers on the overview information display area for a long time, it indicates that the operator may need to obtain more macroscopic contextual information. In this case, the gaze dwell time threshold can be appropriately shortened to speed up the display of the overall overview view. In practical applications, different mapping relationships between interaction behavior types and threshold adjustment rules can be preset. For example, for "detailed zooming" behavior, the threshold is increased by 500 milliseconds; for "no interaction static" behavior, the threshold is decreased by 200 milliseconds.

[0121] This application's solution introduces the detection of the operator's gaze focus, gaze duration, and the type and duration of interaction behavior in the magnified area. Based on this information, it dynamically adjusts the gaze dwell time threshold, enabling the system to more intelligently determine when the operator truly needs to switch to the overall overview view. Because the gaze dwell time threshold can adaptively adjust according to the operator's actual operating context, the timing of displaying the overall overview view better aligns with the operator's cognitive habits and diagnostic process, avoiding interference or delays that might arise from a fixed threshold.

[0122] In some preferred embodiments, a specific example is given below. Suppose an operator is observing a magnified region of an image in detail.

[0123] First, the system uses an eye-tracking device to detect that the operator's gaze remains focused on the overview information display area for an extended period of time and records the duration of the gaze.

[0124] At the same time, the system also captures the operator's interactive behavior in the magnified area.

[0125] Scenario 1: If the system detects that the operator is frequently using the mouse wheel for fine zooming, and the interaction time is relatively long, it indicates that the operator may be delving into the details of the lesion. In this case, the system will extend the preset gaze dwell time threshold based on the interaction behavior type of "fine zooming," for example, from 500 milliseconds to 1000 milliseconds. In this way, even if the operator's gaze briefly lingers on the overview information display area, it will not immediately trigger the display of the overall overview view, thus avoiding interference with the operator's current fine-grained operation.

[0126] Scenario 2: If the system detects that the operator has no interactive behavior in the zoomed-in area and their gaze remains on the overview information display area for an extended period, it indicates that the operator may have finished observing the local details and wants to obtain more macroscopic contextual information. In this case, the system will shorten the preset gaze dwell time threshold based on the "no-interaction stillness" interaction behavior type, for example, adjusting it from 500 milliseconds to 300 milliseconds. Thus, once the operator's gaze time exceeds the adjusted shorter threshold, the system will quickly display the overall overview view, promptly meeting the operator's information needs.

[0127] Ultimately, the system determines whether to display the overall overview view and the identification results corresponding to potential lesion areas based on the operator's gaze focus, fixation time, and the adjusted gaze dwell time threshold. Through this dynamic adjustment mechanism, the system can more intelligently respond to the operator's true intentions, providing a smoother and more efficient assisted diagnostic experience.

[0128] In practical applications, this application further proposes an optimization scheme, which aims to provide a more spatial and interactive diagnostic information display by introducing a three-dimensional spatial model of the digestive tract and combining it with the spatial information of the recognition results, so as to assist doctors in more accurately locating and assessing potential lesions.

[0129] The above methods also include:

[0130] Obtain the spatial position information of the endoscope probe and the probe position coordinates corresponding to the recognition results, and construct a three-dimensional spatial model of the digestive tract.

[0131] The recognition result is projected onto the digestive tract region corresponding to the three-dimensional spatial model;

[0132] Based on the depth information of the recognition results in the three-dimensional spatial model, a depth visual label is assigned to each recognition result;

[0133] The overall overview view displays a three-dimensional spatial model of the digestive tract and the recognition results with depth visual markers.

[0134] The 3D spatial model can be rotated, scaled, and moved using 3D interactive tools;

[0135] When a selection event is received for the recognition result in the three-dimensional spatial model, the recognition result is locally magnified.

[0136] Specifically, acquiring the spatial position information of the endoscopic probe and the corresponding probe position coordinates for the recognition results refers to the real-time acquisition of the probe's three-dimensional spatial coordinates and orientation information within the patient's digestive tract by sensors integrated on the endoscopic probe (such as an inertial measurement unit, optical tracker, or electromagnetic positioning system) during the endoscopic examination. Simultaneously, when the artificial intelligence model identifies a potential lesion area and generates a recognition result, the system records the corresponding endoscopic probe position coordinates. This position information and coordinate data form the basis for constructing the three-dimensional spatial model.

[0137] Furthermore, constructing a three-dimensional spatial model of the digestive tract can be understood as using the probe position information, endoscopic images, and possible pre-defined anatomical models obtained above, and generating a virtual three-dimensional model of the patient's digestive tract through three-dimensional reconstruction algorithms (such as simultaneous localization and mapping, structured light scanning, or multi-view solid geometry). This model can accurately reflect the tortuosity, lumen size, and internal structure of the digestive tract.

[0138] Projecting the recognition results onto the corresponding digestive tract region of the three-dimensional spatial model refers to precisely mapping the two-dimensional image information of the potential lesion region identified by artificial intelligence onto the corresponding surface position on the constructed three-dimensional digestive tract model, based on its corresponding probe position coordinates and the endoscopic image's perspective. This gives the lesion region a clear location in three-dimensional space.

[0139] In practical applications, assigning a depth visual label to each recognition result based on its depth information in a three-dimensional spatial model means that the system assigns different visual representations to the lesion area according to its relative depth in the three-dimensional model (such as its distance from the endoscope probe or its depth from the digestive tract wall). For example, depth can be represented by color gradients, changes in transparency, size adjustments, or special textures, allowing doctors to intuitively perceive the hierarchical relationship of the lesion in three-dimensional space.

[0140] In addition, the overall overview view displays a three-dimensional spatial model of the digestive tract and identification results with depth visual markers. This means that in addition to displaying traditional two-dimensional images or overview information on the monitor, an interactive three-dimensional model of the digestive tract is overlaid, and all identified potential lesion areas are clearly marked on the model. These marks have the aforementioned depth visual markers.

[0141] As a preferred implementation, controlling the rotation, scaling, and movement of the three-dimensional spatial model through a three-dimensional interactive tool means that doctors can freely adjust the perspective, size, and position of the three-dimensional model using three-dimensional interactive devices such as a mouse, touchpad, gesture recognition device, or dedicated joystick, so as to observe the lesion area and its surrounding structures from different angles.

[0142] Therefore, when a selection event is received for the recognition result in the three-dimensional spatial model, the recognition result is magnified locally. This means that when a doctor clicks, selects, or focuses on a lesion marker on the three-dimensional model using interactive tools, the system will automatically magnify the lesion area locally in the three-dimensional model and may simultaneously display detailed image fragments, recognition criteria, patient history information, etc., corresponding to the lesion area, so that the doctor can conduct a more in-depth analysis.

[0143] This application's solution effectively addresses the issue of insufficient spatial information in the aforementioned overall overview view by introducing a three-dimensional spatial model of the digestive tract. Specifically, by acquiring the spatial position information of the endoscopic probe and the probe's position coordinates corresponding to the recognition results, the system can accurately construct a three-dimensional spatial model of the patient's digestive tract. Subsequently, the potential lesion areas identified by artificial intelligence are projected onto this three-dimensional model, integrating the originally scattered two-dimensional recognition results into a unified three-dimensional spatial background. Furthermore, depth visual markers are assigned based on the depth information of the recognition results in the three-dimensional model, enabling doctors to intuitively perceive the relative position and hierarchical relationship of lesions within the digestive tract, compensating for the shortcomings of traditional two-dimensional views in depth perception. Displaying this three-dimensional model and the recognition results with depth visual markers in the overall overview view provides doctors with a global and three-dimensional perspective, enabling them to more comprehensively understand the spatial distribution of lesions. By rotating, scaling, and moving the model using three-dimensional interactive tools, doctors can examine lesions from any angle, greatly enhancing the flexibility and depth of observation. When a doctor selects a recognition result in the 3D model, the system can quickly zoom in on that area and retrieve detailed information, thus achieving a seamless switch from a macroscopic 3D overview to microscopic detailed analysis, significantly improving the efficiency and accuracy of diagnosis.

[0144] In some preferred embodiments, a specific example is illustrated below. Assume a patient is undergoing an endoscopy. The artificial intelligence system acquires endoscopic images in real time and identifies multiple potential lesion areas within the digestive tract. Simultaneously, the positioning module built into the endoscope probe continuously transmits precise three-dimensional position and orientation data of the probe. Using this data and the endoscopic images, the system constructs or updates a three-dimensional spatial model of the patient's digestive tract in real time. When the AI ​​identifies a polyp in the stomach and a duodenal ulcer, the system records the probe's position coordinates corresponding to these two lesion areas. Subsequently, these two identification results are precisely projected onto the surface of the constructed three-dimensional model of the stomach and duodenum. Based on the relative depth of the polyp and ulcer in the three-dimensional model, the system assigns a light red, semi-transparent depth visual marker to the polyp and a dark red, opaque depth visual marker to the ulcer to distinguish their depth levels. On the doctor's diagnostic display screen, in addition to displaying real-time endoscopic images and overview information, an interactive three-dimensional model of the digestive tract is also displayed side-by-side. On this 3D model, doctors can see the overall structure of the stomach and duodenum, as well as polyps and ulcers with visual markers at different depths. Doctors can use the mouse to drag the 3D model to rotate it and observe the relative positions of polyps and ulcers from the side; they can also scroll the mouse wheel to zoom in for a clearer view of lesion details. When a doctor clicks on a polyp marker on the 3D model, the system immediately zooms in on that polyp area and simultaneously displays a detailed image segment corresponding to the polyp, the AI ​​recognition criteria (such as texture and color features), and the patient's past medical history on the side of the screen. This intuitive 3D interactive method allows doctors to quickly and comprehensively understand the spatial distribution and detailed characteristics of lesions, thereby making more accurate diagnostic decisions.

[0145] Secondly, referring to Figure 2 This application further proposes an endoscopic AI-assisted diagnostic system, which includes:

[0146] Image acquisition module 201 is used to acquire endoscopic images;

[0147] The region recognition module 202 is used to identify potential lesion regions in the endoscopic image and generate scoring information corresponding to the potential lesion regions;

[0148] Threshold determination module 203 is used to determine whether the scoring information is greater than a preset score threshold;

[0149] The information recording module 204 is used to generate and store recording information containing the time point, image segment, probe position and recognition result of the corresponding potential lesion area when the scoring information is greater than a preset score threshold.

[0150] The information display module 205 is used to display the recorded information when the inspection is completed.

[0151] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A diagnostic method based on AI under endoscopy, characterized in that, include: Acquire endoscopic images; Identify potential lesion areas in the endoscopic images and generate scoring information corresponding to the potential lesion areas; Determine whether the scoring information is greater than a preset score threshold; When the scoring information is greater than a preset score threshold, a record information containing the time point, image segment, probe position, and recognition result of the corresponding potential lesion area is generated and stored. When the inspection is complete, the recorded information is displayed.

2. The endoscopic AI-assisted diagnostic method according to claim 1, characterized in that, The method further includes: The flatness value of the feature extraction layer in the potential lesion area is calculated to obtain the evaluation result; Under different endoscopic imaging modes, lesion features of potential lesion areas are identified, the consistency between multiple lesion features is verified, and the verification results are obtained. Based on the evaluation results and the verification results, the uncertainty level of the identification result is determined; When the uncertainty level of the identification result is higher than a preset level threshold, feature extraction is performed on the image segment corresponding to the potential lesion area to obtain the feature extraction result; Based on the feature extraction results, various diagnostic information and corresponding probability scores are generated; The system displays various diagnostic information and their corresponding probability scores, and based on the selected events for the diagnostic information, it displays the evidence information corresponding to the diagnostic information, including image features, identification criteria, and patient history information.

3. The endoscopic AI-assisted diagnostic method according to claim 2, characterized in that, The method further includes: The evidence information is weighted and sorted to obtain the weighted and sorted results; Based on the weighted and sorted results, an evidence presentation path is generated; the evidence presentation path represents the display order or display process of the evidence. Based on the evidence presentation path, the image feature regions in the image segments are displayed in layers, and an interactive evidence navigation bar is generated; When a selection event is received for the evidence navigation bar, the image feature area is magnified locally, and the identification basis and patient history information corresponding to the image feature area are displayed; as well as diagnostic information, probability score, and overview information; wherein, the overview information refers to an overview display of the evidence information.

4. The endoscopic AI-assisted diagnostic method according to claim 2, characterized in that, The method further includes: Deformation assessment was performed on the digestive tract mucosa region in the endoscopic images to obtain deformation assessment results; When the deformation assessment results show that there is transient deformation of the digestive tract mucosa, the preset level threshold is adjusted according to the deformation assessment results.

5. The endoscopic AI-assisted diagnostic method according to claim 3, characterized in that, The image feature region includes a highlighted area; the evidence navigation bar includes a feature selection tool; the step of displaying the image feature regions in the image segment in layers according to the evidence presentation path and generating an interactive evidence navigation bar includes: Identify lesion features in highlighted areas of an image segment; wherein the lesion features are correlated with different diagnostic information; Generate visual identifiers for each lesion feature; The lesion features are displayed in layers and overlaid according to the visual identifiers; When a selection event for the lesion feature is received, the display effect of other lesion features is adjusted through the feature selection tool, and the identification basis and patient history information corresponding to the selected lesion feature are displayed at the same time.

6. The endoscopic AI-assisted diagnostic method according to claim 3, characterized in that, When a selection event is received for the evidence navigation bar, the image feature region is partially magnified, and the identification criteria and patient history information corresponding to the image feature region are displayed; as well as diagnostic information, probability scores, and overview information, including: Set the overview information display area; The overview information display area displays diagnostic information, probability scores, and overview information; When an interactive operation is received targeting the feature region of the magnified image, the transparency of the overview information display area is adjusted. When the gaze duration on the overview information display area exceeds the gaze dwell time threshold, the overall overview view and the identification results corresponding to the potential lesion areas are displayed.

7. The endoscopic AI-assisted diagnostic method according to claim 6, characterized in that, When the gaze duration over the overview information display area exceeds the gaze dwell time threshold, the overall overview view and the identification results corresponding to the potential lesion areas are displayed, including: Detect the operator's visual focus and fixation time; Obtain the type and duration of interactive behaviors in the zoomed-in area; Adjust the gaze dwell time threshold according to the interaction type and interaction time; Based on the gaze focus, the gaze duration, and the adjusted gaze dwell time threshold, determine whether the gaze duration for the overview information display area exceeds the adjusted gaze dwell time threshold. When the gaze duration for the overview information display area exceeds the adjusted gaze dwell time threshold, an overall overview view and the identification results corresponding to the potential lesion areas are displayed.

8. The endoscopic AI-assisted diagnostic method according to claim 7, characterized in that, The method further includes: Obtain the spatial position information of the endoscope probe and the probe position coordinates corresponding to the recognition results, and construct a three-dimensional spatial model of the digestive tract. The recognition result is projected onto the digestive tract region corresponding to the three-dimensional spatial model; Based on the depth information of the recognition results in the three-dimensional spatial model, a depth visual label is assigned to each recognition result; The overall overview view displays a three-dimensional spatial model of the digestive tract and the recognition results with depth visual markers.

9. The endoscopic AI-assisted diagnostic method according to claim 8, characterized in that, The method further includes: The 3D spatial model can be rotated, scaled, and moved using 3D interactive tools; When a selection event is received for the recognition result in the three-dimensional spatial model, the recognition result is locally magnified.

10. An endoscopic AI-assisted diagnostic system, characterized in that, The system includes: Image acquisition module, used to acquire endoscopic images; The region recognition module is used to identify potential lesion regions in the endoscopic image and generate scoring information corresponding to the potential lesion regions; The threshold determination module is used to determine whether the scoring information is greater than a preset score threshold; The information recording module is used to generate and store recording information containing the time point, image segment, probe position, and recognition result of the corresponding potential lesion area when the scoring information is greater than a preset score threshold. The information display module is used to display the recorded information when the inspection is completed.