Methods and electronic devices for image processing

The method and device improve endoscopic image processing by predicting endoscope location and identifying target sites within the body, addressing accuracy and speed issues in AI-assisted diagnosis.

JP7835253B2Active Publication Date: 2026-03-25NEC CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing image processing technologies for medical endoscopes lack accuracy and speed in determining the body part being examined, particularly in distinguishing between gastric and intestinal endoscopy, leading to subjectivity and potential misjudgments in AI-assisted diagnosis.

Method used

A method and device that monitor image sequences collected by an endoscope to predict its location within the body, determine a reference image, and identify the target site using trained classification models and feature matching, ensuring accurate and rapid determination of the examination area without external input.

Benefits of technology

Automates the identification of endoscopic examination sites, enhancing accuracy and speed by minimizing noise and algorithmic inaccuracies, enabling timely and precise application of subsequent AI tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835253000001
    Figure 0007835253000001
  • Figure 0007835253000002
    Figure 0007835253000002
  • Figure 0007835253000003
    Figure 0007835253000003
Patent Text Reader

Abstract

To provide a method for image processing and an apparatus.SOLUTION: A process 400 comprises: Step 410 for monitoring a prediction result indicating whether an endoscope is located in the body of a test subject at a time point when a corresponding image is collected, the prediction result being a result corresponding to an image in a first image sequence collected by the endoscope chronologically during a test; Step 420 for determining a reference image corresponding to entry of the endoscope to the body of the test object from the first image sequence, on the basis of the monitoring; Step 430 for acquiring a second image sequence collected by the endoscope after a collection time of the reference image; and Step 440 for determining a target portion of the test subject to be tested by the endoscope, on the basis of the second image sequence. By the process, identification of the target portion can be completed automatically without depending on external input, therefore correctness and a speed of a preliminary step operation of an endoscope test can be secured.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary embodiments of the present disclosure generally relate to the field of computers, and more particularly to methods and devices for image processing.

Background Art

[0002] Image processing technologies are widely used in various fields. In the medical field, the demand for processing and analyzing a large amount of medical image data is increasing. Particularly for medical endoscope images, a solution for doctors to make more accurate and rapid diagnoses is required.

Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for image processing is provided. The method includes monitoring a prediction result for an image in a first image sequence collected over time by an endoscope during an examination, the prediction result indicating whether the endoscope was located inside the body of the subject being examined when the corresponding image was collected; determining, based on the monitoring, a reference image corresponding to the endoscope entering the body of the subject being examined from the first image sequence; obtaining a second image sequence collected by the endoscope after the collection time of the reference image; and determining a target site of the subject being examined by the endoscope based on the second image sequence.

[0004] A second aspect of this disclosure provides an electronic device, which includes at least one processing circuit. The at least one processing circuit is configured to monitor a prediction result of correspondences to images in a first image sequence collected over time by an endoscope during an examination, indicating whether the endoscope is located inside the body of the subject being examined at the time the corresponding image was collected; to determine a reference image from the first image sequence that corresponds to the endoscope entering the body of the subject being examined, based on the monitoring; to acquire a second image sequence collected by the endoscope after the reference image collection time; and to determine a target site of the subject being examined by the endoscope based on the second image sequence.

[0005] In some embodiments of the second aspect, the first image sequence includes a first image, and the corresponding prediction result includes a first prediction result for the first image. At least one processing circuit is further configured to generate a first prediction result for the first image in the first image sequence, based on a trained first classification model, regarding whether the endoscope is inside the body of the subject being examined at the time the first image was collected.

[0006] In some embodiments of the second aspect, the first image sequence includes a second image, and the corresponding prediction result includes a second prediction result for the second image. At least one processing circuit is further configured to perform the following for the second image in the first image sequence: extract image features of the second image; determine a first similarity of the image features to an internal reference feature that characterizes the inside of the body of the reference subject, and a second similarity of the image features to an external reference feature that characterizes the outside of the body of the reference subject; and generate a second prediction result regarding whether the endoscope is located inside the body of the subject being examined at the time the second image was collected, by comparing the first and second similarities.

[0007] In some embodiments of the second aspect, at least one processing circuit is further configured to determine a first number of adjacent images from a first image sequence before determining a reference image based on monitoring, and to set a prediction result for an intermediate image among the first number of adjacent images based on the prediction result of the correspondence of the first number of adjacent images.

[0008] In some embodiments of the second aspect, at least one processing circuit is further configured to: determine from a first number of adjacent images an image in which the corresponding prediction result indicates that the endoscope is located inside the body of the subject being examined; if the number of determined images is greater than a first threshold number associated with the first number, set the prediction result for the intermediate image to indicate that the endoscope is inside the body of the subject being examined; and if the number of determined images is not greater than the first threshold number, set the prediction result for the intermediate image to indicate that the endoscope is not inside the body of the subject being examined.

[0009] In some embodiments of the second aspect, at least one processing circuit is further configured to determine whether all predicted results of the correspondence of a second number of consecutive images in a first image sequence indicate that the endoscope is located inside the body of the subject being examined, and, in response to the determination that all predicted results of the correspondence of a second number of consecutive images indicate that the endoscope is located inside the body of the subject being examined, to determine a reference image based on the second number of consecutive images.

[0010] In some embodiments of the second aspect, the acquisition time span corresponding to the second number of consecutive images is related to the image acquisition frequency of the endoscope.

[0011] In some embodiments of the second aspect, at least one processing circuit is further configured to monitor whether the image quality of individual images collected by the endoscope after the reference image collection time exceeds a threshold image quality, and, in response to monitoring whether the image quality of the third image exceeds a threshold image quality, to make the images sequentially collected by the endoscope within a certain time after the collection time of the third image a second image sequence.

[0012] In some embodiments of the second aspect, the second image sequence satisfies at least one of the following conditions: the acquisition time span corresponding to the second image sequence is equal to the threshold span, or the cumulative image distance between adjacent images in the second image sequence is greater than the threshold, and the image distance is determined based on optical flow.

[0013] In some embodiments of the second aspect, at least one processing circuit is further configured to generate a first identification result regarding which of a plurality of predetermined regions the third image in the second image sequence corresponds to, based on a trained second classification model.

[0014] In some embodiments of the second aspect, at least one processing circuit is further configured to perform the following for a fourth image in a second image sequence: extract image features of the fourth image; determine the similarity of the correspondence of the image features to a plurality of reference features used to characterize each of a plurality of predetermined regions; and generate a second identification result regarding which of the plurality of predetermined regions the fourth image corresponds to by comparing the similarity of the correspondences.

[0015] In some embodiments of the second aspect, at least one processing circuit is further configured to determine that the target site is the first site in response to the number of identification results indicating the first site among a plurality of predetermined sites in the corresponding identification results exceeding a second threshold number, or to determine that the target site is the second site in response to the proportion of identification results indicating the second site among a plurality of predetermined sites in the corresponding identification results exceeding a threshold proportion.

[0016] A third aspect of this disclosure provides an electronic device comprising at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions to be executed by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device causes the device to perform the method described in the first aspect.

[0017] A fourth aspect of this disclosure provides a computer-readable storage medium which stores a computer program which, when executed by a processor, causes the processor to perform the method described in the first aspect.

[0018] It should be understood that the contents described in the summary section of the present invention are not intended to limit the main or important features of the embodiments of this disclosure, nor do they limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description. [Brief explanation of the drawing]

[0019] The above and other features, advantages, and aspects of each embodiment of this disclosure will become more apparent upon further detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals indicate the same or similar elements. [Figure 1] A schematic diagram of an exemplary environment in which the embodiments of this disclosure can be implemented is shown. [Figure 2] A schematic diagram of an exemplary image processing architecture according to some embodiments of this disclosure is shown. [Figure 3] A schematic diagram of an example of determining whether the stomach or intestine is present according to some embodiments of this disclosure is shown. [Figure 4] A flowchart of the image processing process according to some embodiments of this disclosure is shown. [Figure 5] A block diagram of an electronic device capable of carrying out multiple embodiments of this disclosure is shown. [Modes for carrying out the invention]

[0020] The embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While the accompanying drawings show several embodiments of this disclosure, it should be understood that this disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] Please note that the headings of any section / subsection provided herein are not limiting. Various examples are described throughout this specification, and any type of example may be included in any section / subsection. In addition, examples described in any section / subsection may be combined in any way with other examples described in the same section / subsection and / or different sections / subsections.

[0022] In the description of embodiments of the present disclosure, the term "comprising" and its similar terms should be understood as non-limiting inclusion, that is, "including but not limited to them". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "this embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions included hereinafter.

[0023] As used herein, the term "circuit" may refer to a hardware circuit and / or a combination of a hardware circuit and software. For example, a circuit may be a combination of analog and / or digital hardware circuits and software / firmware. As another example, a circuit may be any part of a hardware processor having software, and the hardware processor includes a (plurality of) digital signal processors, software, and a (plurality of) memories, and these cooperate to enable the device to operate so as to execute various functions. In yet another example, a circuit may be a hardware circuit and / or a processor, such as a microprocessor or a part of a microprocessor, which requires software / firmware for operation, but the software may not be present when it is not required for operation. As used herein, the term "circuit" also includes only the implementation of a hardware circuit or processor, or a part of a hardware circuit or processor, and the software and / or firmware associated therewith (or them).

[0024] As briefly described above, in the medical field, the demand for processing and analyzing a large amount of medical image data is increasing. Conventionally, doctors have had to rely on their own experience and visual judgment to analyze endoscopic images, perform corresponding examination items, or make appropriate diagnoses. However, this method has problems such as subjectivity, time-consuming, and possible misjudgments. Therefore, with the development of artificial intelligence (AI), it has come to be applied to the analysis of endoscopic images. Generally, different AI algorithms or models need to be developed and used for each part of the examination target (e.g., the human body). Therefore, correctly determining the part being examined is extremely important for subsequent image analysis.

[0025] For example, endoscopes are often used in the examination of the digestive tract. In the examination of the digestive tract, usually, the same endoscope device is used to switch to different lenses to perform gastric endoscopy and intestinal endoscopy. The lens used for gastric endoscopy has the "mouth" as the entrance from the outside of the patient's body to the inside, and the lens used for intestinal endoscopy has the "anus" as the entrance from the outside of the patient's body to the inside. During a one-day examination, according to the subject's condition, the examiner will occasionally switch between the gastric endoscope lens and the intestinal endoscope lens, and the corresponding examination program will be called after the lens is switched. How to accurately and quickly determine the part currently being examined is an essential task. In particular, in the case of AI-based assisted diagnosis, the automatic discrimination between gastric endoscopy and intestinal endoscopy is the pre-stage operation of subsequent AI tasks.

[0026] Conventionally, some approaches focused on the identification of parts or organs, so they cannot be directly applied to the scene of identifying the stomach or intestine. In some other approaches, for each frame, it is impossible to identify whether it is a certain part or organ and complete the identification task within a certain time period (e.g., within the first time period when the endoscope enters the patient's body). In some other approaches, there is no comprehensive judgment for multiple frames, so it is not suitable for the scene of identifying the stomach or intestine where it is necessary to comprehensively judge some frames collected when the endoscope enters the patient's body.

[0027] The above explains some of the challenges in image processing using endoscopy in the examination of the human digestive tract as an example. Please understand that similar image processing challenges may exist in endoscopic examinations of other subjects or body parts.

[0028] Therefore, embodiments of this disclosure provide an approach for image processing. This approach includes monitoring predictive results for correspondences between images in a first image sequence. Such a first image sequence is collected over time by an endoscope during the examination. The predictive results indicate whether the endoscope is located inside the body of the subject being examined at the time the corresponding image is collected. Based on monitoring the predictive results, a reference image corresponding to the endoscope entering the body of the subject is determined from the first image sequence. Furthermore, a second image sequence is acquired by the endoscope after the time the reference image was acquired. Based on the acquired second image sequence, the target area of ​​the subject being examined by the endoscope is determined.

[0029] Therefore, the system can automatically determine the corresponding body part and examination items from a limited number of collected images. In this way, accuracy and speed of the pre-endoscopic examination process can be ensured without relying on additional external input (such as text, voice, or keyboard input). In particular, in AI-assisted endoscopy, accurately and quickly determining the area to be examined is useful for the accurate and timely application of subsequent AI tasks.

[0030] Exemplary environment Figure 1 shows a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure can be carried out. In environment 100, an electronic device 120 may receive an image sequence 110 and output a determination result. The image sequence 110 includes a set of images 112 sorted by collection time. The images 112 are, for example, images of an object being examined collected by a medical endoscope. The images 112 may be collected in real time or offline. The determination result is, for example, a site examined by the endoscope, also called a target site. For example, in a gastrointestinal examination scene, the determination result may indicate a colonoscopy or a gastroscopy.

[0031] In the embodiments of this disclosure, the object under examination may be any suitable type of object, such as the human body or an animal. The endoscope may be used for a variety of suitable examinations, such as gastrointestinal examinations or pharyngeal examinations. In some of the following examples and embodiments, a distinction will be made between colonoscopy and gastroscopy in the examination of the human gastrointestinal tract. However, it should be understood that this is merely illustrative and not intended to be limiting in any way. The embodiments of this disclosure are not limited to the type of object under examination, the endoscope, or the part of the object under examination.

[0032] In environment 100, electronic device 120 may be any type of device having computing power, including terminal devices. Terminal devices may be any type of mobile terminal, fixed terminal, or portable terminal, and include mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book readers, game devices, or any combination thereof, and further include accessories and peripherals for these devices, or any combination thereof. Electronic device 120 may also be any type of device having computing power, including server devices. Server devices may include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in a cloud environment.

[0033] In particular, in some embodiments, the electronic device 120 may be a computing device that communicates directly with the endoscope. In some embodiments, the electronic device 120 may be implemented to be included in the endoscope, or the individual operations described below with reference to the electronic device 120 may be performed by the endoscope itself.

[0034] The structure and function of environment 100 are described for illustrative purposes only and should not be understood as limiting the scope of this disclosure. In Figure 1 and subsequent descriptions, electronic device 1 2 Device 0 collects the image sequence 110 and determines the result. However, this is illustrative and is not intended to limit the scope of the disclosure. In some embodiments, the collection of the image sequence 110 and the determination of the result may be performed by different devices.

[0035] The following description will continue with illustrative examples of the present disclosure, with reference to the attached drawings.

[0036] Exemplary Architecture Figure 2 shows a schematic diagram of an exemplary image processing architecture according to some embodiments of the present disclosure. Architecture 200 can be implemented in an electronic device 120. Architecture 200 generally comprises a monitoring module 210, a reference image determination module 220, and a target site determination module 235. Additionally or alternatively, architecture 200 may further comprise an endoscope configured to collect an image sequence. Such an image sequence may include images collected by the endoscope during the process of the endoscope entering the body of a subject from outside the body. Architecture 200 will be described below with reference to Figure 1.

[0037] Continuing to refer to Figure 2, the monitoring module 210 receives the first image sequence 205. The first image sequence 205 may be a set of images collected by the endoscope over time during the examination process. It should be understood that images collected by the endoscope over time during the examination process may be considered as video. Therefore, images collected by the endoscope are also called frames, meaning that the first image sequence 205 may contain multiple frames. The monitoring module 210 may monitor multiple frames and make predictions for each of them. Furthermore, the monitoring module 210 may output a prediction result 215 for each frame. Such a prediction result 215 indicates whether the endoscope is located inside the body of the subject being examined at the time the corresponding frame was collected by the endoscope. That is, the prediction result 215 is a prediction for a single frame. To simplify the explanation below, indicating whether the endoscope is located inside the body of the subject being examined will simply be referred to as "in the body".

[0038] In some embodiments, deep learning methods may be employed to perform such single-frame predictions. For example, the monitoring module 210 may use a trained classification model (also called the first classification model) to generate a prediction (also called the first prediction) about whether an endoscope is inside the body at the time the individual image was collected, for each image (e.g., the first image) in the first image sequence 205.

[0039] For example, to train this classification model, internal and external data of the same type of subject being examined may be collected. For instance, if the subject is a person and the area being examined is the stomach, the internal data would include a series of images of the oral cavity, throat, esophagus, stomach, etc. Also, for example, if the subject is a person and the area being examined is the intestines, the internal data would include a series of images of the buttocks, anus, intestines, etc. The internal / external data may then be collected and annotated. Furthermore, the annotated data may be used to construct a binary classification deep learning model. It should be understood that the actions of collecting and annotating data, and training the model, may be performed on devices other than the electronic device 120, and this disclosure is not limiting thereto.

[0040] The monitoring module 210 may use such a deep learning model (e.g., a binary classification model) to analyze the first image sequence 205 and predict whether the endoscope is inside or outside the body. For example, the monitoring module 210 may use such a deep learning model to obtain a prediction score for being inside the body. Based on a preset threshold score for being inside the body, if the prediction score exceeds the threshold score, the monitoring module 210 determines that the prediction result corresponding to the current frame is inside the body; otherwise, it determines that it is outside the body.

[0041] Alternatively or additionally, in some embodiments, a feature matching method may be employed to perform single-frame prediction. For example, for individual images in the first image sequence 205 (e.g., the second image), the monitoring module 210 extracts image features from these images. The monitoring module 210 determines the similarity of the extracted image features to internal reference features (also called the first similarity) and to external reference features (also called the second similarity). Internal reference features characterize the inside of the reference object, and external reference features characterize the outside of the reference object. In other words, internal reference features represent features inside the reference object, and external reference features represent features outside the reference object. Furthermore, by comparing these two similarities, the monitoring module 210 generates a prediction result (also called the second prediction result) of whether the endoscope is inside the body or not. Such a prediction result indicates whether the endoscope is located inside the body of the subject being examined at the time the corresponding image was collected. For example, for a given image, if the first similarity is greater than the second similarity, the prediction result may indicate an internal location; otherwise, the prediction result may indicate an external location.

[0042] For example, in-vivo and in-vivo data may be collected for the same type of subject under examination for feature matching, and image features may be extracted from this data. Feature operators include, but are not limited to, color histograms, texture histograms, and deep network features. Furthermore, the electronic device 120 uses a clustering algorithm (e.g., the K-means algorithm) to spatially cluster the in-vivo and in-vivo features and construct a feature library. The data collection operation may be performed on a device other than the electronic device 120, and this disclosure is not limited thereto.

[0043] The monitoring module 210 may use such a feature library to calculate feature distance ratios for the first image sequence 205 in order to obtain intracellular feature distance ratios. For example, based on a pre-set intracellular threshold feature distance ratio, the monitoring module 210 determines that the current frame corresponds to an intracellular image if the feature distance ratio of the current frame is greater than the threshold feature distance ratio, and otherwise determines that it corresponds to an extracellular image.

[0044] Continuing to refer to Figure 2, the reference image determination module 220 receives the prediction result 215 and determines the reference image 225. Such a reference image 225 corresponds to the endoscope entering the body of the subject being examined. That is, the reference image determination module 220 receives the prediction result 215 for each frame and determines which frame or sequence of frames indicates that the endoscope entered the body of the subject being examined. As can be seen, the determination of the reference image forms the basis for subsequent operations.

[0045] For example, the first image sequence 205 includes 12 frames, for instance, in chronological order: frame 1, frame 2, frame 3, ..., frame 12. Assume that the prediction results corresponding to frames 1 through 5 all indicate an external location, and the prediction result corresponding to frame 6 indicates an internal location. Ideally, since all frames from frame 6 onward can be predicted to indicate an internal location, frame 6 can be determined as the reference image. However, in some cases, the single-frame prediction results described above may be affected by device noise, environmental noise, or algorithmic accuracy, potentially leading to misclassification.

[0046] In some embodiments, the reference image determination module 220 may determine a predetermined number (also called the first number) of adjacent images from the first image sequence 205 before determining the reference image 225. Such adjacent images include, for example, a predetermined number of consecutive images collected before the image currently being determined and a predetermined number of consecutive images collected thereafter. The reference image determination module 220 may set the prediction result of an intermediate image among these images based on the corresponding prediction results of these images. For example, if the corresponding prediction results of these images all indicate that the image is inside the body, the prediction result of an intermediate image among these images may be set to indicate that the image is inside the body. In this way, the prediction results of a single frame can be smoothed to achieve objectives such as noise reduction and removal of the influence of algorithmic accuracy, thereby improving the accuracy of reference image determination.

[0047] Continuing with the above example, let's assume the first number is 5 frames. When determining whether the 6th frame is a reference image, the adjacent frames collected before the 6th frame are the 4th and 5th frames, and the adjacent frames collected afterward are the 7th and 8th frames. The prediction results for single frames from the 4th to the 8th frame could all indicate that the image is inside the body, or some of the prediction results for these frames could indicate that the image is inside the body. The reference image determination module 220 can reduce the risk of misclassification of the reference image by smoothing the prediction results for single frames by combining the prediction results corresponding to these 5 frames and determining whether the 6th frame is inside or outside the body.

[0048] In some embodiments, the reference image determination module 220 may further determine from a first number of adjacent images an image having a corresponding prediction result indicating that the endoscope is located inside the body of the subject being examined, and determine the number of such images. If the number is greater than a threshold for the first number (also called the first threshold number), the prediction result of an intermediate image among these adjacent images is set to indicate that the endoscope is inside the body of the subject being examined. If the number is not greater than the threshold, the prediction result of an intermediate image among these adjacent images is set to indicate that the endoscope is not inside the body of the subject being examined. The first threshold number may be less than or equal to the first number. The method of smoothing the prediction results indicating whether the endoscope is inside or outside the body based on the prediction results for each frame of a series of frames may also be called smoothing or local smoothing. In this way, the accuracy of the reference image determination can be improved.

[0049] Continuing with the above example, let's assume that the first number is 5 and the corresponding first threshold number is 3. When smoothing the prediction result of a single frame in the 6th frame, the adjacent frames collected before the 6th frame are the 4th and 5th frames, and the adjacent frames collected thereafter are the 7th and 8th frames. If three or more prediction results among the images of these five frames indicate that the subject is present in the body, then the prediction result of the 6th frame is set as being present in the subject's body.

[0050] Furthermore, after determining the prediction results (original or after smoothing) for a single frame, the reference image determination module 220 determines the reference image based on these prediction results.

[0051] However, in some cases, if the frame in which the in-body prediction result is first shown is designated as the reference image, it may lead to misidentification of the reference image. For example, the prediction results for frames 3, 9, 10, and 11 of frames 1 through 11 indicate that the object is inside the body. Considering that frame 3 may be affected by noise, the endoscope actually entered the body of the subject from frame 9 onwards. In this case, using frame 3 as the reference image would result in a certain bias.

[0052] Therefore, in some embodiments, a certain amount of time may be set as a buffer when determining the reference image. For example, the reference image determination module 220 may determine whether the corresponding prediction results of a predetermined number (also called the second number) of consecutive images in the first image sequence 205 all indicate that the endoscope is located inside the body of the subject being examined. Furthermore, the reference image determination module 220 may determine the reference image based on that number of consecutive images. For example, the reference image determination module 220 may determine the last frame of the consecutive frames that satisfy the second number as the reference image. As a result, the effect of the time buffer can be obtained. In this way, the accuracy of reference image determination can be further improved. Note that the time buffer is assumed to correspond to K3 frames. If the prediction results of all consecutive K3 frames indicate outside the body, the last frame of these consecutive K3 frames may be determined as the reference image.

[0053] Continuing with the above example, let's assume the second number is 3. When determining the reference image, if the prediction results for the 3rd, 9th, 10th, and 11th frames indicate that the image is inside the body, then the prediction results for consecutive frames that satisfy the second number (i.e., frames 9 through 11) will all indicate that the image is outside the body. The reference image determination module 220 may continue determining subsequent frames until it determines frame 9 as the reference image, when it indicates that the prediction results for consecutive frames that satisfy the second number (i.e., frames 9 through 11) all indicate that the image is inside the body.

[0054] The second number may be set as a fixed value, or it may be set as the number of frames corresponding to a predetermined collection time span, for example, 3 frames or the number of frames corresponding to 1 second. In some embodiments, the collection time span corresponding to the second number of consecutive images is related to the endoscopic image collection frequency. For example, if the collection frequency is 60 fps, the collection time span is 1 second. In another example, if the collection frequency is 15 fps, the collection time span is 4 seconds.

[0055] Continuing to refer to Figure 2, the acquisition time of reference image 225 may be used to determine the start of timing. For example, this acquisition time may be used as the time to start timing, and the image sequence acquired thereafter (also called the second image sequence 230) may be transmitted to the target site determination module 235. Based on such an image sequence, the target site determination module 235 determines the examination site (also called the target site) of the subject to be examined by the endoscope. That is, timing starts from the acquisition time determined based on the first image sequence 205, and the examination site is determined based on the second image sequence 230 acquired thereafter. In this way, the effects of noise and algorithm accuracy can be minimized, and a valid image sequence can be acquired as early as possible.

[0056] In some embodiments, the acquisition time of the reference image may not be used as the start of timing. To obtain a high-quality second image sequence 230, the quality of individual images acquired after the acquisition time of the reference image 225 may be monitored to determine whether the quality exceeds a threshold quality. Furthermore, the acquisition time of an image whose monitored quality exceeds the threshold quality (also called a third image) may be used as the start of timing, and images acquired sequentially by the endoscope within a certain time after that acquisition time may be used as the second image sequence 230. In such embodiments, the quality may be determined by any suitable method (e.g., a supervised learning-based deep learning model, grayscale variance product, etc.).

[0057] The endpoint of the second image sequence 230, i.e., the timing of the end of timing, may be determined in a variety of preferred ways. In some embodiments, the second image sequence 230 satisfies the condition that the corresponding collection time span is equal to a threshold span. Such a threshold span may be a preset fixed value, for example, 5 seconds. That is, timing starts from the effective time of the endoscope's entry into the body and ends after 5 seconds. The second image sequence 230 includes all frames collected during this time. This ensures that a valid image sequence is collected in the shortest possible time. In this way, the examination site can be determined based on a small number of valid images, thereby not hindering subsequent work.

[0058] Alternatively or additionally, in some embodiments, the acquisition of the second image sequence 230 may use adaptive timing. The second image sequence 230 satisfies the condition that the cumulative image distance between adjacent images is greater than a threshold. Such image distances are determined based on optical flow. This ensures the acquisition of a stable image sequence. In this way, the effectiveness of acquiring the second image sequence can be improved.

[0059] For example, after timing has started, a single-frame optical flow field (e.g., a sparse optical flow field) may be extracted for each frame collected, and the distance between each of the two frames may be calculated. Alternatively or additionally, if the image quality of a frame falls below a threshold, the target site determination module 235 may remove that frame and continue tracking the distance for subsequent frames. In this way, images that cannot be tracked due to excessively low image quality can be removed. Furthermore, these distances may be accumulated, for example, from the starting frame collected at the start of timing to each frame collected thereafter. This accumulated distance is compared to the threshold distance. If the accumulated distance is greater than the threshold distance, timing is stopped. The second image sequence 230 includes all frames between the starting frame at the start of timing and the final frame at the end of timing.

[0060] Alternatively or additionally, the timing of the acquisition time for the second image sequence 230 may be stopped if either the acquisition time span reaches a threshold span or the cumulative distance between images is greater than the threshold distance.

[0061] The above embodiment explained how the time period for collecting the second image sequence 230 is determined. Below, we will continue to explain how the target site determination module 235 determines the target site after acquiring the second image sequence 230.

[0062] In some embodiments, the target site determination module 235 may generate a corresponding identification result for the images in the second image sequence 230, i.e., perform single-frame identification. Such an identification result indicates one of a plurality of predetermined sites. Therefore, the target site determination module 235 can determine the target site based on the corresponding identification result. For example, predetermined sites include the esophagus, stomach, and intestines. The target site determination module 235 may identify from each frame of the second image sequence 230 whether or not it contains features corresponding to these sites, and further determine which site of the body to be examined is currently being examined based on the identification results of these frames. In this way, the site to be examined by the endoscope can be automatically determined.

[0063] In some embodiments, single-frame identification of a target region may be performed using a deep learning method. For example, for an image in the second image sequence 230 (e.g., the third image), the target region determination module 235 may generate an identification result (also called a first identification result) based on a trained classification model (also called a second classification model). Such an identification result can indicate which examination region (also called a predetermined region) each individual frame in the second image sequence 230 corresponds to. For example, to obtain such a classification model, past examination data of these examination regions may be collected and annotated. Then, using the annotated past examination data, a deep learning model for multiple classifications (e.g., binary classification) may be constructed. The operations of collecting and annotating data and training models may be performed on devices other than the electronic device 120, and this disclosure is not limited thereto.

[0064] The target site determination module 235 may use such a deep learning model to analyze the second image sequence 230 and predict the current examination site. For example, the examination sites include the stomach and intestines. The target site determination module 235 uses the deep learning model to obtain a prediction score. Based on a preset threshold score for the stomach, if the prediction score exceeds the threshold score corresponding to the stomach, the target site determination module 235 determines that the examination site corresponding to the current frame is the stomach; otherwise, it determines that it is the intestines.

[0065] In some embodiments, single-frame identification of a target region may be performed using feature matching. For example, for an image in the second image sequence 230 (e.g., the fourth image), the target region determination module 235 may extract its image features. Furthermore, the target region determination module 235 determines the corresponding similarity between the extracted image features and several reference features. Each of the several reference features is used to characterize several examination regions. By comparing the corresponding similarities, the target region determination module 235 generates an identification result (also called a second identification result). Such an identification result can indicate which examination region each frame in the second image sequence 230 corresponds to.

[0066] For example, to obtain the reference features to be used, historical inspection data of these inspection sites may be collected, and image features may be extracted from this historical inspection data. Feature operators include, but are not limited to, color histograms and texture histograms. Furthermore, the electronic device 120 uses a clustering algorithm (e.g., the K-means algorithm) to spatially cluster features of different inspection sites and build a feature library. The data collection operation may be performed on a device other than the electronic device 120, and this disclosure is not limited thereto.

[0067] The target site determination module 235 may use such a feature library to calculate the feature distance ratio for the second image sequence 230 and obtain the feature distance ratio. For example, the examination site may include the stomach and intestines. Based on a preset threshold feature distance ratio for the stomach, if the feature distance ratio of the current frame is greater than the threshold feature distance ratio, the target site determination module 235 determines that the examination site corresponding to the current frame is the stomach; otherwise, it determines that it is the intestines.

[0068] Continuing to refer to Figure 2, the target site determination module 235 outputs a determination result regarding the target site based on the inspection site corresponding to each frame in the second image sequence 230. In some embodiments, due to the influence of noise or algorithmic accuracy, these frames may each correspond to different inspection sites. In some embodiments, if the number of identification results indicating the first site among a plurality of predetermined sites in the corresponding identification results exceeds a second threshold number, the target site determination module 235 determines that the target site is the first site.

[0069] As an example, the examination sites include, for instance, the stomach and intestines. Assume that the second image sequence contains a total of 10 frames. The second threshold number is, for example, 5 frames. Suppose that the identification result corresponding to 8 frames in the second image sequence 230 indicates the stomach. If the number of these frames is greater than the second threshold number, the target site determination module 235 determines that the target site is the stomach; otherwise, it determines that it is the intestines.

[0070] In some embodiments, if the proportion of identification results indicating a second of a plurality of predetermined sites in the corresponding identification results exceeds a threshold proportion, the target site determination module 235 determines that the target site is the second site. In the above example, the threshold proportion is, for example, 0.7. Suppose that the identification results corresponding to 8 out of 10 frames of the second image sequence 230 indicate the stomach. If the proportion of these frames to the total number of frames is greater than the threshold proportion, the target site determination module 235 determines that the target site is the stomach; otherwise, it determines that it is the intestine.

[0071] The data processing methods have been described above through various embodiments of this disclosure. Beneficial effects of these embodiments include the ability to automate the identification of the endoscopic examination site without requiring other presentations or physician intervention. Furthermore, such an approach is not limited to a specific device and can determine the examination site simply by collecting a certain sequence of images. Such image sequences may be collected online in real time or offline. Moreover, such an approach allows for rapid determination of the examination site once the endoscope is inside the body of the subject being examined, without hindering the execution of subsequent procedures. As can be seen, this approach ensures accuracy and speed in the pre-endoscopic examination operations without relying on other means (e.g., text, voice, or keyboard input).

[0072] Exemplary implementation of gastrointestinal examination To better understand the embodiments of this disclosure, specific examples will be described below with reference to Figure 3.

[0073] Figure 3 shows a schematic diagram of Example 300 of a gastric or intestinal determination according to some embodiments of the present disclosure. Example 300 can be implemented in an electronic device 120. It is assumed that the examination sites include the stomach and intestines, and the corresponding examination items are, for example, gastroscopy and colonoscopy. The following explanation will be given with reference to Figure 1.

[0074] In box 305, the electronic device 120 acquires video frames related to the subject being examined via the endoscope. Such video frames are, for example, each frame that the endoscope collects in real time after power is turned on.

[0075] In box 310, the electronic device 120 makes a prediction for each frame in the video frame and determines whether the frame corresponds to the endoscope being located inside or outside the body of the subject being examined.

[0076] In box 315, the electronic device 120 determines whether the current frame corresponds to a situation where the endoscope is located inside the body of the subject being examined. If the electronic device 120 determines that the current frame corresponds to the endoscope being located inside the body of the subject being examined, the process proceeds to box 320. That is, the current frame is the reference image described above. If the electronic device 120 determines that the current frame corresponds to the endoscope being located outside the body of the subject being examined, it returns to box 310, and the electronic device 120 continues to make a judgment for the next frame.

[0077] In box 320, the electronic device 120 begins timing based on the determination that the endoscope is positioned inside the body of the subject being examined. Such timing is recorded from the collection time of the effective frame corresponding to the determination that the endoscope has entered the body.

[0078] In box 325, the electronic device 120 acquires, for example, a target video frame related to the subject being examined via an endoscope. Such a target video frame includes a series of consecutive frames collected after the start of timing.

[0079] In box 330, as target video frames are collected, i.e., during the timing process, the electronic device 120 performs a stomach / intestinal prediction for each of the collected frames.

[0080] In box 335, the electronic device 120 terminates the timing and stops collecting the second video frame. The timing may also be terminated in response to reaching a predetermined duration or when the cumulative distance of adjacent frames exceeds a threshold.

[0081] In box 340, the electronic device 120 performs a stomach / intestine prediction for all video frames. For example, the electronic device 120 aggregates the prediction results for each target frame to determine the stomach / intestine prediction result. In another example, the electronic device 120 performs a stomach / intestine prediction for each target video frame and then aggregates the prediction results for these frames to determine the stomach / intestine prediction result. The stomach / intestine prediction for all target video frames is determined based, for example, on the number or percentage of frames that meet the criteria out of all target frames. Furthermore, based on the overall prediction result, the electronic device 120 determines whether the currently detected area of ​​the endoscope is the stomach or the intestine.

[0082] Through the process described above, the electronic device 120 can begin acquiring images from outside the body of the subject to be examined using an endoscope. Once the endoscope enters the body of the subject through the mouth / anus, the electronic device 120 first completes the prediction of being inside / outside the body, and then begins timing. The electronic device 120 can then continue to acquire images inside the body of the subject to be examined using the endoscope. Once the timing is complete, the electronic device 120 completes the determination of whether it is the stomach / intestine based on the identification results of some of the collected images. In this way, since the identification of the target site is completed automatically without relying on other means, the accuracy and speed of the pre-endoscopic examination operations can be ensured.

[0083] Exemplary process Figure 4 shows a flowchart of an image processing process 400 according to some embodiments of the present disclosure. Process 400 can be implemented in an electronic device 120. For the sake of discussion, process 400 will be described with reference to Figure 1.

[0084] In box 410, the electronic device 120 monitors the predicted results of the correspondence between images in the first image sequence. The first image sequence is collected by the endoscope over time during the examination, and the predicted results indicate whether the endoscope was located inside the body of the subject being examined at the time the corresponding image was collected.

[0085] In box 420, the electronic device 120, based on monitoring, determines a reference image from the first image sequence that corresponds to the endoscope entering the body of the subject being examined.

[0086] In box 430, the electronic device 120 acquires a second image sequence collected by the endoscope after the reference image acquisition time.

[0087] In box 440, the electronic device 120 determines the target area of ​​the object to be examined by the endoscope based on the second image sequence.

[0088] In some embodiments, the first image sequence includes a first image, the corresponding prediction result includes a first prediction result for the first image, and the process 400 further includes the electronic device 120 generating a first prediction result for the first image in the first image sequence, based on a trained first classification model, regarding whether the endoscope is inside the body of the subject being examined at the time the first image was collected.

[0089] In some embodiments, the first image sequence includes a second image, the corresponding prediction result includes a second prediction result for the second image, and the process 400 further includes the electronic device 120 extracting image features of the second image for the second image in the first image sequence, determining a first similarity of the image features to an internal reference feature that characterizes the inside of the body of the reference subject, and a second similarity of the image features to an external reference feature that characterizes the outside of the body of the reference subject, and generating a second prediction result regarding whether the endoscope is located inside the body of the subject being examined at the time the second image was collected by comparing the first and second similarities.

[0090] In some embodiments, the process 400 further includes the electronic device 120 determining a first number of adjacent images from a first image sequence before determining a reference image based on monitoring, and setting a prediction result for an intermediate image among the first number of adjacent images based on the corresponding prediction results for the first number of adjacent images.

[0091] In some embodiments, the process 400 further includes the electronic device 120 determining from a first number of adjacent images an image having a corresponding prediction result indicating that the endoscope is located inside the body of the subject being examined; if the number of determined images is greater than a first threshold number associated with the first number, setting the prediction result for the intermediate image to indicate that the endoscope is inside the body of the subject being examined; and if the number of determined images is not greater than the first threshold number, setting the prediction result for the intermediate image to indicate that the endoscope is not inside the body of the subject being examined.

[0092] In some embodiments, the process 400 further includes the electronic device 120 determining whether all of the corresponding prediction results of a second sequence of images in the first image sequence indicate that the endoscope is located inside the body of the subject being examined, and determining a reference image based on the second sequence of images in response to the determination that all of the corresponding prediction results of the second sequence of images indicate that the endoscope is located inside the body of the subject being examined.

[0093] In some embodiments, the acquisition time span corresponding to the second number of consecutive images is related to the image acquisition frequency of the endoscope.

[0094] In some embodiments, the process 400 further includes the electronic device 120 monitoring whether the image quality of individual images collected by the endoscope after the reference image collection time exceeds a threshold image quality, and, in response to monitoring that the image quality of the third image exceeds the threshold image quality, designating images sequentially collected by the endoscope within a certain time after the collection time of the third image as a second image sequence.

[0095] In some embodiments, the second image sequence satisfies at least one of the following conditions: the acquisition time span corresponding to the second image sequence is equal to the threshold span, or the cumulative image distance between adjacent images in the second image sequence is greater than the threshold, where the image distance is determined based on optical flow.

[0096] In some embodiments, the process 400 further includes the electronic device 120 generating a correspondence identification result for a target area with respect to the images in the second image sequence, and determining the target area based on the correspondence identification result, wherein the identification result indicates one of a plurality of predetermined areas.

[0097] In some embodiments, the process 400 further includes the electronic device 120 generating a first identification result regarding which of a plurality of predetermined regions the third image in the second image sequence corresponds to, based on a trained second classification model.

[0098] In some embodiments, the process 400 further includes, with respect to the fourth image in the second image sequence, the electronic device 120 extracting image features of the fourth image, determining the corresponding similarity of the image features to a plurality of reference features used to characterize each of a plurality of predetermined regions, and generating a second identification result regarding which of the plurality of predetermined regions the fourth image corresponds to by comparing the corresponding similarities.

[0099] In some embodiments, the process 400 further includes the electronic device 120 determining that the target site is the first site in response to the number of identification results indicating the first site among a plurality of predetermined sites in the corresponding identification results exceeding a second threshold number, or determining that the target site is the second site in response to the proportion of identification results indicating the second site among a plurality of predetermined sites in the corresponding identification results exceeding a threshold proportion.

[0100] Exemplary equipment Figure 5 shows a block diagram of an electronic device 500 that can carry out one or more embodiments of the present disclosure. It should be understood that the electronic device 500 shown in Figure 5 is merely illustrative and should not constitute any limitation to the function and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 may be used to carry out the electronic device 120 of Figure 1.

[0101] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual processor or a virtual processor and is capable of performing various processes based on a program stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capability of the electronic device 500.

[0102] The electronic device 500 typically includes a plurality of computer storage media. Such media may include, but are not limited to, volatile and non-volatile media, removable and non-removable media, and may be any available media accessible to the electronic device 500. Memory 520 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or a combination thereof. Storage device 530 may be removable or non-removable media, and may include machine-readable media such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and may be accessible within the electronic device 500.

[0103] The electronic device 500 may further include other removable / non-removable, volatile / non-volatile storage media. Not shown in Figure 5, disk drives for reading from and writing to removable non-volatile disks (e.g., “floppy disks”) and disk drives for reading from and writing to removable non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or operations of various embodiments of the present disclosure.

[0104] The communication unit 540 enables communication with other electronic devices via a communication medium. Furthermore, the functionality of the components of the electronic device 500 may be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can use logical connections to one or more other servers, networked personal computers (PCs), or other network nodes to operate in a networked environment.

[0105] The input device 550 may be one or more input devices, such as a mouse, keyboard, or tracking ball. The output device 560 may be one or more output devices, such as a monitor, speaker, or printer. The electronic device 500 may, if necessary, communicate with one or more external devices (not shown), such as a storage device or display device, via the communication unit 540, with one or more devices that enable a user to interact with the electronic device 500, or with any device (e.g., a network card or modem) that enables the electronic device 500 to communicate with one or more other electronic devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0106] According to exemplary embodiments of the present disclosure, a computer-readable storage medium is provided which stores computer-executable instructions, and the computer-executable instructions are executed by a processor to carry out the method described above. According to exemplary embodiments of the present disclosure, a computer program product is also provided which is stored in a non-temporary computer-readable medium as a tangible object and has computer-executable instructions, and the computer-executable instructions are executed by a processor to carry out the method described above.

[0107] Each aspect of this disclosure is described herein with reference to flowcharts and / or block diagrams of methods, apparatus, devices, and computer program products implemented in accordance with this disclosure. It should be understood that each box in the flowcharts and / or block diagrams, and any combination thereof, can be implemented by computer-readable program instructions.

[0108] These computer-readable program instructions, when provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, generate a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data processing device, it produces a device that performs the functions / operations specified in one or more boxes of a flowchart and / or block diagram. Furthermore, by storing these computer-readable program instructions, which cause computers, programmable data processing devices, and / or other devices to function in a particular manner, on a computer-readable storage medium, the computer-readable medium containing the instructions has a product containing instructions that perform each of the functions / operations specified in one or more boxes of a flowchart and / or block diagram.

[0109] When computer-readable program instructions are loaded onto a computer, other programmable data processing device, or other device, a series of operational steps are executed on the computer, other programmable data processing device, or other device to generate a computer implementation process, thereby enabling the instructions executed on the computer, other programmable data processing device, or other device to implement a function / operation specified in one or more boxes of a flowchart and / or block diagram.

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate architectures, functions, and operations that may be implemented in several implemented systems, methods, and computer program products relating to this disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and a module, program segment, or part of an instruction may contain one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions attached to the boxes may occur in a different order than those attached to the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or in reverse order depending on the functions involved. Also note that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0111] The above descriptions of the various implementations of this disclosure are illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and changes will be apparent to an ordinary art artist without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best describe the principles, practical applications, or improvements in the technology in the market of each implementation, or to enable other ordinary art artists in the art to understand the various embodiments disclosed herein.

Claims

1. When the lens of the imaging optical system of the endoscope is switched between a lens for gastric endoscopy and a lens for colonoscopy, the system monitors the predicted result of the correspondence to the images in the first image sequence collected by the endoscope over time during the examination, indicating whether or not the endoscope is located inside the body of the subject being examined at the time the corresponding image is collected. Based on the monitoring, a reference image corresponding to the endoscope entering the body of the subject being examined is determined from the first image sequence, To acquire a second image sequence collected by the endoscope immediately after the time of collection of the aforementioned reference image, Based on the second image sequence, the target area of ​​the object to be examined by the endoscope is determined, A method for image processing performed by an electronic device, including...

2. The first image sequence includes a first image, and the corresponding prediction result includes a first prediction result for the first image. The method further includes generating a first prediction result regarding whether the endoscope is inside the body of the subject being examined at the time the first image is collected, based on a trained first classification model. The method according to claim 1.

3. The first image sequence includes a second image, and the corresponding prediction result includes a second prediction result for the second image. The aforementioned method, Extracting image features from the second image, The first similarity of the image feature to an internal reference feature that characterizes the inside of the reference subject, and the second similarity of the image feature to an external reference feature that characterizes the outside of the reference subject, By comparing the first similarity and the second similarity, a second prediction result is generated regarding whether the endoscope is located inside the body of the subject being examined at the time the second image was collected. The method according to claim 1, further comprising:

4. Before determining the reference image based on the aforementioned monitoring, Determining a first number of adjacent images from the first image sequence, Based on the prediction result of the correspondence between adjacent images of the first number, a prediction result is set for the intermediate image among the adjacent images of the first number. The method according to claim 1, further comprising:

5. Setting the prediction result for the intermediate image among the adjacent images of the first number is: From the adjacent images of the first number, determine the image in which the corresponding prediction result indicates that the endoscope is located inside the body of the subject being examined, If the number of determined images is greater than the first threshold number associated with the first number, the prediction result for the intermediate image is set to indicate that the endoscope is inside the body of the subject being examined. If the number of determined images is not greater than the first threshold number, the prediction result for the intermediate image is set to indicate that the endoscope is not inside the body of the subject being examined. The method according to claim 4, including the method described in claim 4.

6. From the first image sequence, determining a reference image corresponding to the endoscope entering the body of the subject being examined is: To determine whether the prediction results of the correspondence of the second number of consecutive images in the first image sequence all indicate that the endoscope is located inside the body of the subject being examined, In response to the determination that all of the predicted results of the correspondence of the second number of consecutive images indicate that the endoscope is located inside the body of the subject being examined, the reference image is determined based on the second number of consecutive images, The method according to claim 1, including the method described in claim 1.

7. The method according to claim 6, wherein the acquisition time span corresponding to the second number of consecutive images is related to the image acquisition frequency of the endoscope.

8. Acquiring a second image sequence collected by the endoscope immediately after the collection time of the aforementioned reference image means that The process involves monitoring whether the image quality of each image collected by the endoscope immediately after the collection time of the aforementioned reference image exceeds a threshold image quality, In response to monitoring that the image quality of the third image exceeds the threshold image quality, the images sequentially collected by the endoscope within a certain time immediately following the collection time of the third image are designated as the second image sequence. The method according to claim 1, including the method described in claim 1.

9. The second image sequence satisfies at least one of the following conditions: the acquisition time span corresponding to the second image sequence is equal to the threshold span, or the cumulative image distance between adjacent images in the second image sequence is greater than the threshold. The method according to claim 1, wherein the image distance is determined based on optical flow.

10. Determining the target area of ​​the subject to be examined by the endoscope is, To generate an identification result for the image in the second image sequence that corresponds to the target region and indicates one of a plurality of predetermined regions, The target area is determined based on the identification result of the correspondence, The method according to claim 1, including the method described in claim 1.

11. To generate a correspondence identification result for the target region with respect to the images in the second image sequence, Based on the trained second classification model, a first identification result is generated regarding which of the plurality of predetermined regions the third image in the second image sequence corresponds to. The method according to claim 10, including the method described in claim 10.

12. To generate a correspondence identification result for the target region with respect to the images in the second image sequence, To extract the image features of the fourth image in the second image sequence, The degree of similarity of the correspondence between the image features and a plurality of reference features used to characterize each of the plurality of predetermined regions, By comparing the similarity of the aforementioned correspondences, a second identification result is generated regarding which of the plurality of predetermined parts the fourth image corresponds to, The method according to claim 10, including the method described in claim 10.

13. Determining the target site based on the identification result of the correspondence is, In response to the number of identification results indicating the first of the plurality of predetermined parts in the corresponding identification results exceeding the second threshold number, the target part is determined to be the first part, or In response to the proportion of identification results indicating the second of the plurality of predetermined parts in the corresponding identification results exceeding a threshold proportion, the target part is determined to be the second part. The method according to claim 10, comprising one of the following.

14. An electronic device comprising at least one processing circuit configured to perform the method described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Light source device

    JP2001345003A

  • Electronic endoscope system

    JP2006247371A

  • Variable refractive power endoscope based on liquid lens technology

    JP2013545558A

  • Medical image recording apparatus and x-ray imaging apparatus

    JP2021133083A

  • Medical support operation method, device and computer program product

    JP2023509075A