Determination Method, Device, Computer Equipment and Storage Medium for Endoscope Withdrawal Time

By performing frame-by-frame recognition and observation parameter analysis on colonoscopy videos, identifying and adjusting the recoil time, the problem of inaccurate recoil time in the prior art is solved, and higher accuracy and accuracy are achieved.

CN118864338BActive Publication Date: 2025-07-22CHANGZHOU UNITED IMAGING HEALTHCARE SURGICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310469857.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-27
Publication Date
2025-07-22
Estimated Expiration
2043-04-27

AI Technical Summary

Technical Problem

The existing methods of recoil time detection cannot accurately reflect the doctor's true diagnosis time and quality in colonoscopy, resulting in inaccurate calculation results.

Method used

By identifying the target video frame by frame, the image frames of the first target part and the second target part are determined, the total number of frames of the recollection are calculated, and the image frames that are unstable or inconvenient to detection are identified through observation parameters and observation conditions, and the recollection time is adjusted to reflect the actual diagnosis time.

Benefits of technology

It improves the accuracy of the recoil time, can more accurately reflect the doctor's diagnosis time during colonoscopy, and reduces the impact of unstable or inconvenient image frames on the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864338B_ABST
    Figure CN118864338B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for determining the withdrawal time. By frame-by-frame recognition of a target video, a first image frame including a first target part and a second image frame including a second target part in the target video are obtained, so as to determine the total number of frames for withdrawal according to the first image frame and the second image frame; and the observation parameters of the target image frames between the first image frame and the second image frame are determined, and then according to the observation parameters of the target image frames, the observability of the target image frames is obtained, so as to determine the number of image frames whose observability does not meet the observability condition, and according to the total number of frames for withdrawal and the number of image frames whose observability does not meet the observability condition, the withdrawal time that meets the observability condition is determined. The above method realizes the recognition of unstable image frames during the process from the start to the end of withdrawal, making the accuracy of the finally obtained withdrawal time extremely high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical detection technologies, and particularly to a method, apparatus, computer device, and storage medium for determining the withdrawal time. Background Art

[0002] The withdrawal time is an important indicator for the quality control of colonoscopy positioning. According to the 2021 American Gastroenterological Association Colorectal Cancer Screening Guidelines, the recommended withdrawal time is 9 minutes, with a minimum of 6 minutes. However, simply calculating the withdrawal time cannot reflect the doctor's true diagnosis time and quality, that is, this withdrawal time is not accurate.

[0003] The existing detection methods for the withdrawal time include two types: The first method is to first identify the ileocecal valve and record the time, then identify the outside of the body and record the time again, and subtract the two to obtain the withdrawal time; the second method is based on the first method, excluding the frames of "water absorption", "surgery", and "static" to obtain the effective withdrawal time.

[0004] However, the above two methods still have the problem that the calculated withdrawal time is inaccurate. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for determining the withdrawal time, which can improve the accuracy of determining the withdrawal time.

[0006] In a first aspect, this application provides a method for determining the withdrawal time. The method includes:

[0007] Performing frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for withdrawal according to the first image frame and the second image frame; and determining the observation parameter of each target image frame between the first image frame and the second image frame;

[0008] According to the observation parameter of each target image frame, obtaining the observation degree of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0009] According to the total number of frames for withdrawal and the number of image frames whose observation degree does not meet the observation degree condition, determining the withdrawal time that meets the observation degree condition.

[0010] In one embodiment, the observation parameter includes the target information amount of the target image frame relative to the previous adjacent image frame. The determining the observation parameter of each target image frame between the first image frame and the second image includes:

[0011] Input each of the target image frames into an information recognition model for effective information recognition to obtain the first amount of information contained in each target image frame, and input the previous adjacent image frame into the information recognition model for effective information recognition to obtain the second amount of information contained in the previous adjacent image frame;

[0012] Obtain the target amount of information of each target image frame relative to the previous adjacent image frame according to the first amount of information and the second amount of information.

[0013] In one embodiment, the observation parameter includes the change speed of each target image frame relative to the previous adjacent image frame. Determining the observation parameter of each target image frame between the first image frame and the second image frame includes:

[0014] Determine the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame;

[0015] Obtain the Hamming distance between each target image frame and the previous adjacent image frame according to the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame;

[0016] Obtain the change speed of each target image frame relative to the previous adjacent image frame according to the Hamming distance between each target image frame and the previous adjacent image frame.

[0017] In one embodiment, the observation parameter includes the target amount of information of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame. Obtaining the observation degree of each target image frame according to the observation parameter of each target image frame includes:

[0018] Determine the first weight corresponding to the target amount of information and the second weight corresponding to the change speed;

[0019] Perform an accumulation and operation on the target amount of information of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame according to the first weight and the second weight to obtain the observation degree of each target image frame.

[0020] In one embodiment, the observation degree condition includes: the observation degree of the target image frame is less than a preset observation degree threshold.

[0021] In one embodiment, determining the endoscope withdrawal time that meets the observation degree condition according to the total number of endoscope withdrawal frames and the number of image frames whose observation degree does not meet the observation degree condition includes:

[0022] Determine the number of image frames that meet the observation condition based on the total number of frames during the endoscope withdrawal and the number of image frames whose observation degree does not meet the observation degree condition;

[0023] Determine the duration corresponding to the number of image frames that meet the observation condition as the endoscope withdrawal time that meets the observation degree condition.

[0024] In a second aspect, the present application also provides a device for determining the endoscope withdrawal time. The device includes:

[0025] A first determination module, configured to perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames during the endoscope withdrawal according to the first image frame and the second image frame; and determine the observation parameter of each target image frame between the first image frame and the second image frame;

[0026] A second determination module, configured to obtain the observation degree of each target image frame according to the observation parameter of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0027] A third determination module, configured to determine the endoscope withdrawal time that meets the observation degree condition according to the total number of frames during the endoscope withdrawal and the number of image frames whose observation degree does not meet the observation degree condition.

[0028] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] Perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames during the endoscope withdrawal according to the first image frame and the second image frame; and determine the observation parameter of each target image frame between the first image frame and the second image frame;

[0030] Obtain the observation degree of each target image frame according to the observation parameter of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0031] Determine the endoscope withdrawal time that meets the observation degree condition according to the total number of frames during the endoscope withdrawal and the number of image frames whose observation degree does not meet the observation degree condition.

[0032] Fourthly, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0033] Perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for withdrawing the lens according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame;

[0034] According to the observation parameters of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of image frames in which the observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0035] According to the total number of frames for withdrawing the lens and the number of image frames in which the observation degree does not meet the observation degree condition, determine the lens withdrawal time that meets the observation degree condition.

[0036] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0037] Perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for withdrawing the lens according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame;

[0038] According to the observation parameters of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of image frames in which the observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0039] According to the total number of frames for withdrawing the lens and the number of image frames in which the observation degree does not meet the observation degree condition, determine the lens withdrawal time that meets the observation degree condition.

[0040] The above method, device, computer equipment and storage medium for determining the withdrawal time identify each frame of the target video to obtain the first image frame including the first target part and the second image frame including the second target part in the target video, so as to determine the total number of frames for withdrawal according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame, and then obtain the observation degree of each target image frame according to the observation parameters of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame, and determine the withdrawal time that meets the observation degree condition according to the total number of frames for withdrawal and the number of image frames whose observation degree does not meet the observation degree condition. The above method realizes the identification of unstable image frames during the process from the start to the end of the withdrawal, and then removes the duration corresponding to the unstable image frames when determining the withdrawal time, so that the finally obtained withdrawal time that meets the observation degree condition can reflect the accurate diagnosis duration involved in the doctor's withdrawal process during colonoscopy. Therefore, the accuracy of the withdrawal time obtained based on the above method is extremely high. Description of the Drawings

[0041] Figure 1 It is a schematic structural diagram of an endoscope system in an embodiment;

[0042] Figure 2 It is a schematic flowchart of a method for determining the withdrawal time in an embodiment;

[0043] Figure 3 It is a schematic flowchart of a method for determining the withdrawal time in another embodiment;

[0044] Figure 3A It is a schematic flowchart of a training method in an embodiment;

[0045] Figure 4 It is a schematic flowchart of a method for determining the withdrawal time in another embodiment;

[0046] Figure 5 It is a schematic flowchart of a method for determining the withdrawal time in another embodiment;

[0047] Figure 6 It is a schematic flowchart of a method for determining the withdrawal time in another embodiment;

[0048] Figure 7 It is a schematic flowchart of a method for determining the withdrawal time in another embodiment;

[0049] Figure 8 It is a schematic structural diagram of a device for determining the withdrawal time in an embodiment;

[0050] Figure 9Schematic diagram of the internal structure of a computer device in an embodiment. Detailed implementation manners

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] First, the background source of the technical problem to be solved by the present application will be introduced. When a doctor uses a digestive endoscope for colonoscopy, the intestinal mucosa is diagnosed during the process of withdrawing the endoscope. When the withdrawal time is relatively long, the detection rates of polyps, adenomas and other lesions will increase significantly. Therefore, the withdrawal time is defined as an important indicator for colonoscopy quality control. The traditional method for determining the withdrawal time is to determine the theoretical withdrawal time by referring to the relevant parameters in the specified rectal cancer screening guidelines. However, simply calculating the withdrawal time cannot reflect the doctor's real diagnosis time and quality. Therefore, the calculated withdrawal time is not accurate. Based on this, the present application proposes a method for determining the withdrawal time to solve the above technical problems. The following embodiments will specifically illustrate the method for determining the withdrawal time.

[0053] The method for determining the withdrawal time provided by the embodiments of the present application can be applied to an Figure 1 endoscope system as shown. The endoscope system includes a mirror body device 101 and a detection device 102, and the detection device 102 is connected to the mirror body device 101 in a wired or wireless manner. The mirror body device 101 is used to obtain a detection video and transmit the detection video to the detection device 102. The detection device 102 processes the received detection video and then determines the withdrawal time corresponding to the detection video. Among them, the mirror body device 101 can be the mirror body part of the endoscope, and the detection device 102 can be but is not limited to various terminals, for example, devices with image processing functions such as personal computers, laptop computers, smart phones, and tablet computers. For another example, the detection device 102 can directly be a device including an image processor. The detection device 102 can also be a server, or can be implemented by an independent server or a server cluster composed of multiple servers.

[0054] Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the endoscope system to which the solution of the present application is applied. The specific endoscope system may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0055] In one embodiment, as Figure 2 shown, a method for determining the withdrawal time is provided, and this method is applied toFigure 1 Taking the detection device in

[0056] S201, perform frame-by-frame recognition on the target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for withdrawing the endoscope according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame.

[0057] Among them, the target video can be the video of withdrawing the endoscope during colonoscopy. The first target part can be the ileocecal part, which includes the ileocecal valve, the appendiceal orifice, the terminal ileum, etc., which are parts of the intestinal tissue structure. The second target part can be the tissue structure adjacent to the intestinal tissue structure, or any tissue structure other than the intestinal tissue structure, or can also be called the tissue structure outside the body. The first image frame is the endoscope withdrawal image collected at the moment when the endoscope withdrawal starts, and the second image frame is the endoscope withdrawal image collected at the moment when the endoscope withdrawal ends; the target image frame can be any image frame between the first image frame and the second image frame. The observation parameters of the target image frame include the target information amount of the target image frame relative to the previous adjacent image frame and the change speed of the target image frame relative to the previous adjacent image frame. The observation parameters of the target image frame are used to determine the observation degree of the target image frame.

[0058] In the embodiment of the present application, after the endoscope body device obtains the target video, it can transmit the target video to the detection device. When the detection device obtains the target video, it can further use the corresponding image recognition algorithm or image recognition model to recognize each image frame in the target video. First, specifically recognize whether each image frame contains the first target part, so as to detect the first image frame containing the first target part, and recognize whether each image frame contains the second target part, so as to detect the second image frame containing the second target part. Since the first image frame is the endoscope withdrawal image collected at the moment when the endoscope withdrawal starts, and the second image frame is the endoscope withdrawal image collected at the moment when the endoscope withdrawal ends, therefore, the number of image frames included between the first image frame and the second image frame is the total number of frames for withdrawing the endoscope. In the above frame-by-frame recognition process, when the first image frame is detected, the next image frame can be used as the target image frame frame by frame, and then the observation parameters of the target image frame can be determined, so that the detection device can determine the observation degree of the target image frame according to the observation parameters, and thus detect the target image frame that does not meet the observation degree condition between the first image frame and the second image frame.

[0059] It should be noted that when each video frame in the target video frame is recognized and the first target part in the recognized first image frame is the ileocecal part, the above image recognition model can specifically adopt an ileocecal part recognition module, which can be a pre-trained ileocecal part recognition model. In practical applications, assume that when each video frame in the target video frame is recognized, the "ileocecal part" is recognized in 12 consecutive frames (or other number of frames), then it is determined that the ileocecal part has been reached. At this time, the last frame image in the 12 consecutive frames is the first image frame. The ileocecal part recognition model can be a neural network model or other machine learning models. Optionally, the neural network model can include convolutional neural networks such as ResNet, VGG16, and AlexNet. The method for recognizing the ileocecal part includes: recognizing any one of the ileocecal valve, appendix orifice, and terminal ileum can be considered as recognizing the ileocecal part.

[0060] When each video frame in the target video frame is recognized and the second target part in the recognized second image frame is an external part, the above image recognition model can specifically adopt an external recognition module, which can be a pre-trained external recognition module for recognizing whether the endoscope has reached the outside of the body during withdrawal. For example, if the recognition result is "outside the body" in 12 consecutive frames (or other number of frames) when each video frame in the target video frame is recognized, it is determined that the outside of the body has been reached. At this time, the last frame image in the 12 consecutive frames is the second image frame. Optionally, the above image recognition model can specifically adopt convolutional neural networks such as ResNet, VGG16, and AlexNet.

[0061] When the observation parameter of the above target image frame includes the target information amount of the target image frame relative to the previous adjacent image frame, a corresponding information change recognition algorithm can be used to perform information recognition on the target image frame and the previous adjacent image frame respectively, so as to obtain the information amount of the target image frame and the information amount of the previous adjacent image frame, and finally obtain the target information amount of the target image frame relative to the previous adjacent image frame based on these two information amounts; Optionally, the detection device can also adopt a corresponding information change recognition model to recognize the information change situation of the target image frame relative to the previous adjacent image frame, so as to obtain the target information amount of the target image frame relative to the previous adjacent image frame. The information change recognition model therein can be pre-trained based on sample data. When in use, the target image frame and the previous adjacent image frame can be input into the information change recognition model at the same time to directly obtain the target information amount of the target image frame relative to the previous adjacent image frame.

[0062] When the observation parameter of the target image frame includes the change speed of the target image frame relative to the previous adjacent image frame, the detection device may adopt a corresponding speed change recognition model to recognize the picture change speed of the target image frame relative to the previous adjacent image frame, so as to obtain the change speed of the target image frame relative to the previous adjacent image frame, and the speed change recognition model therein may be pre-trained based on sample data.

[0063] S202. According to the observation parameter of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame.

[0064] Among them, the observation degree of the target image frame is used to measure the stability of the target image frame, that is, the higher the observation degree, the stronger the stability of the corresponding target image frame; optionally, the observation degree of the target image frame can also be used to measure the characteristics of the target image frame that are convenient for detection, that is, the higher the observation degree, the more conducive the corresponding target image frame is for doctors to detect, that is, the image frame with a higher observation degree can more accurately reflect the real diagnosis time of the doctor during the withdrawal process corresponding to the target video. If the observation degree of the target image frame meets the observation degree condition, it means that the target image frame is a stable image frame, or the target image frame is an image frame that is convenient for detection. If the observation degree of the target image frame does not meet the observation degree condition, it means that the target image frame is an unstable image frame, or the target image frame is an image frame that is not convenient for detection. The observation degree condition includes that the observation degree of the target image frame is less than a preset observation degree threshold.

[0065] In the embodiment of the present application, when the detection device obtains the observation parameter of the target image frame based on the foregoing steps, the observation parameter of the target image frame can be input into the corresponding observation degree recognition model, and then the observation degree of the target image frame is obtained. The observation degree recognition model therein can be a model pre-trained based on a sample data set; optionally, the detection device can also calculate based on the observation parameter of the target image frame by using a corresponding observation degree recognition algorithm, so as to obtain the observation degree of the target image frame. When the detection device obtains the observation degree of each target image frame included between the first image frame and the second image frame, the observation degree of each target image frame can be compared with a preset observation degree threshold, so as to screen out the number of image frames that do not meet the observation degree condition from the image frames included between the first image frame and the second image frame based on the comparison result, and then obtain the unstable image frames among the image frames included between the first image frame and the second image frame, or obtain the image frames that are not convenient for detection among the image frames included between the first image frame and the second image frame.

[0066] S203. Determine the withdrawal time that meets the observation degree condition according to the total number of withdrawal frames and the number of image frames whose observation degree does not meet the observation degree condition.

[0067] Among them, the withdrawal time that meets the observation condition can reflect the accurate diagnosis duration involved in the doctor's withdrawal process during colonoscopy examination.

[0068] In the embodiment of the present application, since the number of image frames with the observation degree not meeting the observation condition represents the number of unstable image frames, and the withdrawal time not meeting the observation condition represents the time corresponding to the number of unstable image frames, therefore, the difference operation is performed between the number of withdrawal frames included between the first image frame and the second image frame and the number of unstable image frames included therein. The result obtained is the number of stable image frames, and then the number of stable image frames is converted into the corresponding time, and the withdrawal time that meets the observation condition can be obtained.

[0069] Alternatively, since the number of image frames with the observation degree not meeting the observation condition represents the number of image frames that are not convenient for detection, and the withdrawal time not meeting the observation condition represents the time corresponding to the number of image frames that are not convenient for detection, therefore, the difference operation is performed between the number of withdrawal frames included between the first image frame and the second image frame and the number of image frames that are not convenient for detection included therein. The result obtained is the number of image frames that are convenient for detection, and then the number of image frames that are convenient for detection is converted into the corresponding time, and the withdrawal time that meets the observation condition can be obtained.

[0070] In the above method for determining the withdrawal time, by frame-by-frame recognition of the target video, the first image frame including the first target part and the second image frame including the second target part in the target video are obtained, so as to determine the total number of withdrawal frames according to the first image frame and the second image frame; and the observation parameter of the target image frame between the first image frame and the second image frame is determined, and then according to the observation parameter of the target image frame, the observation degree of the target image frame is obtained, so as to determine the number of image frames with the observation degree not meeting the observation condition, and according to the total number of withdrawal frames and the number of image frames with the observation degree not meeting the observation condition, the withdrawal time that meets the observation condition is determined. The above method realizes the recognition of unstable image frames or image frames that are not convenient for detection during the process from the start of withdrawal to the end of withdrawal. Furthermore, when determining the withdrawal time, the duration corresponding to the unstable image frames therein is removed or the image frames that are not convenient for detection are removed, so that the finally obtained withdrawal time that meets the observation condition can reflect the accurate diagnosis duration involved in the doctor's withdrawal process during colonoscopy examination. Therefore, the accuracy of the withdrawal time obtained based on the above method is extremely high.

[0071] In one embodiment, the observation parameter of the target image frame may include the target information amount of the target image frame relative to the previous adjacent image frame. The following embodiments illustrate the method for obtaining the above target information amount, as Figure 3 shown, the method includes:

[0072] S301. Input each target image frame into the information recognition model for effective information recognition to obtain the first amount of information contained in each target image frame, and input the previous adjacent image frame into the information recognition model for effective information recognition to obtain the second amount of information contained in the previous adjacent image frame.

[0073] Optionally, the information recognition model can be used to classify the information level of each image frame. This information level can be divided according to the size of the effective information contained in the image. For example, assuming the information level is divided into levels 0 - 4, from level 0 to level 4 represents the effective information contained in the image increasing in turn. 0 represents the least amount of effective information, and 4 represents the most amount of effective information. Optionally, this information level can be divided according to the type of invalid information contained in the image. For example, the image frame can include any one of the invalid information such as endoscope inspection, biopsy, surgery, and blur. Among them, the blur can include at least one of motion blur, defocus blur, and flushing blur. Assuming the information level is divided into levels 0 - 3, where level 0 represents that the image includes the invalid information of the blur type, level 1 represents that the image contains the invalid information of the endoscope inspection type, level 2 represents that the image contains the invalid information of the biopsy type, and level 3 represents that the image contains the invalid information of the surgery type.

[0074] Optionally, the information recognition model can also be used to directly identify the effective information contained in each image frame to determine the size of the effective information contained in each image frame. If the effective information contained in each image frame is large, it indicates that the stability of the image frame is good or the image frame is easy to detect; if the effective information contained in each image frame is small, it indicates that the stability of the image frame is poor or the image frame is not easy to detect.

[0075] Correspondingly, the above information recognition model can be trained based on sample image frames and standard image frames with annotations of the effective information amount, and the effective information amount can be annotated according to the corresponding level. Optionally, the above information recognition model can be trained based on sample image frames and gold standard image frames with non - information annotations. The non - information annotations include the annotations of any one of the information such as endoscope inspection, motion blur, defocus blur, biopsy, surgery, and flushing. For example, 200 sets of colonoscopy videos can be collected first, and the colonoscopy videos can be segmented. Specifically, the colonoscopy videos can be segmented into a series of 5 - second video segments. Then each video segment is annotated. When specifically annotating, the video segments can be annotated as different levels according to any one of the above annotation methods. For example, the image frame containing endoscope inspection is annotated as level 2, and the image frame containing surgery is annotated as level 3.

[0076] In the embodiments of the present application, when the detection device determines the target image frame and the previous adjacent image frame, these two image frames can be respectively input into the information recognition model for effective information recognition to obtain the first amount of information included in the target image frame and the second amount of information included in the previous adjacent image frame; optionally, the detection device can also input the above two image frames into the information recognition model at the same time for effective information recognition to obtain the first amount of information and the second amount of information.

[0077] Optionally, the present application also provides a training method for the above information recognition model. The training method, as Figure 3A described, may include:

[0078] S3010, obtain a sample data set;

[0079] S3011, perform non-information frame annotation on the sample data to obtain the annotated sample data set;

[0080] S3012, train the initial information recognition model according to the sample data set and the annotated sample data set to obtain the information recognition model.

[0081] Exemplary illustration Figure 3A For the described training method, step 1, obtain a sample data set. For example, obtain more than 10,000 colonoscopy pictures; step 2, perform non-information frame annotation on some of the sample data. The non-information frames include any one of the information such as blur, endoscope inspection, biopsy, surgery, etc. Among them, the blur includes at least one of motion blur, defocus blur, and flushing blur. One annotation method is to divide the level of the amount of non-information contained in the pictures in the sample data, and then perform annotation. For example, if the picture contains more non-information, for example, the picture contains more pixel points of motion blur, then determine that the level of the picture is low. If the image contains less non-information, for example, the picture contains fewer pixel points of defocus blur, then determine that the level of the picture is high. For another example, the level of the corresponding non-information can also be determined according to the type of non-information contained in the picture. For example, the non-information of the type of motion blur or defocus blur is determined as the non-information of the first level, the non-information of the type of endoscope inspection or surgery is determined as the non-information of the second level, the non-information of the type of biopsy is determined as the non-information of the third level, and the non-information of other types is determined as the non-information of the fourth level. Correspondingly, the image frame with a low level is an unstable image frame or an image frame that is not easy to detect, and the image frame with a high level is a stable image frame or an image frame that is easy to detect. It should be noted that the corresponding relationship between the number of non-information contained in the picture and the level can be determined in advance, and the corresponding relationship between the type of non-information contained in the picture and the level can be determined in advance.

[0082] Another annotation method is to annotate according to the number of confirmations of non-information contained in the pictures in the sample data. Here, the number of confirmations refers to the number of doctors who confirm that the information contained in the picture is non-information. Moreover, the number of confirmations corresponds to the level. That is to say, the more the number of confirmations, the higher the corresponding level, and the lower the number of people, the lower the corresponding level. An optional way is that when annotating, if a preset number of doctors determine that the picture in the sample data is non-information, the corresponding level marked as containing non-information is the preset number. For example, if 3 doctors determine that the above picture is non-information and 1 doctor determines that the above picture is information, the level of this picture is marked as 3, indicating that the possibility of this picture being a stable picture or a picture convenient for detection is relatively high; if 1 doctor determines that the above picture is non-information and 3 doctors determine that the above picture is information, the level of this picture is marked as 1, indicating that the possibility of this picture being an unstable picture or a picture not convenient for detection is relatively high. Optionally, in step 3, after annotating the sample data, the sample data can also be preprocessed, where the preprocessing includes data desensitization, cropping black frames, image normalization, etc. Correspondingly, this preprocessing method can also be used to preprocess the sample data to obtain the processed sample data. Step 4, divide the sample data set into a training data set and a test data set, and further divide the training data set into a training set and a validation set. Step 5, input the aforementioned training data set into a convolutional neural network for training, where the convolutional neural network includes two-dimensional convolutional neural networks such as ResNet, VGG16, and AlexNet. And select parameters such as batchsize, epoch, and learning rate. During the training process, the loss can be determined according to the annotated sample data and the result output by the convolutional neural network, and then the parameters of the convolutional neural network can be adjusted according to the loss until the loss meets the preset training conditions to complete the training. Step 5, complete the training to obtain the information recognition model used above.

[0083] S302. Obtain the target information amount of the target image frame relative to the previous adjacent image frame according to the first information amount and the second information amount.

[0084] Among them, the target information amount of the target image frame relative to the previous adjacent image frame can be the information increment of the target image frame relative to the previous adjacent image frame, or the average information amount of the target image frame and the previous adjacent image frame.

[0085] In an embodiment of the present application, when the detection device obtains the first information amount included in the target image frame and the second information amount included in the previous adjacent image frame based on the foregoing steps, the difference operation can be performed on the first information amount and the second information amount to obtain the target information amount of the target image frame relative to the previous adjacent image frame; optionally, the detection device can also perform an average operation on the first information amount and the second information amount to obtain the average information amount of two adjacent image frames, and use the average information amount of two adjacent image frames as the target information amount of the target image frame relative to the previous adjacent image frame.

[0086] In one embodiment, the observation parameter of the target image frame may include the change speed of the target image frame relative to the previous adjacent image frame. The following embodiments illustrate the method for obtaining the above target information amount, as Figure 4 shown, the method includes:

[0087] S401, determine the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame.

[0088] In an embodiment of the present application, the detection device can compress each target image frame and the previous adjacent image frame, compress them into an image frame of 64*64, and convert the compressed target image frame and the previous adjacent image frame into grayscale image frames to obtain the target image frame and the previous adjacent image frame after grayscale conversion. Then, calculate the hash fingerprints of the target image frame and the previous adjacent image frame after grayscale conversion respectively to obtain the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame.

[0089] S402, obtain the Hamming distance between each target image frame and the previous adjacent image frame according to the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame.

[0090] In an embodiment of the present application, the detection device can use the corresponding Hamming distance calculation method to calculate the Hamming distance between each target image frame and the previous adjacent image frame according to the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame. For example, the Hamming distance calculation method can be implemented by the following relational expression (1):

[0091]

[0092] where d(x, y) represents the Hamming distance, x represents the hash fingerprint of the target image frame, and y represents the hash fingerprint of the previous adjacent image frame.

[0093] S403, obtain the change speed of each target image frame relative to the previous adjacent image frame according to the Hamming distance between each target image frame and the previous adjacent image frame.

[0094] In the embodiments of the present application, the detection device may use a corresponding change speed calculation method to calculate the change speed of the target image frame relative to the previous adjacent image frame according to the Hamming distance between each target image frame and the previous adjacent image frame. For example, the change speed calculation method may be implemented by the following relational expression (2):

[0095]

[0096] Wherein, v represents the change speed, and d(x, y) represents the Hamming distance.

[0097] In one embodiment, when the detection device obtains the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of the target image frame relative to the previous adjacent image frame based on the foregoing Figure 3 and Figure 4 embodiments, the observability of the target image frame can be further determined. The following embodiments illustrate this method. As Figure 5 shown, the method includes:

[0098] S501, determine a first weight corresponding to the target information amount, and assign a corresponding second weight to the change speed.

[0099] Wherein, the first weight represents the influence degree of the target information amount on the observability of the target image frame. If the first weight is large, it means that the influence degree of the target information amount on the observability of the image frame is large; if the first weight is small, it means that the influence degree of the target information amount on the observability of the target image frame is small. The second weight represents the influence degree of the change speed on the observability of the target image frame. If the second weight is large, it means that the influence degree of the change speed on the observability of the target image frame is large; if the second weight is small, it means that the influence degree of the change speed on the observability of the target image frame is small.

[0100] In the embodiments of the present application, the detection device may respectively assign a corresponding first weight to the target information amount and a corresponding second weight to the change speed according to the influence degree on the observability of the target image frame, so as to accurately determine the observability of the target image according to the first weight and the second weight later. Regarding the influence degree of the target information amount and the change speed on the observability respectively, it can be determined in advance according to the actual detection scenario, detection requirements or the number of lesion points included in the detection object. For example, for the process of withdrawing the colonoscope, it is determined that the influence degree of the target information amount on the observability is relatively large, and the influence degree of the change speed on the observability is relatively small. Therefore, the corresponding first weight of the target information amount is set larger, and the corresponding second weight of the change speed is set smaller; for the process of withdrawing the gastroscope, it is determined that the influence degree of the target information amount on the observability is relatively small, and the influence degree of the change speed on the observability is relatively large. Therefore, the corresponding first weight of the target information amount is set smaller, and the corresponding second weight of the change speed is set larger.

[0101] For another example, in the process of withdrawing the colonoscope, if the number of lesion points in the intestinal tissue structure being examined is small, that is, the field of view during this withdrawal process is relatively good, in this case, the first weight corresponding to the target information amount and the second weight corresponding to the change speed can both be appropriately reduced; correspondingly, if the number of lesion points in the intestinal tissue structure being examined is large, that is, the field of view during this withdrawal process is relatively poor, in this case, the first weight corresponding to the target information amount and the second weight corresponding to the change speed can both be appropriately increased.

[0102] For another example, if the doctor determines that during the withdrawal process in a corresponding scenario, the influence degree of the change speed on the observation degree is large, while the influence degree of the target information amount on the observation degree is small, then the first weight corresponding to the target information amount can be set to be small, and the second weight corresponding to the change speed can be set to be large.

[0103] S502. According to the first weight and the second weight, perform an accumulation operation on the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame, to obtain the observation degree of each target image frame.

[0104] In the embodiments of the present application, the detection device can use the following relational expression (3) to perform an accumulation operation on the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame according to the first weight and the second weight, to obtain the observation degree of each target image frame. For example, the relational expression (3) can be:

[0105] I = w1 × n + w2 × v (3);

[0106] Wherein, I represents the observation degree of the target image frame, n represents the target information amount of the target image frame relative to the previous adjacent image frame, v represents the change speed of the target image frame relative to the previous adjacent image frame, w1 represents the first weight, and w2 represents the second weight.

[0107] In one embodiment, there is provided Figure 2 An implementation manner of S203 in the embodiment, that is, the above S203 "determine the withdrawal time that meets the observation degree condition according to the total number of withdrawal frames and the number of image frames whose observation degree does not meet the observation degree condition", as Figure 6 shown, includes:

[0108] S601. According to the total number of withdrawal frames and the number of image frames whose observation degree does not meet the observation degree condition, determine the number of image frames that meet the observation condition.

[0109] S602. Determine the duration corresponding to the number of image frames that meet the observation condition as the withdrawal time that meets the observation degree condition.

[0110] Among them, the observability condition includes that the observability of the target image frame is less than a preset observability threshold. If the observability of the target image frame is less than the preset observability threshold, it indicates that the picture stability in the target image frame is good; if the observability of the target image frame is not less than the preset observability threshold, it indicates that the picture in the target image frame is unstable or not easy to detect. Based on this, the image frames whose observability does not meet the observability condition are unstable image frames or image frames not easy to detect, and the image frames whose observability meets the observability condition are stable image frames or image frames easy to detect.

[0111] In the embodiments of the present application, when the detection device obtains the total number of image frames of the image frames included between the inspection start time and the inspection end time in the target video, it can remove the duration corresponding to the unstable image frames or the image frames not easy to detect among them, and the remaining is the duration occupied by the stable image frames or the image frames easy to detect.

[0112] Optionally, the present application provides a warning method for reminding the user of abnormal image frames during the endoscope withdrawal process, that is, unstable image frames or image frames not easy to detect. The method includes: when the detection device obtains the observability of each target image frame based on any of the foregoing embodiments and further determines whether each target image frame meets the observability condition, if it meets the observability condition, it is determined that the picture stability in the target image frame is good or the image frame is easy to detect, and if it does not meet the observability condition, it is determined that the picture stability in the target image frame is poor or the image frame is not easy to detect. The observability condition includes that the observability of the target image frame is less than a preset observability threshold, and the preset observability threshold can also be a warning threshold, that is, when it is detected that the observability of the target image frame is less than the preset observability threshold, it can be prompted that "the picture stability is good or the picture is easy to detect"; when it is detected that the observability of the target image frame is not less than the preset observability threshold, it can be prompted that "the picture stability is poor or the picture is not easy to detect". For example, assuming that the preset observability threshold is 1 or 2, when the observability of the target image frame is less than 1 or 2, a warning message can be initiated, and further, based on the Figure 3 method described in the foregoing Figure 4 embodiments, the effective information amount included in the target image frame can be identified, and based on the

[0113] method described in the foregoing Figure 7 embodiments, the picture change speed of the target image frame can be identified, so as to prompt whether the picture instability or not easy to detect is caused by "too little information amount", or "too fast picture change speed", or both.

[0113] Combining all the above embodiments, a method for determining the endoscope withdrawal time is also provided. As Figure 7 shown, the method includes:

[0114] S701, perform frame-by-frame recognition on the target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video.

[0115] S702, determine the total number of frames for the camera to pull back based on the first image frame and the second image frame.

[0116] S703, input each pair of adjacent target image frames between the first image frame and the second image frame into the information recognition model for effective information recognition, and obtain the effective information amounts contained in each target image frame and its adjacent target image frame, namely the first information amount and the second information amount.

[0117] S704, determine the target information amount of each target image frame relative to the previous adjacent image frame according to the first information amount and the second information amount.

[0118] S705, determine the hash fingerprint of the target image frame and the hash fingerprint of the previous adjacent image frame.

[0119] S706, obtain the Hamming distance between the target image frame and the previous adjacent image frame according to the hash fingerprint of the target image frame and the hash fingerprint of the previous adjacent image frame.

[0120] S707, obtain the change speed of the target image frame relative to the previous adjacent image frame according to the Hamming distance between the target image frame and the previous adjacent image frame.

[0121] S708, assign a corresponding first weight to the target information amount and a corresponding second weight to the change speed.

[0122] S709, perform an accumulation operation on the target information amount of the target image frame relative to the previous adjacent image frame and the change speed of the target image frame relative to the previous adjacent image frame according to the first weight and the second weight, and obtain the observation degree of the target image frame.

[0123] S710, determine the number of image frames that meet the observation condition according to the total number of frames for the camera to pull back and the number of image frames whose observation degree does not meet the observation degree condition.

[0124] S711, determine the duration corresponding to the number of image frames that meet the observation condition as the pull-back time that meets the observation degree condition.

[0125] The above steps are all described in the foregoing embodiments. For detailed content, please refer to the foregoing description and will not be elaborated here.

[0126] The method for determining the endoscope withdrawal time described in the embodiments of the present application realizes determining the accurate endoscope withdrawal time by detecting the stability of the screen for each target image frame between the detection start time and the detection end time. Moreover, in the method for determining whether each target image frame is stable, the embodiments of the present application determine whether the target image frame is stable from two dimensions: the target information amount of the screen and the change speed of the screen, which can further improve the accuracy of the endoscope withdrawal time.

[0127] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0128] Based on the same inventive concept, the embodiments of the present application also provide an endoscope withdrawal time determination device for implementing the method for determining the endoscope withdrawal time described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the endoscope withdrawal time determination device provided below can refer to the limitations on the method for determining the endoscope withdrawal time in the above text, and will not be repeated here.

[0129] In one embodiment, as Figure 8 shown, an endoscope withdrawal time determination device is provided, including:

[0130] A first determination module 10, configured to perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of endoscope withdrawal frames according to the first image frame and the second image frame; and determine the observation parameter of each target image frame between the first image frame and the second image frame.

[0131] A second determination module 11, configured to obtain the observation degree of each target image frame according to the observation parameter of each target image frame, so as to determine the number of image frames whose observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame.

[0132] A third determination module 12, configured to determine a withdrawal time meeting the observation condition according to the total number of frames during withdrawal and the number of frames of images whose observation degree does not meet the observation degree condition.

[0133] In one embodiment, the above-mentioned second determination module 11 is specifically configured to input each of the target image frames into an information recognition model for effective information recognition to obtain the first amount of information included in each of the target image frames, and input the previous adjacent image frame into the information recognition model for effective information recognition to obtain the second amount of information included in the previous adjacent image frame; according to the first amount of information and the second amount of information, obtain the target amount of information of each of the target image frames relative to the previous adjacent image frame.

[0134] In one embodiment, the above-mentioned second determination module 11 is further specifically configured to determine the hash fingerprint of each of the target image frames and the hash fingerprint of the previous adjacent image frame; according to the hash fingerprint of each of the target image frames and the hash fingerprint of the previous adjacent image frame, obtain the Hamming distance between each of the target image frames and the previous adjacent image frame; according to the Hamming distance between each of the target image frames and the previous adjacent image frame, obtain the change speed of each of the target image frames relative to the previous adjacent image frame.

[0135] In one embodiment, the above-mentioned second determination module 11 is specifically configured to assign a corresponding first weight to the target amount of information and a corresponding second weight to the change speed; according to the first weight and the second weight, perform an accumulation and operation on the target amount of information of each of the target image frames relative to the previous adjacent image frame and the change speed of the target image frame relative to the previous adjacent image frame to obtain the observation degree of each of the target image frames.

[0136] In one embodiment, the observation degree condition includes: the observation degree of the target image frame is less than a preset observation degree threshold.

[0137] In one embodiment, the above-mentioned third determination module 12 is specifically configured to determine the number of frames of images meeting the observation condition according to the total number of frames during withdrawal and the number of frames of images whose observation degree does not meet the observation degree condition; determine the duration corresponding to the number of frames of images meeting the observation condition as the withdrawal time meeting the observation degree condition.

[0138] Each module in the above-mentioned device for determining the withdrawal time can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0139] In one embodiment, an image processing device is provided. The image processing device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the method described in any of the foregoing embodiments are implemented.

[0140] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, a method for determining the withdrawal lens time is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or may be a button, a trackball, or a touchpad provided on the casing of the computer device, or may also be an external keyboard, a touchpad, or a mouse, etc.

[0141] Those skilled in the art can understand that Figure 9 the structure shown in

[0142] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0143] Perform frame-by-frame recognition on the target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for withdrawing the lens according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame;

[0144] According to the observation parameters of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of image frames in which the observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame;

[0145] Determine the withdrawal time that meets the observation condition according to the total number of frames of the withdrawal and the number of frames of the image where the observation degree does not meet the observation condition.

[0146] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0147] Perform frame-by-frame recognition on the target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames of the withdrawal according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame;

[0148] According to the observation parameters of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of frames of the image where the observation degree does not meet the observation condition among the image frames between the first image frame and the second image frame;

[0149] Determine the withdrawal time that meets the observation condition according to the total number of frames of the withdrawal and the number of frames of the image where the observation degree does not meet the observation condition.

[0150] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0151] Perform frame-by-frame recognition on the target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames of the withdrawal according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame;

[0152] According to the observation parameters of each target image frame, obtain the observation degree of each target image frame, so as to determine the number of frames of the image where the observation degree does not meet the observation condition among the image frames between the first image frame and the second image frame;

[0153] Determine the withdrawal time that meets the observation condition according to the total number of frames of the withdrawal and the number of frames of the image where the observation degree does not meet the observation condition.

[0154] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0155] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0156] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for determining the withdrawal time, characterized in that The method includes: Performing frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames for pulling back the lens according to the first image frame and the second image frame; and determining the observation parameters of each target image frame between the first image frame and the second image frame; the observation parameters include the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame; the change speed of each target image frame relative to the previous adjacent image frame is obtained according to the Hamming distance between each target image frame and the previous adjacent image frame; Determining a first weight corresponding to the target information amount and a second weight corresponding to the change speed; According to the first weight and the second weight, performing an accumulation and operation on the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame to obtain the observation degree of each target image frame, so as to determine the number of image frames in which the observation degree does not meet the observation degree condition among the image frames between the first image frame and the second image frame; Determining the lens pulling-back time that meets the observation degree condition according to the total number of frames for pulling back the lens and the number of image frames in which the observation degree does not meet the observation degree condition.

2. The method according to claim 1, wherein The observation parameters include the target information amount of each target image frame relative to the previous adjacent image frame. Determining the observation parameters of each target image frame between the first image frame and the second image frame includes: Inputting each target image frame into an information recognition model for effective information recognition to obtain the first information amount included in each target image frame, and inputting the previous adjacent image frame into the information recognition model for effective information recognition to obtain the second information amount included in the previous adjacent image frame; Obtaining the target information amount of each target image frame relative to the previous adjacent image frame according to the first information amount and the second information amount.

3. The method according to claim 1, characterized in that The observation parameters include the change speed of each target image frame relative to the previous adjacent image frame. Determining the observation parameters of each target image frame between the first image frame and the second image frame includes: Determining the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame; Obtaining the Hamming distance between each target image frame and the previous adjacent image frame according to the hash fingerprint of each target image frame and the hash fingerprint of the previous adjacent image frame; Obtaining the change speed of each target image frame relative to the previous adjacent image frame according to the Hamming distance between each target image frame and the previous adjacent image frame.

4. The method according to claim 1, wherein The observation degree of the target image frame is used to measure the stability of the target image frame, or the observation degree of the target image frame is used to measure the characteristics of the target image frame that are convenient for detection.

5. The method according to any one of claims 1-4, characterized in that, The observation degree condition includes: the observation degree of the target image frame is less than a preset observation degree threshold.

6. The method according to any one of claims 1-4, characterized in that, Determining the withdrawal time that meets the observation condition according to the total number of frames of withdrawal and the number of frames of images that do not meet the observation condition includes: Determining the number of frames of images that meet the observation condition according to the total number of frames of withdrawal and the number of frames of images that do not meet the observation condition; Determining the duration corresponding to the number of frames of images that meet the observation condition as the withdrawal time that meets the observation condition.

7. A device for determining the withdrawal time, characterized in that The device includes: A first determination module, configured to perform frame-by-frame recognition on a target video to obtain a first image frame including a first target part and a second image frame including a second target part in the target video, so as to determine the total number of frames of withdrawal according to the first image frame and the second image frame; and determine the observation parameters of each target image frame between the first image frame and the second image frame; the observation parameters include the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame; the change speed of each target image frame relative to the previous adjacent image frame is obtained according to the Hamming distance between each target image frame and the previous adjacent image frame; A second determination module, configured to determine a first weight corresponding to the target information amount and a second weight corresponding to the change speed; according to the first weight and the second weight, perform an accumulation and operation on the target information amount of each target image frame relative to the previous adjacent image frame and the change speed of each target image frame relative to the previous adjacent image frame, to obtain the observation degree of each target image frame, so as to determine the number of frames of images among the image frames between the first image frame and the second image frame whose observation degree does not meet the observation condition; A third determination module, configured to determine the withdrawal time that meets the observation condition according to the total number of frames of withdrawal and the number of frames of images that do not meet the observation condition.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Real-time monitoring method for enteroscope withdrawal time based on random forest algorithm

    CN111767958A

  • Intestinal tract endoscope retreating speed smoothing method

    CN113487553A

  • Effective mirror retreating time evaluation method and device for enteroscopy and storage medium

    CN113962998A