Endoscope device, frame image extraction method, program, and endoscopic system
The endoscope device tags and extracts important frame images based on user operations and feature recognition, addressing the challenge of insufficient image selection in industrial endoscopic examinations by enhancing precision and transparency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- WABTEC INSPECTION TECHNOLOGIES JAPAN CORP
- Filing Date
- 2022-05-20
- Publication Date
- 2026-05-11
AI Technical Summary
Existing endoscopic examination technologies struggle to efficiently extract important frame images in industrial settings due to similarities in internal features of subjects, leading to insufficient image selection and lack of transparency in extraction criteria.
An endoscope device with an insertion unit, operation receiving unit, and control unit that tags frame images based on user operations and specific feature recognition, allowing for precise extraction of important frames using multiple criteria.
Enables the precise narrowing down and extraction of important frame images in endoscopic videos, providing transparency into the extraction criteria used.
Smart Images

Figure 0007856485000001 
Figure 0007856485000002 
Figure 0007856485000003
Abstract
Description
Technical Field
[0001] The disclosure of this specification relates to an endoscope device, a frame image extraction method, a program, and an endoscope system.
Background Art
[0002] In endoscopic examination, there is an operation of taking a video of the inside of a subject and checking it after the examination. This operation becomes less easy as the length of the taken video increases. Therefore, there is a need for a technique that can efficiently check only important parts (important frame images included in the video) in the taken video.
[0003] Patent Document 1 describes an image processing apparatus that processes images acquired in an examination using a capsule endoscope. This image processing apparatus compares the imaging times of a plurality of types of feature images (for example, pylorus images, clip images, Vater's papilla images, and Bauhin's valve images) extracted from an image group (current image group) acquired in the examination with the imaging times of corresponding plurality of types of feature images extracted from an image group (past image group) acquired in a past examination of the same subject. When the time interval between the feature images in the current image group is longer than a reference value or more than the time interval between the corresponding feature images in the past image group, a flag indicating that it is an observation attention image is added to the image captured within the time interval between the feature images in the current image group. Then, when displaying the current image group, the image with the added flag is displayed in a different format from other images in order to arouse the attention of medical staff and make them observe it intensively.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In industrial endoscopic examinations, the internal features of the subject (e.g., turbine blades) are similar. Therefore, if images are extracted based solely on image features, as in the image processing device described in Patent Document 1, the selection of important images will be insufficient. Furthermore, if the device automatically extracts images, as in the image processing device described in Patent Document 1, the user (examiner) cannot know what criteria were used to extract those images.
[0006] One aspect of the present invention is to provide a technique that can narrow down and extract important frame images contained in a video acquired through inspection. [Means for solving the problem]
[0007] An endoscope according to one aspect of the present invention comprises an insertion unit including an image sensor, which is inserted into the body of a subject; an operation receiving unit for receiving operations; and a control unit, wherein the control unit includes, during the recording process of a video including a plurality of frame images generated based on an imaging signal output from the image sensor, a tagging unit that adds information about the operation as a tag to the frame image among the plurality of frame images corresponding to the timing when the operation receiving unit receives an operation, and a tag that adds information about the specific feature image as a tag to the frame image recognized as a specific feature image; and an extraction unit that extracts frame images from the plurality of frame images included in the video based on at least one tag from two or more tags selected from a plurality of types of tags.
[0008] A frame image extraction method according to one aspect of the present invention involves, during the recording process of a video including a plurality of frame images generated based on an imaging signal output from an imaging sensor contained in an insertion unit inserted inside a subject, attaching information about the operation as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation from among the plurality of frame images, attaching information about the specific feature image as a tag to the frame image recognized as a specific feature image, and extracting a frame image from the plurality of frame images included in the video based on at least one of two or more tags selected from a plurality of types of tags.
[0009] A program according to one aspect of the present invention, during the recording process of a video including a plurality of frame images generated based on imaging signals output from an image sensor contained in an insertion unit inserted inside a subject, causes a processor to perform the following processing: attaching information about the operation as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation from among the plurality of frame images; attaching information about the specific feature image as a tag to the frame image recognized as a specific feature image; and extracting a frame image from the plurality of frame images included in the video based on at least one of two or more tags selected from a plurality of types of tags.
[0010] An endoscope system according to one aspect of the present invention comprises an endoscope device and a control device, wherein the endoscope device includes an insertion unit including an image sensor that is inserted into the body of a subject and an operation receiving unit that receives operations, and the control device includes, during the recording process of a video including a plurality of frame images generated based on an imaging signal output from the image sensor, a tagging unit that adds information about the operation as a tag to the frame image among the plurality of frame images corresponding to the timing when the operation receiving unit receives an operation, and a tag that adds information about the specific feature image as a tag to the frame image recognized as a specific feature image, and an extraction unit that extracts frame images from the plurality of frame images included in the video based on at least one tag from two or more tags selected from a plurality of types of tags. [Effects of the Invention]
[0011] According to the above embodiment, it is possible to provide a technology that can narrow down and extract important frame images contained in a video acquired through inspection. [Brief explanation of the drawing]
[0012] [Figure 1] This figure illustrates the external configuration of an endoscope device 1 according to one embodiment. [Figure 2] This figure illustrates the internal configuration of an endoscope device 1 according to one embodiment. [Figure 3] This flowchart illustrates the processing flow performed by the control unit 34 in response to a video recording start operation. [Figure 4] This flowchart illustrates the processing flow performed by the control unit 34 in response to a video selection operation. [Figure 5] This diagram illustrates a condition selection screen where a tag is selected in S230 and displayed in S220. [Figure 6] This diagram illustrates a specific example of extraction in S240. [Figure 7] This figure shows a modified example of the condition selection screen displayed in S220. [Figure 8]It is a diagram showing a modified example of the condition selection screen displayed on S220. [Figure 9] It is a diagram showing an example of tag selection by area specification. [Figure 10] It is a diagram illustrating the hardware configuration of the computer 100.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0014] FIG. 1 is a diagram illustrating the external configuration of the endoscope apparatus 1 according to an embodiment. The endoscope apparatus 1 illustrated in FIG. 1 is used for endoscope inspection in the industrial field and includes an insertion unit 10, an operation unit 20, and a main body unit 30.
[0015] The insertion unit 10 has an elongated shape that can be inserted into the interior of a subject such as a turbine or an engine, and includes a tip portion 11, a bending portion 12 formed to be bendable, and a long flexible tube portion 13 having flexibility.
[0016] The operation unit 20 includes a joystick (bending operation element) 21 that receives an operation (bending operation) for bending the bending portion 12 in a desired direction, a plurality of buttons (not shown) that receive operations for performing various inputs, and the like.
[0017] The main body unit 30 includes a display unit 31, an external interface 32, and the like. The display unit 31 is a display device such as an LCD (Liquid Crystal Display), and performs display of images inside the subject that is the object to be photographed, various screens such as a condition selection screen, playback display of recorded videos, and the like. In addition, the display unit 31 has a touch panel 31a that receives a touch operation for performing various inputs. The touch panel 31a and the above-described operation unit 20 are examples of an operation reception unit that receives operations by a user (examiner). The external interface 32 is connected to an external device such as an external storage device (for example, a USB (Universal Serial Bus) memory).
[0018] FIG. 2 is a diagram illustrating the internal configuration of the endoscope apparatus 1 according to an embodiment. As illustrated in FIG. 2, in the endoscope apparatus 1, the distal end portion 11 included in the insertion portion 10 includes an imaging optical system 11a, an imaging element 11b, a light emitting element 11c, and an illumination optical system 11d.
[0019] The imaging optical system 11a forms a subject image on the imaging element 11b. The imaging element 11b captures (photoelectrically converts) the subject image formed by the imaging optical system 11a to generate an imaging signal, and outputs the imaging signal to the main body portion 30 (image generation portion 33). The imaging element 11b is a CCD (Charge Coupled Device) image sensor, a CMOS (Complementary Metal Oxide Semiconductor) image sensor, or the like.
[0020] The light emitting element 11c emits illumination light for illuminating the subject. The light emitting element 11c is an LED (Light Emitting Diode) or the like. The illumination optical system 11d irradiates the subject with the illumination light emitted by the light emitting element 11c. In the present embodiment, the subject is illuminated with the illumination light emitted by the light emitting element 11c. However, for example, the illumination light emitted by the light source portion included in the main body portion 30 may be guided by a light guide inserted through the insertion portion 10 or the like to illuminate the subject.
[0021] The operation unit 20 receives operations on a joystick 21 and a plurality of buttons (not shown), and outputs a signal corresponding to the operation to the main body portion 30 (control unit 34).
[0022] The main body portion 30 includes an image generation portion 33, a control portion 34, and a recording portion 35 in addition to the above-described display portion 31 and external interface 32. In addition, the touch panel 31a included in the display portion 31 receives a touch operation, and outputs a signal corresponding to the touch operation to the control portion 34.
[0023] The image generation unit 33 generates frame images by applying predetermined signal processing to the imaging signal output from the image sensor 11b, and sequentially outputs the generated frame images to the control unit 34. The image generation unit 33 is composed of, for example, an image generation circuit.
[0024] The control unit 34 controls various parts of the endoscope device 1. For example, the control unit 34 controls the driving of the image sensor 11b, the turning on / off of the illumination (light-emitting element 11c), the bending of the bending section 12, and the display of the display unit 31. The turning on / off of the illumination is controlled in response to operations on the operation unit 20 or the touch panel 31a (illumination operation), and the bending of the bending section 12 is controlled in response to operations on the operation unit 20 (joystick 21) (bending operation).
[0025] The control unit 34 also performs various processes. For example, the control unit 34 performs processes such as displaying frame images sequentially output from the image generation unit 33 on the display unit 31, recording a video containing multiple frame images sequentially output from the image generation unit 33 to the recording unit 35 or an external storage device connected to the external interface 32 (video recording process), and displaying (playback display) the video recorded in the recording unit 35 or the external storage device connected to the external interface 32 on the display unit 31.
[0026] Furthermore, during video recording, the control unit 34 may perform, for example, feature image recognition processing to recognize specific feature images, or, in response to operations on the operation unit 20 or touch panel 31a (measurement operations, comment assignment operations, evaluation result assignment operations, marking assignment operations, etc.), perform measurement processing to measure the length of scratches shown in the frame image, processing to assign comments to the frame image, processing to assign evaluation results to the frame image, processing to assign markings to the frame image, etc. In the processing of assigning comments and evaluation results, arbitrary comments and evaluation results may be input and assigned by operations on the operation unit 20 or touch panel 31a, or comments and evaluation results may be selected from a set of pre-prepared comments and evaluation results.
[0027] Furthermore, the control unit 34 includes a tagging unit 34a and an extraction unit 34b.
[0028] During video recording, the tagging unit 34a adds tags related to the operation to the frame images among the multiple frame images included in the video that correspond to the timing when the operation unit 20 or touch panel 31a receives an operation. For example, the tagging unit 34a adds tags related to the lighting operation to the frame image corresponding to the timing when a lighting on operation is received, and adds tags related to the lighting operation to the frame image corresponding to the timing when a lighting off operation is received. In addition, the tagging unit 34a adds tags related to the bending operation (including information on the bending angle) to the frame image corresponding to the timing when a bending operation is received. If the bending operation is performed continuously, the information related to the bending operation is added as tags to the frame images corresponding to the timing included in the period during which the bending operation was performed continuously. Furthermore, the tagging unit 34a adds information related to the measurement operation to start the measurement process (including information about the measurement target (e.g., scratches)) as a tag to the frame image corresponding to the timing when an operation to start the measurement process is received, and adds information related to the measurement operation to end the measurement process (including information about the measurement target and measurement results (e.g., measured value of scratches)) as a tag to the frame image corresponding to the timing when an operation to end the measurement process is received. Furthermore, the tagging unit 34a adds information related to the comment addition operation as a tag to the frame image corresponding to the timing when a comment addition operation is received. Furthermore, the tagging unit 34a adds information related to the evaluation result addition operation as a tag to the frame image corresponding to the timing when an evaluation result addition operation is received. Furthermore, the tagging unit 34a adds information related to the marking addition operation as a tag to the frame image corresponding to the timing when a marking addition operation is received.
[0029] Furthermore, during video recording, the tagging unit 34a adds tags related to specific feature images to frame images that have been recognized as specific feature images by the feature image recognition process from among multiple frame images included in the video. Specific feature images are, for example, images showing parts of a part of the video that the user is interested in. Specifically, these include images showing gears, blades, elbows, etc. Such specific feature images are determined, for example, before video recording, in response to a feature image selection operation on the operation unit 20 or touch panel 31a.
[0030] The extraction unit 34b extracts frame images from multiple frame images contained in a video recorded by the video recording process, based on at least one of two or more tags selected from a plurality of types of tags. For example, the extraction unit 34b extracts frame images based on an AND condition or OR condition of two or more tags selected from a plurality of types of tags. The two or more tags are selected, for example, in response to a tag selection operation on the operation unit 20 or the touch panel 31a. The tag selection operation is an operation to select two or more tags from the tags attached to the frame images contained in the video.
[0031] Such a control unit 34 is configured to include, for example, a processor such as a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory). The functions of the control unit 34 are realized when the processor executes a program stored in the ROM while using the RAM as a work area. The control unit 34, or the control unit 34 and the image generation unit 33, may be configured using an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), etc.
[0032] The recording unit 35 records video and other data. The recording unit 35 may also record programs executed by the processor of the control unit 34. The recording unit 35 is a non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive).
[0033] In the endoscope device 1 configured as described above, during an endoscopic examination, the user inserts the insertion section 10 into the body of the subject while checking the image displayed on the display unit 31, and when the tip section 11 reaches the desired observation position, the user performs a video recording start operation on the operation unit 20 or touch panel 31a, at which point the control unit 34 starts the process illustrated in Figure 3.
[0034] Figure 3 is a flowchart illustrating the flow of processing performed by the control unit 34 in response to a video recording start operation. In the process illustrated in Figure 3, first, the control unit 34 starts the video recording process (S110). The video recording process is the process of recording a video containing multiple frame images sequentially output from the image generation unit 33 to the recording unit 35 or an external storage device connected to the external interface 32.
[0035] Next, the control unit 34 performs the processes S120 to S123 and the processes S130 to S132 in parallel.
[0036] In the processes S120 to S123, first, the control unit 34 starts the operation reception process (S120). The operation reception process is the process of receiving operations for the operation unit 20 or the touch panel 31a. The processes S121 to S123 following S120 are repeated until a video recording termination operation is received for the operation unit 20 or the touch panel 31a. In the processes S121 to S123, first, the control unit 34 determines whether or not an operation for the operation unit 20 or the touch panel 31a has been received (S121), and if the result of the determination is NO, this determination is repeated. On the other hand, if the result of the determination in S121 is YES, the control unit 34 executes the corresponding control or process according to the received operation (S122), and the tagging unit 34a adds information about the operation as a tag to the frame image corresponding to the timing when the operation was received (S123).
[0037] In the processing of S130 to S132, first, the control unit 34 starts feature image recognition processing (S130). Feature image recognition processing is the process of recognizing specific feature images (more specifically, images that show specific features) in the frame images (which are also frame images included in the video) sequentially output from the image generation unit 33. Note that feature image recognition processing is also the process of recognizing specific features in the frame images sequentially output from the image generation unit 33. Images that show specific features are, for example, images that show parts of interest to the user (images that show gears, blades, or elbows, etc.) and are determined before video recording processing (before the start of video recording processing) according to a feature image selection operation on the operation unit 20 or touch panel 31a. The feature image selection operation may be, for example, an operation to select an image that shows a specific feature from among a plurality of feature images (exemplary feature images) displayed on the display unit 31, or an operation to select a specific part item from among a plurality of part items ("gear", "blade", "elbow", etc.) displayed on the display unit 31. The processes S131 and S132 following S130 are repeated until a video recording termination operation is received on the operation unit 20 or the touch panel 31a. In the processes S131 and S132, first, the control unit 34 determines whether a specific feature (an image showing a specific feature) has been recognized by the feature image recognition process (S131), and if the result of this determination is NO, this determination is repeated. On the other hand, if the result of the determination in S131 is YES, the tagging unit 34a adds information about that specific feature (feature image) as a tag to the frame image in which the specific feature was recognized (the frame image recognized as an image showing a specific feature) (S132).
[0038] Then, when the operation unit 20 or touch panel 31a receives a video recording termination operation, the control unit 34 terminates the video recording process started in S110 (S140), terminates the operation reception process started in S120 (S150), terminates the feature image recognition process started in S130 (S160), and terminates the process illustrated in Figure 3. As a result, for example, a video including frame images to which tags have been attached is recorded in the recording unit 35 or an external storage device connected to the external interface 32.
[0039] After the video is recorded in this manner, in the endoscope device 1, when the user selects a video (a video including tagged frame images) recorded on the recording unit 35 or an external storage device connected to the external interface 32 by performing a video selection operation on the operation unit 20 or touch panel 31a in order to check the recorded video, the control unit 34 starts the process illustrated in Figure 4.
[0040] Figure 4 is a flowchart illustrating the processing flow performed by the control unit 34 in response to a video selection operation. In the processing illustrated in Figure 4, first, the control unit 34 classifies the tags attached to the frame images included in the video selected by the video selection operation by type (S210). Next, the control unit 34 displays a condition selection screen containing the tags classified in S210 on the display unit 31 (S220).
[0041] Next, the control unit 34 selects a tag to be used as the extraction condition on the condition selection screen displayed in S220, in response to a tag selection operation on the operation unit 20 or the touch panel 31a (S230). Here, if multiple tags are selected, the extraction condition can be an AND condition or an OR condition of those multiple tags. Specifically, if an AND condition is set, frame images to which all of the multiple tags are attached are extracted, and if an OR condition is set, frame images to which at least one of the set multiple tags is attached are extracted. When frame images are extracted using an AND condition, it is possible to narrow down the frame images to be more important compared to using an OR condition. On the other hand, when frame images are extracted using an OR condition, in addition to the important frame images, it is possible to obtain frame images before and after those frame images were acquired.
[0042] Next, the extraction unit 34b extracts frame images from among multiple frame images contained in the video based on the tags selected as extraction conditions in S230 (S240). Next, the control unit 34 displays (plays back) the frame images extracted in S240 on the display unit 31 (S250). The processing in S240 and S250 also involves extracting the relevant section (relevant period) in the video based on the tags selected as extraction conditions in S230, and playing back and displaying the partial video of that relevant section.
[0043] Next, the control unit 34 determines whether or not an extraction condition change operation has been performed on the operation unit 20 or the touch panel 31a (S260). If the determination result is YES, the process returns to S220. On the other hand, if the determination result of S260 is NO, the control unit 34 terminates the process illustrated in Figure 4.
[0044] Figure 5 illustrates a condition selection screen displayed in S220 and with tags selected in S230. The condition selection screen illustrated in Figure 5 includes tags categorized in S210, such as "Lighting" (lighting operation), "Bending" (bending operation), "Measurement (Scratch)" (scratch measurement operation), "Feature Image (Gear)" (specific feature image (gear)), "Comment" (comment assignment operation), "Evaluation Result" (evaluation result assignment operation), and "Marking" (marking assignment operation). It also includes checkboxes 41 (41a~41g) corresponding to each tag. The user can select the desired tag by checking the checkbox 41 corresponding to the desired tag. In the condition selection screen illustrated in Figure 5, checkboxes 41c, 41d, and 41e are checked, and the tags "Measurement (Scratch)", "Feature Image (Gear)", and "Comment" are selected. Furthermore, in the condition selection screen illustrated in Figure 5, you can also specify the range of curvature angles in "Curvature" (e.g., "** degrees or more") and the range of measured values in "Measurement (Scratches)" (e.g., "** mm or more"), thereby limiting the range to be extracted.
[0045] Figure 6 illustrates a specific example of extraction in S240. In this example, in the condition selection screen exemplified in Figure 5, tags related to "measurement (scratch)", "feature image (gear)", and "comment" are selected as extraction condition tags, and the AND condition of these tags is set as the extraction condition. In this case, in the extraction of S240, the corresponding section in the video is first identified for each selected tag. Specifically, the measurement processing section is identified based on the frame image to which the tag related to "measurement (scratch)" is attached, the feature image recognition section is identified based on the frame image to which the tag related to "feature image (gear)" is attached, and the comment assignment section is identified based on the frame image to which the tag related to "comment" is attached. The measurement processing section is the section from the frame image to which the tag related to the measurement operation at the start of measurement processing is attached to the frame image to which the tag related to the measurement operation at the end of measurement processing is attached. The feature image recognition section is the section of frame images to which a tag related to a specific feature image (gear) is attached. The comment assignment section is a predetermined time (e.g., 1 minute) section including the time of the frame image to which the tag related to the comment assignment operation is attached. In this specific example, as shown in Figure 6, in a video with a recording time of "01:32:00", the measurement processing intervals are identified as "00:02:00" to "00:03:00" and "01:30:30" to "01:31:30", the feature image recognition intervals are identified as "00:01:00" to "00:03:30" and "01:30:00" to "01:31:30", and the comment assignment intervals are identified as "00:02:00" to "00:03:00" and "01:30:30" to "01:31:30". Once the intervals for each selected tag are identified in this way, the next step is to extract the intervals that satisfy the AND condition of those intervals for each tag. In this specific example, the intervals from "00:02:00" to "00:03:00" and from "01:30:30" to "01:31:30" are extracted. This means that the frame images included in those intervals are extracted. In this specific example, the frame images included in those intervals are frame images in which scratch measurement processing has been performed, frame images in which gears are represented, and frame images within a predetermined time period (e.g., 1 minute) including the time when the comment was added.
[0046] Figure 6 also shows, for reference, the intervals identified based on frame images tagged with "Other" (unselected). Specifically, it also shows the lighting-on interval identified based on frame images tagged with "Lighting," the curvature operation interval identified based on frame images tagged with "Curvature," the evaluation result application interval identified based on frame images tagged with "Evaluation Result," and the marking application interval identified based on frame images tagged with "Marking." The lighting-on interval is the interval from the frame image tagged with lighting operation (lighting on) to the frame image tagged with lighting operation (lighting off). The curvature operation interval is the interval of frame images tagged with curvature operation. The evaluation result application interval is a predetermined time interval (e.g., 1 minute) including the time of the frame image tagged with evaluation result application operation. The marking application interval is a predetermined time interval (e.g., 1 minute) including the time of the frame image tagged with marking application operation. In Figure 6, the lighting section is from "00:00:30" to "01:32:00", the curve operation section is from "00:00:30" to "00:01:30", the evaluation result application section is from "00:02:30" to "00:03:30" and from "01:31:00" to "01:32:00", and the marking application section is from "00:01:30" to "00:02:30" and from "01:30:00" to "01:31:00".
[0047] As described above, the endoscope device 1 according to this embodiment can extract frame images from multiple frame images contained in a video acquired during an endoscopic examination, not limited to extraction conditions based on one type of condition element (one type of tag), but based on extraction conditions based on multiple types of condition elements (multiple types of tags). Therefore, important frame images can be narrowed down and extracted more precisely. In addition, the user can select the condition elements to be used as extraction conditions, so they can understand what criteria were used to extract the images.
[0048] In this embodiment, some tags included in the condition selection screen displayed in S220 may be fixed as selected, as illustrated in Figure 7. Figure 7 shows a modified example of the condition selection screen displayed in S220. In the condition selection screen illustrated in Figure 7, the tag related to "feature image (gear)" is fixed as selected. It is also displayed in a different format from other tags so that it can be identified as being fixed as selected. In this example, it is displayed in a different format from other tags by superimposing a predetermined pattern, but it may also be displayed in a different format from other tags by graying it out. According to the condition selection screen illustrated in Figure 7, the user can further narrow down the feature image recognition section in the video (for example, the section from "00:01:00" to "00:03:30" and the section from "01:30:00" to "01:31:30", which are the feature image recognition sections shown in Figure 6) by selecting tags other than the tag related to "feature image (gear)" based on the selected tags. Furthermore, if the video contains frame images tagged with multiple specific feature images, such as frame images tagged with "feature image (gear)" or frame images tagged with "feature image (blade)," the user may be allowed to specify which feature image tags will be fixed as selected in the condition selection screen exemplified in Figure 7 (for example, by specifying the tag related to "feature image (blade)"). By using such a condition selection screen exemplified in Figure 7, it is possible to pre-set the most important tags among multiple tags, thereby enabling more accurate extraction of the target frame images.
[0049] Furthermore, in this embodiment, the condition selection screen displayed in S220 may further include a video time bar 53, as illustrated in Figure 8, and may also display the tags classified in S210 on the time bar 53 in an identifiable manner. Figure 8 is a diagram showing a modified example of the condition selection screen displayed in S220. The condition selection screen illustrated in Figure 8 includes tags related to "lighting," "measurement (scratches)," "feature image (gears)," and "comments" as tags classified in S210, and includes checkboxes 51 (51a to 51d) and icons 52 (52a to 52d) corresponding to each tag. In addition, the condition selection screen illustrated in Figure 8 further includes a video time bar 53, and icons 52 corresponding to the tags classified in S210 are displayed on the time bar 53 at their corresponding positions. With such a condition selection screen illustrated in Figure 8, the user can select the desired tags by checking the checkboxes 51 of the desired tags while confirming the icons 52 displayed on the time bar 53. Furthermore, on the time bar 53, the position where the icon 52c corresponding to the tag related to the "feature image (gear)" is displayed is, for example, the position at which the feature image began to be recognized. With a condition selection screen like the one exemplified in Figure 8, the timing of tag assignment and the type of tag can be visually understood, making the extraction of frame images more intuitive.
[0050] Furthermore, in the condition selection screen illustrated in Figure 8, tag selection may also be performed by selecting icons 52 displayed on the time bar 53 by specifying an area (e.g., by dragging), as illustrated in Figure 9. Figure 9 shows an example of tag selection by specifying an area. In the selection example shown in Figure 9, the icons 52b corresponding to the tag related to "Measurement (Scratch)" and 52c for the tag related to "Feature Image (Gear)" displayed on the time bar 53 are selected by specifying an area, thereby selecting the tags related to "Measurement (Scratch)" and "Feature Image (Gear)". In this case, the checkboxes 51b and 51c for the tags selected by specifying an area are automatically checked. Tags may be selected in this way by specifying an area.
[0051] Furthermore, in this embodiment, a partial video, which is a part of the video, may be generated using the frame image extracted in S240 (the corresponding section of the video extracted in S240) and recorded in the recording unit 35 or an external storage device connected to the external interface 32. This allows the partial video to be reviewed later.
[0052] Furthermore, in this embodiment, the specific feature image recognized by the feature image recognition process during the video recording process may be determined, for example, based on the content of tags related to the specific feature image added by the tagging unit 34a to frame images included in a video recorded by a video recording process previously performed on the same type of subject under the operation of a skilled inspector. This makes it possible to add tags related to the specific feature image in the same way as when the video recording process is performed under the operation of a skilled inspector, even when the video recording process is performed under the operation of an inexperienced inspector.
[0053] Furthermore, in this embodiment, the specific feature image recognized by the feature image recognition process during the video recording process may be obtained, for example, by machine learning the features of frame images with a specific type of tag attached, which are included in videos recorded by video recording processes previously performed on the same type of subject. For example, if the specific feature image is obtained by machine learning the features of frame images with tags related to scratch measurement operations, then tags can be attached to frame images on which scratch measurement operations should be performed. The tags in this case may include information indicating that the frame image on which the scratch measurement operation should be performed. Alternatively, the specific feature image may be obtained by machine learning the features of frame images with tags related to scratch measurement operations with a limited range of measurement results (for example, scratch measurement operations of "**mm or more"). This allows tags to be attached to frame images on which scratch measurement operations with a limited range should be performed.
[0054] Furthermore, in this embodiment, the main unit 30 may be connected to a network by wire or wireless and may be equipped with a communication interface for communicating with external devices (such as servers) connected to the network. This allows, for example, data (such as video) acquired by the endoscope device 1 to be shared on the cloud.
[0055] Furthermore, in this embodiment, the functions of a part of the main body 30 (for example, the control unit 34, etc.) may be implemented by an external control device, so that the endoscope device 1 is implemented as an endoscope system comprising an endoscope device and a control device. In this case, the control device may be implemented by a computer 100 as illustrated in Figure 10.
[0056] Figure 10 is a diagram illustrating the hardware configuration of computer 100. The computer 100 illustrated in Figure 10 includes a processor 101, memory 102, input device 103, output device 104, storage device 105, portable storage medium drive device 106, communication interface 107, and input / output interface 108, each of which is connected to a bus 109 and can send and receive data from each other.
[0057] The processor 101 is a CPU, and it performs various processes by executing programs such as the OS (Operating System) and applications. The memory 102 includes RAM and ROM. RAM temporarily stores parts of the programs executed by the processor 101. RAM is also used as the work area of the processor 101. ROM stores programs executed by the processor 101 and various data necessary for program execution.
[0058] The input device 103 is a keyboard, mouse, touch panel, joystick, etc. The output device 104 is a display device such as an LCD.
[0059] The storage device 105 is a device that stores data, such as an HDD or SSD. The portable storage medium drive 106 drives the portable storage medium 106a and accesses its contents to read and write data. The portable storage medium 106a is a memory device, flexible disk, optical disk, magneto-optical disk, etc. This portable storage medium 106a also includes CD-ROMs (Compact Disc Read Only Memory), DVDs (Digital Versatile Discs), Blu-ray discs, USB memory, SD card memory, etc.
[0060] The communication interface 107 is connected to a network by wire or wireless connection and is an interface for communicating with external devices connected to the network. The input / output interface 108 is connected to an external device such as an endoscope and is an interface for inputting and outputting data with the external device.
[0061] In such a computer 100, the programs executed by the processor 101 and the various data necessary for program execution may be stored not only in the memory 102, but also in the storage device 105 or the portable storage medium 106a. Furthermore, the programs executed by the processor 101 and the various data necessary for program execution may be stored from an external device via a network or communication interface 107 in one or more of the memory 102, storage device 105, and portable storage medium 106a.
[0062] Furthermore, the computer 100 is not limited to the configuration illustrated in Figure 10; it may be configured with multiple of the components illustrated in Figure 10, or with some components omitted. For example, the computer 100 may have multiple processors.
[0063] Furthermore, the computer 100 may be configured to include hardware such as a microprocessor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field-Programmable Gate Array). For example, the processor 101 may be implemented using at least one of these hardware components.
[0064] Although embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various improvements and modifications are possible without departing from the spirit of the invention. [Explanation of Symbols]
[0065] 1 Endoscopy System 10 Insertion part 11 Tip 11a Imaging optical system 11b Image sensor 11c light-emitting element 11d Illumination optical system 12 Curved section 13 Flexible tube section 20 Control section 21 Joysticks 30 Main body 31 Display section 31a Touch Panel 32 External Interfaces 33 Image generation unit 34 Display section 34a Tagging section 34b Extraction part 35 Records Section 41, 41a, 41b, 41c, 41d, 41e, 41f, 41g checkboxes 51, 51a, 51b, 51c, 51d (community) 52, 52a, 52b, 52c, 52d Icon 53 Time Bar 100 Computers 101 Processors 102 memory 103 Input device 104 Output device 105 Storage device 106 Portable storage medium drive device 106a Portable storage medium 107 Communication Interface 108 Input / Output Interfaces 109 Bus
Claims
1. An insertion part including an image sensor, which is inserted into the inside of a subject that is a machine, An operation reception unit that accepts operations, Control unit and Equipped with, The control unit, During the recording process of a video including multiple frame images generated based on the imaging signal output from the image sensor, a tagging unit adds information about the operation as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation, and adds information about the specific feature image as a tag to the frame image recognized as a specific feature image representing a component contained inside the machine. An extraction unit that extracts frame images from the multiple frame images contained in the video based on at least one of two or more tags selected from the multiple types of tags, An endoscope device for the industrial field, characterized by including the following features.
2. The endoscope apparatus according to claim 1, characterized in that the specific feature image is determined in accordance with a feature image selection operation received by the operation reception unit before the video recording process.
3. The endoscope apparatus according to claim 1, characterized in that the specific feature image is determined based on the content of a tag related to the specific feature image that is added by the tagging unit to a frame image included in a video recorded by a recording process previously performed on the same type of subject.
4. The endoscope apparatus according to claim 1, characterized in that the aforementioned specific feature image is obtained by machine learning the features of frame images to which a specific type of tag has been attached, which are included in a video recorded by a recording process previously performed on the same type of subject.
5. The endoscope apparatus according to claim 1, characterized in that the two or more types of tags are selected in accordance with the tag selection operation received by the operation reception unit.
6. A display unit that displays tags attached to frame images included in the aforementioned video. Furthermore, The endoscope apparatus according to claim 5, characterized in that the tag selection operation is an operation to select two or more types of tags from among the tags displayed by the display unit.
7. The endoscope apparatus according to claim 6, further characterized in that the display unit displays a time bar for the video and displays tags attached to the frame images included in the video in an identifiable manner on the time bar.
8. A display unit that displays the time bar of the aforementioned video and displays the tags attached to the frame images included in the video in an identifiable manner on the time bar. Furthermore, The endoscope apparatus according to claim 5, characterized in that the tag selection operation is an operation to select two or more types of tags by specifying a region from among the tags that are identifiable on the time bar displayed by the display unit.
9. The endoscope apparatus according to claim 1, characterized in that some of the two or more types of tags are pre-selected, and the other types of tags are selected in accordance with the tag selection operation received by the operation reception unit.
10. The endoscope apparatus according to claim 1, characterized in that the extraction unit extracts a corresponding section in the video for each of the two or more types of tags and performs the extraction based on the section.
11. A display unit that displays the frame image extracted by the extraction unit. The endoscopic apparatus according to claim 1, further comprising the following:
12. Using an endoscope device for the industrial field, During the recording process of a video including multiple frame images generated based on imaging signals output from an image sensor contained in an insertion part inserted into the inside of a machine (the subject), information about the operation is added as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation, and information about the specific feature image is added as a tag to the frame image recognized as a specific feature image representing a component contained inside the machine. A frame image extraction method characterized by extracting a frame image from the multiple frame images contained in the video based on at least one tag selected from two or more tags from among the multiple types of tags.
13. During the recording process of a video including a plurality of frame images generated based on an imaging signal output from an imaging sensor contained in an insertion part inserted into a subject which is a machine, information about the operation is added as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation, and information about the specific feature image is added as a tag to the frame image recognized as a specific feature image representing a component contained inside the machine, Based on at least one of two or more tags selected from among multiple types of tags, a frame image is extracted from the multiple frame images contained in the video. A program characterized by causing a processor to perform processing either mounted on or located outside an endoscope device used in the industrial field.
14. An endoscope device for the industrial field, Control device and Equipped with, The aforementioned endoscope device is An insertion part containing an image sensor is inserted into the inside of the subject, which is a machine, An operation reception unit that accepts operations, Includes, The control device is During the recording process of a video including multiple frame images generated based on the imaging signal output from the image sensor, a tagging unit adds information about the operation as a tag to the frame image corresponding to the timing when the operation reception unit receives an operation, and adds information about the specific feature image as a tag to the frame image recognized as a specific feature image representing a component contained inside the machine. An extraction unit that extracts frame images from the multiple frame images contained in the video based on at least one of two or more tags selected from the multiple types of tags, An endoscopic system characterized by including the following:
15. The endoscope apparatus according to claim 1, wherein the extraction unit identifies the interval from the start time to the end time in the video for each selected tag.