Method and device for providing video related to medical condition
The method and device efficiently summarize key sections of a medical situation video by processing object and voice monitoring results, enabling quick and intuitive review of medical situations, thus addressing the challenge of lengthy medical videos.
Patent Information
- Application Number
- PCT/KR2024/018272
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-12
AI Technical Summary
The challenge is to efficiently provide a summary video of key sections from a medical situation video, which is often too long to be conveniently reviewed, especially when a patient is being moved to a hospital.
A method and device that receive a medical situation video, process it to identify important situation images based on object and voice monitoring results, and generate a summary video with a shorter playtime, along with optional annotations and thumbnail images.
This solution allows for intuitive access to key sections of a medical situation video, improving user convenience and satisfaction by providing a concise summary that can be easily reviewed, facilitating quicker assessment of patient medical situations.
Smart Images

Figure KR2024018272_12062025_PF_FP_ABST
Abstract
Description
Method and device for providing video related to medical situations
[0001] The technical field of the present disclosure relates to a method for providing a video related to a medical situation, and a technology is disclosed for providing annotations for a video including a key section based on a monitoring result.
[0002] Recently, endoscopic devices equipped with cameras have been used in various fields. As imaging devices inserted into the body or used to film procedures or surgical sites become smaller and more widespread, the acquisition of videos during internal and external surgical procedures is increasing. Furthermore, with the recent advancements in big data and artificial intelligence, an increasing number of hospitals and physicians are processing videos into medical information content and conducting various research projects based on these videos, including standardization research on technologies and evaluation of tool usability. This has led to a growing demand for video storage and rapid retrieval of key scenes. Furthermore, when viewing medical situations during a patient's transfer to the hospital, the length of the video makes it difficult to view the entire video and selectively review relevant sections. Therefore, a method is needed to facilitate access to key sections and provide summarized information.
[0003] The problem to be solved in the present disclosure is to provide an annotation indicating the characteristics of the main sections of a target video related to a medical procedure, and to provide a method for providing a summary video summarized according to the main sections.
[0004] The problems to be solved in this disclosure are not limited to the technical problems described above, and other technical problems may exist.
[0005] As a technical means for achieving the above-described technical task, a method for providing a video related to a medical situation according to the first aspect of the present disclosure may include: a step of a receiving unit obtaining a medical situation video including a plurality of images and a voice signal; a step of the processor obtaining an object monitoring result for one or more objects included in the plurality of images; a step of the processor obtaining a voice monitoring result for the voice signal; a step of the processor determining a plurality of important situation images corresponding to an important situation among the plurality of images based on the object monitoring result and the voice monitoring result; and a step of the processor providing a summary video including the plurality of important situation images and having a play time shorter than the medical situation video.
[0006] In addition, the step of obtaining the object monitoring result may include a step of the processor obtaining an annotation for the one or more objects; and a step of the processor obtaining the object monitoring result including the annotation.
[0007] In addition, the step of providing the summary video may include a step in which the processor obtains a plurality of image groups by grouping temporally adjacent important situation images among the plurality of important situation images; and a step in which the processor obtains the summary video by arranging the plurality of image groups according to the passage of time.
[0008] In addition, the processor may further include a step of obtaining a thumbnail image corresponding to each of the plurality of image groups; and a step of displaying the thumbnail image by overlapping it on the summary video.
[0009] In addition, the processor further includes a step of obtaining a vital signal synchronized with the medical situation video; and the step of determining the plurality of important situation images allows the processor to determine the plurality of important situation images based on the vital signal.
[0010] In addition, the step of determining the plurality of important situation images may allow the processor to determine the plurality of important situation images based on the importance of a body part corresponding to the object targeted as a result of the object monitoring.
[0011] In addition, the step of providing the summary video may include the processor providing the summary video based on a weight given in the order of importance of the body part corresponding to the voice monitoring result, the vital signal, and the object monitoring result targeted by the processor.
[0012] In addition, the step of providing the summary video may include providing the summary video based on a weight given in the order of the vital signal, the voice monitoring result, the importance of the body part, and the type of medical tool corresponding to the targeted object as a result of the object monitoring, for a section in which the voice decibel as a result of the voice monitoring is higher than a preset risk-predicted decibel.
[0013] In addition, the step of providing the summary video may include providing the summary video based on a weight given in the order of the type of medical tool corresponding to the object targeted by the object monitoring result, the vital signal, the importance of the body part, and the voice monitoring result for a section in which the voice decibel of the voice monitoring result is less than a preset decibel.
[0014] In addition, the step of acquiring the plurality of image groups may update the number of adjacent important situation images to be grouped when the type of medical tool corresponding to the object targeted as a result of the object monitoring is included in a preset type of risk tool.
[0015] In addition, the processor may further include a step of providing a medical situation report according to the annotation of the plurality of important situation images, the object monitoring result, and the voice monitoring result together with the summary video.
[0016] In addition, a device for providing a video related to a medical situation according to a second aspect of the present disclosure may include a receiving unit for obtaining a medical situation video related to a medical practice and including a plurality of images and a voice signal; and a processor for obtaining an object monitoring result for one or more objects included in the plurality of images, obtaining a voice monitoring result for the voice signal, determining a plurality of important situation images corresponding to an important situation among the plurality of images based on the object monitoring result and the voice monitoring result, and providing a summary video including the plurality of important situation images and having a play time shorter than the medical situation video.
[0017] Additionally, the processor can obtain an annotation for the one or more objects and obtain an object monitoring result including the annotation.
[0018] In addition, the processor can obtain a plurality of image groups by grouping temporally adjacent important situation images among the plurality of important situation images, and can obtain the summary video by arranging the plurality of image groups according to the flow of time.
[0019] Additionally, the present disclosure may include a non-transitory computer-readable recording medium having recorded thereon a program for implementing the first aspect according to the third aspect.
[0020] According to one embodiment of the present disclosure, convenience can be improved in that annotations are provided for key sections that are judged to be important in medical situations, so that users can intuitively confirm key sections that are judged to be important.
[0021] Additionally, it can improve user satisfaction by providing a summary video of the medical situation video as a key section of the medical situation, making it easy to check the patient's medical situation in the hospital.
[0022] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0023] FIG. 1 is a diagram illustrating an example of a device or server implemented on a system according to one embodiment.
[0024] Figure 2 is a block diagram schematically illustrating the configuration of a device according to one embodiment.
[0025] Figure 3 is a flowchart illustrating each step of operation of a device according to one embodiment.
[0026] FIG. 4 is a schematic diagram illustrating an example of a device obtaining a medical situation video and obtaining annotations for one or more objects according to an embodiment.
[0027] FIG. 5 is a diagram illustrating an example of a device arranging multiple image groups over time to obtain a summary video according to an embodiment of the present invention.
[0028] FIG. 6 is a diagram illustrating an example of a device according to one embodiment displaying a thumbnail image superimposed on a summary video.
[0029] FIG. 7 is a diagram illustrating an example of a device displaying a summary video according to one embodiment.
[0030]
[0031] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure is complete and to fully inform those skilled in the art of the scope of the present disclosure.
[0032] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present disclosure. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the specification, and "and / or" includes each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present disclosure.
[0033] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those skilled in the art. Furthermore, terms defined in commonly used dictionaries should not be interpreted ideally or excessively unless explicitly defined otherwise.
[0034] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" can be used to easily describe the relationship between one component and other components as depicted in the drawings. Spatially relative terms should be understood to include different orientations of the components during use or operation in addition to the orientations depicted in the drawings. For example, if a component depicted in the drawings were flipped over, a component described as "below" or "beneath" another component could end up "above" the other component. Thus, the exemplary term "below" can include both the above and below orientations. Components can also be oriented in other directions, and thus spatially relative terms can be interpreted accordingly.
[0035] Hereinafter, embodiments are described in detail with reference to the drawings.
[0036]
[0037] FIG. 1 is a drawing showing an example of a device (100) or server implemented on a system according to one embodiment.
[0038] As illustrated in FIG. 1, the medical information system may include an information acquisition device (110), a device (100), an external server (130), a storage medium (140), a communication device (150), a virtual server (160), a user terminal (170), and a network.
[0039] However, those skilled in the art will appreciate that, in addition to the components illustrated in Figure 1, other general-purpose components may be included in the medical information system. For example, the medical information system may further include a blockchain server (not shown) that operates in conjunction with a network. Alternatively, those skilled in the art will appreciate that, in other embodiments, some of the components illustrated in Figure 1 may be omitted.
[0040] According to one embodiment, a device (100) can obtain information related to a medical procedure from an information acquisition device (110). The information acquisition device (110) may include, but is not limited to, a photographing device, a recording device, a biosignal acquisition device, etc. The biosignal may include, without limitation, signals obtained from a living body, such as a body temperature signal, a pulse signal, a respiration signal, a blood pressure signal, an electromyogram signal, and an brain wave signal. An example of the information acquisition device (110), a photographing device, may include, but is not limited to, a first photographing device (e.g., CCTV, etc.) that photographs the entire operating room situation and a second photographing device (e.g., endoscope, etc.) that focuses on the surgical site.
[0041] According to one embodiment, a device (100) may acquire images (videos, still images, etc.) related to medical procedures such as surgery from an information acquisition device (110). The device (100) may perform image processing on the acquired images. Image processing according to one embodiment may include, but is not limited to, naming, encoding, storage, transmission, editing, and metadata generation for each image.
[0042] According to one embodiment, a device (100) can transmit medical practice-related information acquired from an information acquisition device (110) to a network, either as is or in an updated form. The transmission information that the device (100) transmits to the network can be transmitted to an external device (130, 140, 150, 160, 170) via the network. For example, the network can transmit the transmission information that the device (100) transmitted to the network, either as is or in an updated form, to an external server (130), a storage medium (140), a communication device (150), a virtual server (160), a user terminal (170), etc. The device (100) can also receive information (e.g., feedback information, an update request, etc.) received from an external device (130, 140, 150, 160, 170). A communication device (150) can refer to a device used for communication without limitation (e.g., a gateway), and the communication device (150) can communicate with a device that is not directly connected to a network, such as a user terminal (180).
[0043] A device (100) according to an embodiment may include an input unit, an output unit, a processor, a memory, etc., as described below, and may also include a display device (not shown). For example, a user may check a communication status, memory usage status, power status (e.g., battery state of charge, whether external power is supplied, etc.), a thumbnail image for a stored video, the current operating mode, etc., through the display device. Meanwhile, the display device may be a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode, a flexible display, a 3D display, an electrophoretic display, etc. In addition, the display device may include two or more displays depending on the implementation form. In addition, when the touchpad of the display is configured as a touch screen by forming a layer structure, the display can be used as an input device in addition to the output device.
[0044] In addition, the network can perform mutual communication through wired communication or wireless communication. For example, the network can be implemented as a type of server and can include a Wi-Fi chip, a Bluetooth chip, a wireless communication chip, an NFC chip, etc. Of course, the device (100) can perform communication with various external devices using a Wi-Fi chip, a Bluetooth chip, a wireless communication chip, an NFC chip, etc. The Wi-Fi chip and the Bluetooth chip can perform communication using the Wi-Fi method and the Bluetooth method, respectively. When using a Wi-Fi chip or a Bluetooth chip, various connection information such as an SSID and a session key are first transmitted and received, and then communication is established using this, and various information can be transmitted and received. The wireless communication chip can perform communication according to various communication standards such as IEEE, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), etc. NFC chips can operate in the NFC (Near Field Communication) mode using the 13.56MHz band among various RF-ID frequency bands such as 135kHz, 13.56MHz, 433MHz, 860~960MHz, and 2.45GHz.
[0045] An input unit according to one embodiment may refer to a means for a user to input data for controlling a device (100). For example, the input unit may include, but is not limited to, a key pad, a dome switch, a touch pad (contact electrostatic capacitance type, pressure resistive film type, infrared detection type, surface ultrasonic conduction type, integral tension measurement type, piezo effect type, etc.), a jog wheel, a jog switch, etc.
[0046] An output unit according to one embodiment may output an audio signal, a video signal, or a vibration signal, and the output unit may include a display device, an audio output device, and a vibration motor.
[0047] The user terminal (170) according to one embodiment may include various wired and wireless communication devices such as a smartphone, a smart pad, a tablet PC, etc., but is not limited thereto.
[0048] According to one embodiment, the device (100) can update medical practice related information acquired from the information acquisition device (110). For example, the device (100) can perform naming, encoding, storage, transmission, editing, metadata generation, etc. for images acquired from the information acquisition device (110). As an example, the device (100) can perform naming of image files using metadata (e.g., creation time) of the acquired images. As another example, the device (100) can classify images related to medical practices acquired from the information acquisition device (110). The device (100) can classify images related to medical practices based on various criteria, such as the type of medical practice and the medical practitioner, using learned AI.
[0049] Additionally, in FIG. 1, the device (100) may be implemented as a server, and the range of physical devices in which the device (100) may be implemented is not interpreted as being limited.
[0050]
[0051] Figure 2 is a block diagram schematically illustrating the configuration of a device (100) according to one embodiment.
[0052] As illustrated in FIG. 2, a device (100) according to an embodiment may include a receiver (210), a processor (220), an output unit (230), and a memory (240). However, not all of the components illustrated in FIG. 2 are essential components of the device (100). The device (100) may be implemented with more components than the components illustrated in FIG. 2, or may be implemented with fewer components than the components illustrated in FIG. 2.
[0053] For example, a device (100) according to one embodiment may further include a communication unit in addition to a receiving unit (210), a processor (220), an output unit (230), and a memory (240). In addition, a display (not shown) may be included as an example of the output unit (230).
[0054] According to one embodiment, a receiving unit (210) can obtain a medical situation video including a plurality of images and voice signals.
[0055] According to one embodiment, the processor (220) may obtain an object monitoring result for one or more objects included in a plurality of images. In addition, the processor (220) may obtain a voice monitoring result for a voice signal. In addition, the processor (220) may determine a plurality of important situation images corresponding to an important situation among the plurality of images based on the object monitoring result and the voice monitoring result. In addition, the processor (120) may provide a summary video including a plurality of important situation images and having a shorter play time than a medical situation video. In one embodiment, the provision may include an example of a display, and the processor (120) may control the output unit (230) or the display to display the summary video.
[0056] Figure 3 is a flowchart illustrating each step of operation of a device (100) according to one embodiment.
[0057] Referring to step S310, the device (100) may acquire a medical situation video including a plurality of images and audio signals. In one embodiment, the plurality of images may be images representing still images included in a medical situation video played in real time. The plurality of images may be acquired continuously over time. Additionally, the audio signal may refer to a signal recorded in the medical situation video in real time.
[0058] Referring to step S320, the device (100) according to one embodiment may obtain an object monitoring result for one or more objects included in a plurality of images. In one embodiment, the one or more objects may be one of a plurality of preset object targets related to a medical situation. For example, the device (100) may classify one or more objects for each of the plurality of images according to an AI-based image analysis algorithm. Accordingly, the device may obtain an object monitoring result in real time over time for the determined one or more objects. In addition, the one or more objects may further include an object performing a motion. For example, the device may obtain an object monitoring result for an object performing a motion, such as a patient, a medical practitioner, or a guardian. The device (100) according to one embodiment may obtain an annotation for one or more objects. For example, the device (100) may obtain an annotation that has a label for one or more object types and a label for a motion performed by one or more objects. Accordingly, the device (100) can obtain an object monitoring result including an annotation. In this regard, this can be explained with reference to FIG. 4.
[0059] FIG. 4 is a schematic diagram illustrating an example of a device (100) obtaining a medical situation video and obtaining annotations for one or more objects according to one embodiment.
[0060] Referring to FIG. 4, a device (100) according to an embodiment can identify one or more objects for each of a plurality of images included in a medical situation video. Accordingly, labeling can be performed on one or more predetermined regions corresponding to one or more objects. Accordingly, the device (100) can perform image analysis on one or more objects included in one image to obtain annotations for which labeling has been performed. As illustrated in FIG. 4, the device (100) can obtain annotations for which labeling has been performed on each object target (such as a cervical collar or blood pressure monitor) or annotations for which labeling has been performed on the motion of the object (such as a vital monitor touch) by identifying one or more objects and performing labeling.
[0061] Referring to step S330, the device (100) according to one embodiment may obtain a voice monitoring result for a voice signal. In one embodiment, the voice monitoring result may include an example of obtaining an annotation for a corresponding area based on a voice signal of a preset pattern. For example, an annotation for a corresponding area may be obtained according to a pattern in which voice signals corresponding to preset words, such as “It’s 2040,” “Your body temperature is OO degrees,” “Are you male?” are continuously mentioned, or according to the continuity of one or more of the preset words. In addition, the device (100) may obtain a vital signal synchronized with a medical situation video. The device (100) may obtain a vital signal and further obtain an annotation for an area in which the vital signal is determined to fall within an abnormal range.
[0062] Referring to step S340, the device (100) according to one embodiment may determine a plurality of important situation images corresponding to important situations among a plurality of images based on the object monitoring results and the voice monitoring results. The device (100) according to one embodiment may determine a plurality of important situation images corresponding to major sections among a plurality of images using the annotation results included in the medical situation video. The device (100) according to one embodiment may determine a plurality of important situation images further based on vital signals. The device (100) may determine a plurality of important situation images based on the importance of a body part corresponding to a targeted object based on the object monitoring results. In one embodiment, the targeted object may be an object to be analyzed. The device (100) may determine a plurality of important situations based on the importance of a preset body part. For example, when the target of an object is labeled as a heart, the importance may be higher than when the target of an object is labeled as a toe. Accordingly, multiple critical situations can be determined for a body part with high importance based on the importance of the preset body part. The device (100) can also determine multiple critical situation images based on the importance of the type of medical tool corresponding to the targeted object. The device (100) can determine multiple critical situations based on the importance of the type of medical tool set. For example, if the target object is labeled as a cardiopulmonary resuscitator, the importance may be higher than if the target object is labeled as an IV bag. Accordingly, multiple critical situations can be determined for a type of medical tool with high importance based on the importance of the type of medical tool set.
[0063] Referring to step S350, the device (100) according to one embodiment may provide a summary video including a plurality of important situation images and having a shorter play time than the medical situation video. The play time may refer to the playback time of the summary video, and the device (100) may provide a short summary video in which a plurality of important situation images corresponding to the plurality of important situations determined in step S340 are played continuously. For example, a medical situation video including a plurality of images representing a medical situation of one hour or more in which a patient is transferred to a hospital may be updated and provided as a summary video including a plurality of important situation images summarized for a preset time (e.g., one to two minutes). The device (100) according to one embodiment may obtain a plurality of image groups by grouping temporally adjacent important situation images among the plurality of important situation images. The device (100) may obtain a summary video by arranging the plurality of image groups in chronological order. In this regard, a description may be made with reference to FIG. 5.
[0064] FIG. 5 is a drawing for explaining an example of a device (100) arranging multiple image groups over time to obtain a summary video according to one embodiment.
[0065] Referring to FIG. 5, a device (100) according to an embodiment may group one or more important situation images included in a grouping section that further includes a certain time point in both directions from one or more time points that include important situation images with an importance level higher than a preset level. Accordingly, the device (100) may obtain a summary video by arranging multiple image groups over time. For example, a certain time point may be a time point that requires continuity for a plurality of important situation images. Accordingly, multiple important situation images within a time point that requires continuity may be formed into a single group, and time points and sections corresponding to time points that do not correspond to the certain time point and sections that do not include the plurality of important situation images may be determined as unnecessary time points and sections. Accordingly, the device (100) may group and continuously provide one or more important situation images included in an area of a certain time point that requires continuity, and may provide multiple image groups from which images of unnecessary time points and sections are excluded over time. In one embodiment, the important situation images included in a certain time point may be different for each time point. In addition, the device (100) according to one embodiment may update the range of a section at a certain point in time depending on the degree to which the importance of the important situation image exceeds a preset level. For example, if the type of medical tool corresponding to the targeted object as a result of object monitoring is included in the preset type of dangerous tool, the number of adjacent important situation images to be grouped may be updated. As illustrated in FIG. 5, the device (100) may obtain a first image group (511) corresponding to a first section (510), a second image group (521) corresponding to a second section (520), and a third image group (531) corresponding to a third section (530), where the sections at a certain point in time are determined differently.A first image group (511) may include important situation images that are temporally adjacent to a first time, a second image group (521) may include important situation images that are temporally adjacent to a second time, and a third image group (531) may include important situation images that are temporally adjacent to a third time. The predetermined points in time corresponding to the first time, the second time, and the third time may be determined differently depending on the situation. For example, the first time, the second time, and the third time may be determined based on any one image among a plurality of important situation images whose importance level exceeds a preset level. In one embodiment, three times are specified, but this is not limited thereto, and a plurality of sections may be further included. In one embodiment, the importance of an important situation image may be determined based on whether it includes a preset type of risk tool, whether it includes an abnormal range for a vital signal, and whether it includes a preset risk voice word. In addition, when a preset risk tool type is continuously acquired for a preset period of time and / or when the number of preset risk tools is two or more, the interval range at a certain point in time may be updated upward. In addition, when the degree to which the vital signal deviates from the normal range is greater than or equal to a preset difference, the interval range at a certain point in time may be updated upward. In addition, when a preset risk sound word is included, the interval range at a certain point in time may be updated upward. Specifically, in one embodiment, the certain point in time may be a point in time that includes a preset time period in both directions (for example, 3 seconds in both directions) from a point in time that includes any image whose importance level exceeds a preset level. Accordingly, the device (100) may group at least one important situation image that is determined to be an important situation image but has an importance level lower than the preset level included at the certain point in time.That is, when a certain point in time is updated, some of the important situation images included in each group may be increased. If the number of preset dangerous tool types in one image determined as an important situation image exceeds a preset level and the device (100) has two or more, the certain point in time may be updated to increase by 2 seconds. In addition, if dangerous medical tools are continuously acquired at the certain point in time increased by 2 seconds, the certain point in time may be updated to increase by an additional 2 seconds. In another embodiment, the certain point in time may also be updated to increase by 2 seconds if the degree to which the vital signal deviates from the normal range is greater than or equal to a preset difference (e.g., a difference of 10 percent or more). In addition, if a preset dangerous negative word is included, the certain point in time may be updated to increase by 1 second, which is a shorter time than the above-described seconds. If one or more important situation images corresponding to the updated certain point in time include the preset dangerous negative word continuously, the updated certain point in time may be updated to increase by an additional 1 second. However, in one embodiment, the number of consecutive updates may be limited to two. If a certain point in time continuously increases, the length of the summary video may become longer, which may lower the efficiency in obtaining the summary video and lower the accuracy. Therefore, the number of consecutive updates may be limited to a preset number. In addition, if the same criteria are continuously updated (e.g., when the types of dangerous tools are consecutive or when dangerous words appear consecutively) such as whether a predetermined type of dangerous tool is included, whether an abnormal range for vital signals is included, and whether a predetermined type of dangerous voice word is included, the number of updates may be limited to two, and if different criteria overlap and are continuously updated (when a predetermined type of voice word is included and a type of dangerous tool is included), the number of updates may be limited to three. In other words, the number of times a certain point in time can be updated in cases where danger is predicted overlappingly based on different criteria may be increased to make it longer.In addition, the device (100) can provide a rendering so that the length of a certain point in time can be visually easily confirmed, such as the first section (510), the second section (520), and the third section (530) illustrated in FIG. 5. Therefore, there is an effect in that it is possible to visually easily confirm which section contains a large number of important situation images. In another embodiment, when the rendering section (certain point in time) is greater than a preset percentage (e.g., 20 percent) of the medical situation video, a summary video can be obtained by allowing only sections corresponding to 3 seconds each in both directions to be continuous based on the middle point in the rendering section. Since the efficiency of the summary video may be lowered if the rendering section is too numerous or long, only a part of the rendering section may be included in the summary video when the rendering section is greater than a preset percentage. In addition, a group of images corresponding to a certain point in time with a long rendering section can be provided with priority. For example, in general, multiple important situation images included in multiple image groups can be provided sequentially over time, but if the length of the summary video exceeds a preset time (e.g., 2 minutes), multiple important situation images included in corresponding image groups can be provided sequentially in the order of the length of the rendering section.
[0066] According to one embodiment, a device (100) can obtain thumbnail images corresponding to each of a plurality of image groups. Furthermore, the device (100) can display the thumbnail images by overlaying them on a summary video. This will be described with reference to FIG. 6.
[0067] FIG. 6 is a drawing illustrating an example of a device (100) according to one embodiment of the present invention displaying a thumbnail image by overlaying it on a summary video.
[0068] Referring to FIG. 6, a device (100) according to an embodiment may display thumbnail images corresponding to each of a plurality of image groups. In an embodiment, the thumbnail images may be representative images for each of the plurality of image groups. For example, an image corresponding to one image with the highest importance in each group may be determined as the thumbnail image. Accordingly, the device (100) may display one or more determined thumbnail images by overlaying them on a summary video. In addition, when a selection input for a thumbnail image is received, the device (100) may sequentially arrange and display one or more important situation images included in the corresponding group. In addition, an annotation list in which annotation results obtained from one or more important situation images included in the corresponding group are sequentially arranged may be further provided.
[0069] According to one embodiment, the device (100) can provide a summary video based on a weight given in order of importance of a body part corresponding to a targeted object based on a result of voice monitoring, a result of vital signals, and a result of object monitoring.
[0070] For example, in one embodiment, the accuracy of deriving voice monitoring results from voice signals may be the highest based on voice recognition technology, and in general, in cases where a patient is transported to a hospital, the importance may be high because critical situations can be determined based on the pattern of conversation between the patient or guardian and the medical practitioner, preset medical words mentioned by the medical practitioner, and the pattern of sequential listing of preset medical words. In addition, since the accuracy may be the highest because the emergency situation can be intuitively classified based on the frequency, decibel, etc. of the voice signal, the highest weight may be given to the voice monitoring results. In addition, since the vital signals may be judged as a dangerous situation when they fall within an abnormal range, the vital signals may be an important factor in determining a critical situation. However, in cases where the vital signals do not fall within an abnormal range, it may be difficult to determine a critical situation, and thus the vital signals may be given a second highest weight. In addition, in the case of the importance of a body part, it may be a factor that can be considered in that a part with an importance higher than a preset level can include the most important critical situation, but since the range of body parts is wide and diverse, it may be less important than the two factors described above in that it is difficult to confirm critical situations due to motion, etc. Therefore, a high weight may be assigned to the body part in the third rank. In addition, the device (100) according to one embodiment may provide a summary video based on the weights assigned in the order of the type of medical tool corresponding to the targeted object, vital signals, importance of the body part, and voice monitoring result for a section in which the voice decibel is lower than the preset decibel as a result of voice monitoring. For example, a section in which the voice decibel is lower than the preset decibel may be predicted to be a section in which no conversation occurs between the patient or the medical practitioner or a stable section.Therefore, in the case of a section where the voice decibel is below the preset decibel, there is a high probability that the vital signals are within the normal range, and there is a high probability that the body part is at low risk. Therefore, the type of medical tool can be given the highest weight because the importance of the action being taken during patient transport can be highly judged by checking the type of medical tool. In addition, as described above, in the case of a section where the voice decibel is below the preset decibel, there is a high probability that the vital signals are within the normal range, but even if the vital signals are within the normal range, the importance can be highly judged based on the real-time rate of change, and thus the vital signals can be given a second-highest weight. In addition, as described above, in the case of a section where the voice decibel is below the preset decibel, there is a high probability that the body part is at low risk, and there is a third-highest weight because the rate of change is consistent compared to the vital signals. In addition, for a section where the voice decibel is lower than the preset decibel, the voice monitoring result may be given a fourth-highest weight because the importance of the voice signal may be very low. In addition, the device (100) according to one embodiment may provide a summary video based on the weights given in the order of vital signals, voice monitoring results, importance of body parts, and type of medical tools corresponding to the targeted object in the object monitoring result, for a section where the voice decibel is higher than the preset risk-predicted decibel as a result of voice monitoring. In one embodiment, the risk-predicted decibel may correspond to a preset high decibel. For a section where the voice decibel is higher than the preset risk-predicted decibel, it may be predicted to be a risk section.Therefore, in the case of a section where the voice decibel is higher than the preset decibel, it may be desirable to determine whether to respond to an abnormal situation by checking the vital signals first, so the highest weight may be given to the vital signals. In addition, if the voice decibel is higher than the preset risk prediction decibel, the voice monitoring result may be the most important in that it can be predicted as a risk section. However, since the voice monitoring result includes the annotation result for words included in the voice, etc., in the case of a time point after the predicted risk section, the importance of the vital signals may be higher in terms of direct relevance to the patient's life, so the voice monitoring result may be given a second highest weight. In addition, the importance of body parts may be high in the section predicted as a risk section. For example, in the case of high-importance parts such as the head and heart, the need for monitoring is high, so it may affect the decision to use an image as a critical situation. However, when determining multiple important situation images based on voice decibels, the importance of the body part may be given a third-highest weighting because the influence of the patient's vital signs and situations appearing in high-decibel voices may be greater. In addition, since all medical tools used in areas predicted to be at risk are likely to be tools for responding to dangerous situations, the importance of the type of medical tool may be lower than the above-described factors, and therefore the type of medical tool may be given a fourth-highest weighting. In one embodiment, since the voice signal is a highly accurate signal, the device (100) may give different weights to the above-described factors according to each situation based further on the voice decibels. Therefore, in various situations, rather than applying a uniform weighting size, there is an advantage in that more appropriate multiple important situation images can be determined because the weights determined by differently judging the importance according to the situation are applied.
[0071] According to one embodiment, the device (100) can provide a medical situation report based on annotations of multiple critical situation images, object monitoring results, and voice monitoring results, along with a summary video. In one embodiment, the medical situation report can be a report that organizes the annotation results, object monitoring results, and voice monitoring results of multiple critical situation images, such as an emergency record or a surgical record. The device (100) can obtain a medical situation report created in a preset manner based on the annotations, object monitoring results, and voice monitoring results. Accordingly, the device (100) can provide the medical situation report along with a summary video.
[0072] FIG. 7 is a diagram showing an example of a device (100) displaying a summary video according to one embodiment.
[0073] Referring to FIG. 7, a device (100) according to an embodiment may sequentially arrange and display a plurality of thumbnail images corresponding to a plurality of image groups, each with annotations. For example, the device (100) may display a plurality of thumbnail images and annotations corresponding to the images at the bottom of a reservation video. In addition, the device (100) may display an annotation list in which annotations are arranged by section according to object monitoring results and voice monitoring results on the right side of the reservation video. The annotations by section may correspond to a plurality of thumbnail images displayed at the bottom of a summary video. According to the annotations by section illustrated in FIG. 7, the device (100) may classify and display annotations in which labels for one or more object types and motions performed by one or more objects are labeled by section. Therefore, a user can easily identify sections corresponding to important situations in a medical situation video through the annotations by section. Additionally, convenience can be enhanced in that users can receive a summary video based on the annotation results. In another embodiment, the order of annotations for each segment can be updated in descending order of the length of the rendering segment, as described above. Accordingly, multiple important situational images corresponding to segments with longer rendering segments can be displayed preferentially.
[0074] In one embodiment, providing annotations for key sections deemed important within a medical situation can enhance user convenience by allowing users to intuitively identify key sections deemed important. Furthermore, providing a summary video of the medical situation video, highlighting key sections within the medical situation, can enhance user satisfaction by allowing hospitals to easily check the patient's medical condition.
[0075]
[0076] Various embodiments of the present disclosure may be implemented as software comprising one or more instructions stored in a storage medium (e.g., memory) readable by a machine (e.g., a display device or a computer). For example, a processor of the machine (e.g., processor 220) may call at least one instruction from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0077] According to one embodiment, the method according to the various embodiments disclosed in the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0078] Although the present invention has been described with reference to the drawings, it is not limited to the disclosed embodiments and drawings, and a person skilled in the art related to the present embodiment will understand that the present invention can be implemented in a modified form without departing from the essential characteristics of the above-described description. Therefore, the disclosed methods should be considered from an illustrative rather than a restrictive point of view. Even if the operation and effect according to the configuration of the present invention is not explicitly described and described while describing the embodiment, the effect that can be predicted by the configuration can also be recognized. The scope of the present invention is indicated by the claims rather than the foregoing description, and all differences within the equivalent scope should be interpreted as being included in the present invention.
Claims
1. In the method of providing a video related to a medical situation, A step of obtaining a medical situation video including a plurality of images and audio signals by a receiver; A step in which the above processor obtains an object monitoring result for one or more objects included in the plurality of images; A step in which the above processor obtains a voice monitoring result for the voice signal; The step of the processor determining a plurality of important situation images corresponding to the important situation among the plurality of images based on the object monitoring result and the voice monitoring result; and A method comprising: providing a summary video comprising the plurality of critical situation images and having a shorter playing time than the medical situation video; 2. In paragraph 1, The step of obtaining the above object monitoring results is a step of the processor obtaining an annotation for the one or more objects; and A method comprising: a step of obtaining an object monitoring result including the annotation by the processor; 3. In paragraph 1, The steps to provide the above summary video are: The step of the above processor obtaining a plurality of image groups by grouping temporally adjacent important situation images among the plurality of important situation images; and A method comprising: a step of the processor arranging the plurality of image groups over time to obtain the summary video; 4. In paragraph 3, The step of the above processor obtaining a thumbnail image corresponding to each of the plurality of image groups; and A method further comprising: a step of the processor displaying the thumbnail image by overlaying it on the summary video; 5. In paragraph 1, The above processor further comprises a step of obtaining vital signals synchronized with the medical situation video; The step of determining the above multiple important situation images is A method wherein the processor determines the plurality of important situation images based on the vital signals.
6. In paragraph 1, The step of determining the above multiple important situation images is A method wherein the processor determines the plurality of important situation images based on the importance of body parts corresponding to the targeted object as a result of the object monitoring.
7. In paragraph 5, The steps to provide the above summary video are: A method wherein the processor provides the summary video based on weights assigned in order of importance of body parts corresponding to the voice monitoring results, the vital signals, and the targeted object based on the object monitoring results.
8. In paragraph 7, The steps to provide the above summary video are: A method wherein the processor provides the summary video based on weights given in the order of high to low in the vital signal, the voice monitoring result, the importance of the body part, and the type of medical tool corresponding to the targeted object in the object monitoring result, for a section in which the voice decibel of the voice monitoring result is higher than a preset risk expected decibel.
9. In paragraph 7, The steps to provide the above summary video are: A method wherein the processor provides the summary video based on a weight given in the order of the type of medical tool corresponding to the targeted object as a result of the object monitoring, the importance of the vital signal, the body part, and the result of the voice monitoring for a section in which the decibel of the voice monitoring result is less than a preset decibel.
10. In paragraph 3, The step of obtaining the above multiple image groups is A method for updating the number of grouped adjacent important situation images when the type of medical tool corresponding to the object targeted by the object monitoring result is included in a preset type of dangerous tool.
11. In paragraph 2, A method further comprising: providing a medical situation report according to the annotation of the plurality of important situation images, the object monitoring results and the voice monitoring results together with the summary video by the processor; 12. In the method of providing a video related to a medical situation, A receiving unit for acquiring a medical situation video containing multiple images and audio signals; and Obtaining object monitoring results for one or more objects included in the above multiple images, Obtaining voice monitoring results for the above voice signal, Based on the object monitoring result and the voice monitoring result, a plurality of important situation images corresponding to the important situation are determined among the plurality of images, A device comprising a processor comprising the plurality of important situation images and providing a summary video having a shorter play time than the medical situation video.
13. In paragraph 12, The above processor Obtain annotations for one or more of the above objects, A device for obtaining the object monitoring result including the above annotation.
14. In paragraph 12, The above processor obtains a plurality of image groups by grouping temporally adjacent important situation images among the plurality of important situation images, A device for obtaining the summary video by arranging the above-mentioned multiple image groups in time.
15. A computer-readable recording medium having recorded thereon a program for executing the method of any one of claims 1 to 11 on a computer.
Citation Information
Patent Citations
Surgery video transmitting device and method for the same, surgery video reproducing device, and surgery video transmitting and reproducing system
JP2010029511A
A means supporting of curb and construction methode thereof
KR1020210026186A
Apparatus for detecting biological and non-biological particles using reflective optical system
KR1020250043698A
Blockchain-based voting system
KR102160819B1
Forming apparatus, determination method, and article manufacturing method
KR102809595B1