Medical information processing apparatus, medical information processing method, and program
The medical information processing device improves event identification accuracy in clinical procedures by integrating multiple data modalities with a common timeline to classify events, facilitating automated and precise reporting.
Patent Information
- Application Number
- JP2025117450
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-07-11
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods for identifying medical events during clinical procedures, such as manual video labeling and automated video processing, are time-consuming and inaccurate.
A medical information processing device that integrates multiple data modalities (e.g., video, audio, imaging) with a common timeline to identify and classify events by processing data for each modality, generating labels, and outputting higher-level event information based on relative occurrence times.
Enhances the accuracy of event identification and enables automated, real-time reporting and analysis of clinical procedures by leveraging a combination of data modalities to derive higher-level event classifications.
Smart Images

Figure 2026015264000001_ABST
Abstract
Description
[Technical Field]
[0001] The embodiments disclosed in the present specification and drawings relate to a medical information processing device, a medical information processing method, and a program. [Background technology]
[0002] A clinical procedure may include several events that occur throughout the procedure. For example, a surgical procedure may include several stages, such as administering an appropriate injection, e.g., an anesthetic; inserting an instrument, e.g., an endoscope; a cleaning phase; an incision phase; and a suturing phase. Additional events that may be identified throughout the clinical procedure include the presence of all medical personnel and the arrival of a guidewire at its target. It may be desirable to identify the occurrence of such events during the clinical procedure to provide real-time guidance or to provide a record of the procedure for future use.
[0003] One way to identify event occurrences is for a clinician to manually provide an identification of events during a procedure, for example, based on a video recording of the procedure. However, manual labeling of video data would significantly consume clinician time that could be better directed at other efforts. Also, identifying event types based solely on video data would be highly limited. Another proposed approach is to process video data with an automated model. However, the types of events identified based on video data would be highly limited. Also, the evaluation of video data may be inaccurate. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-025590 Summary of the Invention [Problem to be solved by the invention]
[0005] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to improve the accuracy of outputting medical events from a data source. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described below can also be positioned as other problems. [Means for solving the problem]
[0006] An information processing device according to an embodiment includes an acquisition unit, an identification unit, and an output unit. The acquisition unit acquires a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of which is an imaging data type, and the remaining data modalities other than the imaging data type are additional data types different from imaging data, and the plurality of data is provided in association with a common timeline that is common to the time at which the plurality of data were collected. The identification unit processes the data for each of the plurality of data modalities, creates one or more labels occurring at times on the common timeline, and identifies events indicated by each of the processed plurality of data. The output unit processes the labels for each identified event based on the relative occurrence times of the events, and outputs information indicating additional medical events.
[0007] Some embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0008] [Figure 1A] FIG. 1A is a diagram illustrating an example of a medical information processing apparatus according to an embodiment. [Figure 1B] FIG. 1B illustrates a further example of an exemplary medical interrogation device that may be used to train one or more machine learning models. [Figure 2]FIG. 2 is a diagram illustrating an example of a system including a medical image processing device that communicates with multiple devices for collecting data belonging to different modalities. [Figure 3] FIG. 3 shows an example of a number of modules, each processing data belonging to a different channel in order to derive sub-event labels. [Figure 4] FIG. 4 is a diagram showing an example of components of a module for processing a data stream belonging to a certain channel. [Figure 5] FIG. 5 shows an example of using data belonging to three different channels to derive sub-event labels and event labels. [Figure 6] FIG. 6 shows some examples of events that can be identified based on data belonging to different channels. [Figure 7] FIG. 7 is a diagram illustrating an example of a neural network according to an embodiment. [Figure 8A] FIG. 8A illustrates an example of a first portion of an exemplary convolutional neural network for classifying frames of video. [Figure 8B] FIG. 8B illustrates an example of a second portion of an exemplary convolutional neural network for classifying frames of video. [Figure 9A] FIG. 9A is a diagram illustrating an example of a recurrent neural network (RNN). [Figure 9B] FIG. 9B illustrates a further example of a recursive neural network (RNN). [Figure 10] FIG. 10 shows an example of training a 3D convolutional neural network and an example of training a recurrent neural network. [Figure 11] FIG. 11 illustrates an example of training a machine learning model to classify events based on several sub-event labels. [Figure 12]FIG. 12 shows an example of a portion of a training dataset that includes several event labels and associated sub-event labels. [Figure 13] FIG. 13 is a diagram illustrating an example of a method according to an embodiment. [Figure 14] FIG. 14 illustrates an example of content that may be displayed on the device's user interface during or following a clinical procedure. DETAILED DESCRIPTION OF THE INVENTION
[0009] During a clinical procedure, data belonging to one or more modalities may be available to identify the occurrence of specific events occurring as part of the procedure. Two common data modalities include, for example, video and audio, but additional data modalities that may be available depend on the type of procedure. For example, types of data that may be available during a medical procedure in addition to video and audio include imaging data resulting from fluoroscopy, imaging data resulting from other scanning modalities, data obtained from radiofrequency tags, heart rate data, blood pressure data, temperature data, etc.
[0010] One embodiment provides a medical image processing device including a processing circuit configured to: acquire a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, the plurality of data modalities including an additional data type different from imaging data, the plurality of data being provided in association with a common timeline common to times at which the plurality of data were collected; process the data for each of the plurality of data modalities, generate one or more labels occurring at times on the common timeline, identify a plurality of events indicated by each of the processed plurality of data; process the labels for each of the identified events based on relative occurrence times of the plurality of events; and output information indicating a further medical event.
[0011] Here, "multiple data belonging to multiple data modalities" means that each of the multiple data belongs to one of the multiple data modalities. Also, "processing a label for each identified event based on the relative occurrence times of multiple events and outputting information indicating a further medical event" means outputting an event label, i.e., a medical event, based on the relative occurrence times of multiple sub-events.
[0012] By monitoring and processing data belonging to different modalities, sub-events in the data of some of these modalities may be identified. Labels of these sub-events may then be used to derive higher-level event classifications based on the time at which the labeled events occurred. As a result, higher-level event identification may be performed with greater accuracy. The higher-level classifications may then be tagged and added to the overall timeline. This allows for automated reporting and analysis of the procedure workflow. This may be done in real time during the procedure or after the procedure to generate a labeled summary of the past procedure. The identified higher-level events are presented as salient moments in the summarized procedure.
[0013] A common timeline is referenced to provide data belonging to different data modalities. In other words, each item of data (e.g., a frame of video, a segment of audio data, an RF tag measurement) may be associated with a specific time during a clinical procedure. The medical information processor receives this time information, indicating the specific time of each item of data, along with the data itself. The time information may be used to derive timestamps associated with sub-events. The timestamps may then be used to derive time information for higher-level events, which may be placed at appropriate points on the timeline.
[0014] According to an embodiment, a method is provided that includes receiving data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type and another of the plurality of data modalities being an additional data type other than imaging data, the data being provided with reference to a common timeline spanning a period of time over which the data is collected; processing the data for each of the plurality of data modalities to create labels that identify events occurring at times on the common timeline and that are represented by the processed data; and processing the labels for each of the identified events based on the time of occurrence of each of the plurality of events to obtain an output indicative of further medical events.
[0015] According to one embodiment, a computer program product is provided, the computer program product comprising computer-readable instructions that, when executed by at least one processor, cause the at least one processor to perform a method comprising: receiving data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type and another of the plurality of data modalities being an additional data type other than imaging data, the data being provided with reference to a common timeline spanning a period of time over which the data was collected; processing the data for each of the plurality of data modalities to create labels identifying events occurring in time on the common timeline and represented by the processed data; and processing the labels for each of the identified events based on the relative occurrence times of each of the plurality of events to obtain an output indicative of further medical events. According to one embodiment, a non-transitory computer-readable medium is provided having the computer program product stored thereon.
[0016] The embodiments will be described in more detail with reference to the accompanying drawings.
[0017] 1A, which shows a medical information processing device 100 in the form of a computing device. The processing performed to implement a method for determining the occurrence of an event occurring during a clinical procedure is performed by the medical information processing device 100. The device 100 may be a mobile user equipment (UE), a personal computer (PC), a terminal or workstation, a server, or other form of device.
[0018] The device 100 comprises an interface 140 for transmitting and receiving signals. The interface 140 may be a wired or wireless interface. For example, the interface 140 may comprise a wired interface for connecting to a wired network (e.g., a local area network and / or the Internet). Alternatively, or in addition, the interface 140 may comprise a transceiver device configured to transmit and receive communications via a wireless interface. The transceiver device may be implemented, for example, via a radio portion and an associated antenna arrangement. The antenna arrangement may be located internal or external to the device 100.
[0019] The device 100 comprises at least one data processing entity 115, at least one RAM (Random Access Memory) 120, at least one ROM (Read Only Memory) 125, and possibly other components 130 used and designed to perform software- and hardware-assisted tasks, including control, access, and communication with access systems and other communication devices. The at least one RAM 120 and hard drive 125 communicate with the data processing entity 115, which may be a data processor. Medical information processing devices, storage devices, and other associated controls may be provided on a suitable circuit board and / or chipset. A user controls the operation of the device 100 via a suitable user interface, such as a keypad 110, or by voice commands. A display 105 is included on the device 100 for displaying visual content to the user. The device 100 may also include a speaker for providing audio content.
[0020] The memory of device 100 (i.e., random access memory 120 and hard drive 125) may store computer-readable instructions for data processor 115 to perform the data processing functions described herein to be performed by device 100. Alternatively, component 130 may comprise hardware components, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), to perform the operations described herein to be performed by device 100. In some embodiments, the operations described herein to be performed by device 100 may be performed by a combination of hardware components or by a processor executing computer-readable instructions.
[0021] Although device 100 is shown as a single unified device 100, in other embodiments device 100 may include multiple interconnected devices.
[0022] The device 100 may receive data according to a number of different modalities, each of which may be referred to as a different data channel, and the data for each modality may be received from a different data acquisition device.
[0023] 1B, which illustrates an exemplary computing device 150 that may be used to train machine learning models that may be used to perform some of the processing described herein. Device 150 is shown as a single enclosed device. However, in some embodiments, device 150 may be a distributed system having multiple medical information processing devices operative to communicate with each other. Device 150 may comprise a server, a back-end system, etc.
[0024] The device 150 includes at least one random access memory (RAM) 160, at least one hard drive 170, at least one data processing unit 180, 190, and an input / output interface 195. The memory 160, 170 stores data for input to one or more models and data for storing results of processing performed during execution of the one or more models. The memory 160, 170 stores training data applied to train machine learning models. The memory 160, 170 additionally stores computer-executable code that, when executed by the at least one data processing unit 180, 190, provides one or more machine learning models. At least one of the data processing units 180, 190 performs one or more of the processing associated with the one or more models, training the models, and any necessary preprocessing of data used by the models. Via the interface 195, the device 150 receives data items for constructing a training dataset and / or data items for constructing an operating dataset. Device 150 additionally transmits via interface 195 the results generated from running the model on the input data.
[0025] 2, which illustrates an example of a system 200 comprising an apparatus 100 for receiving and processing data according to different modalities, and a plurality of exemplary data collection devices 220a-e, which may communicate with the apparatus 100 via a network 210.
[0026] The exemplary data collection devices 220a-e may include a camera 220a for obtaining images of a clinical procedure. In this embodiment, the image refers to a still image (camera image) and a moving image (video). The camera 220a may be a camera mounted on the ceiling of a clinic or the like to record images within the clinic. In this case, the image obtained by the camera 220a may capture a patient, medical equipment, and / or medical staff members during a clinical procedure. Specifically, the camera 220a may capture the hands of a medical staff member (operator) when the medical staff member is operating a certain medical equipment. For example, the medical equipment may be a contrast agent injection device, an anesthesia machine, a guidewire, an indwelling device (MitraClip, a medical device (e.g., a stent or balloon) used in percutaneous coronary intervention (PCI), a pressure sensor used in fractional flow reserve (FFR) assessment, an artificial valve used in transcatheter aortic valve implantation (TAVI), etc.). The camera 220a may also capture images of a medical staff member (doctor, nurse, etc.) or a patient entering an operating room. The camera 220a may also be part of an endoscope used to obtain video data of the patient's interior. The video data obtained by the camera 220a is transmitted to the apparatus 100. The exemplary data collection devices 220a-e may include a microphone 220b for collecting and acquiring audio data during a clinical procedure. The exemplary data collection devices 220a-e may include an RF detector 220c for detecting the presence of an RF tag 230 in close proximity to the RF detector 220c. Such an RF tag 230 may be attached to a piece of medical equipment and scanned by the detector 220c before a medical staff member uses the medical equipment. The data collection devices 220a-e may include an X-ray detector 220d for obtaining X-ray data of the patient P during a clinical procedure. The X-ray detector 220d may be part of an X-ray imaging device 240, which also includes an X-ray tube 12 for generating X-rays and a collimator 13 for limiting the field of view of the X-ray tube 12. The X-ray imaging device 240 may further include a filter 14 for filtering the X-ray beam output by the X-ray tube 12.
[0027] The data collection devices 220a-e may include a barcode scanner 220e that is used to detect and read barcodes that may be attached to a piece of medical equipment and scanned by the barcode scanner 220e before a medical staff member uses the medical equipment.
[0028] Thus, each of the exemplary data collection devices 220a-e collects data belonging to a different data modality and provides this data to the device 100. In this manner, the device 100 receives data belonging to multiple channels. Each of the data collection devices 220a-e collects data synchronously. In other words, the data collection devices 220a-e collect the data with reference to a common timeline of a clinical procedure, and each item of collected data (e.g., a frame of video) is labeled with time information indicating its position on the timeline. The data collection devices 220a-e provide the time information to the device 100 along with the data belonging to the multiple channels. Each of the devices 220a-e may operate according to a common system clock that is used to provide time information associated with the data recorded throughout the procedure.
[0029] The exemplary data acquisition devices 220a-e are provided by way of example only, and other types of data acquisition devices may be part of the system 200 and provide data of a given data modality to the device 100. For example, other types of data acquisition devices may acquire other types of imaging data, such as Positron Emission Tomography (PET) data or Magnetic Resonance Imaging (MRI) data, instead of or in addition to fluoroscopy. At least one of the data channels includes imaging data, which may take the form of video data acquired by the camera 220a or medical imaging data acquired using an imaging technique, such as fluoroscopy, DA acquisition, CT data acquisition, or MR data acquisition. Also, for example, other types of data acquisition devices may acquire ultrasound image data via ultrasound scanning. Note that ultrasound scanning can be considered an example of an imaging technique. In this case, ultrasound image data is considered an example of medical imaging data. Other types of data collection devices may also collect vital data (e.g., electrocardiogram waveform, heart rate, respiratory waveform, respiratory rate, blood pressure, body temperature, transcutaneous arterial oxygen saturation (SpO2), etc.).
[0030] The device 100 provides multiple modules for processing data belonging to different channels received by the device 100 and extracting labels indicating events occurring in the data belonging to the different channels from the data belonging to the different channels. Each of these modules is referred to herein as a “watcher.” A watcher is an example of an acquirer, an identifyr, an outputter, and a determiner. Such a watcher may be a software module running on the processor 115 of the device 100 or may be implemented in hardware. A watcher identifies basic events, such as the type of anatomical structure in a medical image or the presence of people in a room where a procedure is being performed, or a given type of interaction with medical equipment. Events identified by a watcher are referred to herein as subevents to distinguish them from higher-level events identified based on the labels output by the watcher. A model, which may be a machine learning model or a state machine, uses only the labels and their position in the timeline to perform higher-level classification and understand what is happening in the room.
[0031] See FIG. 3, which illustrates multiple watchers 310a-f, each associated with a different data channel. Each of the watchers 310a-f may receive data belonging to the particular data channel with which it is associated and identify sub-events represented by the received data. Upon identifying a sub-event, each watcher 310a-f outputs a label representing the sub-event. Each such label takes the form of one or more numeric values or a string identifying the type of sub-event and is suitable for input into a model (e.g., a machine learning model or a traditional state machine) to identify the event represented by some of the sub-events. In addition to outputting a label representing the detected sub-event, each watcher 310a-f also outputs a timestamp indicating the time the sub-event occurred. The device 100 may use the timestamp to determine the time of the identified event based on the label.
[0032] In this embodiment, the watchers 310a to 310f may output labels indicating sub-events detected in a rule base (such as a data table) instead of a model.
[0033] The Watcher 310a receives video data. The Watcher 310a may identify sub-events in the video data. A sub-event may be, for example, a specific action of a medical staff member (e.g., preparing a syringe) or a specific structure recognized in the endoscopic video. Based on the labeled sub-events identified in the video data, the Watcher 310a outputs one or more numeric values or strings representing the sub-event labels.
[0034] The watcher 310b receives the audio data. The watcher 310b may identify sub-events in the audio data. The sub-events may include specific spoken words or phrases, or spoken phrases with specific semantic meanings (e.g., phrases such as "insert catheter," "start contrast injection," or "inject anesthetic"). The watcher 310b may employ automatic language recognition to identify words belonging to the audio data. The watcher 310b may additionally apply natural language understanding to determine the semantic meanings of the identified words. The watcher 310b may match the identified words and / or the determined semantic meanings to a specific set of identified words and / or to identify a specific label.
[0035] Watcher 310c receives as input data indicating whether radio frequency (RF) waves were detected due to the presence of an RF tag in close proximity to the detection device, and outputs a label indicating that an RF tag was detected and a timestamp indicating the time of detection.
[0036] The Watcher 310d receives medical imaging data as input. The medical imaging data may include fluoroscopy data. The Watcher 310d may identify sub-events within the medical imaging data. For example, the medical imaging data may be used to identify a region to be treated (e.g., heart, head, abdomen, etc.). Furthermore, for example, the medical imaging data may be used to measure the position of a guidewire or an indwelling device inserted into a patient. In this case, a sub-event may be, for example, the arrival of the guidewire or the indwelling device at a given position within the patient. Based on a labeled sub-event identified in the medical imaging data, the Watcher 310d outputs a label representing the sub-event. The Watcher 310d additionally outputs a timestamp indicating the time at which the sub-event occurred.
[0037] In addition to outputting the label of the sub-event, the watcher 310d may also cause the display 105 to display medical image data (CT data, MR data) captured before surgery as a reference image.
[0038] The Watcher 310d may also receive status information indicating the status of the medical device. For example, if the medical device is an X-ray imaging device 240, the status information of the medical device may include the angle of the C-arm, the position of the X-ray detector 220d, position information of the X-ray irradiation field (for example, identified from the mechanical position of the aperture blades), and position information of the bed. The Watcher 310d may identify a sub-event in the medical imaging data from the status information of the medical device. The Watcher 310d may also identify a sub-event indicating that a specific part of a patient is being imaged based on the collected medical imaging data.
[0039] The Watcher 310e receives ultrasound image data as input. The Watcher 310e may identify sub-events within the ultrasound image data. For example, the ultrasound images may be used to determine the position of a medical device, such as a catheter, inserted into a patient. In this case, a sub-event may be, for example, the arrival of the catheter or the like at a given location within the patient. The Watcher 310e outputs a label representing the sub-event based on the labeled sub-event identified in the ultrasound image data. The Watcher 310e additionally outputs a timestamp indicating the time the sub-event occurred.
[0040] The watcher 310e may cause the display 105 to display ultrasound image data collected before surgery as a reference image, in addition to outputting the label of the sub-event.
[0041] The Watcher 310f receives vital data as input. The vital data may include, for example, an electrocardiogram waveform, a heart rate, a respiratory waveform, a respiratory rate, a blood pressure, a body temperature, and SpO2. The Watcher 310f may identify sub-events in the vital data. Based on the labeled sub-events identified in the vital data, the Watcher 310f outputs one or more numerical values or character strings representing the labels of the sub-events.
[0042] The watcher 310f may display the vital data on the display 105 together with the output of the label of the sub-event.
[0043] The Watcher 310 may also receive status information indicating the status of the medical device. For example, if the medical device is the X-ray imaging device 240, the status information of the medical device may include the angle of the C-arm, the position of the X-ray detector 220d, position information of the X-ray irradiation field (for example, identified from the mechanical position of the aperture blades), and position information of the bed. For example, if the medical device is a contrast agent injection device (injector) that automatically injects a contrast agent, the status information may include information indicating the injection start timing of the contrast agent (for example, identified by receiving an operation (button operation) to instruct the start of injection). For example, if the medical device is an anesthesia machine, the status information may include information indicating the supply start timing of an anesthetic agent (for example, identified by receiving an operation (button operation) to instruct the start of supply). The Watcher 310 may identify a sub-event from the above information.
[0044] Furthermore, the watcher 310 may identify a sub-event indicating the type of X-ray imaging (fluoroscopic imaging, contrast imaging, CAG imaging, etc.) from the various pieces of information received.
[0045] See Figure 4, which illustrates examples of components that may belong to a watcher 310. The watcher 310 may be any of the watchers 310a-f shown in Figure 3.
[0046] The Watcher 310 includes a noise gate module 400. The noise gate 400 monitors the data stream received on the data channel with which the Watcher 310 is associated. The noise gate 400 identifies when changes occur in the data on that channel that may indicate a sub-event. The Watcher 310 also includes a labeling module 410. The labeling module 410 is applied when the noise gate 400 identifies activity in the data on the channel. The labeling module 410 receives the data, and when the noise gate 400 is triggered (i.e., when it identifies activity in the data), it identifies and performs classification of the data to determine whether a sub-event belonging to one of a predetermined set of sub-events defined for the Watcher 310 has occurred. If the labeling module 410 determines that one of these sub-events has occurred, the labeling module 410 outputs a label indicating the sub-event.
[0047] For example, the watcher 310 may be a watcher 310b used to receive audio data. In this case, a noise gate 400 belonging to the watcher 310 may receive an audio data stream and identify changes in the volume of the audio data, which may indicate, for example, speech. Upon identifying points in the audio data where speech occurs, the noise gate 400 provides this indication to the labeling module 410. The labeling module 410 then applies a language recognition model to identify words spoken at those points in the audio data. The labeling module 410 may further apply natural language understanding to the identified words in the audio data to determine their semantic meaning and identify labels based on the identified semantic meaning. As a further example, the noise gate 400 may detect movement in a video data stream and then trigger the labeling module 410 to identify sub-events occurring in the video, such as the presence of one or more people.
[0048] The sub-event labels derived from multiple watchers, and the relative order and timing of the sub-events, are used to identify the event that the multiple sub-events represent. For example, the sub-event labels may be fed into a machine learning model (such as a recurrent neural network) in the order in which they were generated to derive the appropriate event label.
[0049] See FIG. 5, which shows an example of a timeline covering a period during which multiple sub-events are detected in different data channels.
[0050] At 505, the presence of an RF tag is detected. This RF tag is associated with a particular item (item 1) and may be attached to the item's exterior, packaging, etc. In an embodiment, when the item is removed for use by a medical staff member, the tag associated with the item is scanned to indicate that the item has been removed. In this case, detection device 220c sends a signal to device 100 indicating that the RF tag associated with item 1 has been scanned. Device 100 identifies the scan of item 1 as a sub-event and retrieves a label identifying this sub-event along with a timestamp for the sub-event.
[0051] In addition, 505 may detect that a code symbol, such as a one-dimensional code or a two-dimensional code, has been scanned. In this case, the code symbol may be associated with a particular item (item 1) and may be attached to the exterior, packaging, or the like of the item. For example, when the item is removed for use by a medical staff member, the code symbol associated with the item may be scanned by a reading device to indicate that the item has been removed. In this case, the reading device may send a signal to device 100 indicating that the code symbol associated with item 1 has been scanned. Device 100 may identify the scanning of item 1 as a sub-event and retrieve a label identifying this sub-event along with a timestamp for the sub-event.
[0052] At 510, video is recorded on a video data channel of a nurse preparing an injection for injection into a patient. Camera 220a provides a video stream to device 100, which (using watcher 310a) analyzes the video stream to identify a point in the video stream where the nurse prepares the injection. The device may analyze the video footage by applying one or more convolutional neural networks to the video to identify a portion of the video that depicts a person preparing an injection. Upon identifying the portion of the video, device 100 derives a label representing a sub-event (i.e., preparing the injection) and a timestamp representing the time when the sub-event occurred.
[0053] At 515, the radiography data is recorded on the fluoroscopy data channel. Fluoroscopy is an example of an x-ray imaging technique for obtaining an x-ray image stream, although x-ray imaging equipment 240 may operate in other ways to obtain one or more x-ray images. X-ray imaging equipment 240 provides the x-ray imaging data to apparatus 100, which (using watcher 310d) analyzes the imaging data stream to identify points in the imaging data where contrast agent is detected. Upon identifying that portion of the imaging data, apparatus 100 derives a label representing the sub-event and a timestamp representing the time the sub-event occurred.
[0054] Here, the X-ray imaging data includes imaging images, which are high-dose X-ray images, and fluoroscopic images, which are low-dose X-ray images.
[0055] At 520, audio is recorded on the audio data channel, where the words "start injection" are present. Microphone 220b provides the audio data stream to device 100, which (using watcher 310b) analyzes the audio data stream to identify the point in the audio data stream where the words are spoken. The spoken words are analyzed by applying a language recognition algorithm to identify the words. The words may be further analyzed to determine their semantic meaning. Upon determining that the words indicate a command to start or perform an injection, device 100 obtains a label indicating the sub-event where the command to start or perform an injection was spoken. Device 100 further derives a timestamp indicating the time at which the sub-event occurred.
[0056] In this example, device 100 thus obtains various labels that represent different sub-events that occurred on different data channels. Each of these labels is provided in the form of one or more numeric values or text strings suitable for input to model 520. To obtain the labels of the events that they represent, the device provides the sub-event labels as input to model 520. For example, given the four sub-events 505, 510, 515, and 520 shown in FIG. 5 and the corresponding labels derived by watchers 310a-d, model 520 may output a set of values that represent the events of preparing and injecting an injection.
[0057] To ensure that the relative timing of sub-events is taken into account when deriving event labels, the watcher processes each data stream (i.e., data belonging to different modalities) in real time and generates sub-event labels as they occur. The apparatus 100 may be configured to input each sub-event label as it is generated. In this manner, the time at which the sub-event label is input to the model 520 corresponds to the time at which the sub-event occurs in the common timeline. In the example of FIG. 5, when the first sub-event 505 occurs at approximately 2:45 in the timeline, the watcher 310c extracts the appropriate sub-event label, which the apparatus 100 provides as an input to the model 520, thereby updating the model's state. Subsequently, when the second sub-event 510 occurs at approximately 5:55 in the timeline, the watcher 310a extracts the appropriate sub-event label, which the apparatus 100 provides as an input to the model 520, thereby again updating the model's state. Subsequently, when a third sub-event 515 occurs at approximately 7:20 in the timeline, watcher 310d retrieves the appropriate sub-event label, which apparatus 100 provides as an input to model 520, again updating the state of model 520. Subsequently, when a fourth sub-event 520 occurs at approximately 9:05 in the timeline, watcher 310b retrieves the appropriate sub-event label, which apparatus 100 provides as an input to model 520, again updating the state of model 520. As the state of model 520 is updated multiple times with the sub-event label inputs, the output of model 520 represents the event labels of the events represented by the multiple sub-events. In this example, the represented event may be digital angiography (DA) preparation start 620, an example of which is shown in FIG. 6.
[0058] By considering the relative timing of the sub-events, it may be possible to ignore sub-events that occurred long ago and are therefore unrelated to more recent sub-events in identifying the event label.
[0059] Model 520 may take the form of a recurrent neural network (RNN) that remembers state with respect to past sub-events. Such a recurrent neural network may update its state based on the relative timing of the sub-events. For example, the network may "forget" sub-events that occurred very early in the common timeline.
[0060] Model 520 may be a traditional state machine that stores state for past sub-events and updates that state in response to further sub-event labels to derive one or more event labels.
[0061] In some embodiments, the relative timing of sub-events may be taken into account by processing the timestamps generated for the sub-events to derive the event labels. These timestamps, along with the sub-event labels, may be provided as inputs to model 520.
[0062] The different events for which the model 520 outputs labels may include clinical events, stages of a procedure, and / or staff actions. These labels may be used to index and summarize the procedure. See FIG. 6, which shows an example of different event labels that may be output during a procedure. FIG. 6 illustrates how these labels are output at different points during the procedure. Each of these labels is output based on one or more different sub-events.
[0063] 6 , a first label 610 of an event output by model 520 is that all staff for a procedure are present. Model 520 may output this first label 610 in response to receiving as input a label indicating that a staff member has entered the room where the procedure will be performed. Such a label of a sub-event may be output by watcher 310 based on video footage. Additionally, model 520 may output first label 610 in response to receiving as input a label indicating that a sub-event has been detected in audio data received by device 100. Such sub-events detected in the audio data may include detecting utterances by different staff members, the utterance of a given staff member's name, or the utterance of a phrase with a given semantic meaning (e.g., "everyone is here").
[0064] 6, a second label 620 of the event output by model 520 is preparation for obtaining a digital angiography (DA) image. This label 620 may be derived by device 100 based on a label indicating a sub-event in which an RF tag attached to equipment for taking images is scanned by a detector, a label indicating a sub-event in which video data indicating preparation for an injection is recorded by microphone 220b, and a label indicating a sub-event in which audio related to the start of an injection is recorded by microphone 220b.
[0065] 6, a third label 630 of the event output by model 520 is the end of DA image acquisition. This label 630 may be output by model 520 based on a label representing a sub-event of reaching the end of a series of DA images. This sub-event may be detected based on a DA imaging data stream received by an apparatus from an apparatus for performing DA imaging. Label 630 may additionally be output based on a sub-event detected in audio data, for example, a label representing a spoken phrase indicating the end of a DA imaging procedure.
[0066] As a further example, a label indicating DA image acquisition may be output by model 520 based on a first label representing a sub-event of detecting contrast agent diffusion in the vasculature in a clinical image, a second label representing a sub-event of a clinician requesting contrast agent injection, and a third label representing a sub-event of a barcode of new vitals for contrast agent being scanned.
[0067] Although DA scans are an exemplary scan type for outputting event labels at 620 and 630, it will be appreciated by those skilled in the art that other scan types or ultrasound scans may be used to output event labels at 620 and 630.
[0068] 6, a fourth label 640 of the event output by model 520 is the insertion of a guidewire into a patient. Model 520 may output this label 640 based on a label representing the sub-event of fluoroscopy imaging starting, where the label is derived by device 100 based on the fluoroscopy imaging data it receives from device 220d. Label 640 may be output depending on labels representing one or more sub-events detected in the audio data related to the guidewire.
[0069] A fifth label 650 of the event output by model 520 is navigating the guidewire to the target. Model 520 may output this fifth label 650 in response to receiving a label representing the sub-event of the guidewire being in motion, where the label is derived from a fluoroscopy image in which the guidewire is shown. Additional labels of the sub-event determined from additional data channels may also be provided as inputs to model 520 to derive label 650.
[0070] A sixth label 660 of the event output by model 520 is the guidewire reaching the target. This label 660 may be output by model 520 in response to model 520 receiving a label representing the sub-event of the guidewire ceasing to move, where the label is derived from a fluoroscopy image in which the guidewire is shown. Additional labels of the sub-event determined from additional data channels may also be provided as inputs to model 520 to derive label 660.
[0071] As shown in the example of Figure 6, a location on the common timeline is assigned to each event label. The time of each event label may be derived by apparatus 100 from the timestamps of the sub-event labels used to derive each event label. For example, for a particular event label, the most recent (or slightly later) timestamp of the sub-events used to derive the event label may be assigned to that event label. In the example of Figures 5 and 6, the time assigned to event label 620 is slightly after the timestamp of sub-event 520.
[0072] In the example discussed above, model 520 may comprise one or more neural networks that receive sub-event labels as inputs and provide corresponding event labels as outputs, and one or more of watchers 310 may each comprise one or more neural networks that receive sub-event labels as inputs and provide corresponding event labels as outputs.
[0073] FIG. 7 is a schematic diagram of a neural network 700. Neural network 700 includes input nodes 710, hidden nodes 720, and output nodes 730. In practice, network 700 will likely have many more nodes and hidden layers than shown. Each input node 710 receives a single value of input data and generates at its output an activation or node value, which is generated by feeding the input value into an activation function (e.g., a sigmoid). Each input node 710 is connected to each hidden node 720. A matrix of weights defines the connectivity between the input node 710 and the hidden nodes 720. The vector of node values output from input node 710 is scaled by a respective vector of weights at the input of each hidden node 720, with each weight defining the connectivity of one of the input nodes 710 to one of the hidden nodes 720. In FIG. 7, the weights applied at the input of one of the hidden nodes 720 are shown as w0...w3. At each hidden node 720, the input value at that node is given as the dot product of the associated weight vector and the output value of input node 710. An activation function is then applied to the input value at hidden node 720 to provide the output value for that node 720. The output vector of hidden node 720 is fed to each of nodes 730 in the next layer of network 700, which generate output values for the next layer in a similar manner.
[0074] Network 700 may be trained using supervised or unsupervised learning. In one embodiment, network 700 is trained using supervised learning by determining at least one set of output values based on at least one set of input values included in training data. The output values are compared to existing labels in the training data, and an error or loss (i.e., based on the difference between the output value and the label) is calculated. The error or loss is then backpropagated through network 700, and the weights are updated to train network 700 to obtain a better approximation of the labels from the input values. In the next cycle, the revised weights are used with additional training data to further update the weights and more closely generate labels for the additional training data based on the input values for the additional training data. In this manner, network 700 can be trained to perform a specific task.
[0075] Convolutional neural networks may be used when performing video or image classification. A convolutional neural network is a neural network that utilizes convolutional computation in at least one of its layers. Convolutional neural networks are particularly well suited for image analysis and processing because they are shift invariant. For feature recognition in videos or a series of medical images, a 2D convolutional neural network (CNN) may be applied to identify features in individual frames within a video or a series of medical images. Alternatively, for feature recognition in videos, a 3D convolutional neural network (CNN) may be applied to identify features within the video, including determining the temporal relationships between frames.
[0076] Instead of or in addition to using a CNN, traditional image analysis algorithms may be used to perform video or image classification.
[0077] 8A and 8B, which illustrate an example of the operation of a convolutional neural network that may be used to identify given features in a frame of video and perform classification of those features. In the illustrated example, the input image is an X-ray image 805 showing multiple implants being inserted into a patient. The convolutional neural network may be used to identify when each implant is in its final location during a surgical procedure.
[0078] A kernel 810 is applied to determine the convolution of the input image 805 with the kernel 810. The output of this convolution is multiplied by an activation function and added nonlinearly. The activation function used in Figure 8A is a rectified linear activation unit (RELU), which outputs the input if it is positive and zero if it is not positive. Multiple feature maps are generated from the input image by performing convolutions between the input image and different kernels. Each kernel represents a different basic feature, such as a vertical or horizontal line.
[0079] Each feature map generated by the convolution and activation function then undergoes a pooling process to reduce the spatial size of the convolved features. Pooling involves translating a kernel across the feature map to sample pixels and returning the maximum or average value from each sampled pixel in the feature map. The resulting pooled feature maps are then further convolved (with a RELU function applied) using different kernels to generate a further set of feature maps, from which pooling is again performed.
[0080] As shown in FIG. 8B, the pooled feature maps resulting from multiple stages of convolution and pooling are flattened to generate a one-dimensional array (denoted as a flattening layer). The one-dimensional array is provided as a set of input values to a feedforward neural network. The resulting output values represent the state of the implant in the x-ray image. The apparatus 100 may process the output values to infer whether the implant is in its final location within the patient.
[0081] A convolutional neural network may be trained by comparing the output values of different images with the labels of those images and adjusting the weights in the forward propagation portion of the convolutional neural network.
[0082] As described above, model 520 receives as input the different labels of the sub-events in sequence to derive an output representing the event. To ensure that the output of model 520 depends on the relative timing of the sub-events, in some embodiments, model 520 may include at least one recurrent neural network to process the input.
[0083] See Figure 9A, which shows a simple example of a recurrent neural network (RNN) with input nodes 910, hidden layer nodes 920, and output nodes 930. Multiple sets of input data are provided as inputs to the RNN at different times. For the first iteration of the RNN, a first set of input data X t as input at time t, and for the second iteration of the RNN, we use a second set of input data X t+1 as input at a subsequent time t+1, and a third set of input data X t+2 is given as input at a further subsequent time t+2. In this simple example, each set of input data contains only a single value.
[0084] To calculate the activation values at hidden layer nodes 920 during the first iteration of the RNN, the input values X tby a weight W1. A bias b1 is added to the result of this multiplication, and the result of the addition is provided to an activation function (ReLU). The output of the activation function (e.g., a rectified linear unit (ReLU) function) provides the activity for hidden layer node 920. The activity of hidden layer node 920 may be further processed (i.e., multiplied by a weight W3 and added with a bias b2) to produce the activity for output node 930. However, the activity of output node 930 may be ignored and not calculated until the entire set of input values has been processed.
[0085] The activity of node 920 determined during the first iteration constitutes the hidden state, which is the next input value X t+1 is used when processing the next input value X as part of the second iteration of the RNN. t+1 When processed by node 910, it is multiplied by weight W1. The result of this multiplication is added to the hidden state of node 920 determined in the first iteration (shown as "Sum" in Figure 9). The result of this summation is then added to bias b1 and fed into an activation function to determine the activity of hidden layer node 920 for the second iteration. The activity of hidden layer node 920 calculated for the second iteration is then multiplied by the third input value X t+2 to calculate the activations of hidden layer nodes 920 for the third iteration.
[0086] The processing of multiple sets of input values may continue in a similar manner as described above for each ongoing hidden state of the RNN being used for the current iteration, until the final set of inputs has been processed in the final iteration. The activity of output node 930 for the final iteration provides the output of the RNN.
[0087] Figure 9A shows a simplified example RNN with only three nodes 910, 920, and 930. Figure 9B shows a further example where the RNN contains two nodes in each layer. In this further example, the RNN contains two nodes 950 and 955 in the input layer, two nodes 960 and 965 in the hidden layer, and two nodes 970 and 975 in the output layer. In this case, the RNN contains multiple states that are used when processing the next set of input values. The first set of input values X 1,t,X 2,t are provided as inputs to the input layer nodes 950, 955 for processing in the first iteration. The activations of each hidden layer node 960, 965 calculated during this iteration are then calculated based on the second set of input values X 1,t+1 ,X 2,t+1 is used to calculate the activations of each of the hidden layer nodes 960, 965 during the second iteration in which
[0088] As mentioned above, several of the watchers 310 may be implemented using machine learning models that are used to provide labels for sub-events. To provide each of these models, a training process is performed by the device 150 to train the model using several sets of training data.
[0089] Refer to FIG. 10 , which illustrates how two exemplary machine learning models 1100, 1110 can be trained. The machine learning model 1100 is a 3D convolutional neural network for identifying subevents detected in a series of X-ray images / frames. To train the model 1100, multiple sets of X-ray frames, each labeled by a human user, are provided to the device 150. FIG. 10 illustrates a first set of frames labeled to indicate instances where the guidewire is moving, a second set of frames labeled to indicate instances where the guidewire is stationary, and a third set of frames labeled to indicate images without the guidewire. Each of these sets of frames is input by the device 150 into the 3D convolutional neural network 1100, which derives a set of outputs. For each set of frames, a comparison stage 1120 compares the outputs with the corresponding labels to determine an error / loss, which is subsequently used to update the parameters of the model 1100.
[0090] The machine learning model 1110 is a recurrent neural network 1110 used to process text strings to determine the semantic meaning of the text. To train the model 1110, the system uses multiple strings of text strings, each labeled by a human user. FIG. 10 shows a first text string 1140 labeled to indicate that a guidewire has reached a given location, a second text string 1150 labeled to indicate an instruction to end a current action, and a third text string 1160 labeled to indicate an instruction to begin an injection. Each of these text strings is input by the device 150 into the recurrent neural network 1110, which elicits a set of outputs. For each of the text strings 1140, 1150, and 1160, a comparison stage 1130 compares the output with the corresponding label to determine an error / loss, which is subsequently used to update the parameters of the model 1100.
[0091] The apparatus 150 may be used to train a machine learning model 520 that may be used to derive event labels from sub-event labels. See FIG. 11, which illustrates a process for training the machine learning model 520 using multiple sub-event labels and multiple event labels. The event labels are provided by a human user, while the sub-event labels may be provided by a human user or may be derived by applying the watcher 310 to different data channels as described above.
[0092] A set of sub-event labels is input by device 150 to machine learning model 520 to derive an output value. A comparison stage 1120 compares the output value with one or more numeric values representing the event labels provided by a human user to determine a loss / error, which is subsequently used to output parameters for model 520. This process is repeated with multiple sets of sub-event labels, each with a corresponding event label.
[0093] 12 illustrates a portion 1300 of an exemplary training dataset for training the machine learning model 520. The training data 1300 includes three sets of sub-event labels, each with a corresponding event label. The first set of sub-event labels includes a label indicating a first sub-event (denoted as Audio 1) identified in the audio data at time t1, a label indicating a second sub-event (denoted as RF Tag 1) identified by the RF detector at time t2, a label indicating a third sub-event (denoted as Video 1) identified in the video data at time t3, and a label indicating a fourth sub-event (denoted as Audio 2) identified in the audio data at time t3. Each of these sub-events is associated with an event assigned event label 1 and is considered by a human user to represent that type of event. The model 520 may be an RNN, with the first set of sub-event labels input in chronological order, starting with the earliest sub-events, and may derive output values from the RNN that are compared to event label 1 in the comparison stage 1120.
[0094] The training data 1300 also includes a second set of sub-event labels (including audio and fluoroscopic data) and corresponding event labels, and a third set of sub-event labels (including RF tag data, video data, and audio data) and corresponding event labels. The second set of sub-event labels, the third set of sub-event labels, and their corresponding event labels are processed to train the model 520 in the same manner as the first set of sub-event labels.
[0095] Reference is made to FIG. 13, which illustrates a method 1400 implemented in the device 100.
[0096] In S1410, the device 100 receives data of multiple data modalities collected during a clinical procedure.
[0097] In S1420, the apparatus 100 identifies sub-events in the data of at least some of these data modalities, and generates a label for each identified sub-event.
[0098] In S1430, apparatus 100 processes the labels obtained in S1420 depending on the relative timing to obtain an output indicative of a further medical event that is a higher level event than the sub-event identified in S1420. S1430 may include, for each label, providing the label as an input to model 520 upon generation of the respective label, and obtaining an output indicative of the further medical event.
[0099] Apparatus 100 may subsequently perform S1420 again to derive additional sub-event labels corresponding to subsequent times in the common timeline of the clinical procedure, and continue. Apparatus 100 then performs S1430 again to derive an output indicative of another event. In some embodiments, model 520 may be a recurrent neural network that updates its output value for each subsequent sub-event label input and classifies the set of labels as presented.
[0100] Processing to recognize events based on data from multiple channels may be performed on the recorded data online or offline, in other words, in real time, i.e., during the procedure as data becomes available, or after the procedure when the complete set of data belonging to the different channels has been collected. The categorized event sets can be used to generate procedure summaries and reports, and can be individually called up for reference and preparation.
[0101] See FIG. 14, which shows an example of content that may be displayed on the user interface (display 105) of the device 100 during or following a clinical procedure. The content includes a timeline with several labels of events 1500a-d. The content may be part of a procedure summary and report generated following the procedure. Alternatively, the content may be part of a real-time report generated during the procedure, in which case additional labels are added to the timeline in response to events detected based on further collected data.
[0102] Although the above describes an example in which a label is displayed on the display 105 as an output indicating a medical event, the output is not limited to a display. For example, the device 100 may output content corresponding to the label indicating the medical event to an analysis application. Specifically, in a scene in which a cerebral aneurysm is being treated, the output of a label indicating that a catheter has reached a given position in a patient's cerebral blood vessel and that contrast imaging has been performed may be used as a trigger to launch an analysis application that measures the neck size of the cerebral aneurysm through image analysis. In this case, the device 100 may output X-ray imaging data collected by contrast imaging to the analysis application. This allows medical staff to obtain the measurement results of the neck size of the cerebral aneurysm without performing operations such as launching an analysis application or inputting X-ray imaging data.
[0103] Also, for example, device 100 may derive an output indicating an event corresponding to a specific order from medical staff, such as "give me an image with a good view of the blood vessels," based on sub-event labels derived from various data received from various data collection devices.
[0104] Implementations of the subject matter and operations described herein can be realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures and structural equivalents disclosed herein, or in one or more combinations thereof. For example, hardware may include a processor, microprocessor, electronic circuit, electronic component, integrated circuit, etc. Implementations of the subject matter disclosed herein can be realized using one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium and executed by or controlling the operation of a medical information processing device. Alternatively, or additionally, the program instructions can be encoded in an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiving device for execution by the medical information processing device. The computer storage medium can be or include a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Alternatively, if the computer storage medium is not a propagated signal, the computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagated signal, and the computer storage medium may be or include one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0105] According to various embodiments, a method for event recognition is provided, comprising: a) synchronously receiving multiple data sources; b) labeling significant events on individual data channels; c) sending the labels to an event identification algorithm; and d) identifying and labeling the events using individual channel labels that accumulate over time, wherein at least one of the multiple data sources in a) is a medical imaging data source. In some further embodiments, the algorithm in d) is an RNN state machine using multiple states and a multi-head architecture. In some further embodiments, the algorithm in b) analyzes the data for discrete time quantities and triggers a significance detector for the signal. In some further embodiments, the events identified in d) are used to provide a procedure summary. In some further embodiments, the data in a) includes data sources used in medical imaging or interventional procedures. In some further embodiments, the data in a) includes data from one or more clinical imaging channels, video, sound, interactions with equipment, and scanning RF tags or barcodes on consumables. In some further embodiments, the method is applied to live streaming data, and in some further embodiments, the method is applied to recorded data.
[0106] According to one embodiment, there is provided a medical imaging device including a processing circuit configured to: receive data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type and another of the plurality of data modalities being an additional data type other than imaging data, the data being provided with reference to a common timeline spanning a period of time over which the data was collected; process the data for each of the plurality of data modalities to generate one or more labels, each label occurring in time on the common timeline and identifying an event indicated by the processed data; and process the label for each of the identified events based on the relative occurrence time of each of the plurality of events to obtain an output indicative of a further medical event.
[0107] According to an embodiment, processing the label for each of the identified events in each of the plurality of data modalities comprises providing the label as an input to a machine learning model to obtain the output indicative of the further medical event.
[0108] In one embodiment, for each of the plurality of data modalities, processing the data for each of the plurality of data modalities and creating the labels occurs in real time as the data is received, and providing the labels as input to the machine learning model includes, for each of the labels, providing the respective label as input to the machine learning model as the respective label is created.
[0109] According to one embodiment, the machine learning model is a recurrent neural network.
[0110] In one embodiment, the processing circuitry is configured to provide each of the labels as input to the machine learning model in the order in which the corresponding identified events occurred on the common timeline.
[0111] In one embodiment, the processing circuitry is configured to output, for each of the identified events, time information indicating a time at which each of the identified events occurred on the common timeline, and process the time information of the identified events to determine a time associated with the further medical event.
[0112] In one embodiment, the imaging data includes at least one of video data and medical imaging data.
[0113] In one embodiment, the multiple data modalities include one or more of video data, audio data, medical imaging data, and radio frequency tag data.
[0114] According to an embodiment, processing the data of each of the plurality of data modalities to identify, for one or more of the plurality of data modalities, each of the plurality of events occurring during the procedure includes subjecting the data of each of the plurality of data modalities to a further machine learning model to derive the label for each of the identified events.
[0115] In one embodiment, one or more of the plurality of data modalities includes the imaging data, and each of the further machine learning models used to derive the labels for each of the plurality of identified events for the imaging data includes a convolutional neural network.
[0116] In one embodiment, one or more of the plurality of data modalities includes audio data, and each of the further machine learning models used to derive the labels for each of the plurality of identified events for the audio data includes language recognition configured to derive text representing the audio data.
[0117] In one embodiment, the processing circuitry is configured to process the text using a natural language understanding model to derive the labels identifying each of the plurality of events.
[0118] According to one embodiment, the processing circuitry is configured to control a display to provide a visual indication of the further medical event and associated time information indicating when the further medical event occurred in the clinical procedure.
[0119] According to an embodiment, a method is provided that includes receiving data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type and another of the plurality of data modalities being an additional data type other than imaging data, the data being provided with reference to a common timeline spanning a period of time over which the data is collected; processing the data for each of the plurality of data modalities to create labels that identify events occurring in time on the common timeline and that are represented by the processed data; and processing the labels for each of the identified events based on the relative occurrence time of each of the plurality of events to obtain an output indicative of further medical events.
[0120] According to an embodiment, there is provided a computer program comprising computer-readable instructions that, when executed by at least one processor, cause the at least one processor to perform a method comprising: receiving data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type and another of the plurality of data modalities being an additional data type other than imaging data, the data being provided with reference to a common timeline spanning a period of time over which the data was collected; processing the data for each of the plurality of data modalities to create labels identifying events occurring in time on the common timeline and indicated by the processed data; and processing the labels for each of the identified events based on the relative occurrence times of each of the plurality of events to obtain an output indicative of further medical events.
[0121] Although several embodiments have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, modifications, and combinations of embodiments can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0122] 100 devices 12 X-ray tube 13 Collimator 14 Filters 105 Display 110 keypad 115 processors 120 RAM 125 ROM 160, 170 memory 180, 190 Data processing unit 195 Input / Output Interfaces 200 systems 220a~e Data collection equipment 220a camera 220b microphone 220c RF detector 220c Detection device 220d X-ray detector 220e Barcode Scanner 240 X-ray equipment 310, 310a, 310b, 310c, 310d, 310e, 310f Watcher Various aspects of the present disclosure are summarized below as appendices. [Appendix 1] an acquisition unit that acquires a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; an identification unit that processes data of each of the plurality of data modalities, creates one or more labels occurring at times on the common timeline, and identifies a plurality of events indicated by each of the plurality of processed data; an output unit that processes the labels for each of the identified events based on the relative occurrence times of the events and outputs information indicative of a further medical event; A medical information processing device comprising: Here, "multiple data belonging to multiple data modalities" means that each of the multiple data belongs to one of the multiple data modalities. Also, "processing a label for each identified event based on the relative occurrence times of multiple events and outputting information indicating a further medical event" means outputting an event label, i.e., a medical event, based on the relative occurrence times of multiple sub-events. [Appendix 2] the identification unit processes the label for each of the identified events in each of the plurality of data modalities by performing processing including providing the label as an input to a machine learning model to obtain information indicative of the further medical event. 2. A medical information processing device according to claim 1. [Appendix 3] When one of the plurality of data is acquired, the identification unit processes the data of each of the plurality of data modalities and creates the label in real time for each of the plurality of data modalities; Providing the labels as inputs to the machine learning model includes, for each of the labels, providing each of the labels as inputs to the machine learning model at the time of creation of each of the labels. 3. The medical information processing device according to claim 2. [Appendix 4] the machine learning model is a recurrent neural network; 4. The medical information processing device according to claim 2 or 3. [Appendix 5] the identification unit provides each of the labels as an input to the machine learning model in the order in which each of the corresponding identified events occurred on the common timeline; 4. The medical information processing device according to claim 2 or 3. [Appendix 6] the output unit outputs, for each of the identified events, time information indicating a time at which each of the identified events occurred on the common timeline; a determination unit that processes the time information of each of the identified events to determine a time associated with the further medical event. 4. The medical information processing device according to claim 2 or 3. [Appendix 7] The imaging data type includes at least one of video data, X-ray image data, and ultrasound image data. 4. The medical information processing device according to claim 2 or 3. [Appendix 8] the plurality of data modalities include at least one of video data, audio data, X-ray image data, ultrasound image data, radio frequency tag data, and vital data; 4. The medical information processing device according to claim 2 or 3. [Appendix 9] For one or more of the plurality of data modalities: processing the data from each of the plurality of data modalities to identify the plurality of events occurring during the procedure includes providing the data from each of the plurality of data modalities to a further machine learning model to derive the label for each of the plurality of identified events. 4. The medical information processing device according to claim 2 or 3. [Appendix 10] one or more of the plurality of data modalities includes the imaging data type; each of the additional machine learning models used to derive the label for each of the identified events for the imaging data type comprises a convolutional neural network; 10. The medical information processing device according to claim 9. [Appendix 11] one or more of the plurality of data modalities includes audio data; each of the further machine learning models used to derive the label for each of a plurality of identified events in the audio data includes language recognition configured to derive text representing the audio data; 11. The medical information processing device according to claim 9 or 10. [Appendix 12] the identifying unit processes the text using a natural language understanding model to derive the labels that identify the plurality of events. 12. The medical information processing device according to claim 11. [Appendix 13] Further comprising a display control unit that controls the display, the display controller controls the display to display visual information indicating the further medical event and associated time information indicating when the further medical event occurred in the clinical procedure. 4. The medical information processing device according to claim 2 or 3. [Appendix 14] further comprising an application execution unit that executes an application corresponding to the further medical event in response to output of the information indicating the further medical event. 4. The medical information processing device according to claim 2 or 3. [Appendix 15] A medical information processing method by a medical information processing device, comprising: Acquire a plurality of data collected during a clinical procedure, the plurality of data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the plurality of data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; For each of the plurality of data modalities, processing data for each of the plurality of data modalities to generate one or more labels occurring at times on the common timeline and identifying events indicated by each of the plurality of processed data; Each of the plurality of events processing the labels for each of the identified events based on the relative occurrence times of the events and outputting information indicative of a further medical event; Medical information processing method. [Appendix 16] On the computer, Acquire a plurality of data collected during a clinical procedure, the plurality of data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the plurality of data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; For each of the plurality of data modalities, processing data for each of the plurality of data modalities to generate one or more labels occurring at times on the common timeline and identifying events indicated by each of the plurality of processed data; processing the labels for each of the identified events based on the relative occurrence times of each of the identified events to output information indicative of a further medical event; A program that executes a process.
Claims
1. an acquisition unit that acquires a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; an identification unit that processes data of each of the plurality of data modalities, creates one or more labels occurring at times on the common timeline, and identifies a plurality of events indicated by each of the plurality of processed data; an output unit that processes the labels for each of the identified events based on the relative occurrence times of the events and outputs information indicative of a further medical event; A medical information processing device comprising:
2. the identification unit processes the label for each of the identified events in each of the plurality of data modalities by performing processing including providing the label as an input to a machine learning model to obtain information indicative of the further medical event. The medical information processing device according to claim 1 .
3. When one of the plurality of data is acquired, the identification unit processes the data of each of the plurality of data modalities and creates the label in real time for each of the plurality of data modalities; Providing the labels as inputs to the machine learning model includes providing, for each of the labels, each of the labels as inputs to the machine learning model at the time of creation of each of the labels. The medical information processing device according to claim 2 .
4. the machine learning model is a recurrent neural network; The medical information processing device according to claim 2 or 3.
5. the identification unit provides each of the labels as an input to the machine learning model in the order in which each of the corresponding identified events occurred on the common timeline; The medical information processing device according to claim 2 or 3.
6. the output unit outputs, for each of the identified events, time information indicating a time at which each of the identified events occurred on the common timeline; a determination unit that processes the time information of each of the identified events to determine a time associated with the further medical event. The medical information processing device according to claim 2 or 3.
7. the imaging data type includes at least one of video data, X-ray image data, and ultrasound image data; The medical information processing device according to claim 2 or 3.
8. the plurality of data modalities include at least one of video data, audio data, X-ray image data, ultrasound image data, radio frequency tag data, and vital data; The medical information processing device according to claim 2 or 3.
9. For one or more of the plurality of data modalities: processing the data from each of the plurality of data modalities to identify each of the plurality of events occurring during the procedure includes providing the data from each of the plurality of data modalities to a further machine learning model to derive the label for each of the identified events. The medical information processing device according to claim 2 or 3.
10. one or more of the plurality of data modalities includes the imaging data type; each of the additional machine learning models used to derive the label for each of the identified events for the imaging data type comprises a convolutional neural network; The medical information processing device according to claim 9 .
11. one or more of the plurality of data modalities includes audio data; each of the further machine learning models used to derive the label for each of the plurality of identified events in the audio data includes language recognition configured to derive text representing the audio data; The medical information processing device according to claim 9 or 10.
12. the identifying unit processes the text using a natural language understanding model to derive the labels identifying each of the plurality of events. The medical information processing device according to claim 11 .
13. the output unit controls a display to display visual information indicative of the further medical event and associated time information indicating when the further medical event occurred in the clinical procedure. The medical information processing device according to claim 2 or 3.
14. further comprising an application execution unit that executes an application corresponding to the further medical event in response to output of the information indicating the further medical event. The medical information processing device according to claim 2 or 3.
15. A medical information processing method by a medical information processing device, comprising: Acquire a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; For each of the plurality of data modalities, processing data for each of the plurality of data modalities to generate one or more labels occurring at times on the common timeline and identifying events indicated by each of the plurality of processed data; processing the labels for each of the identified events based on the relative occurrence times of each of the identified events to output information indicative of a further medical event; Medical information processing method.
16. On the computer, Acquire a plurality of data collected during a clinical procedure, the data belonging to a plurality of data modalities, at least one of the plurality of data modalities being an imaging data type, and the data modalities other than the imaging data type being a further data type different from imaging data, the plurality of data being provided in association with a common timeline that is common to the time at which the plurality of data were collected; For each of the plurality of data modalities, processing data for each of the plurality of data modalities to generate one or more labels occurring at times on the common timeline and identifying events indicated by each of the plurality of processed data; processing the labels for each of the identified events based on the relative occurrence times of each of the identified events to output information indicative of a further medical event; A program that executes a process.
Citation Information
Patent Citations
Medical information processing device, x-ray diagnostic device, and medical information processing program
JP2022025590A