Generating training data for a machine learning model and using a machine learning model

Automatic annotation of sensor data using situational information in clinical settings addresses the inefficiencies of manual training data generation, improving the accuracy and efficiency of machine learning model training in clinical environments.

DE102024132853A1Pending Publication Date: 2026-05-13KARL STORZ SE & CO KG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
KARL STORZ SE & CO KG
Filing Date
2024-11-11
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Generating training data for machine learning models in clinical settings is time-consuming and prone to errors due to manual sorting and annotation, which is inefficient and inconsistent.

Method used

Automatically annotate sensor data using situational information from clinical settings, such as changes in device states or actions by personnel, to generate training data for machine learning models, allowing for autonomous, local, and anonymous data generation.

Benefits of technology

Simplifies the generation of training data, reduces manual effort, and enhances data accuracy by leveraging situational context for efficient and reliable machine learning model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In a computer-implemented method for generating training data for a machine learning model for use in a clinical setting, sensor data from at least one sensor present in the clinical setting is received. Situational information is received, which depends on a change in the state of a device in the clinical setting and / or an action by a person in the clinical setting. At least a portion of the sensor data is automatically annotated using this situational information to generate the training data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method and a device for generating training data for a machine learning model in a clinical setting. The present invention further relates to a method and a device for using a machine learning model in a clinical setting. Finally, the present invention relates to a system comprising a device for generating training data for a machine learning model and a device for using the machine learning model in a clinical setting.

[0002] In the medical field, systems are known that can provide support in clinical situations, such as during certain procedures. For example, a system-assisted system for minimally invasive surgery is known from WO 2024 / 008854 A1.

[0003] The application of machine learning or artificial intelligence (AI) methods in clinical situations has the potential to improve support during interventions, diagnostic accuracy, efficiency and patient care.

[0004] In endoscopy, for example, AI algorithms can be used to analyze endoscopic images in real time and detect anomalies. AI systems can also support the user in making a diagnosis, as well as provide real-time feedback and recommendations to highlight or confirm potential pathological findings.

[0005] Machine learning models are trained on large amounts of data to recognize patterns and features that are characteristic of certain disease states.

[0006] In the current stage of machine learning, generating training data requires considerable effort. Data must first be carefully sorted, evaluated, and weighted manually before it can be integrated into the training process. These manual steps are not only time-consuming but also carry the risk of errors and inconsistencies. Summary of the invention

[0007] It is therefore an object of the present invention to provide devices and methods for generating training data for a machine learning model in a clinical setting, in order to simplify the generation of training data. Furthermore, methods and devices for using a machine learning model in a clinical setting are to be provided. Finally, a system that combines the two aforementioned devices is to be provided.

[0008] This problem is solved by the subject matter of the independent claims of the present invention. Advantageous embodiments are the subject matter of the dependent claims.

[0009] According to a first aspect, a computer-implemented method for generating training data for a machine learning model for use in a clinical setting is provided. Sensor data is received from at least one sensor present in the clinical setting. Situational information is received, which depends on a change in the state of a device in the clinical setting and / or an action by a person in the clinical setting. At least a portion of the sensor data is automatically annotated using the situational information to generate the training data.

[0010] A fundamental idea of ​​the present invention is that the sensor data can be at least partially automatically annotated (i.e., labeled). This eliminates the need for time-consuming manual annotation.

[0011] Situational information is taken into account for this purpose. Within the scope of this invention, situational information can be understood as information that depends on or describes states, changes in state, actions, events, or the like in the clinical situation.

[0012] For example, a specific action by a person (such as a treating physician or assistant) can indicate the circumstances or events present in the clinical situation, and the sensor data can be labeled accordingly. Similarly, a change in the state of a device used in the clinical setting can indicate the occurrence of an event or the presence of a specific situation, allowing the sensor data to be labeled accordingly. This enables more efficient generation of training data for the machine learning model.

[0013] Within the scope of this invention, a clinical situation can be understood to mean an operation, a surgical procedure, the treatment of a patient, an external examination of a patient, or the like. In particular, medical or surgical instruments may be used.

[0014] A change in the state of a device can refer, in particular, to switching it on, off, or changing its operating parameters. The device could be, for example, a medical instrument or system. In an endoscopy, for instance, a smoke extraction system could be activated or deactivated.

[0015] According to one embodiment of the method for generating training data for the machine learning model, the training data can be generated autonomously, locally and anonymously, which is particularly advantageous with regard to data protection considerations.

[0016] According to one embodiment, the method is further provided for generating the machine learning model. For this purpose, the machine learning model is trained using the annotated sensor data. The trained machine learning model can then be output and used in the clinical setting.

[0017] According to one embodiment of the method for generating training data for the machine learning model, the system determines, based on situational information, whether sensor data should be automatically annotated or discarded. This allows for the automatic detection of whether sensor data is relevant to the training process. If so, it is automatically annotated; otherwise, it is discarded. For example, if a doctor generates and saves a still image during an examination or operation, it can be inferred that the still image is important or contains helpful information. The still image can then be used for training and is automatically annotated.

[0018] According to one embodiment of the method for generating training data for the machine learning model, situational information is determined based on sensor data and / or other sensor data. For example, gestures or actions of people in a clinical situation can be recognized based on sensor data, from which conclusions about the situation can be drawn.

[0019] According to one embodiment of the method for generating training data for the machine learning model, the sensor data includes image data from at least one image sensor, audio data from at least one audio sensor, and / or video data from at least one video sensor. The sensor can, in particular, detect a person undergoing treatment.

[0020] According to one embodiment of the method for generating training data for the machine learning model, a person's action in the clinical situation includes activating or deactivating a device used in the clinical situation. For example, the person might activate or deactivate a medical or surgical device or an auxiliary device. The auxiliary device could be, for instance, a smoke extraction system during an endoscopy.

[0021] The action of the person can also be a specific movement of a piece of equipment used in the clinical situation. The progress of an operation or procedure can be inferred from this movement. The action can also be the person uttering a specific word or sentence, such as an instruction, from which the situation can be determined. The person in this context can be, in particular, a treating physician, assistant, or similar professional.

[0022] According to one embodiment of the method for generating training data for the machine learning model, the clinical situation includes video endoscopy. The sensor data can be image data and / or video data. The machine learning model can be trained for use in smoke detection, progress monitoring of the video endoscopy, and / or energy monitoring of a high-frequency device used in the video endoscopy.

[0023] According to a second aspect, the present invention provides a device for generating training data for a machine learning model for use in a clinical setting. The device comprises an interface that receives sensor data from at least one sensor present in the clinical setting. The interface receives situational information that depends on a change in the state of a device in the clinical setting and / or an action by a person in the clinical setting. A computing unit generates the training data by automatically annotating at least a portion of the sensor data using the situational information.

[0024] According to a third aspect, the invention provides a method for using a machine learning model in a clinical setting. Sensor data is received from at least one sensor present in the clinical setting. A machine learning model is used to generate a recommendation for action and / or to automatically perform an action, with the sensor data serving as input data for the machine learning model. Reaction information is determined based on a person's response to the recommendation for action and / or to the action performed.

[0025] Within the scope of the invention, a recommendation for action can be understood as the suggestion of one of several possibilities, or the suggestion to perform or not perform a specific action. For example, the user can be advised to activate a flue gas extraction system.

[0026] The automatic execution of an action could be, for example, the automatic activation of a device in a clinical situation, such as activating a smoke extraction system.

[0027] Reaction information can be understood as information that depends on how a person acts after receiving the recommended course of action or after the action has been carried out automatically.

[0028] According to one embodiment of the method for using the machine learning model in a clinical setting, the machine learning model is adapted based on the response information. For example, if the person accepts the recommended action or does not interrupt the automatically performed action, it can be recognized that the machine learning model correctly predicted the action. The machine learning model can be retrained, for example, by generating new training data based on the sensor data and the results of the machine learning model.

[0029] According to one embodiment of the method for using the machine learning model in a clinical situation, the machine learning model is first trained on the basis of training data which have been generated according to a method according to the first aspect.

[0030] According to one embodiment of the method for using the machine learning model in a clinical situation, the response information includes information on whether the person acted in accordance with the recommended course of action. This allows actions by experts, such as a treating physician, to be taken into account in order to improve the machine learning model.

[0031] According to one embodiment of the method for using the machine learning model in a clinical situation, the response information includes whether the person canceled or reversed the automatic execution of the action. If so, the prediction was likely incorrect; otherwise, it was correct.

[0032] According to one embodiment of the method for using the machine learning model in a clinical setting, an algorithm is used to calculate a reliability metric that indicates how reliable the recommended action and / or the automatic execution of the action is. Thus, a metric or hit rate can be specified in advance, quantifying the expected reliability. The algorithm can be adapted depending on the response information. For example, if the user does not perform the recommended action or interrupts the automatic execution of the action, the reliability metric can be reduced by adjusting the algorithm. Conversely, if the user performs the recommended action or does not interrupt the automatic execution of the action, the reliability metric can be increased by adjusting the algorithm.

[0033] According to one embodiment of the method for using the machine learning model in a clinical situation, the machine learning model is adapted depending on a user profile of the responding person. For example, the initially pre-trained model can be specifically adapted for each user and thus better support the user taking into account their typical behavior.

[0034] According to one embodiment of the method for using the machine learning model in a clinical situation, the machine learning model is retrained during operation. In particular, new training data can be generated.

[0035] According to a fourth aspect, the invention provides a device for using a machine learning model in a clinical setting. An interface receives sensor data from at least one sensor present in the clinical setting. A computing unit uses a machine learning model to generate a recommendation for action and / or to automatically execute an action. The sensor data are used as input data for the machine learning model. A response information acquisition unit determines response information based on a person's reaction to the recommendation for action and / or to the action performed.

[0036] According to a fifth aspect, the invention provides a system for use in a clinical setting. The system comprises a device for generating training data for a machine learning model according to the second aspect and a device for using the machine learning model according to the fourth aspect.

[0037] Although some functions are described here, in the foregoing and below, as being performed by "devices," "interfaces," or "modules," it should be understood that this does not necessarily mean that such devices, interfaces, or modules are provided as separate units. In cases where one or more devices, interfaces, or modules are provided wholly or partially as software, the devices, interfaces, or modules may be implemented by sections or snippets of program code that are distinct from one another but may also be intertwined.

[0038] Similarly, where one or more devices, interfaces, or modules are provided as hardware, the functions of one or more devices, interfaces, or modules may be provided by one and the same hardware component, or the functions of one device, interface, or module, or the functions of several devices, interfaces, or modules, may be distributed across several hardware components, which need not necessarily correspond one-to-one with the devices, interfaces, or modules. Therefore, any device, system, method, etc., that possesses all the features and functions attributed to a particular device and / or interface and / or module is to be understood as constituting, comprising, or implementing the device and / or interface and / or module.

[0039] In particular, it is possible that all facilities, interfaces, or modules are implemented by program code that is executed by a computing facility.

[0040] The computing device can be implemented as any device or means for performing calculations, in particular for executing software, an application, or an algorithm. For example, the computing device can include at least one processor, such as at least one central processing unit (CPU), and / or at least one graphics processing unit (GPU), and / or at least one field-programmable gate array (FPGA), and / or at least one application-specific integrated circuit (ASIC), and / or any combination thereof. The computing device can further include main memory operationally connected to the at least one processor, and / or non-volatile memory operationally connected to the at least one processor and / or the main memory. The computing device can be implemented partially and / or entirely in a local device and / or partially and / or entirely in a remote system, such as a remote system.be implemented through a cloud computing platform.

[0041] According to a sixth aspect, the invention provides a computer program product comprising executable program code which, when executed by a computing device, is configured to perform the method according to an embodiment of the first aspect or the third aspect of the present invention.

[0042] According to a seventh aspect, the invention provides a non-volatile, computer-readable data storage medium comprising executable program code which, when executed by a computing device, is configured to carry out the method according to an embodiment of the first aspect or the third aspect of the present invention.

[0043] The non-volatile, computer-readable data storage medium can include or consist of any type of computer memory, in particular semiconductor memory, such as solid-state memory. The data carrier can also include or consist of a CD, DVD, Blu-ray disc, USB flash drive, or the like.

[0044] According to an eighth aspect, the invention provides a data stream comprising executable program code or configured to generate executable program code which, when executed by a computing device, is set up to perform the method according to an embodiment of the first aspect or the third aspect of the present invention.

[0045] Further advantageous variants, options, embodiments, and modifications will become apparent from the following figures and the accompanying detailed description, as well as from the claims. It is understood, however, that while the detailed description and specific examples indicate preferred embodiments of the invention, they are provided for illustrative purposes only, since various changes and modifications within the scope of the invention are obvious to the person skilled in the art. Brief description of the characters

[0046] Individual embodiments of the present disclosure will be explained in detail with reference to the following figures. The components in the drawings are not necessarily to scale, but serve to illustrate the principles of the present invention. The numbering of process steps initially serves only to distinguish them and does not necessarily imply a corresponding sequence; however, it is one option to carry out the steps in the order of their numbering. Several steps can also be carried out overlapping or simultaneously. The figures show: Fig. 1 a schematic block diagram of a system for use in a clinical situation according to an embodiment of the invention; Fig. 2. A flowchart of a procedure for generating training data for a machine learning model for use in a clinical situation; Fig. 3. A flowchart of a procedure for using a machine learning model in a clinical situation; Fig. 4. A schematic block diagram of a computer program product; and Fig. 5 A schematic block diagram of a non-volatile, computer-readable data storage medium. Detailed description of the figures

[0047] Fig. Figure 1 shows a schematic block diagram of a System 300 that can be used in a clinical situation. The clinical situation can refer to a situation in a building or vehicle dedicated to medical purposes, for example, a medical research institute, a laboratory, a hospital, a medical university, a doctor's private practice, or the interior of an ambulance.

[0048] The system 300 includes a device 100 for generating training data for a machine learning model and a device 200 for using the machine learning model.

[0049] The device 100 for generating training data for the machine learning model comprises a first interface 101, which receives sensor data from one or more sensors. The first interface 101 can be a wired or wireless interface, such as a USB interface, optical interface, or the like.

[0050] The data is stored in a first storage device 103, such as a semiconductor memory, memory stick or the like.

[0051] The sensor data can include image data received from at least one image sensor. The sensor data can also additionally or alternatively include audio data from at least one audio sensor and / or video data from at least one video sensor. The sensor data can be recordings of the environment surrounding a patient being treated, or recordings of the patient's internal body, for example, using a probe, X-rays, ultrasound images, or similar methods.

[0052] The first interface 101 continues to receive situational information, which depends on a change in the state of a facility in the clinical situation.

[0053] The situational information may additionally or alternatively depend on an action taken by a person in the clinical situation. This action may consist of activating or deactivating a piece of equipment used in the clinical situation. It may also consist of performing a specific movement of a piece of equipment used in the clinical situation, such as a surgical instrument. Finally, it may consist of the person uttering a specific word or phrase.

[0054] Another example of an action is the user pressing a foot switch to generate an image. This suggests that corresponding sensor data is important and should therefore be considered when training the machine learning model.

[0055] More generally, situational information can encompass several pieces of information. In the case of smoke detection during an endoscopy, this can include whether the endoscope is inside the body, whether the procedure is in a phase where smoke can be generated, whether a high-frequency function is activated that could contribute to smoke development, and / or whether image processing can determine whether optical flow is occurring, whether the sensor image is generally blurry, or whether it has sharp segments in the focus area.

[0056] Situational information can be received externally or determined from received sensor data. Furthermore, it may be possible to consider additional sensor data to determine the situational information.

[0057] The device 100 for generating training data for the machine learning model further comprises a first computing unit 102, which includes, for example, a processor, an application-specific integrated circuit, or the like. The computing unit 102 automatically annotates at least a portion of the sensor data using the situation information. The training data is then generated using the annotated sensor data.

[0058] Depending on the situation information, it can be determined whether the sensor data is automatically annotated or discarded.

[0059] The training data can be used for supervised learning of a machine learning model.

[0060] The training data generally comprises a variety of data pairs, each consisting of an input and a corresponding target. The input can be generated from sensor data. For example, the input could be a camera image, a video sequence, or an audio sample. The target corresponds to the appropriate label, which is determined through automatic annotation. For example, the input can be classified, meaning the label corresponds to its assignment to a specific class.

[0061] For example, the target variable can correspond to a system state. In an endoscopy, for instance, a distinction can be made between two system states, one where flue gas is present and the other where it is not. The system state thus describes the presence of flue gas. The machine learning model can then be trained to detect flue gas.

[0062] More generally, automatic annotation can be used to identify a dataset that is then divided into training data and test data. After the machine learning model is trained with the training data, the test data is used to evaluate the performance of the machine learning model.

[0063] The device 100 for generating training data for the machine learning model can also be configured to generate the machine learning model itself. In this configuration, the device 100 trains the machine learning model using the annotated sensor data.

[0064] The clinical situation could be, for example, a video endoscopy. The machine learning model can then be trained for use in smoke detection, progress tracking during video endoscopy, or energy monitoring of a high-frequency device used in video endoscopy.

[0065] The device 200 for using the machine learning model includes a second interface 201, which receives sensor data from at least one sensor present in the clinical situation. The second interface 201 can be a wireless or wired interface.

[0066] The sensor data is stored in a second memory 203, such as a semiconductor memory, memory stick or the like.

[0067] A second computing unit, 202, uses the machine learning model to generate a recommended course of action. Alternatively or additionally, an action can be carried out automatically. The sensor data is used as input data for the machine learning model.

[0068] For example, during an endoscopy, the second computing unit 202 can use the machine learning model to decide whether to activate smoke extraction. For this purpose, the machine learning model can be trained as a classifier that detects the presence of smoke.

[0069] The device 200 for using the machine learning model further comprises a reaction information determination unit 204, which determines reaction information depending on a person's reaction to the action recommendation and / or to the action performed. The reaction information determination unit 204 can be identical to the second computing unit 202 or be a separate component.

[0070] The response information can include whether the person acted in accordance with the recommended course of action. It can also include whether the person canceled or reversed the automatic execution of the action. For example, if the user frequently performs a recall (i.e., interrupts the action), this may indicate that the action should not have been performed. The machine learning model can then be adapted based on the response information, for example, by automatically adjusting the weights of layers in a neural network of the machine learning model.

[0071] This allows for continuous improvement of the machine learning model without the need for manual retraining.

[0072] The response information can be used to evaluate the machine learning model. Every single action can be traced. Because the process can be executed autonomously, self-learning, and locally, all values ​​and parameters are available during runtime. This allows all actions to be evaluated and documented retrospectively without gaps. A chain of evidence can thus be generated for quality assurance and potential regulatory approval.

[0073] The machine learning model can also be adapted depending on the user profile of the responding person.

[0074] Furthermore, it may be provided that an algorithm is used to calculate a reliability metric, indicating how reliable the recommended course of action and / or the automatic execution of the action is. The algorithm is adapted depending on the response information. For example, a specific action can only be executed automatically if the reliability metric exceeds a predefined threshold, which may depend on the type of action.

[0075] Furthermore, it can be provided that if the reliability metric falls below a predefined threshold, the user is given a recommendation for action. The user can trigger this recommendation automatically, for example, by activating a control element (such as a button). Depending on whether the user performs the action or not, the calculation of the reliability metric is adjusted accordingly and / or the machine learning model can be retrained. For instance, the sensor data can be automatically annotated based on the user's response. This allows the machine learning model to be improved.

[0076] Fig. Figure 2 shows a flowchart of a procedure for generating training data for a machine learning model for use in a clinical setting. The procedure can be performed using the device 100 described above for generating training data for a machine learning model.

[0077] In step S101, sensor data is received from at least one sensor present in the clinical setting. This sensor data can include image data received from at least one image sensor. The sensor data can also additionally or alternatively include audio data from at least one audio sensor and / or video data from at least one video sensor. The sensors can be located in the vicinity of a patient. However, the sensors can also be located on or in a device used during an examination or surgery.

[0078] In step S102, situation information is received, which depends on a change in the state of a device in the clinical situation and / or an action by a person in the clinical situation. The action of a person in the clinical situation could consist of activating or deactivating a device used in the clinical situation, such as a smoke extraction system.

[0079] A person's action may consist of performing a specific movement of a device used in the clinical situation, such as moving a medical instrument in a specific direction.

[0080] The action may also include the person uttering a specific word or sentence, such as giving an instruction to an assisting person or during an assessment.

[0081] In step S103, at least some of the sensor data is automatically annotated using the situation information, thereby generating training data.

[0082] For example, several scenarios can occur in a clinical situation. The sensor data can be used to determine which scenario is most likely, and the sensor data is annotated accordingly. Another example involves progress detection. Here, the sensor data is used to determine how far an examination or operation has progressed, and the sensor data is annotated accordingly.

[0083] Fig. Figure 3 shows a flowchart of a procedure for using a machine learning model in a clinical setting. The machine learning model can be based on the one described in Fig. The procedures described in section 2 are used to build the machine learning model. In particular, training data for the machine learning model can be generated according to the procedure described there. The machine learning model can then be trained based on the generated training data. The machine learning model thus provided is then used in the procedure described below.

[0084] In step S201, an interface 201 receives sensor data from at least one sensor present in the clinical situation, such as a sensor attached to a medical device or a sensor permanently installed in the environment.

[0085] In step S202, a computing device 202 uses a machine learning model to generate a recommendation for action and / or to automatically execute an action. The sensor data is used as input data for the machine learning model. Based on the sensor data, the machine learning model determines which action should be recommended or executed.

[0086] In step S204, a reaction information determination unit 204 determines reaction information depending on a person's reaction to the action recommendation and / or to the action performed. For example, it can be determined whether the person accepts the action recommendation or accepts the automatically executed action, i.e., does not interrupt it.

[0087] Fig. Figure 4 shows a schematic block diagram of a computer program product 400. The computer program product 400 comprises executable program code 401, which, when executed (e.g., by a computing device), is configured to perform the method according to an embodiment of the present invention, for example, according to Fig. 2 or Fig. 3.

[0088] Fig. Figure 5 shows a schematic block diagram of a non-volatile, computer-readable data storage medium 500 according to an embodiment of the present invention. The data storage medium 500 comprises executable program code 501, which, when executed (e.g., by a computer), is configured to perform the method according to an embodiment of the present invention, for example, according to Fig. 2 or Fig. 3.

[0089] The non-volatile, computer-readable data storage medium 500 can, for example, be designed as or comprise a semiconductor memory, e.g., an SSD. The data storage medium 500 can also comprise or comprise a CD, DVD, Blu-ray disc, or a magnetic storage device.

[0090] The foregoing description of the disclosed embodiments contains only examples of possible implementations, which are described to enable a person skilled in the art to manufacture or use the present invention. Various variations and modifications of these embodiments are readily apparent to a person skilled in the art – upon knowledge of the present invention – and the general principles defined herein can be applied to other embodiments without departing from the scope of this disclosure.

[0091] Therefore, the present invention is not to be limited to the specific embodiments shown herein, but is to be granted the broadest scope consistent with the principles and novel features disclosed herein. Therefore, the present invention is to be limited only in accordance with the following claims.

[0092] The invention can be roughly summarized as follows: Sensor data is acquired and automatically labeled or annotated (i.e., without human intervention), taking situational information into account. This simplifies the generation of training data. Furthermore, actions can be automatically recommended or carried out, and the user's reaction is recorded. Reference symbol list 100 Device for generating training data for a machine learning model 101 first interface 102 first computing facility 103 first memory 200 Device for using a machine learning model 201 second interface 202 second computing facility 203 second storage 204 Response Information Determination Unit 300 System 400 computer program product 401 Program Code 500 data storage medium 501 Program Code S101 to S103, S201 to S203 Procedure steps QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] WO 2024 / 008854 A1

[0002]

Claims

Computer-implemented method for generating training data for a machine learning model for use in a clinical situation, comprising the steps of: Receiving (S101) sensor data from at least one sensor present in the clinical situation; Receiving (S102) situation information which depends on a change in the state of a facility in the clinical situation and / or an action of a person in the clinical situation; and Generating (S103) the training data by automatically annotating at least some of the sensor data, using the situation information. Method according to claim 1, wherein the machine learning model is trained using the annotated sensor data. Method according to claim 1 or 2, wherein, depending on the situation information, it is determined whether the sensor data is automatically annotated or discarded. Method according to one of the preceding claims, further comprising the step: determining the situation information based on the sensor data and / or based on further sensor data. Method according to any of the preceding claims, wherein the action of a person in the clinical situation comprises at least one of activating or deactivating a device used in the clinical situation by the person, of a certain movement of a device used in the clinical situation by the person, and of uttering a certain word or phrase by the person. Method according to one of the preceding claims, wherein the clinical situation comprises video endoscopy, wherein the sensor data comprises image data and / or video data, and wherein the machine learning model is trained for use in smoke gas detection, progress detection of the video endoscopy and / or energy monitoring of a high-frequency device used in the video endoscopy. Method for using a machine learning model in a clinical situation, comprising the steps of: Receiving (S201) sensor data from at least one sensor present in the clinical situation; Using (S202) a machine learning model to generate a recommendation for action and / or to automatically perform an action, using the sensor data as input data to the machine learning model; and Determining (S203) a response information depending on a person's response to the recommendation for action and / or to the action performed. Method according to claim 7, wherein the machine learning model is adapted depending on the reaction information. Method according to claim 7 or 8, wherein the response information includes information on whether the person acted in accordance with the recommended course of action. Method according to any one of claims 7 to 9, wherein the response information includes information on whether the person has cancelled or reversed the automatic execution of the action. Method according to one of claims 7 to 10, wherein a reliability parameter is further calculated using an algorithm, which indicates how reliable the recommendation for action and / or the automatic execution of the action is, and wherein the algorithm is adapted depending on the reaction information. Method according to one of claims 7 to 11, wherein the machine learning model is adapted depending on a user profile of the responding person. Device (100) for generating training data for a machine learning model for use in a clinical situation, comprising: an interface (101) configured to: receive sensor data from at least one sensor present in the clinical situation, and to receive situation information which depends on a change in the state of a device in the clinical situation and / or an action of a person in the clinical situation; and a computing device (102) configured to automatically annotate at least part of the sensor data using the situation information in order to generate training data. Device (200) for using a machine learning model in a clinical situation, comprising: an interface (201) configured to receive sensor data from at least one sensor present in the clinical situation; a computing unit (202) configured to use a machine learning model to generate a recommendation for action and / or to automatically perform an action, wherein the sensor data are used as input data for the machine learning model; and a response information determination unit (204) configured to determine response information depending on a person's response to the recommendation for action and / or to the action performed. System (300) for use in a clinical situation, comprising: a device (100) for generating training data for a machine learning model according to claim 13; and a device (200) for using the machine learning model according to claim 14.