Method and voice control device for voice control of a medical device

By employing dual voice analysis and redundant verification methods, the problem of preventing first-time failure in voice control of medical devices was solved, ensuring the reliability and security of voice commands and enabling the reliable execution of safety-critical operations.

CN115938357BActive Publication Date: 2026-03-31SIEMENS HEALTHINEERS AG
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The existing voice control methods for medical devices lack reliable first-failure protection, making them unusable in safety-critical applications.

Method used

Through a dual speech analysis process, first and second computational linguistics algorithms are used to identify and verify speech commands respectively, ensuring the consistency of speech commands. Control signals are only executed after confirmation that there are no errors, including redundant verification using independent hardware and software systems.

Benefits of technology

This invention achieves first-failure protection for voice control methods in medical devices, ensuring operational reliability and safety, and avoiding potential dangers caused by voice recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938357B_ABST
    Figure CN115938357B_ABST
Patent Text Reader

Abstract

A method for voice control of a medical device (1) having the steps of: detecting a voice input intended to control the device containing an operator; analyzing the audio signal to provide a first voice analysis result; identifying a first voice command based on the first voice analysis result; determining a verification signal for the first voice command, including: analyzing the audio signal to provide a second voice analysis result; identifying a second voice command of the operator based on the second voice analysis result; comparing the first and second voice commands, wherein the verification signal confirms the first voice command only when the first voice command and the second voice command meet a consistency criterion; generating a control signal for controlling the medical device based on the verification signal if the first voice command is confirmed according to the verification signal, wherein the control signal is adapted to control the medical device according to the first voice command; and inputting the control signal into the medical device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for voice control of a medical device by processing audio signals, said audio signals containing voice input from an operator intended to control the device. In particular, this invention relates to a method for voice control of a medical device to prevent first-time failure. This invention also relates to a corresponding medical system having the medical device. Background Technology

[0002] Medical devices are typically used to treat and / or examine and / or monitor patients, such as imaging modalities. Such as magnetic resonance imaging (MRI) equipment, computed tomography (CT) equipment, PET (positron emission tomography) equipment, or interventional and / or therapeutic equipment, such as radiotherapy or radiation therapy equipment. Patient treatment and / or examinations are typically performed with the assistance of an operator.

[0003] Before and during treatments and / or examinations performed on patients using such medical devices, the devices are typically configured in various ways, such as inputting patient data and setting different device parameters. These steps are performed by an operator, and the configuration of the medical device is usually done via a physical user interface provided on the device, through which the operator can input information.

[0004] For the economical operation of such medical devices, a smooth workflow or process is desired. In particular, setup should be designed to be as simple as possible. Voice control is especially suitable for this, where the operator transmits control commands to the medical device via natural language signals. In this regard, DE 10 2006 045719B4 describes a medical system with a voice input device, in which specific functions of the system can be activated and deactivated by means of voice control. Here, the audio signal detected by the voice input device is processed by a voice analysis module to determine the operator's voice commands.

[0005] In voice control, i.e., analyzing or recognizing user intentions or voice commands expressed in natural language, artificial intelligence algorithms, particularly neural networks, are preferred. These AI algorithms are particularly well-suited for mapping a high-dimensional input space comprising a large number of different speech sequences corresponding to natural speech inputs to a target space comprising multiple defined control commands.

[0006] Furthermore, for approval, multiple medical devices must meet first-failure protection or functional safety requirements at least for selected operating procedures or maneuvers to ensure patient and operator safety at all times during the increasingly automated operation of medical devices. First-failure protection means that no single first failure will render the use of the medical device unsafe for its entire lifespan.

[0007] Particularly safety-critical control commands may involve triggering / initiating X-ray radiation during image data inspection or radiotherapy. Another example of safety-critical control commands involves the (autonomous) controlled movement of a medical device or one of its components, such as a robotic arm, in space. Unauthorized or unverified radiation triggering or device movement can directly endanger the health of the patient or operator.

[0008] For example, in medical procedures within the scope of interventional radiology, where X-ray images must be recorded at different times using X-ray radiation, the physician performing the intervention, due to their aseptic nature, cannot effectively operate, for example, manually. The procedure must either be interrupted to initiate X-ray image recording, or at least one other operator must be present besides the physician, who is responsible for inputting the corresponding control signals.

[0009] To ensure functional safety and prevent first-time failure in configuring the hardware and / or control software of medical devices for processing and translating control commands detected via any user interface, it is currently common practice to require manual authorization of the recognized control commands by the operator. The so-called dead man grip is used as an example here. This switch / lever / handle must be continuously operated by the operator to perform automatic adjustment movements on the medical device. When the operator releases the dead man grip, the adjustment movement automatically stops. In this way, accidental or unintended input of control commands is largely avoided.

[0010] Alternatively, the medical device can run a redundant second software system on independent hardware, i.e., its own processor or its own memory, to protect the control software. The medical device only executes the initially identified control commands if the redundant software system verifies their credibility; otherwise, the control commands are discarded.

[0011] However, in the field of voice control, there remains a fundamental lack of established and reliable methods to verify the quality of speech recognition algorithms required for functional or first-failure-proof security. This is particularly due to the fact that known methods for speech processing remain error-prone and nondeterministic. Consequently, there is a lack of universally accepted standards for the training datasets of (AI) speech recognition algorithms (AI = Artificial Intelligence) that ensure first-failure-proof security, or universally effective robustness measures for (AI) speech recognition algorithms to correctly classify speech inputs altered, for example, due to distortion or background noise. This is the subject of current research. Methods for verifying recognized speech commands based on existing security standards are also insufficiently secure or impractical, as there is also a lack of classically defined requirements / standards for AI speech recognition algorithms to verify their implementation. This would be a mandatory prerequisite for the approval or certification of (AI) speech recognition algorithms. Furthermore, verifying recognized speech commands by repeating the same redundant speech recognition algorithm with the first failure if in doubt does not guarantee first-failure-proof security.

[0012] Therefore, voice control of medical devices has been used only in applications where safety is a critical concern until now. Summary of the Invention

[0013] The object of this invention is to solve this problem and provide a mechanism for voice control of medical devices, which allows for the determination of operator voice commands from audio signals in an improved and more reliable manner. In particular, the object of this invention is to provide a mechanism for ensuring first-failure protection through voice control.

[0014] According to the present invention, the objective is achieved by a method for voice control of a medical device, a corresponding voice control device, a medical system including the voice control device, a computer program product, and a computer-readable storage medium, as described in embodiments of the present invention. Advantageous designs are the subject of the embodiments.

[0015] The solution according to the purpose of the invention is described below with respect to the claimed method and the claimed device. The features, advantages, or alternative embodiments mentioned herein are equally applicable to other claimed subjects, and vice versa. In other words, the entity claims (e.g., relating to voice control devices) can also be improved by incorporating the features described or claimed in conjunction with the method. The corresponding functional features of the method are here constituted by the corresponding entity features of one of the devices, such as modules or units.

[0016] In a first aspect, the present invention relates to a method for voice control of a medical device. In an embodiment, the method is configured as a computer-implemented method. The method includes several steps.

[0017] One step involves detecting a voice signal, the voice signal containing voice input from an operator relating to control of the device. One step involves analyzing an audio signal to provide a first voice analysis result. One step involves recognizing a first voice command based on the first voice analysis result. One step involves determining a verification signal to confirm the first voice command.

[0018] Determining the verification signal includes sub-steps. One sub-step involves analyzing the audio signal to provide a second voice analysis result. Another sub-step involves recognizing a second voice command from the operator based on the second voice analysis result. A third sub-step involves comparing a first voice command and a second voice command. The verification signal confirms the first voice command only when the first voice command and the second voice command meet a consistency criterion.

[0019] Another method step involves generating a control signal for controlling a medical device based on a verification signal and a possible first voice command. This step is performed if the first voice command has been confirmed according to the verification signal. Here, the control signal is adapted or configured to control the medical device according to the first voice command. Another step involves inputting the control signal into the medical device.

[0020] In the context of this invention, audio signals can particularly contain sound information. Audio signals can be analog, digital, or digitized signals. Digitized signals can be generated from analog signals, for example, via an analog-to-digital converter. Accordingly, the detection step can include providing a digitized audio signal based on the received audio signal or digitizing the received audio signal. In embodiments, audio signal detection can include recording the audio signal using an input device comprising a suitable sensor, such as an acoustic sensor in the form of a microphone, which can be part of the user interface of a medical device. Audio signal detection can also include providing digital or digitized signals for other analytical steps.

[0021] Audio signals can include communication from the user (hereinafter also referred to as the operator), such as instructions to be performed or information about those instructions, such as voice control information. In other words, audio signals can include the operator's speech input in the form of natural language. Natural language is typically language spoken by humans. Natural language can include tone and / or intonation to communicate modulation. Unlike formal language, natural language can have structural and lexical ambiguity.

[0022] In the context of this invention, analyzing audio signals involves inferring or recognizing the content of an operator's speech input. The analysis may include recognizing human voice. The analysis may include analyzing values ​​representing human language in the speech input, such as frequency, amplitude, pitch modulation, etc. The results of the analysis are provided, at least in part, in the form of a first speech analysis result.

[0023] The analysis according to the invention can be applied to methods for processing natural language. In particular, according to a first aspect of the invention, a first computational linguistics algorithm can be applied to audio signals or speech input. One possibility for processing speech input expressed in natural language to provide a first speech analysis result is to convert the speech input expressed in natural language into text, i.e., structured language, by means of a speech-to-text (software) module. Subsequently, further analysis to provide the first speech analysis result, for example by means of latent semantic indexing (LSI), can assign meaning to the text. In particular, in this case, the audio signal is analyzed as a whole. It is possible to identify individual words or phrases and / or syntax. It is also possible to consider the relationship / position of words or phrases with words or phrases that appear before or later in the audio signal. The analysis of the audio signal can particularly include grammatical analysis, for example using a language or grammatical model. Additionally or alternatively, the operator's intonation can be evaluated. Thus, the invention achieves natural language understanding (NLU) contained in speech input.

[0024] The step of recognizing a first voice command based on the first speech analysis result involves assigning a specific voice command to the meaning of the text identified in the speech input based on the speech analysis result, or associating it with a predefined group of possible voice commands. Here, in the step of recognizing the first voice command based on the speech analysis result, any number of different voice commands can be assigned to the speech input. Each voice command here represents an instruction performed to execute a work step, a defined action, operation, or movement, which is automatically executed by the medical device based on the voice command. The recognition of the first voice command can include classifying the speech input into a predefined group of command categories based on the speech analysis result. The predefined group of possible command categories can exist, for example, in the form of a command library. In particular, multiple different speech inputs, or the speech analysis results of their individual components, can be assigned to the same command category and thus the same voice command based on their meaning content.

[0025] This approach considers the dimensions of natural language, particularly how different word choices or intonations can express the same user intent.

[0026] The instruction library can also include instruction categories for unrecognized voice instructions, which are always assigned when the audio signal has acoustic features that are not specific to a particular voice instruction or cannot be assigned to another instruction category.

[0027] The step of recognizing the first speech instruction can also be performed by applying a first computational linguistics algorithm. Preferably, the step of recognizing the first speech instruction follows either the creation of a textual record of the speech signal (structured language) or the creation of a semantic analysis result based on the textual record. Both can be included in the first speech analysis result.

[0028] In the implementation scheme, multiple instruction categories are at least partially specific to the medical device and particularly relate to the functionality provided by the medical device. These instruction categories are adaptable in the implementation scheme and can be adjusted, in particular, through the selection, configuration, or training of computational linguistic algorithms used during analysis. For example, additional instruction categories can be added.

[0029] In another step, a verification signal is determined to confirm the first voice command. This step checks whether the identified first voice command actually corresponds to the user intent or expected command anticipated by the operator. The verification signal here indicates a measure of consistency between the anticipated user intent and the first voice command. In embodiments, the verification signal can include verification information. The verification information can have at least two values, such as "1" representing complete consistency and "0" representing inconsistency; in embodiments, it can also have multiple different discrete values, such as between "1" and "0". Some of these discrete values ​​can correspond to different degrees of consistency, which can respectively correspond to fractional or incomplete consistency. In embodiments of the invention, these values ​​can also correspond to sufficient consistency between the voice command and the anticipated user intent.

[0030] According to the present invention, the step of determining the verification signal is performed in the following manner:

[0031] Analyze the audio signal to provide a second speech analysis result.

[0032] The analysis also identifies or infers the content of the operator's speech input. This analysis can also include human voice recognition. The analysis can include analyzing values ​​representing human language in the speech input, such as frequency, amplitude, and pitch modulation. The results of the analysis are provided, at least in part, as a second speech analysis result.

[0033] Methods for processing natural language can also be used here. In particular, a second computational linguistics algorithm can be applied to the audio signal or speech input. Specifically, the analysis can also include: processing the speech input expressed in natural language to provide a second speech analysis result, converting the speech input into text, i.e., structured language, using a speech-to-text (software) module. Subsequently, further analysis to provide the second speech analysis result, such as using latent semantic indexing (LSI), can assign meaning to the text. It is also preferable to analyze the audio signal as a whole. Individual words or phrases and / or syntax can be identified. The relationship / position of words or phrases with words or phrases appearing previously or later in the audio signal can also be considered. The analysis of the audio signal can particularly include grammatical analysis, such as using a language or grammar model. Additionally or alternatively, the operator's intonation can be evaluated.

[0034] The step of recognizing the second voice command based on the second voice analysis result also involves: assigning a specific voice command to the meaning of the text identified in the voice input, or associating it with a predefined group of possible voice commands, based on the second voice analysis result. Here, in the step of recognizing the second voice command based on the second voice analysis result, it is possible to assign a defined set of different voice commands to the voice input. Each voice command here represents a specific instruction for performing a work step, a defined action, operation, or movement, which is automatically executed by the medical device based on the voice command.

[0035] The step of recognizing the second speech instruction can also be performed by applying a second computational linguistics algorithm. Preferably, the step of recognizing the second speech instruction follows either the creation of a textual record of the speech signal (structured language) or the creation of a semantic analysis result based on the textual record. Both can be included in the first speech analysis result.

[0036] According to the present invention, a second analysis of the audio signal is performed, particularly independently of the first analysis step as described in detail below, and a second speech analysis result is generated to provide a second speech instruction, which is then compared with the first speech instruction. The first and second speech analysis results may differ in embodiments, although both are generated based on the same audio signal. The first and second speech instructions can also differ from each other. This is due to the different configurations of the corresponding analysis steps used to identify the speech instruction, as will be described in more detail below.

[0037] The first voice command and the second voice command are then compared, wherein if the first voice command and the second voice command meet a consistency criterion, a verification signal confirms the first voice command. According to the invention, the verification signal is therefore based on a comparison between the first voice command and the second voice command. If consistency is determined between the first voice command and the second voice command, verification information confirming the first voice command is generated for the verification signal. Determining consistency between the first voice command and the second voice command here corresponds to a test step or confirmation cycle for the first voice command according to the invention.

[0038] A consistency criterion can be configured as a threshold representing the equality or similarity between a first voice command and a second voice command. The threshold can have a value representing 100% identity between the first and second voice commands. However, the threshold can also be less than 100%. The threshold of the consistency criterion to be followed can, for example, depend on a security category or security requirement, to which the first and / or second voice commands are assigned, or where the security category or security requirement provides corresponding voice commands based on the work steps to be performed. Accordingly, according to the invention, it is particularly possible to store thresholds associated with second voice commands for a set of possible second voice commands, wherein the threshold can represent a corresponding security level according to the security category of the control command.

[0039] The safety categories of control commands represent safety measures or levels that must be followed when a medical device performs the corresponding work steps, operations, or maneuvers. Control commands may also include voice commands, which must also meet preset safety requirements. The corresponding safety measures can be those determined by normalization. According to the invention, each voice command or command category can be assigned to a safety category. In other words, according to the invention, there is a predetermined set of different safety categories, each representing a different safety measure. Therefore, each possible voice command, which can be identified as a first voice command and / or a second voice command, can be assigned its safety requirements according to one of the safety categories. This assignment rule can be predetermined based on rules or by means of a lookup table. Multiple voice commands can be assigned to a single safety category. In particular, at least one safety category can be provided, which includes safety-critical voice control commands, especially those exhibiting protection against first-time failure. However, multiple safety categories can also be provided for safety-critical voice commands based on different safety-related levels. In particular, at least one safety category can also be provided for non-safety-critical voice commands or general voice input. It is possible to pre-determine the possible classification of voice commands based on the medical device, its function, and / or the corresponding safety specifications for the operating procedures to be controlled by voice commands.

[0040] According to the invention, after confirming the first voice command by means of a verification signal, the verification signal is provided for further processing, for example, to the control unit of the voice control device according to the invention. In this regard, the step of generating a control signal based on the verification signal and the possible first voice command can include: providing the control signal. The verification signal indicates whether the required equivalence or similarity measure for the first voice command has been achieved. Upon corresponding confirmation of the first voice command by means of the verification signal, in a subsequent step, a control signal for controlling the medical device is generated based on the first voice command and the verification signal, and input into the medical device.

[0041] In embodiments of the invention, the steps involving the derivation of the second voice command and the comparison of the first and second voice commands can constitute a P-path, which advantageously suffices without further user interaction. The steps involving the derivation of the second voice command from the same audio signal and the comparison of the first and second voice commands can be performed at least partially in parallel with the derivation of the first voice command in time. For example, in embodiments, the detected audio signal can be simultaneously fed to the first analysis step and the second analysis step. Furthermore, the first and second voice commands can be identified simultaneously because the first and second computational linguistic algorithms are independent of each other and are implemented, in particular, in different software / hardware environments. Alternatively, these steps for determining the second voice command can be performed in time after the derivation of the first voice command.

[0042] By utilizing the step of determining the verification signal, the present invention advantageously applies voice commands, i.e., the monitoring or reliability verification of control commands obtained by means of machine speech analysis methods. In other words, the present invention, through the aforementioned steps, enables the implementation of a test (P-protect) path for voice commands according to a system designed to prevent first-time failure. If the first voice command is confirmed through the reliability verification step (the verification signal indicates consistency), the first voice command is further processed and executed by the medical device. If the first voice command is not confirmed through the reliability verification step (the verification signal indicates deviation), the process is aborted, and the first voice command is discarded and not executed.

[0043] By using a verification signal, the present invention advantageously prevents the execution of user input that has been misidentified or misinterpreted by machine speech recognition. In particular, this method enables the execution of control commands in a security-critical or first-failure-prevention manner. The first voice command is executed only after it has been confirmed by a reliability verification step.

[0044] In embodiments of the invention, the determination of the verification signal is therefore advantageously implemented in a separate, independent hardware and / or software system. In particular, the determination of the verification signal is performed independently of the remaining process steps. The algorithm used to determine the verification signal differs, in particular, from the computational linguistics algorithms previously used in the method. In embodiments of the invention, the step of determining the verification signal is performed in a real or virtual computing unit different from the previous process steps. Thus, according to the invention, it is possible to establish a system that prevents first-failure testing, comprising a control path (C-path, C-control) and a test path (P-path, P-protect), wherein the step of determining the verification signal within the protected P-path is performed. In particular, the steps of detecting audio signals containing voice input from an operator relating to control of the device, performing initial analysis of the audio signals to provide a first voice analysis result, and recognizing a first voice command based on the first voice analysis result form an important component of the C-path.

[0045] As mentioned at the beginning, in another preferred embodiment of the invention, the analysis for providing a first speech analysis result includes applying a first computational linguistics algorithm. The first computational linguistics algorithm preferably includes a first training function applied to the audio signal. The analysis for providing a second speech analysis result further includes applying a second computational linguistics algorithm, which includes a second training function applied to the audio signal. The first training function and the second training function are different from each other.

[0046] In a preferred improvement, the second training function is configured to recognize only safety-critical voice commands.

[0047] Safety-critical voice commands here represent defined work procedures, actions, operations, or movements that must be automatically executed by the medical device based on voice commands and must comply with defined safety requirements, particularly standard safety requirements. Safety-critical voice commands also include, in particular, voice commands designed to prevent first-failure. Voice commands designed to prevent first-failure should be understood as safety-critical voice commands with the highest safety requirements. In implementation schemes, safety-critical voice commands can be assigned to one or more safety categories, wherein, in the latter case, safety-critical voice commands can also have different safety levels as a subset of all possible voice commands.

[0048] The second computational linguistics algorithm should identify security-critical speech commands that typically possess a combination of representative features, namely, representative acoustic signal variation curves, frequency patterns, amplitudes, pitch changes, etc. They can also possess representative and easily distinguishable letter and / or word orders. Security-critical, especially first-failure-resistant speech commands can be characterized, for example, by a minimum number of syllables, such as three or more. Alternatively, they possess one-to-one, i.e., non-confusional speech features, minimizing from the outset the similarity to other speech commands or generally other speech inputs and the resulting risk of confusion. The aforementioned characteristics or features of the audio signal can be detected in the steps of analyzing the audio signal according to the invention and provided as a speech analysis result.

[0049] Accordingly, in a preferred embodiment of the present invention, the second training function is configured to identify structurally and / or lexically explicit speech instructions as security-critical speech instructions based on second tokenized information and / or second semantic information, typically based on the results of second language analysis.

[0050] According to the present invention, at least the structural and / or lexical characteristics of the audio signal are determined by a second computational linguistic algorithm. Therefore, according to the present invention, representative, explicit structural and / or lexical characteristics are associated with safety-critical speech instructions. According to the present invention, these explicit characteristics are thus pre-assigned to safety-critical speech instructions so that the safety-critical speech instructions themselves can be easily and reliably identified. In other words, according to the present invention, based on representative, structural, and lexical characteristics / features, an audio signal corresponding to a reliably identifiable safety-critical speech instruction can be pre-defined.

[0051] In this sense, in a particularly preferred embodiment, the second training function is configured to identify voice commands having at least three, preferably four syllables, as safety-critical voice commands. Here, the invention assumes that the risk of confusion with voice commands decreases with the number of syllables they contain. Therefore, particularly preferably, safety-critical operational steps of the device are assigned voice commands comprising at least three syllables, and particularly preferably, voice commands comprising five to seven syllables.

[0052] Therefore, the first computational linguistics algorithm includes a first training function applied to the audio signal for analysis. The training function should be understood as a trained machine learning algorithm. Preferably, recognizing the first voice command based on the first speech analysis result and assigning the first voice command to a security category further includes applying the first computational linguistics algorithm.

[0053] Preferably, the training function or trained machine learning algorithm comprises a neural network, preferably a convolutional neural network. The neural network is constructed substantially similarly to a biological neural network, such as the human brain. Preferably, the artificial neural network comprises an input layer and an output layer. Multiple intermediate layers can be included in between. Each layer comprises one, preferably multiple, nodes. Each node is here considered a biological computing unit or switching point, for example, a neuron. In other words, a single node corresponds to a specific computational operation applied to the input data. Nodes in one layer can be connected to each other and / or to nodes in other layers via corresponding edges or connections, particularly via directed connections. The edges or connections define the data flow of the network. In a preferred embodiment, the edges / connections are equipped with parameters, also referred to as "weights." These parameters adjust the influence or weight of the output data of the first node on the input of a second node connected to the first node via a connection.

[0054] According to an embodiment of the invention, the neural network is a trained network. Training of the neural network is preferably performed in the sense of "supervised learning" based on training data from a training dataset, i.e., known input-output data pairs. Here, known input data is transmitted as input data to the neural network, and the output data of the neural network is compared with the known output data of the training dataset. The artificial neural network then learns independently and adjusts the weights of individual nodes or connections until the output data of the output layer of the neural network is sufficiently similar to the known output data of the training dataset. In this context, "deep learning" is also mentioned in convolutional neural networks. The terms "neural network" and "artificial neural network" can be understood as synonymous.

[0055] According to the present invention, in one embodiment, the training function of the convolutional neural network, i.e., the first computational linguistics algorithm, is trained during the training phase to analyze the detected audio signals and identify a first voice command or a command category corresponding to the first voice command. The first voice command and / or command category then corresponds to the output data of the training function.

[0056] The speech instructions corresponding to the set of control commands to be identified by the first computational linguistics algorithm can have a large number of acoustic, structural, and / or lexical feature combinations, i.e., a large number of different combinations of frequency patterns, amplitudes, modulations, words, word order, syllables, etc. Accordingly, the training function learns during the training phase to assign one of the possible speech instructions based on the feature combinations extracted from the audio signal in the sense of the first speech analysis result, or to classify the audio signal according to one of the instruction categories. The training phase can also include manually assigning training input data in the form of speech input to various speech instructions or instruction categories. The first training function is trained to recognize the widest and most diverse set of audio signals as speech instructions based on a predefined, especially large, set of instruction categories, which includes human language.

[0057] The first set of neural network layers is capable of extracting or determining acoustic features of an audio signal, i.e., providing a first speech analysis result comprising a combination of acoustic features specific to the audio signal. The first speech analysis result can be provided in the form of an acoustic feature vector. In this regard, a speech data stream, preferably entirely comprising the audio signal, is used as input data to the neural network. The first speech analysis result can also be used as input data for a second set of neural network layers, also referred to as a "classifier." The second set of neural network layers is used to assign at least one speech instruction or instruction category to the extracted feature vector. The set of instruction categories can also include, in particular, instruction categories for unrecognized speech instructions. If a speech instruction cannot be explicitly identified based on the features, the neural network can be trained accordingly to assign the audio signal to that category.

[0058] The analysis steps or functions can also be performed by multiple, especially two or three, independent neural networks. Therefore, a first computational linguistics algorithm can include one or more neural networks. For example, feature extraction can be performed using a first neural network, and classification can be performed using a second neural network.

[0059] Classifying audio signals into instruction categories based on the initial speech analysis results involves comparing the extracted feature vectors of the audio signals with instruction category-specific feature vectors stored in an instruction library. For each instruction category, one or more feature vectors can be stored to account for the multidimensional nature of human language and to identify specific speech instructions based on a large number of different types of speech input.

[0060] The comparison of feature vectors can include a comparison of individual features, prioritizing all features included by the feature vectors. Alternatively or additionally, the comparison can be based on feature parameters derived from the feature vectors, which consider individual features. A consistency metric derived during the comparison for the feature vectors or feature parameters indicates which voice command or command category was assigned. Command categories are assigned based on the highest similarity or similarity exceeding a defined threshold.

[0061] The threshold used to determine the similarity measure can be automatically preset or preset by the operator. It can also depend on a specific combination of features identified for the audio signal. The threshold can represent a large number of individual thresholds for individual features in the feature vector or a general threshold that takes into account a large number of individual features contained in the feature vector.

[0062] Further analysis of the audio signal can now include applying a second computational linguistics algorithm, including a second training function, to the audio signal. Preferably, recognizing a second speech instruction based on the second speech analysis result further includes applying the second computational linguistics algorithm. The comparison of the first and second speech instructions can also be performed using the second training function. The first and second training functions are different training functions according to the invention.

[0063] As described at the beginning, the second training function in the implementation is configured to identify only security-critical voice commands in the audio signal. In this sense, the second training function or the second computational linguistics algorithm is specific to security-critical voice commands, which also include voice commands designed to prevent first-failure.

[0064] The step of recognizing a second voice command based on the results of the second voice analysis specifically involves associating the results of the second voice analysis with predefined groups of security-critical, particularly first-failure-prevention, voice commands. Therefore, recognizing a second voice command includes recognizing security-critical voice commands, particularly first-failure-prevention voice commands.

[0065] The second training function or the trained second machine learning algorithm also includes a neural network, which can be configured substantially as described with reference to the first training function.

[0066] According to an embodiment of the invention, the second neural network is also a trained network. As described above, the training of the neural network is preferably performed in the sense of "supervised learning" based on training data of the training dataset, i.e., known input and output data pairs. According to the invention, in an embodiment, the second neural network, i.e., the second training function of the second computational linguistic algorithm, is trained during the training phase to analyze the detected audio signals and identify the second voice command as a safety-critical voice command, particularly a first-failure-proof voice command in the set of safety-critical voice control commands. In an embodiment, the second voice command can correspond to the output data of the second training function. In other embodiments, the second training function also undertakes the step of comparing the first voice command and the second voice command.

[0067] The second computational linguistics algorithm seeks to identify security-critical speech instructions that possess representative combinations of features as described at the beginning. These feature combinations are particularly different from those of the speech instructions that the first computational linguistics algorithm seeks to identify. Security-critical, especially first-failure-resistant speech instructions, possess distinct and unambiguous phonetic, lexical, or structural features that are difficult to confuse with the features of other speech instructions.

[0068] Correspondingly, a second training function is also learned during the training phase, assigning one of the security-critical speech commands based on the feature combinations extracted from the audio signal, in the sense of the second speech analysis result. The training phase can also include manually assigning training input data, presented as speech input, to each speech command.

[0069] Similar to the case of the first training function, in the case of the second training function, the first set of neural network layers can also be involved in extracting or determining the acoustic features of the audio signal and providing a second speech analysis result. Refer to the description above, which can be reused here. The second speech analysis result can also be used as the input dataset for the second set of neural network layers, i.e., the "classifier," which assigns one of the safety-critical speech commands to the second speech analysis result. If no safety-critical speech command is recognized, the second training function can be advantageously configured to assign an error output to the second speech command. If the second training function does not recognize a safety-critical speech command, the first speech command is always discarded in the sense of human and device safety.

[0070] Here, the functions can also be executed by multiple, especially two, independent neural networks.

[0071] Classifying audio signals into safety-critical, particularly first-failure-prevention, voice commands based on the second speech analysis results can also be based on comparisons of the extracted second feature vector of the audio signal with a large number of specific and stored feature vectors for each safety-critical voice command. In embodiments of the invention, exactly one feature vector is stored for each safety-critical, particularly first-failure-prevention, voice command to take into account the safety requirements of safety-critical actions or movements for use in medical devices. The comparison of feature vectors can include individual comparisons, prioritizing all features included by the feature vectors. Alternatively or additionally, the comparison can be based on feature parameters derived from the corresponding feature vectors, which consider subsets or all individual features. The consistency metric for the feature vectors or feature parameters derived during the comparison indicates which safety-critical voice command was assigned or which was not identified. Safety-critical voice commands with the highest similarity or similarity above a determined threshold are assigned to the second voice command. The remainder refers to embodiments with respect to the first training function.

[0072] If the second training function does not identify a safety-critical voice command, the first and second voice commands are not compared, and a verification signal is created indicating that the second voice command does not exist. The method is aborted, and the first voice command is discarded. If the second training function identifies a safety-critical voice command, the first and second voice commands are compared. If the comparison shows that the two voice commands are inconsistent, a verification signal is generated that does not acknowledge the first voice command, and the first voice command is also discarded. A verification signal acknowledging the first voice command is generated only if the comparison shows that the two voice commands are sufficiently consistent.

[0073] It is also possible to train a second training function, particularly using a second set of neural network layers, to derive keywords, word orders, or instruction triggers from the feature vectors. Therefore, the second set of neural network layers is used to assign explicit keywords to the acoustic feature vectors. For this purpose, one or more keywords can be stored regarding safety-critical voice commands, for example, one or two keywords can be explicitly assigned to a safety-critical voice command.

[0074] Keywords are especially short words or word orders with, for example, three or four syllables. Specific acoustic feature vectors can also be stored for keywords of various instruction categories. A neural network for a second computational linguistics algorithm can be trained accordingly, particularly by means of a second set of neural network layers, performing a comparison between the feature vectors extracted from the audio signal and the stored feature vectors based on the assigned keywords. This comparison can be performed as described with reference to the first training function. If a keyword is identified in the audio signal, a corresponding safety-critical speech instruction is assigned. If the second training function determines that it does not match one of the stored keywords, no second speech instruction is recognized.

[0075] The first training function is trained to identify a wide and diverse set of audio signals as speech commands based on a predefined, particularly large, set of command categories, including human language. Other training functions are trained to identify keywords in the audio signals.

[0076] Therefore, the significant difference between the first training function and the second training function can be particularly evident in the type or range of the training data. The first training function is trained using a first training dataset, which, in an embodiment, comprises a broad and diverse set of audio signals, including human language and various speech instructions assigned according to a large number of different instruction categories. The first training dataset can also include speech input not assigned to any instruction category. The second training function is trained using a second training dataset, which is limited to specific, particularly security-critical, and first-failure-resistant speech instructions, i.e., speech, vocabulary, and / or structurally clear and unambiguous speech instructions. Therefore, in an embodiment, the second training dataset comprises different audio signals and corresponding security-critical speech instructions. In particular, in an embodiment, the second training dataset is a subset of the first training dataset, wherein the subset can only involve first-failure-resistant speech instructions. In this regard, according to the invention, the second training function can be trained using a small training vocabulary, and conversely, the first training function can be trained using a large training vocabulary.

[0077] Another difference between the first and second training functions lies in the construction of the validation or classification function, which is preferably performed using a second set of neural network layers. The first training function is configured to classify the speech input into a large number of different speech commands based on a large number of different categories, while the second training function is configured to classify the speech input into only a small number of safety-related speech commands based on a small number of categories.

[0078] In particular, according to the present invention, the first training function and the second training function differ in the manner or type of neural networks they use. In this way, for example, the risk of similar system errors occurring when performing speech command recognition using the first training function and the second training function can be reduced.

[0079] Particularly preferably, in the case of the first training function, a speech recognition algorithm that is known and available per se can be employed. In embodiments of the invention, the second training function corresponds to a speech recognition algorithm generated during a secure, for example, software development process at the manufacturer's end.

[0080] The first training function can be configured as a feedforward network, and the second training function can be configured as a recursive or feedback network, in which nodes of a layer are also linked to themselves or to other nodes in the same layer or at least one of the previous layers.

[0081] According to some implementation schemes, the first computational linguistics algorithm can be implemented as a so-called front-end algorithm, which acts as the master in, for example, the local computing unit or local speech recognition module of a medical device. As a front-end, processing can be performed particularly well in real time, allowing results to be obtained virtually without significant time delay. Correspondingly, the second computational linguistics algorithm can be implemented as a so-called back-end algorithm, which acts as the master in, for example, a remote computing device, such as a real server-based computing system or a virtual cloud computing system. In the back-end implementation, complex analysis algorithms requiring high computational power can be used in particular. Accordingly, the method can include transmitting audio signals to the remote computing device and receiving one or more analysis results from the remote computing device. In alternative implementations, the second computational linguistics algorithm can also be implemented as a front-end algorithm. Conversely, the first computational linguistics algorithm can also be implemented as a back-end algorithm.

[0082] According to several preferred embodiments of the present invention, the analysis of the audio signal includes tokenization for segmenting letters, words, and / or sentences within the audio signal using a first computational linguistics algorithm and a second computational linguistics algorithm. Here, a first speech instruction and a second speech instruction are identified based on a first speech analysis result including first tokenized information and a second speech analysis result including second tokenized information.

[0083] In embodiments of the present invention, only the first computational linguistics algorithm or the second computational linguistics algorithm can perform tokenization. In particular, the first network layers of the first training function and / or the second training function can be configured to tokenize audio signals.

[0084] In computational linguistics, tokenization refers to segmenting text into units at the letter, word, and sentence levels. According to some implementations, tokenization can include converting speech contained in an audio signal into text. In other words, it is possible to create a text record and then tokenize it. A wide range of methods known per se can be used for this, such as those based on formant analysis, hidden Markov models, neural networks, electronic dictionaries, and / or language models. Preferably, as described at the outset, this analysis step is performed using a first training function, a second training function, and / or other training functions.

[0085] The first speech analysis result and / or the second speech analysis result can include first or second tokenized information. By using the tokenized information, the structure of the speech input used to determine the speech command can be considered, or the user's intent can generally be considered.

[0086] According to some particularly preferred embodiments, analyzing the audio signal includes performing semantic analysis on the audio signal using a first computational linguistics algorithm and a second computational linguistics algorithm. Here, a first speech instruction and a second speech instruction are identified based on a first speech analysis result including first semantic information and a second speech analysis result including second semantic information.

[0087] In embodiments of the present invention, only the first or second computational linguistics algorithm can perform semantic analysis of the audio signal. Specifically, the first network layers of the first training function and / or the second training function can respectively constitute the audio signal for semantic analysis.

[0088] Semantic analysis is suitable for inferring the meaning of operator speech input. In particular, semantic analysis can include upstream speech recognition steps (speech-to-text) and / or tokenization steps.

[0089] According to some embodiments, semantic information indicates whether the audio signal contains one or more user intentions. User intentions can, in particular, be voice input from an operator relating to one or more voice commands under consideration. Voice commands can, in particular, be voice commands related to controlling a medical device. According to some embodiments, semantic information indicates or includes at least one characteristic of the user intention contained in the audio signal. Through semantic analysis, specific acoustic features or characteristics are thus extracted from the voice input, which can be considered for determining the voice command or would be relevant to determining the voice command.

[0090] In other embodiments, the method according to the invention further includes classifying at least a first voice instruction into one of a plurality of security categories, wherein at least one of a plurality of security categories is provided for safety-critical voice instructions. Here, when the first voice instruction has been assigned to a security category of a safety-critical voice instruction, determination of a verification signal for the first voice instruction is performed. Therefore, when the first voice instruction is classified according to the security category of a safety-critical voice instruction, a verification step is preferably performed. As mentioned at the beginning, in embodiments, safety-critical voice instructions include those executed by a medical device, and in particular, must be configured to prevent first-time failure.

[0091] In one implementation, voice commands or command categories are assigned to safety categories based on pre-determined assignment rules that consider the risk measure of the type of voice command or the operation, action, or movement of a medical device triggered by a voice command for the patient and / or operator and / or medical device or for other medical facilities. In another implementation, the assignment is performed via a pre-determined lookup table. In a further implementation, predefined keywords can be provided, which can be identified in the audio signal by a first computer linguistics algorithm, and each keyword is linked to one of the safety categories. Here, one or more keywords can be assigned to a safety category. In another implementation, if one or more keywords are identified in the audio signal by the first computer linguistics algorithm, a safety level is assigned to the first voice command. Alternatively, when recognizing the first voice command, the safety category can be automatically determined using predefined assignment rules.

[0092] The step of assigning the first voice instruction to a safety category therefore involves identifying the first voice instruction as a safety-critical voice instruction, particularly one that prevents first-time failure, or a non-safety-critical voice instruction. In an implementation, this step is also performed by a first computational linguistics algorithm, preferably by a first training function. Thus, the first training function can be trained to assign safety categories to the first voice instruction. Consequently, the output data of the training function of the first computational linguistics algorithm can also include safety categories. In particular, the third set of neural network layers can be configured to assign safety categories based on the instruction category and / or the identified first voice instruction, wherein the determined instruction category and / or the first voice instruction is used as input data for the third set of neural network layers. The third set is now configured to classify the instruction category and / or the identified voice instruction as a safety-critical command, particularly one that prevents first-time failure, or a non-safety-critical command.

[0093] In an embodiment of the present invention, recognizing the first voice command as a security-critical voice command is a prerequisite for determining the verification signal.

[0094] In other preferred embodiments, detecting the audio signal includes detecting the audio signal using a first input device and a second input device, wherein the first input device and the second input device are different from each other. Detection also includes providing the audio signal detected by means of the first input device for analysis and providing a first speech analysis result, and providing the audio signal detected by means of the second input device for analysis and providing a second speech analysis result.

[0095] In this implementation, audio signals are detected using two redundant input devices, such as two independent microphones. Errors that may occur during signal detection, i.e., during recording, digitizing, and / or storing voice input in one of the input devices, can only propagate to the derivation of the first or second voice command, but are unlikely to propagate to either.

[0096] According to another aspect, the present invention provides a voice control device for voice control of a medical device. The voice control device includes at least one interface for detecting audio signals containing voice input from an operator relating to controlling the device. The voice control device also includes a first evaluation unit configured to analyze the audio signal and provide a first voice analysis result, and to identify a first voice command based on the first voice analysis result. The first evaluation unit is further configured to assign the first voice command to a safety category, wherein at least one safety category is provided for safety-critical voice commands. The voice control device also includes a second evaluation unit configured to determine a verification signal for the first voice command and provide it for further processing. The second evaluation unit is specifically configured to analyze the audio signal to provide a second voice analysis result, identify a second voice command from the operator based on the second voice analysis result, and compare the first and second voice commands, wherein if the first and second voice commands meet a consistency criterion, the verification signal confirms the first voice command. The voice control device also includes a control unit configured to control the medical device based on the verification signal and, if necessary, generate the first voice command, provided that the first voice command has been confirmed according to the verification information. Here, the control signal is applicable to controlling the medical device according to the first voice command. The voice control device also includes an interface for inputting control signals into the medical device. The two interfaces can be combined into a single input and output interface.

[0097] Medical devices, in the sense of this invention, are particularly physical medical devices. These devices are typically used to treat and / or examine patients. Medical devices are especially capable of being configured to perform and / or assist medical procedures. Medical procedures can include imaging and / or interventional and / or therapeutic procedures, but also include patient monitoring. In particular, medical devices can include imaging modalities such as magnetic resonance imaging (MRI), single-photon emission computed tomography (SPECT), positron emission tomography (PET), computed tomography (CT), ultrasound, X-ray equipment, or X-ray equipment configured as a C-arm. Imaging modalities can also be combined medical imaging devices, comprising any combination of multiple imaging modalities mentioned above. Furthermore, medical devices can have interventional and / or therapeutic devices, such as biopsy devices, radiotherapy or radiation therapy devices for irradiating patients, and / or interventional devices for performing interventions, especially minimally invasive interventions. According to another embodiment, the medical device additionally or alternatively includes a patient monitoring module, such as an ECG device, and / or patient care equipment, such as ventilation equipment, infusion equipment, and / or dialysis equipment. In this embodiment, the treatment and / or examination and / or monitoring of patients by means of the medical device is typically assisted or controlled by an operator, such as a nurse, technician, X-ray assistant, or physician.

[0098] The first and second evaluation units and the control unit can be configured as one or more central and / or distributed computing units. Each computing unit can have one or more processors. The processors can be configured as central processing units (CPU / GPU). In particular, the first and / or second evaluation units and the control unit can each be implemented as part or module of a medical device to be controlled via voice input. In embodiments, the evaluation units can each be configured as submodules of the control unit, or vice versa. Alternatively, at least one evaluation unit can be implemented as a local or cloud-based processing server. Furthermore, at least one evaluation unit can include one or more virtual machines.

[0099] In a preferred embodiment, the first and second evaluation units are configured as separate computing units. That is, the two evaluation units are implemented on separate hardware or in different software modules, as described at the beginning. In this way, the voice control device is configured to represent a classic control protection (CP) structure of a control system for first-failure prevention in response to control commands in the form of voice input, wherein the first evaluation unit constitutes part of the control path (C path) and the second evaluation unit constitutes the test path (P path).

[0100] The interface of a voice control device is typically configured for data exchange between the voice control device and other components and / or for data exchange between components or modules of the voice control device. In this regard, the interface can be implemented as one or more separate data interfaces, which can have hardware and / or software interfaces, such as a PCI bus, USB interface, FireWire interface, ZigBee, or Bluetooth interface. The interface can also have an interface to a communication network, wherein the communication network can be a local area network (LAN), such as an intranet or a wide area network (WAN). Correspondingly, one or more data interfaces can have a LAN interface or a wireless LAN interface (WLAN or Wi-Fi). The interface can also be configured for communication with an operator via a user interface. Accordingly, the interface can be configured to display voice commands via the user interface and receive associated user input via the user interface. In particular, the interface can include an acoustic input device for recording audio signals, and in embodiments, includes an acoustic output device for outputting an audio signal including a request to acknowledge a first voice command.

[0101] In another preferred embodiment, the interface includes first and second input devices for simultaneously detecting audio signals. Systematic or random errors during audio signal detection are limited to one of the evaluation paths used to determine the first and second voice commands.

[0102] The advantages of the proposed device substantially correspond to the advantages of the proposed method. Features, advantages, or alternative embodiments / aspects can also be applied to other claimed subjects, and vice versa.

[0103] According to another aspect, the present invention provides a medical system comprising a voice control device according to the invention and a medical device for performing medical procedures.

[0104] In another aspect, the present invention relates to a computer program product comprising a program and capable of being directly loaded into the memory of a programmable computing unit and having program mechanisms, such as libraries and auxiliary functions, so that when the computer program product is executed, a method for voice control of a medical device, particularly according to the above embodiments / aspects, is performed.

[0105] In another aspect, the invention also relates to a computer-readable storage medium on which readable and executable program segments are stored, so that when the program segments are executed by a computer, all steps of the method for voice control of a medical device according to the above embodiments / aspects are performed.

[0106] The computer program product herein can include software having source code or executable software code, where the source code still needs to be compiled and linked, or the source code only needs to be interpreted, and the executable software code only needs to be loaded into a computing unit for execution. The computer program product enables the rapid, repeatable, and robust execution of the method according to the invention. The computer program product is configured such that it can execute the method steps according to the invention by means of a computing unit. Here, the computing unit must have preconditions, such as corresponding working memory, corresponding processor, or corresponding logic unit, to enable efficient execution of the corresponding method steps.

[0107] Computer program products, for example, are stored on computer-readable storage media or stored on a network or server, from which they can be loaded into the processor of a corresponding computing unit, the processor being directly connected to or constituting part of the computing unit. Furthermore, control information of the computer program product can be stored on a computer-readable storage medium. The control information of the computer-readable storage medium can be configured such that, when a data carrier is used in the computing unit, the control information executes the method according to the invention. Examples of computer-readable storage media are DVDs, magnetic tapes, or USB sticks on which electronically readable control information, particularly software, is stored. If the control information is read from the data carrier and stored in the computing unit, all embodiments / aspects of the method according to the invention described above can be executed. Therefore, the invention can also be based on the computer-readable medium and / or the computer-readable storage medium. The advantages of the proposed computer program product or the associated computer-readable medium substantially correspond to the advantages of the proposed method. Attached Figure Description

[0108] Other features and advantages will become apparent from the following description of the embodiments, as illustrated in the diagrams. The modifications mentioned herein can be combined with each other to form new embodiments. In the different drawings, the same reference numerals are used for the same features.

[0109] Figure 1 A schematic block diagram of a system for controlling a medical device according to one embodiment is shown;

[0110] Figure 2 Another block diagram of a system for controlling a medical device according to another embodiment is shown;

[0111] Figure 3 A schematic flowchart illustrating a method for controlling a medical device according to one embodiment is shown;

[0112] Figure 4Another schematic flowchart of a method for controlling a medical device according to one embodiment is shown; and

[0113] Figure 5 A neural network of a computational linguistics algorithm according to the present invention is shown in one embodiment. Detailed Implementation

[0114] Figure 1 A functional block diagram of a system 100 for controlling a medical device 1 is schematically shown. The system 100 includes the medical device 1, configured to perform medical procedures on a patient. The medical procedures may include imaging and / or interventional and / or therapeutic procedures. The system also includes a voice control device 10.

[0115] Medical device 1 can include an imaging modality. The imaging modality is typically configured to image anatomical regions of a patient when the patient is brought to the detection area of ​​the imaging modality. Examples of imaging modalities include magnetic resonance imaging (MRI), single-photon emission computed tomography (SPECT), positron emission tomography (PET), computed tomography (CT), ultrasound, X-ray equipment, or X-ray equipment configured as a C-arm. The imaging modality can also be a combined medical imaging device comprising any combination of multiple imaging modalities mentioned above.

[0116] Furthermore, medical device 1 can have interventional and / or therapeutic devices. Interventional and / or therapeutic devices are generally configured to perform interventional and / or therapeutic medical procedures on a patient. For example, interventional and / or therapeutic devices can be biopsy devices for obtaining tissue samples, radiotherapy or radiation therapy devices for irradiating a patient, and / or interventional devices for performing interventions, especially minimally invasive interventions. According to embodiments of the invention, interventional and / or therapeutic devices can be automated or at least partially automated, and especially robotically controlled. Radiotherapy or radiation therapy devices can, for example, have a medical linear accelerator or other radiation source. For example, interventional devices can have catheter robots, minimally invasive surgical robots, endoscopic robots, etc.

[0117] According to another embodiment, the medical device 1 may additionally or alternatively have units and / or modules for assisting in the execution of medical procedures, such as patient support equipment that can be at least partially automated and / or monitoring equipment for monitoring the patient's condition, such as an ECG device, and / or patient care equipment, such as ventilation equipment, infusion equipment and / or dialysis equipment.

[0118] According to an embodiment of the invention, one or more components of the medical device 1 should be operable via one or more voice inputs from an operator. For this purpose, the system 100 includes a voice control device 10, which includes an interface with an acoustic input device 2.

[0119] Acoustic input device 2 is used to record or detect audio signal E1, that is, to record the speech produced by the operator of system 100. Input device 2 can be implemented as a microphone, for example. Input device 2 can be fixedly installed on medical device 1 or in another location, such as in a remote control room. Alternatively, input device 2 can also be implemented portablely, for example, as a microphone in an earphone that can be carried by the operator. In this case, input device 2 advantageously has a transmitter 21 for wireless data transmission.

[0120] The voice control device 10 has an input terminal 31 for receiving signals and an output terminal 32 for providing signals. The input terminal 31 and the output terminal 32 can form an interface device for the voice control device 10. The voice control device 10 is generally configured to perform data processing procedures and to generate electrical signals.

[0121] For this purpose, the voice control device 10 can have at least one computing unit 3. The computing unit 3 can include, for example, a processor, such as a CPU. The computing unit 3 preferably has first and second central evaluation units, for example, evaluation units each having one or more processors. The computing unit 3 can preferably include a control unit configured to generate control signals for the medical device.

[0122] The computing unit 3 can be at least partially configured as a control computer (Systemcontrol) or a part thereof for the medical device 1. According to the invention, the computing unit 3 includes units or modules configured to perform safety functions, particularly standardized safety functions (also known as SIL (safety integrity level) units). These standardized safety functions are used to minimize operational risks to the medical device 1 caused by executing incorrectly recognized control commands. In particular, the safety functions can protect against the first failure when automatically recognizing control commands from voice input. According to another embodiment, the functions and components of the computing unit 3 can be distributed across multiple computing units or control modules of the system 100.

[0123] Furthermore, the voice control device 10 has at least one data storage device 4, and more specifically, it can have a non-volatile data storage device that can be read by the computing unit 3, such as a hard disk, CD-ROM, DVD, Blu-ray disc, floppy disk, flash memory, etc. Typically, software P1 and P2 can be stored on the data storage device 4, which is configured to cause the computing unit 3 to execute the steps of the method.

[0124] As in Figure 1 As schematically shown, the input terminal 31 of the voice control device 10 is connected to the input device 2. The input terminal can also be connected to the medical device 1. The input terminal 31 can be configured for wireless or wired data communication. For example, the input terminal 31 can have a bus terminal. Alternatively, or in addition to a wired connection terminal, the input terminal 31 can also have an interface, such as a receiver 34 for wireless data transmission. For example, as in... Figure 1 As shown, the receiver 34 can communicate with the transmitter 21 of the input device 2. For example, the receiver 34 can be equipped with a WIFI interface, a Bluetooth interface, etc.

[0125] On one hand, the output terminal 32 of the voice control device 10 is connected to the medical device 1. The output terminal 32 can be configured for wireless or wired data communication. For example, the output terminal 32 can have a bus terminal. Alternatively, or in addition to the wired connection terminal, the output terminal 32 can also have an interface for wireless data transmission, such as an interface with the online module OM1, such as a WIFI interface, Bluetooth interface, etc.

[0126] The voice control device 10 is configured to generate and provide one or more control signals C1 at its output terminal 32 to control the medical device 1. The control commands C1 cause the medical device 1 to perform specific operating steps or sequences of steps. Taking an imaging mode configured as an MR device as an example, such steps could involve performing a specific scanning sequence by exciting a magnetic field in a specific manner through the generator circuitry of the MR device. Furthermore, such steps could involve the movement of movable system components of the medical device 1, such as the movement of a patient support device or the movement of the emission or detector components of the imaging mode. These steps could also, in particular, involve the triggering or initiation of X-ray radiation.

[0127] To provide the control signal C1, the computing unit 3 can have different modules M1-M3. The first module M1, corresponding to the first evaluation unit, is configured to analyze the audio signal E1 and, based on this, provide a first speech analysis result, and identify a first speech command SSB1 based on the first speech analysis result. In an embodiment, module M1 can also be configured to assign the first speech command SSB1 to the security category SK. For this purpose, module M1 can be configured to apply a first computational linguistic algorithm P1 to the audio signal E1. In particular, module M1 is configured to execute method steps S20 to S40.

[0128] The first voice command SSB1 can then be input to another module M2 corresponding to a second evaluation unit independent of the first evaluation unit. Specifically, when the first voice command SSB1 is identified as a security-critical voice command, particularly one designed to prevent first-time failure, the first voice command SSB1 is then forwarded to module M2 (e.g., by assigning it to a security category SK that includes the security-critical voice command). Module M2 is configured to determine a verification signal VS for confirming the first voice command SSB1. To this end, module M2 can be configured to apply a second computational linguistics algorithm P2 to the audio signal E1. The second computational linguistics algorithm P2 can be configured to identify the second voice command SSB2 in the audio signal E1 and perform a comparison between the first voice command SSB1 and the second voice command SSB2. The second computational linguistics algorithm P2 can also be configured to identify at least one keyword or specific word order in the second audio signal E2 that can be linked to a security-critical voice command.

[0129] The second module M2 is specifically configured to generate a verification signal VS that includes verification information confirming the first voice command SSB1, based on a comparison between the first voice command SSB1 and the second voice command SSB2. The verification signal VS may preferably also include the first voice command SSB1.

[0130] A verification signal VS is provided to a third module M3 corresponding to the control unit of the voice control device 10. Module M3 is configured to provide one or more control signals C1 based on the verification signal VS and a possible first voice command SSB1, the control signals being adapted to control the medical device 1 according to the first voice command SSB1.

[0131] If the first voice command SSB1 belongs to security category SK, which involves voice commands that are not security-related, then module M2 is not activated in the implementation scheme. Module M1 then inputs the recognized first voice command SSB1, for example, directly into module M3. This eliminates the need for verification according to the invention.

[0132] The subdivision into modules M1-M3 is used here only to more briefly illustrate the operation of computing unit 3 and should not be construed as restrictive. Modules M1-M3 can also be understood here as computer program products or computer program segments that, when executed in computing unit 3, perform one or more of the following functions or method steps.

[0133] Preferably, at least module M2 is configured as a SIL unit for performing security functions. According to the invention, the security functions involve checking or verifying the identified first voice command SSB1 in at least one independent verification loop.

[0134] According to the classic CP structure, module M1 and module M3 together constitute part of the C path, while module M2 constitutes part of the P path.

[0135] Figure 2 A functional block diagram of a system 100 for performing medical procedures on a patient, according to another embodiment, is shown schematically.

[0136] exist Figure 2 The implementation shown in the document is the same as that in the document. Figure 1 The difference in the implementation shown is that, in one respect, the function of module M1 is at least partially controlled within the online module OM1. In other respects, the same reference numerals denote the same components or components with the same function.

[0137] The online module OM1 can be stored on server 61, and the voice analysis device 10 can exchange data with server 61 via an internet connection and an interface 62. Accordingly, the voice control device 10 can be configured to transmit an audio signal E1 to the online module OM1. The online module OM1 can be configured to determine a first voice command SSB1 based on the audio signal E1, and optionally determine the relevant security category and return it to the voice control device 10. Accordingly, the online speech recognition module OM1 can be configured to make a first computational linguistics algorithm P1 available in a suitable online storage. The online module OM1 can be understood here as a centralized device, where multiple, particularly local clients, provide speech recognition services (in this sense, the voice control device 10 can be understood as a local client). The advantage of using a central online module OM1 is that it allows the application of more powerful algorithms and consumes more computing power.

[0138] In an alternative implementation, the online speech recognition module OM1 can also return "only" the first speech analysis result SAE1. The first speech analysis result can then, for example, contain machine-usable text converted from the audio signal E1. Based on this, module M1 of computing unit 3 can recognize the first voice command SSB1. This design is advantageous when the voice command SSB1 depends on the condition of medical device 1, and the online module OM1 cannot access the medical device and / or is not prepared to consider it. In this case, the performance of the online module OM1 is used to create the first speech analysis result, but otherwise the voice command is determined within computing unit 3.

[0139] However, according to another modification not shown, other functions of the speech analysis device 10 can also be executed in the central server. Therefore, it is conceivable that the second computational linguistics algorithm P2 also plays a controlling role in the online module.

[0140] On the other hand, the implementation shown here is similar to that in Figure 2 The difference in the embodiment shown lies in the additional acoustic input device 20. This acoustic input device is also used to record or detect the audio signal E1. Detection here occurs simultaneously or in parallel with the detection of the audio signal E1 by means of input device 2. Input device 20 can also be implemented as a microphone, which can be directly and fixedly mounted separately from input device 20 within the housing of medical device 1 or in other locations, such as in an operating room. Input device 20 can also be configured in embodiments, as in input device 2, for wireless data communication with input terminal 31. Input terminal 31 is here configured to provide the audio signal detected by means of input device 2 for analysis and providing a first speech analysis result SAE1, and to provide the audio signal detected by means of input device 20 for analysis and providing a second speech analysis result SAE2. Errors in signal detection, i.e., errors in the recording, digitization, and / or storage of the speech input, are thus advantageously limited to the evaluation path of the audio signal.

[0141] exist Figure 1 In the system 100 illustrated in the example, the medical device 1 is controlled via... Figure 3 The method is illustrated as a flowchart. The order of the method steps is not limited by the order shown or the selected numbering. Therefore, the order of the steps can be reversed if necessary. The steps can be performed in parallel in the implementation. Individual steps can also be omitted.

[0142] Typically, it is proposed here that the operator of medical device 1 issues commands via voice or speech, for example, by the operator uttering a sentence such as "Start X-ray recording" or "Take the patient to the initial position," input device 2 and, if necessary, input device 20 detect and process the relevant audio signal E1, and voice control device 10 analyzes the detected audio signal E1 and generates a corresponding control command C1 to operate medical device 1. One advantage of this method is that the operator can also perform other tasks while speaking, such as preparing the patient. This advantageously speeds up the workflow. Furthermore, medical device 1 can therefore be controlled at least partially "non-contactly," thereby improving the hygiene of medical device 1.

[0143] A method for voice control of medical device 1 includes steps S10 to S70. These steps are preferably performed using a voice control device 10. In step S10, an audio signal E1 is detected using an input device 2, the audio signal containing voice input from an operator relating to controlling the device 1. In an embodiment, the audio signal E1 is detected simultaneously using an input device 20 in step S10.

[0144] Audio signal E1 is provided to voice control device 10 via input terminal 31. Step S20 includes analyzing audio signal E1 to provide a first speech analysis result SAE1. For this purpose, the audio signal detected by means of input device 2 can be provided in particular. Therefore, step S20 includes providing a language expression related to the control of the medical device as the first speech analysis result SAE1 from audio signal E1. Generating speech analysis result SAE1 can include multiple sub-steps.

[0145] The sub-step may involve: firstly converting the sound information contained in the audio signal E1 into text information, i.e., generating a text record. The sub-step may involve: tokenizing the audio signal E1 or the operator's speech input or text record T. Tokenization here means segmenting the speech input, i.e., dividing the spoken text into units at the word or sentence level. Accordingly, the first speech analysis result SAE1 may include first tokenized information, such as indicating whether the operator has finished speaking the current sentence.

[0146] The sub-step can involve, additionally or alternatively, performing semantic analysis of the audio signal E1 or the operator's voice input or text record T. Accordingly, the first voice analysis result SAE1 can include first semantic information of the operator's voice input. Semantic analysis here involves assigning meaning to the voice input. For this purpose, for example, it can be compared word-by-word or word-by-word groups with a general or medical device 1-specific database of words and / or voice commands 50 of the medical device 1 or the system 100 according to the invention. In particular, in this step, for example, one or more statements contained in the command library 50 of the medical device 1 can be assigned to the text record T as different voice commands. Thus, the user's intent regarding the voice command can be identified.

[0147] The described sub-steps can be executed by the language understanding module included in module M1, and in particular, the sub-steps can be executed by the online module OM1.

[0148] In step S30, a first voice instruction SSB1 is identified based on the first speech analysis result SAE1. Here, semantic information representing the user's intent is used in particular to assign the corresponding voice instruction as the first voice instruction SSB1. In other words, the identified voice instruction is assigned to one of several possible instruction categories, where different expressions, words, word orders, or word combinations are stored for each instruction category, representing the user's intent corresponding to the voice instruction. Step S30 may also include comparing the first speech analysis result SAE1 with the instruction database 5 or the instruction library 50.

[0149] In optional step S40, the first voice command SSB1 is assigned to the safety category SK. Here, according to the invention, at least one safety category is provided for safety-critical voice commands to which the first voice command SSB1 belongs. Here, each voice command can be assigned to the safety category SK by means of a classic lookup table or other assignment rules. According to the invention, at least one of the possible safety categories is now provided for safety-critical voice commands whose execution by the medical device 1 must meet specific, particularly standardized, safety requirements. In particular, this includes voice commands whose execution must be implemented in a manner that prevents first-time failure.

[0150] Steps S20 to S40 can be implemented, for example, by software P1, which is stored in data storage 4 and causes computing unit 3, particularly module M1 or online module OM1, to execute the steps. Software P1 can include a first computational linguistics algorithm, which includes a first training function applied to an audio signal, particularly for analyzing audio signal E1, which will refer to... Figure 4 and Figure 5This will be explained in more detail. Thus, this invention implements the classic C path of the CP architecture.

[0151] In step S50, a verification signal VS is determined to confirm the first voice command SSB1. The verification signal VS is determined based on the detected audio signal E1. Preferably, the audio signal detected by means of the input device 20 can be provided for step S50. According to the present invention, the verification signal VS can be determined in multiple sub-steps, as referenced... Figure 4 As explained in more detail, sub-steps S51 to S54 of S50 involve analyzing the audio signal to provide a second speech analysis result SAE2, identifying the operator's second speech command SSB2 based on the second speech analysis result, and comparing the first and second speech commands.

[0152] According to the present invention, the verification signal is determined using a separate module M2 of the computing unit 3. Therefore, the present invention implements a separate P-path for protecting security-critical voice commands, wherein the corresponding security functions are based on voice commands obtained by means of machine language processing methods according to the present invention.

[0153] During step S50, steps S20 to S40 are performed on the audio signal detected by means of input device 2.

[0154] In the implementation, step S50 can be executed in parallel, i.e., simultaneously with steps S20 to S40. Alternatively, step S50 can be executed after steps S20 to S40. In the implementation, step S50 can be executed for each control command identified as the first voice command SSB1 in step S40. In a preferred implementation, step S50 is executed according to step S40, i.e., only when the first voice command SSB1 is identified as a security-critical voice command by being assigned to a security category SK that includes security-critical voice commands, particularly as a voice command designed to prevent first-time failure.

[0155] In step S60, if the first voice command SSB1 is confirmed in step S50, a control signal C1 for controlling the medical device is generated based on the verification signal VS or the first voice command SSB1. The control signal C1 is then suitable for controlling the medical device 1 according to the first voice command SSB1. The first voice command SSB1 can be transmitted to module M3 of the computing unit 3 (or, for example, corresponding software stored in the data memory 4) in the form of an input variable or as part of the verification signal VS, and at least one control signal C1 can be derived from it. In step S70, the control signal C1 is input to or transmitted to the medical device 1 via output terminal 32.

[0156] Figure 4A flowchart illustrating the verification signal VS for determining the first voice command SSB1 in one embodiment of the invention is shown. Here, the first voice command SSB1 is confirmed without further user interaction. In other words, the present invention enables particularly simple and rapid operation of the medical device 1 using only an initial voice input in the sense of an audio signal E1.

[0157] exist Figure 5 The steps shown are not necessarily in a pre-defined order; they can be customized. Figure 3 This is performed during step S50. Accordingly, step S50, which determines the verification signal VS, includes the following steps.

[0158] Step S51 involves analyzing the audio signal E1 to provide a second speech analysis result SAE2, while step S52 involves recognizing a second speech command SSB2 based on the second speech analysis result SAE2.

[0159] Steps S51 to S54 are implemented by software P2, which is stored in data storage 4 and causes computing unit 3, especially module M2, to execute the steps. Software P2 may include a second computational linguistics algorithm P2, which includes a second training function. The second training function, like software P1, is also applied to audio signal E1, and is particularly used for analyzing audio signal E1.

[0160] Step S51 can also include the tokenization of audio signal E1 and / or semantic analysis of audio signal E1. Accordingly, the second speech analysis result SAE2 can also include second tokenized information or second semantic information.

[0161] The second computational linguistics algorithm P2 can advantageously include a second training function. A key feature of the invention is that the first and second training functions are different from each other. In particular, the second training function is configured to recognize only safety-critical voice commands in the audio signal E1, especially voice commands designed to prevent first-failure.

[0162] The first training function is configured to recognize a very broad vocabulary of various voice inputs from the operator, and especially of voice commands not related to safety. In this sense, the second training function is specific to voice commands related to safety. Therefore, the second training function does not recognize voice commands not related to safety or general voice input. In this regard, at least step S52 is specific to recognizing voice commands related to safety as the second voice command SSB2.

[0163] Safety-critical voice commands are characterized by a combination of representative features, such as frequency patterns, amplitude, and modulation. Furthermore, safety-critical voice commands, especially those designed to prevent first-time failures, typically have a minimum of three or more syllables. The more syllables a voice command has, the easier it is to distinguish it from other voice commands, thus minimizing the risk of confusion.

[0164] The special features of steps S51 and S52 are achieved by constructing a second training function that differs from the first training function. A significant difference between the first and second training functions may lie in the training data used. The first training dataset used for the first training function includes a broad and diverse vocabulary concerning various different speech commands or general speech inputs. The second training dataset used for the second training function, however, is limited to vocabulary relating to specific, particularly clearly defined, speech commands, in order to enable the precise identification of safety-critical, especially first-failure-prevention speech commands with a low error rate. In this respect, according to the present invention, the second training function can be trained using a small training vocabulary, and conversely, the first training function can be trained using a large training vocabulary.

[0165] For example, a first training dataset for two functions or operating steps of medical device 1 can include a large number of voice inputs. The first function can involve the safety-critical initiation of X-ray image recording. The input data of the first training dataset can include, for example, voice inputs such as: “New Scan!”, “Please start a new scan!”, “Let’s do a new acquisition now!”, “Will you perform a scan from me?”, etc. The corresponding output data of the first training dataset is limited to the voice command “Perform acquisition!”. Therefore, the first training function is trained such that it classifies all these voice inputs into the same command category, which includes the voice command “Perform acquisition!”. The first training function can also be configured here for further generalization such that it also assigns, for example, the voice input “Let’s do a new performance for me!” to the aforementioned voice command.

[0166] The second function, for example, can relate to the non-safety-critical activation of the ventilator of medical device 1. The input data of the first training dataset can, for example, include the following voice inputs: “Start fan!”, “Switch on the fan!”, “OK, activate the fan!”, etc. The corresponding output data of the first training dataset is limited to the voice command “Start fan!”. The first training function is also trained to classify all these voice inputs into another command category, which includes the voice command “Start fan!”.

[0167] Accordingly, the first training data can include a large amount of additional input data and output data involving additional, and especially arbitrary, voice commands and other inputs.

[0168] The second training function is specifically configured to identify structurally and / or lexically explicit speech instructions as safety-critical speech instructions based on second tokenized information and / or second semantic information. In sub-step S51, structural and / or lexical characteristics of the audio signal E1 are determined. Representational, explicit structural and / or lexical characteristics are associated with safety-critical speech instructions so that they can be easily and reliably identified.

[0169] In this sense, in a particularly preferred embodiment, the second training function is configured to identify voice commands having at least three, preferably four syllables, as safety-critical voice commands, thereby reducing the risk of confusion with safety-critical voice commands. Therefore, particularly preferably, the safety-critical operating steps of the device are assigned voice commands comprising at least three syllables, and more preferably, voice commands comprising five to seven syllables.

[0170] Therefore, the second training function is trained to classify only specific speech inputs as safety-critical speech commands. Specifically, the second training function can be trained to recognize one or more keywords or word sequences.

[0171] Therefore, the second training dataset includes exactly one voice input corresponding to the input data for the output data corresponding to the safety-critical voice commands. According to the invention, the voice input is now selected as structurally, lexically, and / or phonetically defined to avoid confusion with other voice inputs.

[0172] In the example above regarding the safety-critical function of initiating X-ray image recording with medical device 1, the second training dataset now includes "Perform Acquisition" as voice input and output data corresponding to the safety-critical voice command "Perform Acquisition". Particularly preferably, the voice input can include trigger or initiation terms, such as in the sense of "Siemens, perform Acquisition", where Siemens represents the specific trigger term.

[0173] An alternative, feasible voice command to trigger X-ray radiation could be: “Start Scan!” However, this is highly susceptible to confusion due to its acoustic similarity to “Start fan!” Another alternative voice command might be “New Scan!” However, this is also easily confused with two syllables and could be recorded as part of general speech input and incorrectly identified as a voice command, for example, as part of the speech input “When the patient is in we're doing a new scan immediately.”

[0174] The voice command "Perform acquisition" is particularly well-suited as a safety-critical voice command due to its acoustic characteristics and the fact that it is difficult to confuse due to its seven syllables.

[0175] According to the present invention, the training functions themselves can also be different, particularly in constituting the verification function or the classification function. The first training function has a large number of categories corresponding to a wide variety of voice commands. The second training function, however, is limited to a smaller number of categories corresponding to small, especially medical device-specific, safety-critical voice commands. The present invention reduces the risk of similar system errors occurring when recognizing voice commands using the first and second training functions by using different types of training functions.

[0176] If in step S52 the second training function identifies a voice instruction as a second voice instruction SSB2 derived from the second language analysis result SAE2, then it is itself a voice instruction of security concern.

[0177] By constructing the first and second training functions in different types, in this implementation scheme, the risk of executing the first voice command SSB1 that is incorrectly recognized in module M1 is minimized, because the system voice command recognition error of the first training function is unlikely to be repeated by the second training function.

[0178] In step S53, now according to the consistency criteria The consistency between the first voice command SSB1 and the second voice command SSB2 is checked. Here, a consistency measure between the two voice commands is compared with a predetermined threshold specific to the first and / or second voice commands SSB1 / SSB2. The threshold, or the required similarity measure, can vary significantly depending on the voice command. According to the invention, the threshold for voice commands with a high security level, i.e., the threshold for voice commands that particularly demonstrate protection against first-time failure, is set to be greater than that for voice commands with a lower security level.

[0179] If test step S53 determines that the first voice command SSB1 and the second voice command SSB2 have a consistency measure that is near or above the threshold, then the first voice command SSB1 is confirmed.

[0180] In step S54, module M2 of the voice control device 10 also generates a verification signal VS based on the test results in step S53. The verification signal VS is then transmitted to module M3 corresponding to the control unit of the voice control device 10, which (in...) Figure 3 In step S60, a control signal C1 is generated based on the first voice command SSB1 confirmed by the verification signal VS.

[0181] In the example above of the safety-critical voice command used to initiate X-ray image recording, the medical device will therefore only trigger X-ray radiation when the operator says "Perform acquisition" as voice input.

[0182] Figure 5 Showing the ability to be based on Figures 3 to 4 The method uses an artificial neural network 400. Specifically, the shown neural network 400 can be a first or second training function for a first or second computational linguistic algorithm P1, P2. The neural network 400 operates on multiple input nodes x. i 410 responds to the input values ​​applied to it in order to produce one or more outputs. j In this embodiment, the neural network 400 adjusts the weighting factor w of each node based on the training data. i Learning is done using weights. Input node x iPossible input values ​​for 410 could be, for example, speech input or audio signals from a first or second training dataset. The neural network 400 weights 420 on the input values ​​410 based on a learning process. In one embodiment, the output value 440 of the neural network 400 corresponds to the first or second speech command SSB1, SSB2. In other embodiments, the output value 440 of the neural network can also include an indication of the security category SK of the first speech command SSB1 or a check result of a consistency criterion between the first speech command SSB1 and the second speech command SSB2. The output 440 can be transmitted via one or more output nodes. j conduct.

[0183] The artificial neural network 400 preferably includes multiple nodes h j The hidden layer is 430. Multiple hidden layers h can be configured. jn Hidden layer 430 uses the output value of another (hidden) layer 430 as its input value. Nodes in hidden layer 430 perform mathematical operations. Node h j The output value here corresponds to its input value x. i and weighting factor w i The nonlinear function f. When the input value x is obtained... i After that, node h j For each input value x i The weighted factor w i The weighted product is summed, as determined by the following function:

[0184] h j =f(∑ i x i ·w ij )

[0185] In particular, node h j The output value is formed as a function f for node activation, such as a sigmoid function or a linear ramp function. Output value h j Transmitted to one or more output nodes j The node activation function f is used to recalculate the value h for each output value. j The sum of weighted products:

[0186] o j =f(∑ i h i ·w′ ii )

[0187] The neural network 400 shown here is a feedforward neural network, as is preferably used in the first computational linguistics algorithm P1, wherein all nodes 430 process the output values ​​of the previous layer into input values ​​in the form of a weighted sum. According to the invention, other neural network types are particularly capable of being used in the second computational linguistics algorithm, such as feedback, e.g., recurrent neural networks, wherein nodes h j The output value can also be its own input value.

[0188] The neural network 400 is preferably trained using a supervised learning method to recognize patterns. A known method is backpropagation, which can be applied to all embodiments of the invention. During training, the neural network 400 is applied to training input values ​​and must produce corresponding, pre-known output values. The mean square error (MSE) between the calculated output value and the expected output value is iteratively calculated, and the various weighting factors 420 are adjusted until the deviation between the calculated output value and the expected output value is below a predetermined threshold.

[0189] If it has not yet clearly occurred, but is meaningful and in the sense of the invention, the various embodiments, their various sub-aspects or features can be combined or interchanged with each other without departing from the scope of the invention. Where applicable, the advantages described with reference to one embodiment of the invention also apply to other embodiments unless explicitly mentioned.

[0190] This invention enables the use of computational linguistic algorithms, including neural networks, for speech recognition and voice command derivation in security-related applications, particularly those designed to prevent first-time failures. Testing demonstrates that the described solution is more reliable than conventionally implemented speech recognition algorithms lacking security features. In embodiments, the invention is characterized by scalability regarding user-friendliness and / or security (failure detection). On one hand, this allows for the appropriate protection of voice commands at different security levels. On the other hand, it is feasible to improve operability over time if experience shows that voice command recognition or verification in a defined target environment corresponds to the desired security specifications and can demonstrate sufficient reliability. Alternatively, it is feasible to improve security if it is recognized that, for example, due to specific environmental conditions (e.g., background noise), reliable recognition of commands transmitted via voice input does not work sufficiently well.

[0191] Scalability typically enables gradual changes to the safety aspects of the implementation, both in conventional and machine learning-based parts. This provides a migration path toward the use of safety-tolerant computational linguistic algorithms, including neural networks.

[0192] Finally, it should be noted that the construction scheme according to the invention is capable of applying verification mechanisms for misinterpretations of control commands derived by means of speech recognition (incorrect voice commands are assigned to voice input including user intent) or for untimely, i.e., slow, recognition of voice commands. However, according to the invention, failures to recognize voice commands in voice input despite including user intent cannot be prevented. Therefore, the invention is not applicable to protection, for example, of emergency stop functions.

Claims

1. A method for voice control of a medical device (1), the method having the steps of: - detecting (S10) an audio signal (El) containing a voice input of an operator intending to control the device; - analyzing (S20) the audio signal for providing a first speech analysis result (SAE1); - identifying (S30) a first speech command (SSB1) of the operator based on the first speech analysis result; - determining (S50) a verification signal (VS) for the first speech command, comprising: - analyzing (S51) the audio signal for providing a second speech analysis result (SAE2); - identifying (S52) a second speech command (SSB2) of the operator based on the second speech analysis result; - comparing (S53) the first speech command and the second speech command, wherein the verification signal confirms the first speech command when the first speech command and the second speech command comply with a consistency criterion, - generating (S60) a control signal (Cl) for controlling the medical device based on the verification signal, if the first speech command is confirmed according to the verification signal, wherein the control signal is adapted to control the medical device according to the first speech command; and - inputting (S70) the control signal into the medical device, wherein - the analyzing (S20) for providing a first speech analysis result comprises applying a first computer linguistics algorithm (Pl) to the audio signal, the first computer linguistics algorithm (Pl) comprising a first training function, and - the analyzing (S51) for providing a second speech analysis result comprises applying a second computer linguistics algorithm (P2) to the audio signal, the second computer linguistics algorithm (P2) comprising a second training function, wherein the first training function and the second training function differ from each other, wherein the second training function is configured to identify only safety-critical speech commands.

2. The method of claim 1, wherein, The analyzing of the audio signal comprises tokenization for segmenting letters, words and / or sentences within the audio signal by means of the first computer linguistics algorithm and the second computer linguistics algorithm, and identifying the first speech command and the second speech command based on the first speech analysis result comprising first tokenization information and the second speech analysis result comprising second tokenization information.

3. The method of claim 1 or 2, wherein analyzing the audio signal comprises: The audio signal is semantically analyzed by means of the first computer linguistics algorithm and the second computer linguistics algorithm, and the first speech command and the second speech command are recognized based on the first speech analysis result comprising first semantic information and the second speech analysis result comprising second semantic information.

4. The method according to claim 2, wherein the second training function is configured to recognize as safety-critical speech commands speech commands that are structurally and / or lexically unambiguous according to the second tokenization information.

5. The method according to claim 3, wherein the second training function is configured to recognize speech commands which are structurally and / or lexically unambiguous as safety-relevant speech commands depending on the second semantic information.

6. The method according to claim 3, wherein the second training function is configured to recognize speech commands having at least three syllables as safety-relevant speech commands.

7. The method according to claim 1 or 2, further comprising the steps of: - classifying (S40) the first speech command into one of a plurality of safety categories, wherein at least one of the plurality of safety categories is provided for safety-relevant speech commands; - determining a verification signal for the first speech command if the first speech command is assigned to a safety category of safety-relevant speech commands.

8. The method according to claim 1 or 2, wherein detecting the audio signal comprises the steps of: - detecting the audio signal by means of a first input device (2) and a second input device (20), wherein the first input device and the second input device are different from each other; - providing the audio signal detected by means of the first input device for analyzing and providing the first speech analysis result; and - providing the audio signal detected by means of the second input device for analyzing and providing the second speech analysis result.

9. A speech control device (10) for speech control of a medical device (1), the speech control device comprising: - an interface (31) for detecting an audio signal (El) containing a speech input of an operating person intended to control the device; - a first evaluation unit (Ml) configured to - analyze the audio signal and provide a first speech analysis result (SAE1), - identify a first speech command (SSBl) based on the first speech analysis result; - a second evaluation unit (M2) configured to - determine a verification signal (VS) for the first speech command by: - analyzing the audio signal (El) to provide a second speech analysis result (SAE2); - identifying a second speech command (SSB2) of the operating person based on the second speech analysis result; - comparing the first speech command and the second speech command, wherein the verification signal confirms the first speech command if the first speech command and the second speech command comply with a conformity criterion (UK); - a control unit (M3) configured to - generate a control signal (Cl) for controlling the medical device based on the verification signal if the first speech command is confirmed according to the verification signal, wherein the control signal is adapted to control the medical device depending on the first speech command; and - an interface (32) for inputting the control signal into the medical device, wherein - the first speech command is a speech command of the operating person intended to control the medical device. ​ - the analysis (S20) for providing the first speech analysis result comprises applying a first computer linguistics algorithm (P1) to the audio signal, the first computer linguistics algorithm (P1) comprising a first trained function, and - the analysis (S51) for providing the second speech analysis result comprises applying a second computer linguistics algorithm (P2) to the audio signal, the second computer linguistics algorithm (P2) comprising a second trained function, wherein the first trained function and the second trained function are different from each other, wherein the second trained function is configured to identify only safety-critical speech commands.

10. The speech control device according to claim 9, wherein the first evaluation unit and the second evaluation unit are configured as separate computing units from each other.

11. The speech control device according to claim 9 or 10, wherein the interface (31) comprises a first input device and a second input device (2, 20) for simultaneously detecting the audio signal.

12. A medical system (1), comprising: - a speech control device (10) according to any one of claims 9 to 11; and - a medical device (1) for performing a medical procedure.

13. A computer program product comprising a program and being directly loadable into the memory of a programmable computing unit, the computer program product having program means to perform the method according to any one of claims 1 to 8 when the program is executed.

14. A computer-readable storage medium on which a readable and executable program section is stored to perform all steps of the method according to any one of claims 1 to 8 when the program section is executed. ​

Citation Information

Patent Citations

  • Medical system with a voice input device

    DE102006045719B4

  • Voice recognition method and device

    CN101807399A

  • Voice recognition method and device and electronic equipment

    CN111816165A