Method for controlling a machine or a machine element, and control arrangement

EP4609364A1Pending Publication Date: 2025-09-03SENSOR TECHN WIEDEMANN GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023794033
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-24
Filing Date
2023-10-20
Publication Date
2025-09-03

AI Technical Summary

Technical Problem

Current control systems for complex machines face delays in recognizing human gestures, leading to unreliable and inefficient human-machine interaction, as users receive no immediate feedback on whether their commands are interpreted correctly.

Method used

A method utilizing an image capture system, processing device, and neural network to capture and analyze multiple images, determine image and overall confidence values, and output signals for recognized gestures, providing feedback to users and ensuring control commands are issued only when confidence thresholds are met, thereby improving recognition speed and reliability.

Benefits of technology

This approach enhances the speed and reliability of gesture recognition, allowing users to adjust their gestures based on immediate feedback and ensuring accurate and timely execution of control commands, improving overall human-machine interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method (S1-S8) for controlling a machine or a machine element (6) by means of predefined objects (11, 11a, 11b) or gestures detected by a processing device (17), in which, during the identification and categorisation of a gesture, the progress is signalled to a user in the form of an overall confidence value, wherein in some cases, a predefined confirmation object must be additionally detected for the execution of a control command (23). The invention also relates to a control arrangement (1) for carrying out such a method, a computer arrangement and a computer program product. The progress of detecting objects, in particular gestures, which trigger certain control commands of the machine, can be transmitted for example by a display in a display apparatus (5, 5a, 5b), or alternatively or additionally by acoustic or haptic signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR CONTROLLING A MACHINE OR A MACHINE ELEMENT AND CONTROL ARRANGEMENT

[0002] The present invention is in the field of control technology and electronics and relates to a method for controlling a machine or a

[0003] The invention further relates to a control arrangement, a computer arrangement, and a computer program product. In the field of robotics, but also generally for controllable machines that have a certain complexity or are used in a complex environment, an efficient design of human-machine interaction is required for

[0004] REPLACEMENT SHEET (RULE 26) Control of complex processes and devices is becoming increasingly important. Modern user interfaces include not only keyboard inputs and inputs from other hardware elements such as joysticks and pads, but also additional input channels. These include commands that a user transmits to a machine not via manual input, but rather with gestures or vocalizations.

[0005] For this purpose, machines should recognize such commands as autonomously as possible, for example when there is imminent danger or when a user is performing additional activities at the same time.

[0006] From the perspective of a user interacting with a machine, the main focus is on ensuring that his commands, such as gestures, are recognized reliably and as quickly as possible.

[0007] A common problem with control via optical signals, such as gestures, is the time delay between the human signal and the recognizable execution of a corresponding control command. Depending on the algorithm used, the time delay between the gesture and the corresponding machine action can range from several hundred milliseconds to a few seconds.

[0008] This delay can lead to the user aborting or changing the gesture because he or she does not receive feedback from the machine and cannot therefore judge whether the signal was not interpreted or was interpreted incorrectly.

[0009] The present invention is therefore based on the object of improving the speed and reliability of command recognition when controlling a machine or a machine element and thus the interaction between man and machine.

[0010] The problem is solved with the features of the invention by the subject matter of the independent patent claims, i.e., by a method, a computer arrangement, a computer program product, and a control arrangement. The dependent patent claims specify possible embodiments.

[0011] REPLACEMENT LEAF (RULE 26) Thus, the invention relates to a method for controlling a machine or a machine element, comprising:

[0012] Capturing a plurality of images sequentially by means of an image capturing system; in particular, outputting all or some of the plurality of captured images by an image display device;

[0013] capturing at least a subset of the plurality of captured images by a processing device;

[0014] Identifying at least one first predetermined object on a plurality of successively recorded images of the subset of captured images, in particular by a network trained by machine-based learning,

[0015] Classifying the or an identified first predetermined object for a plurality of successively recorded images into one of a plurality of categories, each of which is assigned to a predetermined object or a group of predetermined objects, wherein each category is assigned a defined control command for controlling the machine or the machine element, in particular via the network;

[0016] Outputting a signal associated with a first predetermined object identified and / or categorized, in particular an optical, acoustic or haptic signal, further in particular displaying an image or a symbol in an image display device;

[0017] Determining an image confidence value for a plurality of consecutively taken images, which indicates the certainty or probability with which a first predetermined object was identified in the respective image and assigned to a category;

[0018] Repeatedly determining an overall confidence value for the first predetermined object from the image confidence values ​​of several consecutively recorded images, in particular taking into account their temporal arrangement;

[0019] Outputting an overall confidence value, in particular by means of an optical, acoustic or haptic signal, in particular by means of a display in an image display device in which the image or symbol associated with the categorised first predetermined object is also displayed,

[0020] Comparing the overall confidence value with a threshold and

[0021] REPLACEMENT SHEET (RULE 26) Generating and issuing the control command associated with the category into which the identified first predetermined object has been classified, under the condition that the overall confidence value exceeds a predetermined threshold value, wherein in particular the signal associated with the identified and / or categorized first predetermined object is changed upon issuing the control command.

[0022] For reliable detection of optical signals in the form of recognizable objects and, consequently, correct activation of control commands, several criteria must be met. For example, the respective object should be recognized across several consecutively recorded images / frames, and the position, shape, and size of the recognized object should only change by a certain amount at most. Objects can be predetermined gestures, represented by images or image sequences, that a user can perform, for example, with one or both hands.

[0023] The method in question transmits information about the identification of the gesture and the corresponding categorization to the user while they are performing the gesture. In this way, the user can either adapt or slightly modify their gesture and also maintain it for the necessary period of time until the machine starts executing the control command intended by the gesture. Accordingly, in one method, a plurality of images are recorded using an image recording system, from which the user's gestures are identified and categorized. In addition, the recorded images are optionally also output again, e.g. on an image display device, for example on a stationary screen, on a screen of a mobile device or in a heads-up display device, so that the user can see their gesture and also adapt it if necessary.This improves gesture recognition by providing feedback to the user.

[0024] The image acquisition system may comprise multiple cameras or scanners, which may also be spaced apart and may capture the user from different viewing angles. The different images captured simultaneously by different cameras may be used for a

[0025] REPLACEMENT SHEET (RULE 26) Three-dimensional analysis of a gesture can be combined before, during, or after identification or categorization to achieve greater recognition reliability or to enable identification of certain gestures in the first place. The images captured in the context of 3-dimensional capture can also be understood as 3-dimensional representations of objects / gestures.

[0026] All captured images, or a subset thereof, are fed to a processing device, which may contain an analysis device with a trained algorithm, for example, a trained neural network. For example, a selection of the captured images can be fed to the analysis device that exhibits particularly good image quality and / or corresponds to an image set that can be easily processed by the analysis device without delay due to its capacity. For this purpose, for example, only a specified portion of the captured images can be fed to the processing device.

[0027] If the analysis device includes a neural network for identifying and categorizing graphic objects, this can advantageously be a CNN, a so-called convolutional neural network, since such networks are particularly suitable for identifying objects in image files and subsequent categorization due to the selection, design and number of their layers layered between the input side and the output side.

[0028] The analysis device or a self-learning element within the analysis device can be trained with a large number of images containing gestures of varying clarity and image quality. The intended gesture and a confidence value are then provided for each image in the analysis device's training mode. The network trained in this way can then recognize and categorize objects / gestures. The confidence value can also be determined by the neural network mentioned above or by a second neural network trained with images of objects / gestures that are known to it and that are represented in images in a more or less distorted manner by specific image processing tools in a controlled manner.

[0029] REPLACEMENT SHEET (RULE 26) During use, the analysis device first detects that a gesture is present, for example, when a user pauses their natural movement and raises at least one hand. Identification can be achieved through pattern recognition.

[0030] In this context, gestures are primarily intended to be stationary signs that a user can form with body parts, especially the hands, but also, for example, with the head. However, the method can also refer to certain elementary movements that are recognizable as moving signs or gestures (so-called moves). These can also be identified using optical pattern recognition, whereby the movement patterns in the form of several consecutively recorded images are compared with corresponding image sequences from a storage device, or corresponding short movement sequences are trained in a neural network.

[0031] Once the analysis device has identified a gesture, it compares it with predetermined objects / categories of objects stored in a memory device and determines similarities with these objects. If the similarity of an identified object to a predetermined stored object or category clearly outweighs the similarity to other predetermined objects, the identified object is assigned to the respective category, to which a control command is then linked.

[0032] Depending on the trained algorithm / network used, the steps of identifying an object and classifying it into a category can be performed jointly and simultaneously or sequentially using a CNN, a convolutional neural network, which, due to its number of neuron layers, can in many cases be capable of "deep learning."

[0033] The similarity metrics used for such comparisons are well-known from image processing. Size and orientation deviations—that is, rotations of the identified object relative to a stored representation—are allowed within a predetermined range and compensated for in the similarity analysis. CNN

[0034] REPLACEMENT SHEET (RULE 26) have proven to be translation invariant in this context and can compensate for shifts in the recognized objects particularly well. From the similarity or degree of agreement between a recognized and identified object / gesture, which is determined from one or more simultaneously recorded images, and a stored reference object of a specific category into which the recognized object is assigned, an image confidence value is determined. This value can be proportional to the similarity or directly identical to the similarity. When determining the image confidence value, image quality that is independent of the object (e.g. the clarity of the gesture) can also be taken into account; this quality can depend on weather conditions and brightness, for example. An image confidence value refers only to one image or a group of images recorded simultaneously or almost simultaneously.The confidence value reflects the probability determined by the network that—based on the ground truth—the categorized object corresponds to the true object and thus to the correct gesture. CNNs have a layer in which the analysis results are categorized, and the assignment to the possible categories is expressed in the weights of the individual elements of the layer. Thus, the reliability of the categorization, related to an image, can be directly measured in such a network based on the initial position of nodes. In this context, the English term "layer" is used as an established technical term for the position of nodes in a network.

[0035] In parallel to the actual analysis process, the user can also be continuously signaled that and which predetermined gesture is currently being recognized and / or when the image confidence value, i.e. the reliability value of the recognition of the gesture in relation to an image or several images taken simultaneously, exceeds a predetermined threshold for the first time, in particular the threshold set for issuing a control command.

[0036] A further embodiment of the method can provide that an overall confidence value is determined taking into account the number of images on which the first predetermined object was categorized with an image confidence value above a predetermined threshold, taking into account

[0037] REPLACEMENT SHEET (RULE 26) counting the number of further images between these images in which the first predetermined object was not categorized with an image confidence value above the threshold.

[0038] To incorporate the consistency of a gesture or a recognized object into the assessment of recognition reliability, an overall confidence value is determined from several image confidence values ​​relating to consecutively captured images, which links the image confidence values ​​together. Since image confidence values ​​are continuously determined while images are being captured, it is useful to repeatedly evaluate the most recently determined image confidence values ​​and to continuously generate updated overall confidence values ​​from these in order to monitor when an overall confidence value is high enough to trigger a control command. In one embodiment, the overall confidence value reflects the probability determined by the trained network that the identified and categorized object consistently corresponds to a shown gesture / object.

[0039] The determination of the overall confidence value can be performed in an analytical step that further processes the results of the neural network's output according to a fixed algorithm. However, in another implementation, the determination of the overall confidence value can also be performed within the neural network. This would require an internal state memory within the neural network in which data is temporarily stored for further processing by the network.

[0040] To determine an overall confidence value, for example, a specified number of the most recently determined image confidence values ​​of captured images can be linked, or the image confidence values ​​determined in a sliding time window can be linked. A threshold can be set for the image confidence values, above which the image confidence values ​​are processed. If a certain number or a certain proportion of image confidence values ​​above the threshold are reached in a specified time unit or within a specified number of consecutive image confidence values, the overall confidence value exceeds the threshold required for issuing a control command. The temporal arrangement of the image confidence values ​​that determine the

[0041] REPLACEMENT SHEET (RULE 26) threshold are taken into account for determining the overall confidence value. For example, in order to reach the threshold of the overall confidence value, the condition can be set that between the number of image confidence values ​​that are each above the required image confidence threshold, there must be fewer than a predetermined number or a predetermined proportion of image confidence values ​​below the threshold. Individual image confidence values ​​below the threshold may be permissible, for example, to compensate for incorrect measurements. Time-related conditions can also be set for reaching the threshold of the overall confidence value, for example the condition that a certain number or a certain proportion of image confidence values ​​must be above the image confidence threshold within a predetermined time.

[0042] If the determined overall confidence value reaches or exceeds the threshold required for issuing a control command, a signal can be sent to the user informing them that the object / their gesture has been reliably recognized. The signal can be provided, for example, via a display device or an acoustic signal generator, or in haptic form, for example through a vibration signal on a wearable, a glove equipped with vibration elements, or a control element, such as a joystick. The recognized gesture can also be transmitted, for example displayed on an image display device in the form of a standardized symbol for the recognized gesture or in the form of the captured images in which the recognized gesture is marked by a frame or information symbols, or transmitted via an announcement.When signaling is provided in acoustic or haptic form, distraction of the user's gaze, which may be directed toward a working situation of the machine being controlled, is avoided. When using an image display, this can be implemented, for example, on a mobile device or in the form of a head-up display to avoid distraction of the user's gaze.

[0043] This provides the user with feedback about the identified object, which in some cases has already been categorized. Additionally, the current image confidence value or an overall confidence value can be signaled, for example, via an image or a

[0044] REPLACEMENT SHEET (RULE 26) acoustic or haptic signal. The user receives information about which of the objects depicted in the images the network has identified as a gesture / object, and which gesture this object should represent after categorization. This allows the user to intervene early if, for example, the object is incorrectly identified, the wrong gesture is categorized, or identification repeatedly fails due to external parameters (e.g., hand position, lighting, etc.). As soon as the overall confidence value exceeds the threshold for issuing a control command, a separate signal can be displayed or transmitted acoustically or haptically.

[0045] In some embodiments, a symbol associated with the issued control command is displayed on the screen. This can be, for example, the gesture that was recognized or a label or a brief description of the control command to be executed. Likewise, the symbol associated with the identified object can be changed on the screen after a defined period of time immediately upon issuing the control command and / or removed again after the assigned control command has been issued. This indicates to the user that the system is available for a new gesture and / or that the previously recognized gesture for initiating the previously executed control command is not currently available.

[0046] Furthermore, it can be provided that if the total confidence values ​​assigned to the categorized object fall below the threshold for a certain number of the captured images, the symbol associated with the identified object is removed from the screen. However, in some cases, it may not be practical to do this if the total confidence value falls below a threshold in a single instance. Rather, the proposed method can exhibit a certain tolerance for miscategorization or categorization with low image confidence values.

[0047] The number of images from the subset of captured images where an identified and categorized object with the image confidence value must be above the threshold in order to have a sufficient control command output.

[0048] REPLACEMENT SHEET (RULE 26) Achieving the overall confidence value may in some cases depend on the category of the object or on the categorized object itself. Accordingly, it is possible that certain categorized objects trigger an associated control command with a higher priority or more quickly than others. For example, for a categorized object to which a stop control command is assigned, the number of images in which the identified object must be categorized with a sufficiently high image confidence value for a control command to be issued may be significantly smaller (e.g., at least 30% or at least 60% smaller), or the time over which a certain image confidence value must be achieved may be significantly shorter (e.g., at least 30% or at least 60% shorter), than for gestures linked to other control commands.

[0049] In addition, when detecting an object / gesture in a single image, a tolerance can be provided that allows the object to be detected even if the object is optically magnified or reduced by a certain factor by changing the distance from an image acquisition system or by a lateral shift.

[0050] The value of the still acceptable magnification or reduction and the value of the acceptable displacement can be proportional to the size of the object or the size of a frame placed around the object. These tolerances and controls are important so that the gesture recognition has an indication of whether a single input or multiple or changing users is being input. In some cases, objects, such as hand gestures, that are in the center of the captured image can be preferentially recognized, i.e., identified, or even categorized. In these cases, the neural network is primarily trained to identify and categorize objects in the center of the captured images.Therefore, a subset of the recorded images that are captured by the processing device and further processed by the neural network can be understood not only as a selection of images but also as a selection of sections from all images or from a subset of images.

[0051] Another aspect concerns blocking further control commands until the previous control command, whose gesture was recognized, has been processed.

[0052] REPLACEMENT SHEET (RULE 26) also provides for indicating to the user that the gesture has been blocked. Accordingly, in some cases, the symbol associated with the identified or categorized object changes in the image display immediately after the control command is issued, and the object / gesture in question is subsequently blocked from recognition, at least until the control command is processed.

[0053] The term "lock" means that further identification or categorization of a specific gesture is not displayed and cannot lead to the issuance of a control command. In this way, the user is informed that the recognized control command corresponding to this gesture has not yet been completed. Accordingly, in these cases, a further identification or categorization of an object is not displayed, even in cases where this is still being executed in the background. In other cases, gesture recognition can also be interrupted for one or more gestures until the control command has been processed.

[0054] The block can be communicated to the user in various ways, e.g., by displaying another character on a display device or by further changing the image associated with the identified and categorized object. For example, after correct recognition, an icon of the recognized gesture can be displayed, whose color changes after the control command is issued. The changed color then remains as long as the control command is executed.

[0055] In some cases, however, gesture identification and categorization can also continue in the background. This is useful when there are objects or gestures that are associated with, for example, the "Stop" control command or an equivalent abort command. This ensures that a user can interrupt the machine's movement or a function being performed at any time with another gesture.

[0056] In one embodiment of the method, it can be provided that a control command is only issued under the additional condition that the processing device, in addition to the first predetermined

[0057] REPLACEMENT SHEET (RULE 26) Object, a confirmation object different from the first predetermined object is recognized with at least a predetermined confirmation total confidence value.

[0058] This sufficiently ensures that a control command is only issued when intended and confirmed by a confirmation gesture following the originally recognized and categorized gesture. A control command cannot then be issued accidentally simply because a certain gesture is unintentionally held. This reliably prevents the risk of incorrect operation and allows compliance with safety regulations.

[0059] In addition, it can be provided that the recognition, specifically the identification and classification of a confirmation object into a category / categorization, is only permitted after the output of a signal associated with the identified and / or categorized first predetermined object, in particular only under the additional condition that the overall confidence value for the first predetermined object / the first predetermined gesture exceeds a predetermined threshold value.

[0060] This prevents a confirmation gesture from being made prematurely, for example at the same time as a gesture to select a control command, and thus negligently, or from leading to identification and classification in a category.

[0061] If the threshold set for issuing a control command has already been exceeded by the overall confidence value of the first predetermined object, the time at which the control command is issued can be determined using the confirmation command—i.e., by pointing a confirmation object or a confirmation gesture. This can be particularly convenient if the threshold of the overall confirmation confidence value—i.e., the overall confidence value set for the confirmation object / confirmation gesture linked to a confirmation command—is lower than the thresholds for other predetermined objects. This makes it likely that the time delay in recognizing a confirmation object will be short.

[0062] REPLACEMENT SHEET (RULE 26) If a first object / gesture has already been detected and the overall confidence value of the first predetermined object exceeds the threshold set for issuing a control command, the analysis device can also switch to an operating mode in which it exclusively attempts to identify or recognize a confirmation object or a confirmation gesture. This reduces the processing effort and response time for processing subsequent images. The identification of a confirmation object and its classification into a category—in this case, specifically the category of confirmation objects—is simplified and accelerated.

[0063] Furthermore, it can be provided that, after outputting a signal associated with the identified and / or categorized first predetermined object or after it has been determined that an overall confidence value determined for the first predetermined object has exceeded a predetermined threshold value, a request signal for showing a confirmation object / a confirmation gesture is output, wherein the request signal is output in particular as a character in an image display device and is selected in particular depending on the control command to be issued or the categorized first predetermined object.

[0064] Such an option also increases operating safety and protects against confusion between different gestures / objects.

[0065] It can also be provided that the request signal is displayed on the same image display device in which the recorded images are displayed, in particular on a partial area of ​​the image display device separated from the display of the recorded images.

[0066] This eliminates the need for the user to look away from a display device, where they might see their gesture or an associated control command, to recognize the prompt. Another option might be to provide the prompt signal as an acoustic or haptic signal, so in these cases, there's no need to look away from the observed event or a display device.

[0067] REPLACEMENT LEAF (RULE 26) Furthermore, a method can be provided in which the image on the screen associated with the identified predetermined object has a surrounding frame or a strip, in particular in the form of a progress bar, wherein optionally the indexing of the overall confidence value is carried out by displaying at least one of the following features:

[0068] A number, especially a percentage, the size of which depends on the current overall confidence level;

[0069] A width of the frame that depends on the overall confidence value;

[0070] A color or brightness of the frame that depends on the overall confidence value;

[0071] Fading the frame in and out at a variable frequency depending on the overall confidence value;

[0072] The length or width, color, brightness, or a frequency of change of the display of a stripe that depends on the overall confidence value;

[0073] A combination of several of the above features.

[0074] The frame can have a rectangular, round or oval shape and, in particular, a closed shape, and the strip can be straight or curved and, for example, when it has reached its full length, completely surround the object depicted.

[0075] This provides transparent display options for the user, allowing them to intuitively perceive the current overall confidence value.

[0076] In the case of signaling via acoustic or haptic signals, these can be switched on and off with a modulatable frequency depending on the overall confidence value, or the frequency of a vibration or an emitted signal tone can be changed depending on the overall confidence value.

[0077] It may further be provided that an identified object is categorised the same in two consecutive images if the position of the identified object in the second image does not change by more than a specified number or

[0078] REPLACEMENT LEAF (RULE 26) a number of pixels depending on the extent of the representation of the first object in the image display device deviates from the position in the first image.

[0079] This simplifies the operation of the analysis system, as it does not have to identify and categorize objects in each new image. Instead, it can first search for and locate a previously detected predefined object in a subsequent image, even if the object has moved slightly. This can be implemented, among other things, by using an algorithm for determining image correlations in addition to the processing by a neural network. This algorithm can be placed upstream of the neural network to save analysis effort.

[0080] Furthermore, it can be provided that the signal associated with an identified and / or categorized first predetermined object remains changed in the form of an image in an image display device after the output of the associated control command at least until the control command has been processed or until the same control command can be meaningfully issued again.

[0081] This measure ensures that a user can recognize that a specific control command has been issued and cannot or does not need to be issued again. This prevents the user from repeating the gesture because they cannot yet recognize that a control command has been executed or at least already started, and thus potentially causing incorrect operation of the machine. The corresponding image can be displayed on a display device, for example, transparently, with a special color, dashed lines, or grayed out.

[0082] A further embodiment of a method can provide that the issuing of a control command or the execution of a control command is interrupted if the processing device detects a demolition object in captured images with at least a predetermined total demolition confidence value.

[0083] REPLACEMENT SHEET (RULE 26) Abort objects can be provided and stored as a category of predetermined objects in a memory device of the analysis device for rapid identification and categorization. Likewise, a neural network can be trained with one or more abort objects / abort gestures. This allows a user to potentially stop or change a given control command if they detect an error during operation or in categorizing a previously identified object. A quick abort is also possible if the work situation requires it.For this purpose, it may be provided that for an abort object or abort gesture which, if identified and categorized, leads to the issuance of an abort command or to the stopping of an issued control command or a machine activity triggered by this command, a lower threshold for the image confidence values ​​and / or for an overall confidence value for issuing the abort command is set than the threshold that applies to the other possible and identifiable and categorizable objects and gestures, including the confirmation objects. This means that an abort object is recognized more quickly than other objects, albeit with less reliability. This may be permissible in scenarios where a system can be brought into a safe state by stopping all activities. In systems where hazardous situations could arise from aborting ongoing actions, such a design may be discouraged.

[0084] In addition to a method of the type described above, the invention also relates to a computer arrangement comprising: a memory with a program stored therein; at least one processor connected to the memory; wherein the program is designed such that upon execution of the program instructions by the processor, a method of the type described above is carried out.

[0085] The computer arrangement can, for example, implement one or more neural networks, in particular at least one CNN, by means of a program and contain further program parts which, outside the neural network, implement the input data and / or the output control commands and signals to the user.

[0086] REPLACEMENT SHEET (RULE 26) Furthermore, the invention also relates to a computer program product which contains program instructions stored in a memory, the execution of which by a processor carries out a method of the type described above.

[0087] Furthermore, the invention relates to a control arrangement for a machine or a machine element with an image recording system for recording a plurality of images, with an image display device and a processing device, wherein the control arrangement is set up to carry out the following method steps:

[0088] Capturing a plurality of images sequentially by means of an image capturing system; in particular, outputting all or some of the plurality of captured images by the image display device;

[0089] capturing at least a subset of the plurality of captured images by a processing device;

[0090] Identifying at least one first predetermined object on a plurality of successively recorded images of the subset of captured images, in particular by a network of the processing device trained by machine-based learning,

[0091] Classifying the or an identified first predetermined object for a plurality of successively recorded images into one of a plurality of categories, each of which is assigned to a predetermined object or a group of predetermined objects, wherein each category is assigned a defined control command for controlling the machine or the machine element, in particular via the network;

[0092] Outputting a signal associated with a first predetermined object identified and / or categorized, in particular an optical, acoustic or haptic signal, further in particular displaying an image or a symbol in an image display device;

[0093] Determining an image confidence value for a plurality of consecutively taken images, which indicates the certainty or probability with which a first predetermined object was identified in the respective image and classified into a category;

[0094] REPLACEMENT SHEET (RULE 26) Repeatedly determining an overall confidence value for the first predetermined object from the image confidence values ​​of several consecutively recorded images, in particular taking into account their temporal order;

[0095] Outputting an overall confidence value, in particular by means of an optical, acoustic or haptic signal, in particular by means of a display in an image display device in which the image associated with the categorised first predetermined object is also displayed,

[0096] Comparing the overall confidence value with a given threshold and

[0097] Generating and outputting the control command associated with the category into which the identified first predetermined object has been classified under the condition that the overall confidence value for the first predetermined object exceeds the predetermined threshold value, wherein in particular the signal associated with the identified and / or categorized first predetermined object is changed upon outputting the control command.

[0098] Such a control arrangement is capable of enabling a user to operate a machine or a machine element in a simple, transparent, fast and reliable manner.

[0099] Such a control system can be used to control vehicles or work machines, for example, when they are used in difficult environments. It could also be used by people with disabilities who are prevented from using keyboards or similar input devices.

[0100] Furthermore, it can be provided that the control arrangement is configured to output a control command only under the additional condition that, in addition to the first predetermined object, the processing device identifies a confirmation object different from the first predetermined object with at least a predetermined overall confirmation confidence value and assigns it to a confirmation category, or is generally configured to carry out a method of the type described above.

[0101] REPLACEMENT LEAF (RULE 26) In addition, the control arrangement can be provided with a display device for displaying recorded images and images associated with first predetermined objects, as well as at least one output device for acoustic signals or an output device for haptic signals.

[0102] In this way, at least two different signal channels are available for the operation of the machine or the machine part for interaction and communication between the machine and the user, either an optical and an acoustic channel or a channel for optical signals and a channel for haptic signals or all three of the above-mentioned signal channels.

[0103] The machine user can thus, for example, use a head-up display while performing gestures to monitor both the machine and the display of recognized gestures. They can also receive signals, for example via beeps or vibration detectors, that inform them whether the overall confidence value is rising or falling, or whether it has already exceeded the threshold for issuing a control command. For this purpose, mobile acoustic and haptic signaling devices can also be connected to the machine's control system via a radio link. Haptic signaling devices can be wearables in the form of wristbands or gloves equipped with actuators, particularly vibration elements.

[0104] BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Further aspects and embodiments of the described method and the control arrangement are shown in figures of a drawing and described below.

[0106] Figures 1A to 1D: show a time sequence diagram for a display of identified objects in the context of the method,

[0107] Figures 2A to 2D: show another timing diagram for a display

[0108] REPLACEMENT SHEET (RULE 26) of identified objects in the context of the procedure,

[0109] Figure 3: shows the consideration of size and position changes of an object,

[0110] Figure 4: shows a screenshot of a screen showing the confidence value, as well as a frame around an identified object in the form of a gesture,

[0111] Figure 5: shows a first flowchart of an embodiment of the

[0112] procedure,

[0113] Figure 6: shows a second flowchart of an embodiment of the

[0114] procedure,

[0115] Figure 7: shows an embodiment of a control arrangement according to the

[0116] Invention with variants, and

[0117] Figure 8: shows an embodiment of various gestures as objects that can be classified into categories with associated control commands.

[0118] For technical reasons, the first two illustrations are divided into Figures 1A, 1B, 1C, ID on the one hand, and 2A, 2B, 2C, 2D on the other. The legend for the position of each of Figures 1A, 1B, 1C, ID and 2A, 2B, 2C, 2D in a composite figure is shown below each figure.

[0119] The embodiments and examples shown in the figures, which depict representations of objects / gestures, are not necessarily to scale. Likewise, various elements may be shown enlarged or reduced in size to emphasize individual aspects.

[0120] The proportions between the individual elements do not necessarily have to be realistic. Terms such as "top," "above," "below," "un-

[0121] REPLACEMENT LEAF (RULE 26) However, terms such as "below", "larger", "smaller", "right" and "left", and the like are correctly represented with respect to the elements in the figures. Thus, it is possible to deduce such relationships between the represented elements from the illustrations.

[0122] In Figures 1A to 1D, representations of a gesture 11a, which can be recognized as a graphic object, are shown in an upper row in the form of a raised open hand with the palm facing the viewer.

[0123] The representations are displayed in chronological succession on a display device, for example on a screen, starting on the left side of Figure 1A and progressing to the right side up to Figure 1D, one after the other at the same location on the screen. This representation can originate from the images acquired directly by an image recording system, but it can also originate from a library of stored objects and represent the image / object that is closest to the identified image object. The chronological sequence shows that the certainty or probability of correct recognition, i.e., the successful and reliable assignment of a category to the imaged object, increases over time.This certainty of recognition can correspond exactly to the probability of correct or appropriate categorization, that is, the probability that an assigned category is correctly assigned. This probability can correspond to a determined overall confidence value as a quantity based on multiple image confidence values.

[0124] The increasing probability of an accurate categorization is signaled by frames 13, 13a, 13b, and 13c, also called "bounding boxes," which surround the image of the hand and become wider and darker over time as the probability of correct or accurate categorization increases. This represents the respectively determined image confidence values. Thus, frame 13c is wider and darker than frame 13b, and this is wider and darker than frame 13a. For reasons of reliable and reproducible representation, this is shown in the figures by a dashed representation of the frames, which may be more or less interrupted.

[0125] REPLACEMENT SHEET (RULE 26) The leftmost image may not yet have a frame at all, since at this point the object / gesture may already be identified but not yet linked to a category, or the categorization may still be very unreliable.

[0126] Below the upper row of representations, Figures 1A to 1D show a lower row of representations corresponding to another gesture 11b, which is recognizable as a graphic object, namely a closed hand with two spread fingers. This row also shows that the reliability of a correct or appropriate categorization reaches a required threshold after a certain period of time, which is indicated by the strong frame 13c. Before this, a weak frame 13a appears in representations / frames no. 2 and 3 due to a low image confidence value, and thereafter, the reliability of a categorization initially decreases before increasing again until the threshold is reached. This temporary decrease in the image confidence value is tolerated when determining an overall confidence value.In this case, too, the respective images of the hand are shown next to each other for the purpose of representation on paper, whereas on a screen they can be displayed one after the other in the same place. Here, too, 13c denotes the thickest frame around the representation of the gesture. This also expresses that at the time corresponding to this representation, the probability of an accurate categorization is highest or at least exceeds a required threshold to trigger an action. In addition, a percentage 12 is shown above each individual representation to indicate the reliability of the categorization. The representations 13c in the upper as well as in the lower row of representations bear the textual annotation that a control command has already been triggered that corresponds to the respective recognized gesture.This can optionally also be indicated by a graphic symbol, for example in Figure 1C, ID by a bar 28 above the representations.

[0127] In addition, a circle is shown above the penultimate representations with the frame 13b in the upper and lower rows, indicating that reliable categorization is imminent because a

[0128] REPLACEMENT SHEET (RULE 26) sufficient or nearly sufficient probability of correct categorization has been achieved. This allows the user to see that their gesture has been successfully recognized, allowing them to, for example, lower their hand.

[0129] Previously, the progressive recognition and categorization process was already signaled to the user by the strengthening frame, so that he was encouraged to maintain his gesture until the end.

[0130] The circular symbol 29 shown in the illustrations above the frames 13b can, in some embodiments, also serve as a prompt symbol for entering a confirmation of the gesture. In this case, after this symbol appears, the user should perform a confirmation gesture different from the first, already recognized gesture to signal to the control device that they confirm the execution of the control command. If successful, the confirmation gesture is then recognized by the control device in the same way as the original gesture, and only then is the control command triggered and sent.

[0131] For successful recognition of a confirmation gesture, the control system may have lower requirements than for the recognition and categorization of another gesture, since at the time of a pending and expected confirmation, other gestures do not need to be recognized, and the control system may be focused on recognizing a confirmation gesture. Thus, the threshold for the required image confidence value and / or the threshold for the required overall confidence value may be lower for a confirmation gesture than for other gestures.

[0132] Figures 2A, 2B, 2C, and 2D show, in a similar manner to Figures 1A, 1B, 1C, and 1D, a gesture that was merely identified in the first illustration, beginning on the left side of Figure 2A, and is categorized as safe in the second illustration further to the right in Figure 2A and is accordingly surrounded by a thick frame 13c. Subsequently, the control command is triggered and sent. The initially recognized gesture is then further displayed, as shown in Figure 2B, but the display is visibly different from the display shown before execution.

[0133] REPLACEMENT BLADE (RULE 26), for example, may be color-changed or grayed out. This is indicated in Figure 2B by reference numeral 13d.

[0134] The changed representation 13d is displayed in the image display device, for example, if the gesture displayed or linked to the representation can no longer be identified and categorized by the control arrangement, at least for a while.

[0135] For this to happen, it is not necessary for the user to still perform the gesture at the beginning of this display. Figures 2A to 2D show that after the control command is triggered and the strong frame 13c is displayed, no further gestures are initially performed or detected by the user. After the control command is executed, the display, for example, shown in gray, is then displayed.

[0136] In Figure 3, a first representation 13e shows a gesture in the form of an open hand surrounded by a bounding box 22. A second representation 13f shows a bounding box 22a / a frame that is tilted by a certain angle relative to the first bounding box 22. Simultaneously or alternatively, the bounding box 22a can also be shifted relative to the first bounding box 22 in the image display in which the bounding boxes 22, 22a are displayed one after the other. Additionally or alternatively, the representation of the gesture can also be slightly different in size in the various representations, which is evident in the second bounding box 22a being slightly smaller than the first bounding box 22. The latter situation can occur, for example, if the user approaches or moves away from the camera while performing the gesture.Such deviations should not result in different objects / gestures being recognized in the different representations. Such deviations are relatively easy to compensate for using suitable analysis tools, for example, in the form of a convolutional neural network, since such types of neural networks exhibit a relatively high translation invariance and thus recognize such representations that are merely rotated, shifted, or scaled relative to one another as similar. Rules can also be established for the acceptance of such equivalence of representations.

[0137] REPLACEMENT SHEET (RULE 26) can be set. For example, fixed amounts of translations, rotations, or scale changes can be permitted, or the amount of permissible translations can be determined dynamically depending on the size of the object representation, measured by the length of the diagonals of the bounding box.

[0138] Figure 4 shows a representation of an image captured by an image recording system in the form of a camera in the form of a black and white photograph, which was captured by the processing device and fed into an analysis. The analysis device has identified two objects in the form of gestures and determined the categories assigned to them, or classified the objects into categories. The two identified gestures 12, 12a are labeled "Cylinder Down" and "Stop" and are each surrounded by a frame 13, 13a. For the "Cylinder Down" gesture, an image confidence value of 71% is given in the representation. For the "Stop" gesture, an image confidence value of 56% is given. Thus, none of the gestures has yet been reliably categorized, and it is more likely that a "Cylinder Down" gesture will be conclusively categorized than that a gesture / object in the image will be classified into the "Stop" category.This situation demonstrates that multiple objects can be identified and categorized by the analysis system simultaneously. The decisive factor for generating a control command is then which of the objects is categorized with the highest overall confidence value and whether this value is sufficiently high.

[0139] Figure 5 shows a schematic representation of an embodiment of the method according to the invention.

[0140] A first method step S1 involves capturing images. One or more cameras are directed at a user of a machine who wishes to control it with pointable objects, such as gestures. The user could, for example, also point at objects or signs for this purpose. When using multiple cameras, the images can be merged and linked to obtain 3-dimensional information and / or generally achieve better image quality. Two different cameras with different strengths can also be used for this purpose.

[0141] REPLACEMENT SHEET (RULE 26) TI ken, such as a color still camera and an infrared camera. The recorded images are optionally displayed in a method step S2 on an image display device, for example, on a screen, allowing the user to observe their own gestures.

[0142] In a third method step S3, all or part of the recorded images are fed to a processing device, which, in a fourth method step S4, identifies one or more objects in each image. For this purpose, an analysis device within the processing device may comprise a convolutional neural network, which is particularly suitable, for example, for such an identification step but also for a subsequent method step S5, the categorization of the identified objects. However, other self-learning devices can also be used for this purpose. Categorization is understood to mean the classification of the identified objects / gestures for each image into predetermined categories or the assignment of the categories to the identified objects.After method step S5, in a subsequent method step S6, the identified objects and / or the categories into which they have been classified are displayed in an image display device, as well as a current overall confidence value.

[0143] For this purpose, an image confidence value is determined for each identified and categorized object for each image captured in the processing unit. This value indicates the probability / reliability with which the category assigned to the identified object was correctly determined in the individual image. The image confidence values ​​from several consecutively acquired images are combined to form an overall confidence value. The overall confidence value is determined by taking into account both the number and values ​​of the image confidence values ​​(e.g., the proportion of images with an image confidence value above a certain threshold among all images in a current time unit), as well as any deviations due to images with lower image confidence values.

[0144] This allows the user to see after process step S6 to what extent his gesture was correctly identified and evaluated with a sufficiently high overall confidence value.

[0145] REPLACEMENT LEAF (RULE 26) In a further method step S7, the continuously updated overall confidence value is compared with a predefined threshold. After a decision step S7a, if the threshold is reached or exceeded, the control command linked to the reliably recognized gesture is generated and sent to the machine in a subsequent method step S8. At the same time, in a method step S9, the issued control command or the gesture linked to it is displayed on an image display device. This informs the user that their gesture was successfully implemented.

[0146] Figure 6 schematically shows a further embodiment of the method according to the invention, which up to method step 7a has the same shape as the embodiment shown in Figure 5.

[0147] According to the method illustrated in Figure 6, however, after the overall confidence value reaches or exceeds a threshold, a control command is not immediately triggered and output in step S8. Rather, after reliably categorizing a first gesture linked to a control command, the system waits for the recognition of a confirmation gesture performed by the user. This confirmation gesture must be identified by the analysis device in a method step S11 and reliably categorized with a sufficient overall confidence value in a method step S12. After comparing it with a threshold and the associated decision step S12a, the control command is finally triggered and sent to the machine.

[0148] Optionally, for this purpose, in a method step S10, the user can be prompted by means of a request signal after a successful categorization of the first gesture to show a confirmation object in the form of a confirmation gesture in order to finally trigger the control command.

[0149] If, after process step S12, an overall confidence value is reached that reaches or exceeds the threshold value for confirmation objects, the control command is triggered in the next process step S8 after decision step S12a. At the same time, in a further process step

[0150] REPLACEMENT SHEET (RULE 26) S13 the execution of the control command is displayed after confirmation. Afterward, the symbol for the executed control command or the symbol for the associated gesture may be displayed in a modified form, for example, grayed out, to indicate that this command has just been executed and may no longer be available.

[0151] In Figure 6, process loop 15 also illustrates that new images are continuously being captured and processed. Furthermore, it is clear that the analysis device, in particular a neural network, can also identify and categorize multiple objects simultaneously. In process step S6, different objects belonging to different categories are then output with different weights, which correspond to the different probabilities of reliable detection and classification.

[0152] Among the detected objects, a termination criterion can also be included if a termination object / gesture has been identified and assigned to the termination category. To ensure a sufficient overall confidence level, a different, particularly lower, threshold can be determined when a termination object / gesture is detected, just as for a confirmation gesture, than for other control commands and the associated gestures. The triggering of a termination command is illustrated in Figure 6 by method step S14.

[0153] Figure 7 shows a control arrangement 1 according to the invention.

[0154] This comprises an image recording device with two cameras 4a, 4b directed at a user 16. The user is located in the area of ​​a machine or machine element 6, for example, in a driver's cab. The user operates the machine and issues control commands at least partially via gestures or, more generally, via graphic objects recognizable in an image processing system, which are recorded by the cameras 4a, 4b.

[0155] For this purpose, many images are taken one after the other, for example at a rate of more than 10 or more than 20 images per second.

[0156] REPLACEMENT SHEET (RULE 26) An image can also be understood as two combined images taken simultaneously by different cameras. The captured images or a subset of the images are fed to an analysis device 3 within a processing device 17. At the same time, the captured images can be fed to an image display device 5 in the form of a screen or, for example, a head-up display and displayed there. The analysis device comprises a neural network, in particular a CNN (convolutional neural network). In its input layers 3a, the identification of predetermined graphic objects in the images takes place for each image. The neural network has been previously trained with corresponding objects for this purpose.In further layers, the identified objects are then assigned to pre-formed categories, and in an output layer 3b, the weights / probabilities are then output with which one or more objects are assigned to the various predetermined categories 18a, 18b, 18c.

[0157] This output is generated for each of the consecutively acquired images. The resulting consecutively generated output vectors 3c, 3d, 3e are continuously combined into an overall confidence value in the processing element 19 according to predefined rules for a specific number of recently processed images or for a specific past time period.

[0158] Optionally, a quality analysis of the image quality can also be considered for determining the overall confidence value. For this purpose, the images from the image acquisition device 4a, 4b are sent in parallel to an image analysis unit 20, and the image quality is evaluated in an image evaluation unit 21. These units 20, 21 can, for example, be implemented as a separate, trained neural network trained to evaluate image quality. This allows for the evaluation of any image deterioration caused by environmental influences that could negatively impact the overall confidence value.

[0159] The determined overall confidence value is checked in the processing element 22 to determine whether it meets or exceeds a threshold value, which may depend on the category of the respective identified object. If this condition is met, a command element 23 is used to issue a control command.

[0160] REPLACEMENT BLADE (RULE 26) to the machine or a machine element 6 activated.

[0161] At the same time, the recognized and categorized gesture or the control command associated with it, along with the currently applicable overall confidence value, is sent to an image display device 5a, which may be identical to the display device 5 but may also be different from it. Additionally or alternatively, signals relating to the recognized gesture or the initiation of the control command can also be sent to additional actuators 7, which can transmit optical, acoustic, or haptic signals to the user 16. Such actuators can be, for example, lights, loudspeakers, headphones, or vibration elements in wearables or in machine control elements.

[0162] Optionally, after categorizing an object and reaching a total confidence value sufficient for issuing a control command, and before issuing the control command, in a further step indicated by the dashed line 24, a confirmation request can be issued to the user 16 in an image display unit 5b or by another signal-generating element, which can be designed similarly to an actuator 7. The image display unit 5b can also be combined with the image display unit 5 and / or 5a.

[0163] If a confirmation to issue a control command is requested from the user, the processing device 17 initially does not trigger a control command after the processing element 22 has identified and categorized the object / gesture in question, but rather waits for the user 16 to issue a confirmation gesture, for this gesture to be identified, and for it to be categorized by the analysis device 3 with a specific overall confidence value, wherein the threshold which is determined as sufficient for the overall confidence value of a confirmation gesture can be lower or higher than the threshold which is determined for the overall confidence value of other gestures / objects.

[0164] The dashed line 25 indicates this confirmation analysis process, which is performed in parallel with or as part of the ongoing image processing.

[0165] REPLACEMENT SHEET (RULE 26) In Figure 8, five symbols are depicted in an upper row (27) and the same number of gestures in a lower row (26), with each gesture in the lower row being assigned a control command from the upper row. Thus, the gesture of the closed hand with two spread fingers in the lower row at the left edge of Figure 8 is assigned a start command. The upward-pointing thumb to the right of it is assigned the command "piston up," which is depicted in the second position, from the left, in the upper row in Figure 8. Further to the right, this is followed by the command "piston down," which is assigned to the gesture "thumb down." Next follows the control command for "forward movement," and finally, the raised hand with the flat palm facing the viewer represents the control command "stop."This last control command and the gesture associated with it could also represent an abort command, which does not mean the orderly stopping of a work process, but rather an emergency abort intended to bring the machine into a safe state in the shortest possible time.

[0166] REPLACEMENT SHEET (RULE 26) LIST OF REFERENCE SYMBOLS

[0167] 1 tax order

[0168] 3 Analysis unit, neural network

[0169] 3a Input layer

[0170] 3b Output layer

[0171] 4a, 4b Camera

[0172] 5, 5a, 5b screen

[0173] 6 Machine / Machine Element

[0174] 10 image, subset

[0175] 11, 11a, 11b identified object

[0176] 12, 12a associated sign

[0177] 13, 13a frame

[0178] 13b frame, indicating that gesture was correctly recognized over a longer period of time

[0179] 13c Frame after starting the control command

[0180] 15 Process loop

[0181] 16 users

[0182] 17 Processing unit

[0183] 18a-c Categories with assigned probabilities for each image

[0184] 19 Processing element for overall confidence value

[0185] 20 Image analysis unit

[0186] 21 Image Evaluation Unit

[0187] 22 processing element

[0188] 23 Command element

[0189] 24 Procedural step Request for a confirmation order

[0190] 25 Image processing to detect a confirmation gesture

[0191] 26 rows of categorized objects / gestures

[0192] 27 control commands assigned to the 26 gestures

[0193] 28 bars in the display

[0194] 29 Circle in the display

[0195] 51 Taking pictures

[0196] 52 Outputting images

[0197] 53 Capturing images

[0198] 54 Identifying objects

[0199] REPLACEMENT SHEET (RULE 26) 55 Classify in category

[0200] 56 Outputting a signal

[0201] 57 Compare overall confidence value

[0202] S7a Decision step

[0203] 58 Issuing a control command

[0204] 59 Outputting a signal for the overall confidence value

[0205] 510 Prompting for a confirmation gesture

[0206] 511 Identifying a confirmation object

[0207] 512 Categorizing a confirmation object

[0208] S12a Decision step

[0209] 513 Displaying the output of a control command after confirmation

[0210] 514 Abort detection

[0211] REPLACEMENT SHEET (RULE 26)

Claims

Claims Method for controlling a machine or a machine element (6), comprising: Capturing (S1) a plurality of images successively by means of an image recording system; in particular, outputting (S2) all or some of the plurality of captured images by an image display device (5, 5a, 5b); capturing (S3) at least a subset (10) of the plurality of recorded images by a processing device (17); Identifying (S4) at least one first predetermined object (11, 11a, 11b) on a plurality of successively recorded images of the subset of captured images, in particular by a network (3) trained by machine-based learning, Classifying (S5) the or an identified first predetermined object (11, 11a, 11b) for a plurality of successively recorded images into one of a plurality of categories, each of which is assigned to a predetermined object or a group of predetermined objects, each category being assigned a defined control command (23) for controlling the machine or the machine element, in particular by the network (3); Outputting (S6) a signal associated with a first predetermined object (11, 11a, 11b) identified and / or classified in a category, in particular an optical, acoustic or haptic signal, further in particular displaying an image or a symbol (12, 13, 13a) in an image display device (5a); Determining an image confidence value for several consecutively taken images, which indicates the certainty or probability with which a first predetermined object on the respective image was identified and assigned to a category; repeatedly determining an overall confidence value for the first predetermined object from the image confidence values ​​of several consecutively recorded images, in particular taking into account their temporal order; Outputting an overall confidence value, in particular by means of an optical, acoustic or haptic signal, in particular by means of a display in an image display device (5a) in which the image or symbol associated with the categorized first predetermined object is also displayed; Comparing (S7) the overall confidence value with a threshold value, and generating (S8) and outputting the control command (23) associated with the category into which the identified first predetermined object (11, 11a, 11b) was assigned, under the condition that the overall confidence value exceeds a predetermined threshold value, wherein, in particular, the signal (13c) associated with the identified and / or categorized first predetermined object (11, 11a, 11b) is changed upon output of the control command. The method according to claim 1, characterized in that a control command (23) is output only under the additional condition that, in addition to the first predetermined object (11, 11a, 11b), the processing device (17) recognizes a confirmation object different from the first predetermined object with at least a predetermined overall confirmation confidence value.Method according to claim 2, characterized in that the identification and classification of a confirmation object into a category is permitted only after the output (S6) of a signal associated with the identified and / or categorized first predetermined object (11, 11a, 11b), in particular only under the additional condition that the overall confidence value for the first predetermined object exceeds a predetermined threshold value. Method according to one of claims 1 to 3, characterized in that, after outputting (S6) a signal associated with the identified and / or categorized first predetermined object (11, 11a, 11b) or after it has been determined that an overall confidence value determined for the first predetermined object has exceeded a predetermined threshold value, a request signal (S10) for showing a confirmation object / a confirmation gesture is output, wherein the request signal is output in particular as a character in an image display device (5a) and is selected in particular as a function of the control command (23) to be issued or the categorized first predetermined object.Method according to claim 4, characterized in that the request signal (S10) is displayed on the same image display device (5) in which the recorded images are displayed, in particular on a partial area of ​​the image display device separated from the display of the recorded images. Method according to one of claims 1 to 5, characterized in that the overall confidence value is determined taking into account the number of images in which the first predetermined object (11, 11a, 11b) was categorized with an image confidence value above a predetermined threshold, and taking into account the number of further images lying between these images in which the first predetermined object was not categorized with an image confidence value above the threshold.Method according to one of claims 1 to 6, wherein the image on the screen (5, 5a) associated with the identified predetermined object (11, 11a, 11b) has a surrounding frame (13, 13a, 13b, 13c) or a strip, in particular in the form of a progress bar, wherein optionally the indexing of the overall confidence value is carried out by displaying at least one of the following features:. A number, especially a percentage, the size of which depends on the current overall confidence level; A width of the frame that depends on the overall confidence value; A color or brightness of the frame that depends on the overall confidence value; Fading the frame in and out at a variable frequency depending on the overall confidence value; The length or width, color, brightness, or a frequency of change of the display of a stripe that depends on the overall confidence value; A combination of several of the aforementioned features. Method according to one of claims 1 to 7, characterized in that an identified object (11, 11a, 11b) is categorized identically in two consecutive images if the position of the identified object in the second image deviates from the position in the first image by no more than a specified number or a number of pixels dependent on the extent of the representation of the first object in the image display device. Method according to one of claims 1 to 8, characterized in that the signal associated with an identified and / or categorized first predetermined object (11, 11a, 11b) remains changed in the form of an image in an image display device (5, 5a) after the output of the associated control command (23) at least until the control command has been processed or until the same control command can be meaningfully issued again.Method according to one of claims 1 to 9, characterized in that the issuing of a control command (23) or the execution of a control command is interrupted if the processing device (17) detects a termination object in captured images with at least a predetermined overall termination confidence value. A computer arrangement (1) comprising: a memory with a program stored therein; at least one processor connected to the memory; wherein the program is designed such that, upon execution of the program instructions by the processor, a method according to one of claims 1 to 10 is executed. A computer program product containing program instructions stored in a memory, upon execution of which by a processor, a method according to one of claims 1 to 10 is executed. A control arrangement for a machine or a machine element (6) with an image recording system for recording a plurality of images, with an image display device (5, 5a, 5b) and a processing device (17), wherein the control arrangement is configured to carry out the following method steps: Capturing (S1) a plurality of images successively by means of an image recording system (4a, 4b); in particular, outputting (S2) all or some of the plurality of captured images by the image display device (5, 5a, 5b); Capturing (S3) at least one subset (10) of the plurality of recorded images by a processing device; identifying (S4) at least one first predetermined object (11, 11a, 11b) on a plurality of successively recorded images of the subset of recorded images, in particular by a network (3) of the processing device trained by machine-based learning, Classifying the or an identified first predetermined object (11, 11a, 11b) for a plurality of successively recorded images into one of a plurality of categories, each of which is assigned to a predetermined object or a group of predetermined objects, each category being assigned a defined control command (23) for controlling the machine or of the machine element (S5), in particular by the network; Outputting (S6) a signal associated with a first predetermined object (11, 11a, 11b) identified and / or classified in a category, in particular an optical, acoustic or haptic signal, further in particular displaying an image or a symbol (12, 13, 13a) in an image display device (5, 5a, 5b); Determining an image confidence value for a plurality of consecutively recorded images, which indicates the certainty or probability with which a first predetermined object was identified in the respective image and classified into a category; repeatedly determining an overall confidence value for the first predetermined object from the image confidence values ​​of a plurality of consecutively recorded images, in particular taking into account their temporal arrangement; Outputting an overall confidence value, in particular by means of an optical, acoustic or haptic signal, in particular by means of a display in an image display device in which the image associated with the categorised first predetermined object is also displayed, Comparing (S7) the overall confidence value with a given threshold and Generating (S8) and outputting the control command (23) associated with the category into which the identified first predetermined object (11, 11a, 11b) was classified, under the condition that the overall confidence value for the first predetermined object exceeds the predetermined threshold value, wherein in particular the signal (13c) associated with the identified and / or categorized first predetermined object (11, 11a, 11b) is changed upon outputting the control command.

14. Control arrangement according to claim 13, which is configured to output a control command (23) only under the additional condition that, in addition to the first predetermined object (11, 11a, 11b), the processing device (17) identifies a confirmation object different from the first predetermined object at least with a predetermined overall confirmation confidence value and assigns it to a confirmation category, or which is configured to carry out a method according to one of claims 1 to 11.

15. Control arrangement according to claim 13 or 14, characterized in that it comprises both a display device (5, 5a, 5b) for displaying recorded images and images assigned to first predetermined objects (11, 11a, 11b), and at least one output device (7) for acoustic signals or an output device (7) for haptic signals.