Measuring system for recording object parameters of an object and device

The optical measuring system with a compensation signal and gesture recognition engine addresses the limitations of existing systems by providing device- and speaker-independent gesture recognition with reduced resource consumption and improved robustness.

DE102012025977B4Active Publication Date: 2026-04-23ELMOS SEMICON AG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
ELMOS SEMICON AG
Filing Date
2012-05-23
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing gesture recognition systems require significant computing resources, are sensitive to ambient light and interference, and lack device and speaker independence, leading to suboptimal recognition performance and high resource consumption.

Method used

A multidimensional optical measuring system with a transmitter and receiver system that uses a compensation signal to mitigate environmental influences, combined with a gesture recognition engine that performs feature extraction and emission calculation, allowing for device- and speaker-independent gesture recognition.

Benefits of technology

The system achieves robust and efficient gesture recognition with reduced resource consumption, enabling low-power operation and adaptable recognition across different devices and speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Measurement system for recording object parameters of an object (O), where the object (O) can perform a z-movement, an x-movement and a y-movement during the capture, wherein the measuring system has a physical interface (23) and where the physical interface (23) - at least one controller (C ij ) and - at least one transmitter (H i ) for light and - at least one receiver (D j ) for the transmitter that has at least one transmitter (H i ) emitted signal and wherein the measurement system includes at least one first unit in the form of an HMM gesture sequence recognition engine and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to perform the feature extraction (11) and emission calculation (12) steps, and where at least one controller (C) ij ) the physical interface (23) is set up for this purpose, - that of at least one transmitter (H i ) emitted signal and - with at least one recipient (D) j ) to compare the received signal and - to determine the comparison result as a multidimensional signal (24), and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to perform the feature extraction (11) and emission calculation (12) steps on the basis of this multidimensional signal (24), and wherein the feature extraction step (11) includes at least one of the following sub-steps: - Dividing the multidimensional signal (24) into individual frames of defined length, - Filtering of the multidimensional signal (24), - Normalization of the multidimensional signal (24), - Orthogonalization of the multidimensional signal (24), - nonlinear mapping of the multidimensional signal (24), - Formation of derivatives of the values ​​of the multidimensional signal generated in this way (24) and wherein the first unit in the form of the HMM gesture sequence recognition engine has a sequence lexicon (16) and wherein the sequence lexicon (16) contains sequence prototypes and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to evaluate the temporal sequence of hypothesis lists (39) of the successive frames and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to determine from this sequence of hypothesis lists (39) the most probable of the predefined sequence prototypes for a sequence of recognized prototypes.
Need to check novelty before this filing date? Find Prior Art

Description

Subject matter of the invention

[0001] The invention relates to a measuring system for acquiring object parameters of an object (O), wherein the object can perform a z-movement, an x-movement, and a y-movement during the acquisition. The measuring system has a physical interface (23). The physical interface (23) has at least one controller (C). ij ) and at least one transmitter (H i ) for light and at least one receiver (D j ) for the transmitter that has at least one transmitter (H i ) emitted signal. The measurement system has at least one first unit in the form of an HMM gesture sequence recognition engine. The first unit in the form of the HMM gesture sequence recognition engine is configured to perform the feature extraction (11) and emission calculation (12) steps. The at least one controller (Ci) j) the physical interface (23) is configured to receive the signal from at least one transmitter (H i ) emitted signal and to be connected to at least one receiver (D j ) to compare the received signal and determine the comparison result as a multidimensional signal (24). The first unit, in the form of the HMM-Gesture-Sequence-Recognition Engine, is set up to perform the feature extraction (11) and emission calculation (12) steps based on this multidimensional signal (24). The feature extraction step includes at least one of the following substeps: - Dividing the multidimensional signal (24) into individual frames of defined length, - Filtering of the multidimensional signal (24), - Normalization of the multidimensional signal (24), - Orthogonalization of the multidimensional signal (24), - nonlinear mapping of the multidimensional signal (24), - Formation of derivatives of the values ​​of the multidimensional signal thus generated (24).

[0002] The first unit, in the form of the HMM gesture sequence recognition engine, comprises a sequence lexicon (16). The sequence lexicon contains sequence prototypes. The first unit, in the form of the HMM gesture sequence recognition engine, is configured to evaluate the temporal sequence of the hypothesis lists of successive frames. For this purpose, the most probable of the predefined sequence prototypes for a sequence of recognized prototypes is determined from this sequence of hypothesis lists. Introduction

[0003] The continuous recording of gestures and their parameterization and assignment parameters is an important basic function in many, especially mobile, applications such as the control of computers, robots, mobile phones and smartphones.

[0004] A major problem with many gesture recognition systems is that they rely on image signal processing. This requires cameras and complex image processing equipment. These systems demand significant resources such as computing power, storage capacity, and electrical energy. However, these resources are very limited in a mobile phone. This particularly affects power consumption, storage space, and the lifespan of the processing units.

[0005] Alternative systems have lower sensitivity, which degrades recognition performance. Furthermore, these systems are not application-independent, necessitating the re-collection of gesture databases, a time-consuming and therefore very expensive process. In addition, most systems are sensitive to ambient light and other interference.

[0006] Finally, most application systems have additional sensors or sensory capabilities that should be combined with the operation of a gesture recognition system.

[0007] The following section discusses the state of the art: Several disclosures from the prior art are known that implement gesture recognition for a touchpad. For example, US6677932B1 describes a "Touch Data Structure" (designation 79, Fig. 3 in US6677932B1). It is therefore a gesture recognition device that requires touching the touchscreen. Consequently, only the x and y coordinates are disclosed in this disclosure” (Designation 80,

[0008] Fig. 3 in US6677932B1). Determining movement in the Z-direction is not possible. Crucially, however, this finding does not provide feedback to the physical sensors. The physical recognition method is not adapted to the already recognized sequence and to the two possible gestures to be distinguished (follow-gesture hypothesis). As a result, the recognition result for the subsequent gesture is not optimized.

[0009] The same applies to US20060026535A1. Here too, a "touchscreen" (designation 102) is mentioned. Fig. 2, US20060026535A1). As with US6677932B1 before it, this is only a two-dimensional gesture recognition method, which naturally offers far fewer artistic design possibilities. Here too, the gesture recognition configuration, and thus the recognition result, is not optimized for the next gesture in a gesture sequence.

[0010] Furthermore, US20060026535A1 assumes a capacitive method, which is not claimed here.

[0011] Both US20060026535A1 and US6677932B1 are sensitive to interference, such as sunlight, because they lack a compensation signal to equalize the signal being measured at the sensor. This makes these methods susceptible to contamination and sensor drift.

[0012] Such a compensation signal is described in US20080158178A1 (Claim 1, US20080158178A1). However, this document also describes a touchscreen. Therefore, only two-dimensional recognition is possible. Furthermore, the recognized gestures are not used to precondition the system for optimal recognition of the next gesture.

[0013] In the article "Multimodal Emotion Recognition in Audiovisual Communication" by Schuller, B.; Lang, M.; Rigoll, G., published in ICME '02 2001, pp. 745-748, a DTW recognition system is described (section 2 of the article). This system is known to be speaker- and device-dependent, meaning it must be trained using test datasets. This necessitates device- and speaker-specific gesture databases, which either prevent the industrial production of such devices or require users to undergo time-consuming training before using them. This is generally an obstacle that the average consumer is unwilling to overcome. Furthermore, training performed by untrained individuals is typically flawed.

[0014] Section 2 of this article also mentions a Z-component. However, since the sensor used is a SAW sensor, the Z-component is detected by the contact pressure. Therefore, it is not possible for the user's hand to detach from the touchpad, and thus true 3D recognition is not possible.

[0015] While the article does analyze the dialogue history (section 3.4), it does not provide feedback to the physical level of the recognition system through reconfiguration of the sensor system. Therefore, the sensor system is only optimized for an average value across all these handwriting samples. No extrapolation of the optimal parameter settings for distinguishing the most likely next gestures takes place.

[0016] All of the writings specifically cited above have in common that they require direct contact between the operating hand and the sensor interface and are therefore not 3D-capable and do not allow the detection of movement in the Z direction.

[0017] From DE10300223B3, an optoelectronic measuring arrangement is known, comprising at least two first light sources that emit light in a time-sequentially pulsed, phase-wise manner, and at least one receiver for receiving at least the pulse-synchronous alternating light component originating from the first light sources, as well as a device for compensating for ambient light by regulating the light intensity emitted into the measuring arrangement by at least one light source, such that the pulse-synchronous component occurring between different phases is reduced to zero. The special feature of this device is that the at least one further light source used for regulation is a light source independent of the first light sources, is assigned to the receiver, and its light intensity is adjustable in amplitude and polarity.DE10300223B3 thus discloses a technical teaching which, unlike the aforementioned documents, no longer requires contact between the user's hand and the device surface. However, the performance of this method is still limited insofar as recognized gestures do not influence the measurement itself. Therefore, the false acceptance rate (i.e., the rate of incorrectly accepted gestures) and the false rejection rate (i.e., the rate of incorrectly rejected gestures) are unnecessarily high. DE10300223B3 is only capable of determining a distance. DE10300223B3 discloses neither a method nor a device for associating the measured distances with gestures and identifying them. Object of the invention

[0018] The invention aims to provide a speaker- and largely device-independent gesture interface for human-machine interfaces while simultaneously requiring low resources and being highly robust against ambient light.

[0019] This problem is solved by a measuring system according to claim 1 and a device according to claim 9. Description of the basic idea behind the proposal

[0020] The proposal is divided into the following parts: - The actual detection system, in particular a multidimensional feedback optical measuring system - A smartphone or robot suitable for use in the proposed data collection system - A processing method and system for determining a gesture recognition result, which can be executed fully automatically by the recognition system. - A method for adapting the proposed system and procedure and the associated databases to a specific hardware platform - A method for recording device- and speaker-independent gesture databases - The basic structure of an associated gesture language for controlling a smartphone and similar applications where a system is to be influenced by gestures.

[0021] These are described in the following sections. Exemplary recognition system

[0022] A first exemplary proposed measurement system consists primarily of a typically triangulation-capable optical measurement system with typically a transmitter system and a receiver system, or a combined transmitter / receiver system that can switch very quickly between transmit and receive phases. Each transmitter system can consist of multiple transmitters. Likewise, each receiver system can consist of multiple receivers. The same applies to the combined receiver / transmitter systems.

[0023] To illustrate the operating principle, the device is first demonstrated using a system with only one transmitter and one receiver, with the help of the schematic diagram. Fig. 1 explained. At this point, it should be noted that all figures in this document contain only such details as a knowledgeable person needs to understand and follow the proposed idea.

[0024] The exemplary device consists of at least one first transmitter (3) which transmits into a first transmission path consisting of transmission sections (4) and (6). This transmission path (4, 6) can, for example, be divided by an object into the aforementioned first transmission section (4) and the second transmission section (6). Further serial and parallel divisions into further serial and parallel transmission sections are conceivable. (This is equivalent to the possibility of a unidirectional transmission network with inputs and outputs.)

[0025] At the end of the first transmission path (4) is an interaction zone (2) in which the current of the physical quantities (1), for example, the location, orientation, and surface properties of an object, in particular, for example, a hand, interacts (5) with the signal current of the transmitter (3) exiting the first transmission path (4). This interaction (5) modifies the properties of the transmitter's transmission signal (3). These properties can be, for example, amplitude, phase, spectrum, etc. Thus, the object imprints traces of its properties onto the transmitter's signal, which will later be used to draw conclusions about the object.

[0026] After this modification (5), the signal from the transmitter (3) typically passes through the second transmission path (6) and is received by a receiver or sensor (7).

[0027] The problem of independence from environmental influences remains. To address this, the received signal at the receiver (7) is typically augmented within the physical medium by a compensation signal to achieve a nearly constant overall signal. For this purpose, a compensation transmitter (9) sends a corresponding compensation signal via a second transmission path. This signal is superimposed at the receiver (7) on the signal from the transmitter (3) after it has passed through the transmission paths (4, 6). For the sake of completeness, it should be mentioned that the compensation signal can be generated within the medium, for example as light, or it can be fed directly into the amplifier's circuit as a signal, for example as an electrical current signal emulating the photocurrent of an LED.

[0028] For this purpose, a controller (8) (labeled "Controller") controls a compensation transmitter (9), which, via a predefined and essentially unaffected by the current of physical influences (1), typically separate transmission path (10), also feeds a signal into the receiver (7), where the two signals superimpose, as mentioned, linearly or non-linearly. Ideally, however, the receiver (7) is a linear receiver in which the signals from the transmitter (3) (after modification (5)) and the signal from the compensator (9) superimpose linearly. If the receiver (7) is non-linear, a component is added that is proportional to the product of the two superimposed signals. Although this case can be taken into account in the control engineering, it is not explained in detail here. However, it is easily possible for a person skilled in the art to modify the controller described here so that such non-linearities can be taken into account at any time.

[0029] The controller (8) also generates the transmit signal S5 of the transmitter (3). It is important that the transmit signal S5 is orthogonal to the transmit signals of other controllers with respect to the scalar product performed in this controller (8). This applies if there are other controllers in the system.

[0030] The controller (8) provides a vectorial output signal (24) which is analyzed below. The dimension and content are chosen to optimize gesture recognition.

[0031] Various methods and the devices or system components suitable for these methods can be used for the analysis of the vectorial signal (24) output by the controller (8) as well as for the synthesis of the test signal.

[0032] The result of this analysis when controlled by the controller (8) is a set of parameters that can be supplemented by further parameters obtained, for example, from optional additional sensors (37), touchpads, switches, buttons, and other measurement systems (37). The result is a so-called quantization vector, whose components, the measured parameters, will generally not be completely independent of each other. Each parameter on its own usually exhibits insufficient selectivity for precise gesture discrimination in complex contexts. By creating one or more such quantization vectors from the continuous stream of analog physical parameter values ​​(24) at typically regular time intervals through the physical interface (23) (see Fig. 1) using suitable methods, a temporally and value-quantized multidimensional parameter data stream is created (24).

[0033] This multidimensional signal (24), thus obtained in the form of one or more streams of quantization vectors, is first divided into individual frames of defined length, filtered, normalized, then orthogonalized, and, if necessary, appropriately distorted by a nonlinear mapping—e.g., logarithmization and cepstrum analysis. Furthermore, derivatives of the values ​​thus generated are also calculated. This method has long been known in pattern recognition, particularly speech recognition and material recognition, as feature extraction (11). The possibilities are manifold and can be found in the literature. Finally, the multidimensional quantization sector is multiplied by a so-called LDA matrix.

[0034] The next step of detection can be carried out using two different methods, for example: a) through a neural network or b) by an HMM detector c) by a Petri net

[0035] First, the HMM detector ( Fig. 1) described: Using the aforementioned predefined LDA matrix (14), the modified data stream (38) is mapped from the multidimensional input parameter space to a new parameter space, thereby maximizing its selectivity. The components of the resulting new transformed feature vectors are selected not according to real physical or other parameters, but according to maximum significance, which results in the aforementioned maximum selectivity.

[0036] The LDA matrix was typically calculated offline beforehand using example data streams from the database (18) with known gesture datasets through training (17). In the case of device-independent gesture recognition, the procedure now deviates from the standard procedure in one crucial step: ( Fig. 2)

[0037] If care is taken to ensure that all elements of the feature extraction (11) consist of at least locally reversible functions, deviations in geometry etc. can be taken into account in the form of an approximately linear transformation function.

[0038] Therefore, if the LDA matrix (14) has been determined for a geometric arrangement of transmitters and receivers in a well-defined application, it can be adapted to the specific geometry of the device by simple matrix multiplication with a device / application matrix (26). Fig. 2) In this way, a device-independent and speaker-independent gesture recognition system is obtained, which can easily be transported from device to device - for example, from a type of game console to a smartphone.

[0039] This matrix (26) can be determined, for example, by having a number of test users perform predefined and standardized test gestures for the specific gesture recognition application to be recognized. Since the gestures are known, a simple linear transformation between the new vectors and the vectors in the database (18) previously recorded on another device can now be calculated using a suitable training tool (25).

[0040] To make the test gestures reproducible worldwide, it makes sense to have them performed in a standardized way using suitable devices (e.g., motorized dolls or robots). Fig. 9) This allows these artificial gestures to be precisely compared with the stored prototypes for these robot-based reference gestures. This is done using Fig. Figure 11 explains. It shows an example mobile phone (33) above which a model hand (32) is located. This model hand could, for example, be the hand of a mannequin. The hand is moved by a segmented robot arm (35). Joints (35) are located between the segments (36). Fig. Figure 9 does not show exactly how the hand is moved, i.e., by which motor, but only how the recording of standard gestures / normalization gestures works. If the size of the hand (32) and its shape, color, etc., as well as its orientation and distance to the mobile phone (33), the lighting, the reflective properties of the container and the hand itself in which the entire device and the mobile phone (33) are located, and the movement, indicated here by a swiping motion (34), are defined, then the gesture recognition result of such a defined movement (34) can be predicted and thus compared with the obtained result.

[0041] The result of this calculation is an application-specific LDA matrix (26). (see also Fig. 2) The advantage is now that the data stream (38) leaving the feature extraction is speaker-independent and largely device-independent.

[0042] The key economic advantage of this method for transferring a gesture database from one device type to another is that the expensive recording of the gesture databases only needs to be performed once using actual gesture speakers. These costs can amount to up to €1 million per application.

[0043] The other corresponding methods are known from speech and material recognition. Simultaneously, the prototypes are calculated from these example data streams in the database (18) in the coordinates of the new parameter space and stored in a prototype database (15). In addition to this statistical data, this database can also contain instructions for a computer system regarding what should happen in the event of successful or unsuccessful recognition of the respective prototype. Typically, the computer system will be the computer system (e.g., a mobile phone) whose human-machine interface (hereinafter referred to as HMI) is intended to represent the proposed gesture recognition.

[0044] The pattern vectors (38) thus obtained, which are output by the feature extraction (11), are now compared with these pre-stored, i.e., learned, gesture prototypes (15), for example, by calculating the Euclidean distance between a quantization vector in the coordinates of the new parameter space and all these previously stored prototypes (15) in the emission calculation (12). At least two recognitions are performed in this process: 1. Does the detected quantization vector (38) correspond to one of the pre-stored quantization vector prototypes (= previously known gestures) or not, and with what probability and reliability? 2. If it is one of the gestures already stored, which one is it and with what probability and reliability?

[0045] To perform initial detection, dummy prototypes are typically stored in the prototype database (15), which should cover as many parasitic parameter combinations as possible that occur during operation. These prototypes are stored in a database (15), which will also be referred to as the code book (15) below.

[0046] If the minimum half-known Euclidean distance between two quantization vectors of two different prototypes of elementary gestures in new coordinates is also stored in the database / codebook (15) and falls below a prototype from the codebook, then this prototype is considered recognized. From this point on, it can be ruled out that further distances calculated during a continued search to other prototypes in the codebook (15) could yield even smaller distances.

[0047] The minimum Euclidean distance is calculated using the following formula: DistFV_CbE=MinCb_cnt=Cb_anz1[∑dim_cnt=dim1(FVdim_cnt−CbCb_cnt,din_cnt)2]

[0048] Here, dim_cnt represents the dimension index that traverses up to the maximum dimension of the feature vector (38) dim. FV dim_cnt stands for the component of the feature vector (38) corresponding to the index dim_cnt. Cb_cnt stands for the number of the entry in the Code-Book (15), i.e. for the elementary gesture corresponding to Cb_cnt. CB CB_cnt,dim_cnt Accordingly, dim_cnt represents the component of the prototypical code book feature vector entry corresponding to the elementary gesture corresponding to Cb_cnt. Dist FV_CbE Therefore, Cb_cnt represents the minimum Euclidean distance obtained. When searching for the smallest Euclidean distance, the number Cb_cnt that produces the smallest distance is noted.

[0049] To illustrate this, an example of assembly code is provided:

[0050] The confidence measure for correct detection is derived from the dispersion of the underlying basic data streams for a prototype and the distance of the quantization vector from its center of gravity.

[0051] Fig. Figure 10 demonstrates various detection scenarios. For simplicity, the representation is chosen for a two-dimensional feature vector so that the methodology can be depicted on a two-dimensional sheet of paper. In reality, the feature vectors (38) are typically always multidimensional.

[0052] The focal points of various prototypes (41, 42, 43, 44) are shown. As described above, half the minimum distance between these prototypes can now be stored in the code book (15). This would then be a global parameter that would be valid for all prototypes. However, this decision using the minimum distance presupposes that the variances of the elementary gesture prototypes (41, 42, 43, 44) are more or less identical. If the feature extraction is optimal, this is indeed the case. This would correspond to a circle around each of the prototypes with a gesture database-specific radius.

[0053] In reality, this is rarely achievable. Therefore, recognition performance can be improved by storing the spread for each elementary gesture prototype. This would correspond to a circle around each prototype with a prototype-specific radius. The disadvantage is an increase in computing power.

[0054] A further improvement in recognition performance can be achieved if the scatter width for the prototype of each elementary gesture is modeled by an ellipse. For this, instead of the radius as before, the principal axis diameters of the scattering ellipse and their tilt relative to the coordinate system must now be stored. The disadvantage is a further, massive increase in computing power and storage requirements.

[0055] Of course, the calculation can be made even more complicated, but this usually only massively increases the effort and does not significantly improve the recognition performance.

[0056] Compact detectors for mobile devices will therefore typically resort to the simplest of the described options.

[0057] The positions of the gesture vectors determined by the emission calculation can vary considerably. For example, it is conceivable that such a gesture vector (46) lies too far from any gesture. This distance threshold could, for instance, be the minimum halfway distance between the prototypes. It is also possible that the scatter ranges of the prototypes (43, 42) overlap, and a determined gesture vector (45) lies within this overlap. In this case, a hypothesis list would contain both prototypes with different probabilities due to their different distances.

[0058] In the best case, the determined gesture vector (48) lies within the scattering range (threshold ellipsoid) (47) of a single prototype (41) which is thus reliably detected.

[0059] To improve the modeling of the range of variation of a single gesture, it is conceivable to model it using multiple prototypes, each with its own range of variation. In other words, several prototypes can represent the same gesture. The risk here is that, due to the distribution of the probability of a gesture across several such sub-prototypes, the probability of each individual sub-prototype gesture may become lower than that of another gesture whose probability was lower than that of the original gesture. Thus, this other gesture could potentially become the dominant one, albeit erroneously.

[0060] A significant problem, therefore, is the computing power required to reliably recognize elementary gestures. This will be discussed further: A crucial point is that the computational effort increases with Cb_anz * dim.

[0061] For a non-optimized HMM recognizer, the number of assembly instructions required to calculate a vector component is approximately 8 steps.

[0062] The number A_Abst of the necessary assembler steps to calculate the distance of a single code book entry (CbE) to a single feature vector (FV) is calculated approximately as follows: A_Abst=FV_Dimension*8+8

[0063] This results in the number A_CB of assembler steps for determining the code book entry with the smallest difference: A_CB=Cb_anz*(A_Abst)+4=Cb_anz*(FV_Dimension*8+8)+4

[0064] Using the example of a medium HMM recognizer with 50000 CbE (code-book entries = CB_ance) and 24 FV_dimensions (feature vector dimensions = FV_dimension), the number of steps is: 50000*(24*8+8)+4~10 million operations per feature vector

[0065] At a relatively low sampling rate of 8kHz = 8000 FV per second (feature vectors (24) per second) a computing power of 8GIps (8 billion instructions per second) is required.

[0066] Especially for energy-autonomous and / or mobile systems, this simple example already leads to a resource consumption that is unsustainable.

[0067] In the case of an optimized HMM recognizer, as already mentioned, the smallest codebook distance between two prototypes is pre-calculated and stored in the codebook (15). This has the advantage that the search can be aborted if a distance is found that is less than half of this minimum codebook distance. This halves the average search time. Further optimizations can be made by sorting the codebook according to the statistical occurrence of the elementary gestures. This ensures that the most frequent elementary gestures are found immediately, which further reduces the computation time.

[0068] For such an optimized HMM detector, the computing power requirement is now as follows:

[0069] Again, there are 8 steps to calculate a vector component. The steps for calculating the distance A_Abst of a code book entry (CbE) to the feature vector (FV) are again: A_Abst=FV_Dimension*8+8

[0070] The number of steps required to determine the code book entry with the smallest gap A_CB using optimization is slightly higher: A_CB=Cb_anz*(A_Abst)+4=Cb_anz*(FV_Dimension*8+10)+4

[0071] The two additional assembler instructions are necessary to check whether the distance is less than half the smallest code book entry distance.

[0072] Furthermore, the number of CB_Anz code-book entries for mobile and energy-autonomous applications is limited to 4000 code-book entries or even less.

[0073] Furthermore, the number of feature vectors per second is reduced by lowering the sampling rate.

[0074] This will be explained using a simple example: The aforementioned medium HMM detector is now operated with a codebook containing only a tenth of the entries, e.g., with 4000 CbE, and still with 24 FV dimensions.

[0075] The number of steps is now 4000*(24*8+10)+4∼808004_ operations per feature vector

[0076] Reducing the feature vector rate to 100 FV per second (feature vectors per second) extracted from 10ms parameter windows over, for example, 80 samples each, and aborting the search when the distance is less than half the smallest codebook entry distance, results in at least a halving of the effort with suitable codebook sorting.

[0077] The required computing power then drops to <33 DSP MIPS (33 million operations per second). In reality, codebook sorting leads to even lower computing power requirements, for example, 30 MIPS. This makes the system real-time capable and integrable into a single IC.

[0078] The search space can be narrowed down through preselection. A prerequisite for this is an even distribution of the data, meaning the centers of gravity of the quadrants are located in the geometric center of the quadrant.

[0079] The necessary reduction of the codebook (15) has advantages and disadvantages: Reducing the codebook (15) increases both the false acceptance rate (FAR), i.e., the incorrect gestures that are accepted as gestures, and the false rejection rate (FRR), i.e., the correctly executed gestures that are not recognized.

[0080] On the other hand, this reduces the resource requirements (computing power, chip area, memory, power consumption, etc.).

[0081] Furthermore, the history, i.e., previously observed gestures, can be used in hypothesis formation. A suitable model for this is the so-called Hidden Markov Model (HMM).

[0082] For each gesture prototype, a confidence measure and a distance to the measured quantization vector can thus be derived, which can be used directly by the system (21) and can also be further processed. It is also useful to output a hypothesis list for the recognized prototypes of each frame, containing, for example, the ten most probable gesture prototypes of that frame with their respective probability and reliability of recognition.

[0083] With this data (21), instructions or executable code correlated with the respective prototype can also be output to the said user computer system (e.g. a mobile phone), which were stored, for example, in the prototype database (21) or otherwise previously.

[0084] Since not every temporal and spatial sequence of gestures is meaningful, it is possible to evaluate the temporal sequence of the hypothesis lists of the successive frames.

[0085] The goal here is to find the elementary gesture sequence path through the successive lists of hypotheses that has the highest probability.

[0086] Here too, at least two detections are performed: 1. Is the most likely elementary gesture sequence one of the already stored sequence prototypes or not, and with what probability and reliability? 2. If it is one of the already stored sequences, which one is it and with what probability and reliability?

[0087] This can be achieved either by using a training program (20) to automatically enter a sequence prototype into a sequence dictionary (16), or manually via a type-in ​​tool (19) that allows these sequence prototypes to be entered via the keyboard. It is also conceivable, with suitable standardization, to download individual gesture sequences, for example, via the internet.

[0088] Using a Viterbi decoder (13), the most probable of the predefined sequences for a sequence of elementary gestures, more precisely the recognized gesture prototypes, can be determined from the sequence of hypothesis lists. This is particularly true even if individual elementary gestures were incorrectly recognized due to measurement errors. Therefore, it is very useful to transfer sequences of gesture hypothesis lists (39) from the emission calculation (12), as described above, to the Viterbi sequence recognition (13). The result is the most probable recognized elementary gesture sequence (22) or, analogous to the emission calculation, a sequence hypothesis list. Here, too, instructions can be linked to the most probable recognized gesture sequence in the sequence database (16) or otherwise to the aforementioned computing unit, for example, a mobile phone, specifying what action should be taken. If necessary, for example, a warning message is sent to the user.

[0089] Finally, we will consider the functional software components of the gesture sequence recognition engine. This is located in Fig. 1 is entered as a Viterbi search (13). This search (13) accesses the sequence lexicon (16). The sequence lexicon (16) is fed by a learning tool (20) and a tool (19), in which these elementary gesture sequences can be defined by textual input. The possibility of a download has already been mentioned.

[0090] The basis of sequence recognition (19) is the hidden Markov model. The model is built up from different states. In the Fig. In the example given in section 23, these states are symbolized by numbered circles. In the aforementioned example in Fig. The 23 circles are numbered from 1 to 6. There are transitions between the states. These transitions are described in the Fig. 23 is designated by the letter a and two indices i and j. The first index i denotes the number of the source node, and the second index j the number of the destination node. Besides transitions between two different nodes, there are also transitions a ii or a jj which lead back to the starting node. Furthermore, there are transitions that allow nodes to be skipped; from the sequence, a probability of finding a k-th observable b is thus derived. k actually observable.

[0091] This results in sequences of observables that have predictable probabilities b k can be observed.

[0092] It is important to note that every hidden Markov model consists of unobservable states q i exists. Between two states q i and q j The transition probability a exists ij .

[0093] Thus, the probability p for the transition from q can be determined. i on q j write as: p(qnj|qn−1i)≡aij

[0094] Here, n represents a discrete point in time. The transition therefore occurs between step n with state q. i and step n+1 with state q j instead of.

[0095] The emission distribution b i (Ge) depends on the state q i ab. As already explained, this is the probability of observing the elementary gesture Ge (the observable) when the system (hidden Markov model) is in state q. i is located: p(Ge|qi)≡bi5(Ge)

[0096] In order to start the system, the initial states must be defined. This is done using a probability vector π. i . It can then be stated that a state q i with the probability π i An initial state is: p(qi1)≡πi

[0097] It is important that a new model be created for each gesture sequence. In a model M, the observation probability for an observation sequence of elementary gestures should be defined. Ge→=(Ge1,Ge2,...GeN) will be determined

[0098] This corresponds to a state sequence that is not directly observable and follows the sequence below: Q→=(q1,q2,...qN)

[0099] The probability p of observing the state sequence Q, which depends on the model M, the state sequence Q and the observation sequence Ge, is: p(Ge→|Q→,M)=p(Ge1,Ge2,...GeN|q1,q2,....qN)=p(Ge1|q1)⋅p(Ge2|q2)⋅.......p(GeN|qN)=∏n=1Np(Gen|qn)=∏n=1Nbn(Gen)

[0100] This results in the probability of a state sequence Q→=(q1,q2,......qN) Model M: p(Q→|M)=p(q1,q2,.....qN|M)=p(q1)⋅p(q2|q1)⋅p(q3|q1,q2)⋅.... .p(qN|q1,q2,.....qN−1)=p(q1)∏n=2Np(qn|qn−1)=π1∏n=2Na(n−1)n

[0101] Thus, the probability of recognizing a gesture word is equal to that of a gesture sequence: p(Ge→|Mj)=∑all Qkp(Ge→|Qk,Mj)p(Q→k|Mj)=∑all Qk(∏n=1Nbn(Gen))(π1∏n=2Na(n−1)n)

[0102] The most probable word model (gesture sequence) for the observed emission Ge is determined by summing the individual probabilities over all possible paths Q. k , which lead to this observed gesture sequence Ge. p(Ge→|Mj)=∑all Qkp(Ge→|Qk,Mj)p(Q→k|Mj)=∑all Qk(∏n=1Nbn(Gen))(π1∏n=2Na(n−1)n)

[0103] Summing over all possible paths Q is problematic due to the potential computational effort. Therefore, the process is usually terminated very early. Some predictors only use the most probable path Q.k This will be discussed below.

[0104] The calculation is performed recursively. The probability a n (i) at time n the system is in state q i The observed effect can be calculated as follows: αn(i)=p(Ge1,Ge2,......Gen;qn=qi)≡p(Gein,qni) αn+1(j)=[∑i=1Sαn(i)⋅aij]bj(Gen+1)

[0105] Here, a sum is taken over all S possible paths that lead to state q. i+1 lead in

[0106] It is assumed that the total probability in state q i n+1 The best path is dominant. Then the sum can be simplified with minimal error. αn+1*(j)=[maxi(αn*(i)⋅aij)]bj(cn+1)

[0107] By tracing back from the last state, the best path can now be obtained.

[0108] The probability of this path is a product. Therefore, a logarithmic calculation reduces the problem to a pure summation problem. Here, the probability of recognizing a word corresponds to the probability of recognizing a model M. j This corresponds to the determination of the most probable word model for the observed emission X. This is now done exclusively via the best possible path Q. best p(Ge→|Mj)=∑all Qk(∏n=1Nbn(Gen))(π1∏n=2Nq(n−1)n)

[0109] This will thus become p(Ge→|Mj)=p(Ge→|Qbest,Mj)p(Q→best|Mj)=exp(ln(π1)+ln(b1(Ge1))+∑n=2Nln(bn(Gen))+ln(a(n−1)n))

[0110] It is now of particular importance that the Prototype Book (15) contains only elementary gestures.

[0111] The Fig. Figures 12 to 22 show some of the possible exemplary elementary gestures over a sketched smartphone as an example device.

[0112] The figures depict gestures performed with one hand. These can, of course, also be combined into two-handed gestures. It is certainly conceivable that these gestures could also be performed with other body parts. For example, the circular hand movement over a smartphone could correspond to a circular foot movement in front of a car trunk. The recognition system would then be directed downwards and, upon detecting this foot gesture in combination with other factors, such as the proximity of the correct car key, would open the trunk.

[0113] For simplicity, gestures can be broken down into different hierarchies. First, there are the orientations of the hand or object. Then there are the transitions between these orientations. These are typically rotations, followed by positioning in space and its transitions, which correspond to the hand or object's path in space above the device.

[0114] Let us first turn to the elementary object orientations. These are stored in the Prototype Book (15), just like the movements. As a first approximation, we can, for example, specify the following orientations of a hand. For an object, an analogous basic orientation must first be defined in order to define an analogous orientation for each object. a. A palm or underside of an object for recognition b. A hand from the side with little finger towards the recognizer (object from the right side) c. A back of the hand or the top of an object for recognition d. A hand from the side with thumb towards the recognizer (object from the left side) e. A hand or object oriented in one of the following directions: i. North, ii. North East iii. East iv. South East v. South vi. South West vii. West viii. North West f. One hand tilted with the tip pointing downwards (An object tilted downwards at the front) g. One hand tilted with the tip pointing upwards (An object tilted forwards and upwards) h. One hand vertically with the point facing upwards (Object pointing vertically upwards) i. One hand vertically with the point downwards (Object pointing vertically downwards) j. One hand bent (object rotated) k. To the right l. To the left m. The fingers of one hand spread n. The number of spread fingers of one hand is 1, 2, 3, 4, 5 o. A finger (especially the index finger) of one hand pointing forward, and so on. Regarding the last points j, k, l, and m, their applicability to another object depends heavily on that object's capabilities. Completely different properties may be important for a given object. If multiple objects / hands are present, combinations of these basic orientations are also possible. The above list naturally applies to objects other than hands, such as feet or signaling devices. This also applies to the following sections.

[0115] The distinguishability of the orientations ultimately depends only on the number of senders T_Anz and receivers S_Anz. The number of parameters P_Anz that can be extracted without differentiation is then proportional to P_Anz=T_Anz*S_Anz

[0116] However, this requires a correspondingly comprehensive regulation.

[0117] The basic gestures performed by a hand or an object can be divided into • those without an object (hand), • those with an object (hand) but without movement (e.g., no gesture, remaining still), • Change in the number of objects (hands) in front of the recognizer (none, one, two) typically with a spatial orientation (entry from left, right, top, bottom and corresponding intermediate values ​​such as diagonally above etc.), • those in which the object (the hand) itself is rotated around one of three spatial axes, either clockwise or counterclockwise ( Fig. 12, Fig. 13, Fig. 14), in particular reorientations of the object / hand (e.g. edge of hand, palm, top of hand towards the detector), • those in which the object (the hand) performs circular, elliptical, or hyperbolic movements clockwise or counterclockwise on a coordinate direction of a polar coordinate system, a spherical coordinate system, or another orthogonal coordinate system. Fig. 21 and Fig. Figures 22 show circular movements. The third axis orientation corresponds to this. Fig. Figure 21 is not shown because only the smartphone would have been rotated. Of course, limiting the display to curve segments is conceivable. • those in which the object (the hand) is moved translationally forward or backward in one of the three spatial directions (x, y, z, and diagonal intermediate values) Fig. 18, Fig. 19, Fig. 20) • those in which the structure of the hand (the object) is changed (extending one, two or more fingers, e.g. counting, Fig. 15, where Fig. 15, for simplicity, does not show the numbers 4 and 5; or fingers spreading or closing, Fig. 17; or, for example, opening or closing a hand Fig. 16) • Combinations of these gestures or movements (e.g., circular movement with opening and closing of a hand, whereby the hand is always closed at one pole of the circle and open at the other.)

[0118] Every gesture has an inverse gesture. This reverses the result of a previously performed movement in a geometric sense. For example, the inverse of a clockwise circular motion is a counterclockwise circular motion, and the inverse of closing the fingers is opening them ( Fig. 17) etc.

[0119] It is important to note that combinations of these gestures are conceivable. For example, opening and closing the fingers can occur in six spatial directions (x, y, z, -x, -y, -z) relative to the finger's line of movement. The hand can also be oriented in six directions (x, y, z, -x, -y, -z). (Intermediate diagonal orientations are, of course, also possible.) In total, the basic elementary gesture can therefore be executed in 36 configurations. These can be transformed into one another through rotations. However, not all gestures are equally recognizable. In the case of opening the fingers in the z-direction, the lower finger casts a shadow over the upper one. A skilled practitioner would therefore typically not use such gestures.

[0120] An elementary gesture sequence “swipe in X direction” ( Fig. 19) can therefore be composed of, for example, the following actions: a) Elementary gesture “Idle” = no object b) Start of entry (= hand object appears at the edge of the measuring area) c) Entry process (= object essentially does not change size and distance, but performs a translational movement in the X direction, thus appearing at different times in different places) d) Remaining still (= object does not substantially change size or distance, does not perform any translational movement in the X and Y directions, thus appears at the same locations at different times) (= end of entry process) e) Swiping (= object essentially does not change size and distance, but performs a translational movement in the X direction, thus appearing at different times in different places) f) Static (= object does not substantially change size or distance, does not perform any translational movement in the X and Y directions, thus appears at the same locations at different times) (= end of swipe) g) Exit process (= object does not essentially change size and distance, but performs a translational movement in the X direction, thus appearing at different times in different places) h) End of entry (= hand object disappears at the opposite edge of the measuring range) i) Elementary gesture “Idle” = no object

[0121] A “gesture word” defined in this way, which is composed of the elementary gestures described above, can be recognized with high accuracy by the combined HMM-Viterbi recognizer (13) described above.

[0122] Basically, two recognition scenarios are now possible: Single-gesture word recognition ( Fig. 24) and continuous gesture word recognition ( Fig. 25).

[0123] With single-gesture word recognition, the control system is triggered to start gesture recognition by another event, such as pressing a button. The advantage of this approach is that it saves processing power and therefore energy. This can be beneficial for mobile devices.

[0124] In continuous gesture word recognition ( Fig. 25) The gesture recognition is permanently active. It is in at least one sleep state that corresponds to the permanent recognition of the gesture "no gesture". Experience has shown that this is represented by a large number of different prototypes. In addition, it is also possible to put the gesture recognition into a special power-saving mode. For example, it is conceivable that an energy-autonomous device is powered by a solar cell and the actual gesture recognition is switched off. The energy output of the solar cell can be used as a sensor signal (37, Fig. 1) be used. The recognizer then jumps from the "sleep" state to the "wake up" state. This means that the transmitters and the rest of the system are started, which massively increases power consumption. If no allowed elementary gesture is found, the system returns to "sleep" mode. However, if a allowed elementary gesture is found (In Fig. If the gestures "no gesture" and "bcd" are used, the system transitions to the respective states. If the system remains in the "no gesture" state for too long, it reverts to the "sleep" state to conserve energy. As long as gestures are executed consecutively within a predefined timeframe, the system will always return to the "no gesture" state. In many cases, it is advisable to always conclude the gesture words with a special "end gesture" state, which terminates the gesture words in a system-specific manner.

[0125] The gesture words, whose recognition was described above, can now be combined to form more complex “gesture sentences”.

[0126] It is useful to group gesture words into gesture word classes. Examples of such classes could be classes that group gesture words that, for example... • Subjects • Objects (e.g., items and people) or • Properties of these objects or • Changes in properties or • Properties of the changes themselves • Causes of these changes or • Describe the relationships between these objects.

[0127] The sequence of gesture words can be defined depending on the gesture word class. This enables subsequent HMM-Viterbi recognizers to correctly identify the content of such a defined gesture message, a gesture sentence, despite potentially incorrect gesture word recognition from the previously described combined HMM-Viterbi recognizer. Since the process is analogous to the recognition of elementary gesture sequences (gesture words) based on a hypothesis list (29), it will not be explained further and is also omitted from the figures for the sake of simplicity.

[0128] A gesture grammar defined in this way allows for a significantly improved transmission of compact command sequences, similar to a gesture language.

[0129] In order to make such a gestural language efficient, it is useful to design the transitions between the gesture words in such a way that the end of the preceding gesture and the beginning of the following gesture can be coordinated.

[0130] Gesture words can contain elementary gestures and elementary gesture sequences that express dependencies on other gesture words or convey other information such as temporal references (like past and future, etc.). Elementary gesture sequences can also consist of only a single elementary gesture.

[0131] At this point, we should mention the diverse possibilities of linguistics.

[0132] It is now of particular importance that the gestures performed have consequences and produce an effect. Typically, such gestures will be used to control a computer. This could also be the computer of a robot. Since the applications are usually not developed by the device manufacturer, it is advisable that the gestures or gesture sequences are, or can be, assigned to text strings or other symbol strings. This creates a link between the gestures or gesture sequences on the one hand and the actions of the actuators of the computer and / or robot system on the other. Such a link is best transferred via a data transfer protocol, such as HTTP or XML.

[0133] When using a neural network detector ( Fig. 3) A neural network (27) is fed with the determined quantization vectors (38) of the feature extraction (11). However, a neural network (27) must first have been conditioned using example data (18) and a training procedure (28). Here, too, feedback of the recognition result (29) or (38) to the controller (8) is conceivable. Fig. 8)

[0134] Neural network recognition systems are significantly more compact and therefore easier to integrate than HMM recognition systems. However, their ability to distinguish gestures is limited due to internal quantization noise. They are also more sensitive to interference. Therefore, it is not recommended to attempt to distinguish more than 20 gestures, including parasitic ones, with a single neural network. Furthermore, neural network recognition systems typically produce speaker- and device-dependent results. This means they must be trained by the user, which severely limits their applicability.

[0135] This type of recognition system is state of the art and has long been known for pattern recognition.

[0136] However, in order to supply the system with a particularly efficient set of feature vectors, it is advantageous to drive the transmitter (3) and the compensation transmitter (9) with an optimal set of signals. This concept of feeding a recognition result back into the physical interface (23) is novel.

[0137] For this purpose, the Viterbi recognizer (13) can output a list of the likely gestures (Gn) that will follow next, based on the sequence lexicon (16). This information can now be used to control the transmitters (3) so that the received vectors (24) enable optimal gesture discrimination. Fig. 7)

[0138] This procedure will be explained using a specific example: Fig. Figure 4 shows a schematic example of a mobile device (mobile phone) with eight combinations of transmitter diode (H1-H8), receiver diode (D1-D8) and compensation diode (K1-K8).

[0139] After generating a forecast of the most likely subsequent gestures, the transmitter sends a temporal and, if necessary, spatial pattern that is best suited to distinguishing the likely alternatives.

[0140] This pattern can be pre-calculated using the Database (18).

[0141] The Fig. Figures 27 to 58 show all possible activity patterns of the various transmitters of the example mobile phone. Fig. 4.

[0142] For the expert, it is clear that, assuming linear addition of the activity patterns is possible, the patterns can form a Hilbert space. Thus, by orthogonalizing the patterns, simple basic patterns can be generated from which all other patterns are composed.

[0143] These elementary activity patterns can, in the simplest case, be represented by the activity of a single diode.

[0144] Irradiation with an optimized transmission pattern maximizes the significance of the quantization vectors.

[0145] The last problem is deriving the optimal transmission sequence.

[0146] For this purpose, the optimal stimulation pattern for the next elementary gesture (Gn) identified by the Viterbi recognizer (13) as the most probable next gesture is extracted from the database (18) for verification by the controller (8). Therefore, when creating this database (18), a recognition performance must already be stored for every conceivable pattern. This can be done, for example, using a matrix (LDA_B).

[0147] Finally, the control algorithm of the physical interface (23) will be shown using Fig. 5 will be illustrated with the transmitter (the light-emitting diode H) i ), compensation transmitter (the light-emitting diode K i ) and the sensor (the photodiode D) i ) be linked together: Each of the transmitters H1 to H8 sends a signal into a separate transmission channel (I1) depending on the desired transmission pattern. i ) to the object to be measured (O), typically a hand. There, the light is reflected and transmitted via a further transmission channel (I3). j ) to each receiver diode K1 to K8.

[0148] For each pair of receiving diode D i and transmitting diode H j A controller C will be used ij planned. Since this would result in a large hardware requirement, it makes sense in our example to use only eight such controllers C. ij to provide and always only one transmitting diode H j to operate. This means that between the transmitting diodes H iThe switching occurs using time-division multiplexing. The corresponding feature vectors are therefore typically buffered and combined into a feature vector (24) that is eight times wider in this example, and passed to feature extraction in this example at one-eighth of the original feature vector frequency.

[0149] Each of the receiving diodes D i Each will have a signal generator G i This generator produces a signal orthogonal to the other generators. Orthogonality is explained in more detail below.

[0150] The signal from generator G i , which is the receiver diode D i The assigned value is converted into a short initial pulse S5 by an orthogonalization unit. oi and a trailing signal S5 di split up in such a way that their sum again gives the signal S5 ithis would result and that these signals are orthogonal with respect to the scalar product explained later.

[0151] The signal from the receiving diode D i is amplified by a preamplifier and combined with the signal S5 di to signal S10 dij in the multiplier M 1ij The signal is multiplied. Then the signal passes through filter F. 1ij filtered. The preceding multiplication in the multiplier M 1ij and this filtering in filter F 1ij define the scalar product of the output signal of the receiver D 1j with signal S5 di . When orthogonality of further signals was mentioned in this description, it meant that these two signals, when multiplied by a scalar in this way, result in zero.

[0152] The filter output signal is then amplified again by amplifier V1. ij to signal S4 dijamplified. The inverse transformation is performed by remultiplying the signal S4. dij with signal S5 di to signal S6 dij .

[0153] The same operation is performed with the other signal S5. oi This is carried out. This results in the signal S6 in an analogous manner. oij .

[0154] To now determine the compensation signal of the compensation diode K 1j To obtain the signals S6 dij and S6 oij for all generator signals i added together and after adding a suitable offset B1 j the compensation diode K 1j supplied.

[0155] Each controller C ij In this example, it outputs two values: A ij and D eij In this example, these form a 16-dimensional analog vector together with the seven other pairs of the other seven active controllers. Since only one transmitter H is used at a time due to time-division multiplexing, iWhen active, eight of these vectors are generated within a measurement cycle in which all transmitters H i Once activated, these are temporarily stored as described and assembled into a 128-dimensional feature vector (24). This 128-dimensional feature vector (24) forms the input signal for feature extraction (11). Since usually several transmitters H i and several D1 sensors j The number of components present can be used to create a compensation signal S6. j contributing controller C ij be greater than one. This is in Fig. 5 indicated by the fact that the signal from sensor D1 j is led downwards out of the figure and on the other side an analog signal from below into the summation point with the signal S6 oij is being led. This is to be understood as meaning that the next controller C is located there. kj is connected, where k≠i. This controller C kjwill then of course be connected to the S5 signal k analogous to the Controller C ij operated. So that the two controllers C ij and C kj The signals must not interfere with each other (S5). i and S5 k with respect to the scalar product by multiplication in the multiplier M 1ij (or M1) kj ) and the subsequent filtering in filter F 1ij (or F1) kj ) be orthogonal.

[0156] The compensation by the compensation transmitters K j The compensation takes place within the medium and is therefore largely independent of problems such as sensor contamination and aging. However, a compensation transmitter is always required. In the case of an optical system and an LED as the compensation transmitter K j However, this is not absolutely necessary. It is also conceivable to use the system as in Fig. 6 shown how to build it. Here is what's on Fig. 5 not shown, a base current of the photodiode D1 i shown. This is ensured by the resistor R. The compensation transmitter is the one shown. Fig. 6. Drawn power source K j This emulates the compensation luminous flux I2. j out of Fig. 5 by forcing a current through the resistor R, which corresponds to the current produced by the luminous flux I2 j in Fig. 5 in the D1 photosensor j generated electric current.

[0157] Since the input of the sensors D j Since the signal should always be at the same operating point, experience has shown that it makes sense to use a signal S5. i to define for an i that is a constant signal. In this case, the value of the compensation current K is jAn equivalent value that only changes when the equivalent value of the signal received by Dj changes. Typically, this results in chromatic aberration compensation in optical systems. Therefore, the associated transmitter H i This equivalent value is not typically used and can therefore be omitted.

[0158] With regard to compensation, both systems are therefore the same. The system in Fig. However, version 5 has the advantage that the compensation signal I2j and the measurement signals I1i, I3j are subject to the same influences and are therefore typically affected in the same way. The system of Fig. Therefore, version 5 is technically better, which is made up of Fig. The 6 is the cheaper option, as it has one less transmitter.

[0159] Considering what has been said so far, it becomes clear that a proposed measurement system can be constructed vectorially. This is in Fig. 9 drawn.

[0160] The generator bank G produces a vectorial signal S5, whose components are the already known S5 i These are signals. From these, the vector signal S5o is generated by a logic L and a vector delay unit Δt. Its components are the already known S5. di and S50 i .

[0161] The vectorial generator signal S5 is provided with a vectorial offset B1, if necessary, and fed to the transmitter bank H. Therefore, the eight transmitters from Fig. 4 can be understood as a physical representation of a vector. The elements of this transmitter vector H are the already known transmitters H i Each of the transmitters H i The transmitter vector H sends a signal into a transmission channel I1 i into this I1 iThese in turn form a vectorial transmission channel I1. This leads to at least one object O. (The discussion of multiple objects is not pursued here, as this is readily apparent to a person skilled in the art.) From this object O, a transmission channel I3 leads to each of the following: j to one sensor D1 each j The I3 j The D1j form the vectorial transmission channel I3. The D1j form the vectorial sensor bank D1. Its vectorial output signal is amplified and multiplied in the multiplier M by each signal of the vectorial S5o signal. It is readily apparent that at this point the system expands considerably in terms of the necessary complexity. A reduction to a subset of multiplications and / or time-division multiplexing therefore seems appropriate here. In this respect, it is not a mathematically ideal multiplication. The resulting signal is the vectorial signal S10, whose components are the already known signals S10o.ij and S10d ij These are each filtered individually in the vectorial filter bank F, where this filter consists of the known filters F1. ij and F2 ij is composed of the following: For each of the filtered signals in a vectorial amplifier V, which is again composed of the known amplifiers V1, the following applies: ij and V2 ij out of Fig. 5. The vectorial output signal S4 of the vectorial amplifier bank V is composed of the signals S4o. ij and S4 dij out of Fig. 5 together. Each element S4o is multiplied again. ij and S4 dij of the signal vector S4 with the associated signals S5 di and S50 i The resulting signal S6d is then calculated for all i. ij and S6 oij to the respective component S6 jThe vectorial signal S6 is summed. If necessary, the vectorial bias signal B2 is added to this to form the vectorial compensation signal S3. This drives the vectorial compensation transmitter bank K. This transmitter then feeds the signal back into the vectorial sensor bank D via the vectorial transmission path I2. As previously discussed, the system is designed so that fluctuations in the signal are compensated for via the transmission path I3.

[0162] The recognition result (S4) is passed as a data stream to the feature extraction unit (FE). This unit generates a feature vector pFV which is not yet maximized with respect to selectivity. This vector is maximized with respect to selectivity by multiplication with the LDA matrix (LDA_F) (14). The result is the feature vector FV (38). From this, the emission calculation unit (EC) (12) generates the hypothesis list HL (39). Using the sequence lexicon SL (16), the Viterbi search unit (VS) (13) generates the list of recognized gesture sequences G (22). Simultaneously, a prediction for the next expected gesture Gn is created. This vector is modified, for example, by multiplication using the illumination matrix (LDA_B) (40) and used to modify the generator signals S5. Fig. 9 This is done by multiplying the vector of generator signals signal by signal by signal by signal by signal by vector thus obtained. Advantages of the proposal

[0163] The proposal enables the immediate differentiation of gesture sequences while simultaneously improving immunity to ambient light and improving the suppression of recognition errors through redundancies.

[0164] It is particularly suitable for mobile systems and internet use. Figures Fig. 1: Exemplary functional sequence of speaker-independent gesture or gesture sequence recognition with HMM recognition Fig. 2: Exemplary functional sequence of speaker- and device-independent gesture or gesture sequence recognition with HMM recognition Fig. 3: Exemplary functional sequence of speaker-dependent gesture or gesture sequence recognition with a neural network recognizer Fig. 4: Positioning of the transmitters (especially transmitter diodes) H i , the receiver (especially receiver diodes) D jand the compensation transmitter (especially compensation diodes) D j using the example of a mobile phone with eight exemplary transmitters, compensation transmitters and receivers. Fig. 5: Example of control loop C ij for the compensation regulation of the compensation transmitter K j Fig. 6: Example of control loop C ij for the compensation regulation of the compensation transmitter K j where the compensation transmitter is a power source. Fig. 7: Exemplary functional sequence of speaker- and device-independent gesture or gesture sequence recognition with HMM recognition, influencing the transmitter's transmission pattern depending on the recognition result Fig. 8: Exemplary functional sequence of speaker-dependent gesture or gesture sequence recognition with a neural network recognizer, influencing the transmitter transmission pattern depending on the recognition result Fig. 9: Exemplary vectorial overall control loop for the compensation control of the compensation transmitters K and the adaptation of the transmission pattern of the transmitters T to the detection result. Fig. 10: Diagram illustrating the emission calculation process Fig. 11: Schematic representation of the principle of a mechanically automated calibration device for the reproducible execution of standardized gestures Fig. 12: Rotating the hand around the longitudinal axis Fig. 13: Rotating the hand around the axis perpendicular to the palm Fig. 14: Rotating the hand around the axis perpendicular to the longitudinal axis in the palm Fig. 15: Spreading a different number of fingers (0, 1, 2, 3). 4 and 5 fingers are not drawn for clarity. Fig. 16: Changing the hand shape from a fist to a flat hand Fig. 17: Spread fingers Fig. 18: Moving the hand up and down (Z-movement) Fig. 19: Translational movement perpendicular to the device (x-movement) Fig. 20: Translational movement parallel to the device (y-movement) Fig. 21: Circular movement around the y-axis above the device (x-axis would be analogous) Fig. 22: Circular motion around the z-axis above the device Fig. 23: HMM Model Fig. 24: Single word recognition (single models) Fig. 25: Continuous Gesture Recognition Fig. 26-57: Examples of possible activity patterns List of designations 1 Physical Value Stream / Stream of physical quantities 2 Interaction Region / Area of ​​Interaction 3 First Transmitter / First Sender 4 First transmission path 5. Interaction or modification 6 Second transmission line 7 Sensor or receiver 8 Controllers / Regulators 9 Compensation Transmitter 10 Transmission path from the compensation transmitter (9) to the sensor (7) 11 Feature Extraction 12 Emission Computation / Emission Calculation 13 Viterbi Search / Viterbi decoder 14 LDA Matrix 15 Prototype Book Code Book or CB) 16 Sequence Lexicon / Sequence Lexicon 17 Training tool for creating the LDA matrix (14) and the prototype book (15) 18 Database of previously recorded gesture feature vectors of known elementary gestures in known elementary gesture sequences of a statistical selection of gesture speakers in a statistical selection of gesture situations, preferably including norm gestures and norm gesture sequences for recalibrating the gesture recognizer when the physical interface changes (23) 19 Type-In tool for manually entering recognizable elementary gesture sequences (gestures) 20 Online training tool for inputting new elementary gesture sequences (gestures via the physical interface (23) (decomposition into previously known elementary gestures) 21 Recognized Gesture Parameters 22 Recognized Gesture Sequences 23 Physical Interface 24 Stream of physical values ​​from the Physical Interface (23) This is transformed in the Feature Extraction (11) into a stream of gesture feature vectors. 25 Application-specific training 26 Application-specific LDA matrix (This is always changed when the physical interface (23) is changed.) 27 Neural network 28 Training tools for neural networks 29 Gesture parameters detected by the neural network 30 Example Mobile Phones 31-inch screen 32 Exemplary Artificial Robot Hands for Defining and Demonstrating a Standard Gesture 33 Example Device 34 Exemplary Elementary Gestures 35 Example joints with three rotational degrees of freedom 36 exemplary arms, each with one translational degree of freedom (length) 37 Signals from other sensors and measuring systems 38 Modified Feature Vector Data Stream List of 39 Elementary Gesture Hypotheses 40 Illumination Matrix (LDA B) 41 Prototype 1 42 Prototype 2 43 Prototype 3 44 Prototype 4 45 Unidentified gesture 46 Unregistered gestures 47 Threshold ellipsoid 48 Recognized Gestures A ij Amplitude level of the radiation from transmitter H i into receiver D j B1 Constant vector signal that is added to the vector signal S5. The result of this addition powers the transmitter bank H. B2 Constant vector signal that is added to the vector signal S6. The result of this addition, S3, powers the compensation transmitter bank K. CB Code Book (Prototype database (15)) CBE 1 Code Book entry (= entry in the prototype database (15)) Cij Controller for regulating the signal of the compensation diode K j depending on the transmitter H i Each transmitter must have its own S5 signal in relation to all other signals. iorthogonal S5 i Signal output. K1 Compensation transmitter 1 (Compensation LED 1) j=1 K2 Compensation transmitter 2 (Compensation LED 2) j=2 K3 Compensation transmitter 3 (Compensation LED 3) j=3 K4 Compensation transmitter 4 (Compensation LED 4) j=4 K5 Compensation transmitter 5 (Compensation LED 5) j=5 K6 Compensation transmitter 6 (Compensation LED 6) j=6 K7 Compensation transmitter 7 (Compensation LED 7) j=7 K8 Compensation transmitter 8 (Compensation LED 8) j=8 D1 Sensor 1 (Photodiode 1) j=1 D2 Sensor 2 (Photodiode 2) j=2 D3 Sensor 3 (Photodiode 3) j=3 D4 Sensor 4 (Photodiode 4) j=4 D5 Sensor 5 (Photodiode 5) j=5 D6 Sensor 6 (Photodiode 6) j=6 D7 Sensor 7 (Photodiode 7) j=7 D8 Sensor 8 (Photodiode 8) j=8 Di Sensor j (Photodiode j) j=j The ijDelay level of the radiation from transmitter H i into receiver D j DSP Digital Signal Processor Δt Vectorial delay unit. This serves the logic L to generate the vectorial signal S5o from the vectorial signal S5. F Vectorial filter bank consisting of all filters F1 ij and F2 ij F1 ij Filter 1 of controller C ij F2 ij Filter 2 of controller C ij FV Feature Vector (38) G Vectorial Generator Bank. GS Recognized Gesture Sequence (22) Gn Expected next gesture (vector) H1 Transmitter 1 / Transmitter 1 / Transmitting diode 1 i=1 H2 transmitter 2 / transmitter 2 / transmitting diode 2 i=2 H3 transmitter 3 / transmitter 3 / transmitting diode 3 i=3 H4 transmitter 4 / transmitter 4 / transmitting diode 4 i=4 H5 transmitter 5 / transmitter 5 / transmitting diode 5 i=5 H6 transmitter 6 / transmitter 6 / transmitting diode 6 i=6 H7 transmitter 7 / transmitter 7 / transmitting diode 7 i=7 H8 transmitter 8 / transmitter 8 / transmitting diode 8 i=8 Hi Transmitter i / Transmitter i / Transmit diode ii=i HL Elementalgesten Hypothesis List (39) I1 Vectorial first transmission step from the vectorial transmitter bank H to the object O. I2 Vectorial second transmission path from the vectorial compensation transmitter bank K to the object O. I3 Vectorial third transmission path from object O to vectorial sensor bank D. I1i First steep transmission section from the transmission H i to object O. I2 j Second transmission path from compensation transmitter K j to object O. I3 j Third transmission path from object O to sensor D j . K1 compensation transmitter 1 (compensation LED 1) j=1 K2 compensation transmitter 2 (compensation LED 2) j=2 K3 compensation transmitter 3 (compensation LED 3) j=3 K4 compensation transmitter 4 (compensation LED 4) j=4 K5 compensation transmitter 5 (compensation LED 5) j=5 K6 compensation transmitter 6 (compensation LED 6) j=6 K7 compensation transmitter 7 (compensation LED 7) j=7 K8 compensation transmitter 8 (compensation LED 8) j=8 K j Compensation transmitter j (compensation LED j) j=j The logic L generates the vector signal S5o from the vector signal S5 using the delay unit Δt. LDA F LDA Matrix (14) LDA B Illumination Matrix (40) M1 ij Multiplexer 1 of controller C ij M2 ij Multiplexer 2 of controller C ij O object pFV Feature vector (e.g. after logarithmization, framing and filtering etc.) before multiplication with the LDA matrix (14) S3 Vectorial transmit signal for the vectorial compensation transmitter bank K. The signal is generated by vectorial addition of vector B2 with S6. S3 j Transmission signal for the compensation transmitter K j The signal is created by adding signal B2. j with S6 j . S4d ij Amplitude level of the radiation from transmitter H i into receiver D j S4o ij Delay level of the radiation from transmitter H i into receiver D j S5 Vectorial transmit signal consisting of the signals S5 i S5o Vectorial signal consisting of all signals S5o i and S5di S5o i Gate signal for measuring the reception delay of signal S5 i of the sender Hi S5di gate signal for receiving amplitude measurement of the signal S5 i of the sender Hi S5 i Transmitter signal of transmitter H i . Here, this signal S5 i to all other signals S5 i be orthogonal. S6 Vectorial signal consisting of all signals S6o i and S6di S6o i Back-transformed reception delay measurement of signal S5 i of the sender Hi S6di Back-transformed signal amplitude measurement of the S5 signal i of the sender Hi S6 j Compensation pre-signal for controlling the compensation transmitter K j This signal may be followed by a constant B1. j added up. S10 Vectorial output signal of the multiplexer M consisting of all signals S10d ij and S10o ij S10d ij Output signal of the multiplexer M1 ijof the controller C ij S10o ij Output signal of the multiplexer M2 ij of the controller C ij SL Sequence Lexicon (16) V amplifier bank consisting of all amplifiers V1 ij and V2 ij V1 ij Amplifier 1 of Controller C ij V2 ij Amplifier 2 of Controller C ij

Claims

[1] Measurement system for recording object parameters of an object (O), where the object (O) can perform a z-movement, an x-movement and a y-movement during the capture, wherein the measuring system has a physical interface (23) and where the physical interface (23) - at least one controller (C ij ) and - at least one transmitter (H i ) for light and - at least one receiver (D j ) for the transmitter that has at least one transmitter (H i ) emitted signal and wherein the measurement system includes at least one first unit in the form of an HMM gesture sequence recognition engine and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to perform the feature extraction (11) and emission calculation (12) steps, and where at least one controller (C) ij ) the physical interface (23) is set up for this purpose, - that of at least one transmitter (H i ) emitted signal and - with at least one recipient (D) j ) to compare the received signal and - to determine the comparison result as a multidimensional signal (24), and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to perform the feature extraction (11) and emission calculation (12) steps on the basis of this multidimensional signal (24), and wherein the feature extraction step (11) includes at least one of the following sub-steps: - Dividing the multidimensional signal (24) into individual frames of defined length, - Filtering of the multidimensional signal (24), - Normalization of the multidimensional signal (24), - Orthogonalization of the multidimensional signal (24), - nonlinear mapping of the multidimensional signal (24), - Formation of derivatives of the values ​​of the multidimensional signal generated in this way (24) and wherein the first unit in the form of the HMM gesture sequence recognition engine has a sequence lexicon (16) and wherein the sequence lexicon (16) contains sequence prototypes and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to evaluate the temporal sequence of hypothesis lists (39) of the successive frames and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to determine from this sequence of hypothesis lists (39) the most probable of the predefined sequence prototypes for a sequence of recognized prototypes. [2] Measuring system according to claim 1, including at least one compensation transmitter (K j ) includes and where at least one compensation transmitter (K j ) is set up to do so, into a transmission channel (I2) j ) between this compensation transmitter (K j ) and at least one recipient (D j ) to feed in a signal, - that this corresponds to the signal of at least one transmitter (H i ) superimposed in such a way that at least a predefined part of the transmitter's signal (H) is cancelled out. i ) at the recipient (D j ) results in and where the controller (C ij ) and the compensation transmitter (K j ) are set up in this way, - the signal of the compensation transmitter (K j ) using the controller (C ij ) to generate and regulate in such a way that said annihilation occurs. [3] Measuring system according to one of the preceding claims, the first unit, in the form of the HMM gesture sequence recognition engine, is set up to - those from the controllers (C ij ) supplied data stream in the form of the multidimensional signal (24) in the further course of processing with a speaker- and object-independent and device-dependent LDA matrix (26). [4] Measuring system according to one of the preceding claims, the first unit, in the form of the HMM gesture sequence recognition engine, is set up to - those from the controllers (C ij ) supplied data stream in the form of the multidimensional signal (24) to be multiplied by an application-independent matrix in the further course of processing. [5] Measuring system according to any one of the preceding claims, the first unit in the form of the HMM gesture sequence recognition engine comprises a prototype database (15) and where this prototype database (15) contains prototypes from sample data streams and the first unit, in the form of the HMM gesture sequence recognition engine, is set up to - To compare pattern vectors (38) output by feature extraction (11) within the framework of emission calculation (12) by distance calculation, in particular Euclidean distance calculation, with the distances to prototypical vectors, the said prototypes. [6] Measuring system according to one of the preceding claims, the first unit in the form of the HMM gesture sequence recognition engine comprises the prototype database (15) and where this prototype database (15) contains prototypes from sample data streams and wherein the first unit in the form of the HMM gesture sequence recognition engine is set up to output at least one hypothesis list (39) for the recognized prototypes of each frame as a result of the emission calculation (12), and where the hypothesis list (39) contains the most probable of these prototypes of a frame with the respective probability of recognition and the first unit, in the form of the HMM gesture sequence recognition engine, is set up to form the hypothesis list (39), and where the hypothesis list (39) can only contain one element. [7] Measuring system according to one of the preceding claims and claim 6, wherein the measurement system is set up such that a hypothesis list (39) includes at least one of at least one transmitter (H i ) the transmitted signal is affected. [8] Measuring system according to one of the preceding claims, that the first unit in the form of the HMM gesture sequence recognition engine includes a device-independent prototype database (15). [9] Device, in particular a robot or computer or smartphone, wherein the device has a measuring system in the form of a measuring system according to one or more of the preceding claims and wherein the device is configured to be able to control at least one actuator by means of at least one elementary gesture which is stored as a prototype in the prototype database (15) of the first unit in the form of the HMM gesture sequence recognition engine.

Citation Information

Patent Citations

  • Mode-based graphical user interfaces for touch sensitive input devices

    US20060026535A1

  • Front-end signal compensation

    US20080158178A1

  • System and method for recognizing touch typing under limited tactile feedback conditions

    US6677932B1

  • Optoelectronic measuring device with stray light compensation provided by intensity regulation of additional light source

    DE10300223B3