Measuring system for recording object parameters of an object and device

DE102012025973B4Active Publication Date: 2025-08-07ELMOS SEMICON AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102012025973
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2012-05-23
Publication Date
2025-08-07
Estimated Expiration
2032-05-23

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Measuring system for recording object parameters of an object (O), where the object can perform a z-movement and an x-movement and a y-movement during the capture, wherein the measuring system has a physical interface (23) and where the physical interface (23) - at least one controller (C ij ) and - at least one transmitter (H i ) for light and - at least one recipient (D j ) for the signal from the at least one transmitter (H i ) emitted signal and wherein the measuring system comprises at least a first unit in the form of an HMM gesture sequence recognition engine and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction (11) and emission calculation (12), and wherein the at least one controller (C ij ) the physical interface (23) is arranged to - that of the at least one transmitter (H i ) emitted signal and - with at least one by the at least one recipient (D j ) received signal and - to determine the comparison result as a multidimensional signal (24) and that the first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction (11) and emission calculation (12) on the basis of this multidimensional signal (24), and that the feature extraction (11) step comprises at least one of the following sub-steps: - dividing the multidimensional signal (24) into individual frames of defined length, - filtering of the multidimensional signal (24), - Normalization of the multidimensional signal (24), - Orthogonalization of the multidimensional signal (24), - non-linear mapping of the multidimensional signal (24), - Formation of derivatives of the thus generated values of the multidimensional signal (24) and wherein the first unit in the form of the HMM gesture sequence recognition engine comprises a prototype database (15) and wherein prototypes from example data streams are stored in this prototype database (15) and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to - to compare pattern vectors (38) output by the feature extraction (11) with the distances to prototypical vectors, the said prototypes, by means of distance calculation, in particular Euclidean distance calculation, within the framework of the emission calculation (12).
Need to check novelty before this filing date? Find Prior Art

Description

Subject of the proposal

[0001] The proposal relates to a measuring system for detecting object parameters of an object (O), wherein the object can perform a z-movement and an x-movement and a y-movement during detection. The measuring system has a physical interface (23). The physical interface (23) has at least one controller (C ij ) and at least one transmitter (H i ) for light and at least one receiver (D j ) for the signal from the at least one transmitter (H i ) emitted signal. The measuring system has at least a first unit in the form of an HMM gesture sequence recognition engine. The first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction (11) and emission calculation (12). The at least one controller (C ij) of the physical interface (23) is designed to receive the signal from at least one transmitter (H i ) and to generate a signal emitted by at least one receiver (D j ) and to determine the comparison result as a multidimensional signal (24). The first unit in the form of the HMM Gesture Sequence Recognition Engine is configured to perform the steps of feature extraction (11) and emission calculation (12) based on this multidimensional signal in the form of one or more streams (24). The feature extraction step comprises at least one of the following substeps: - dividing the multidimensional signal in the form of one or more streams (24) into individual frames of defined length, - filtering the multidimensional signal in the form of one or more streams (24), - normalization of the multidimensional signal in the form of one or more currents (24), - orthogonalization of the multidimensional signal in the form of one or more streams (24), - non-linear mapping of the multi-dimensional signal in the form of one or more currents (24), - Formation of derivatives of the thus generated values of the multidimensional signal in the form of one or more currents (24).

[0002] The first unit, in the form of the HMM gesture sequence recognition engine, comprises a prototype database (15). Prototypes from example data streams are stored in this prototype database (15). The first unit, in the form of the HMM gesture sequence recognition engine, is configured to compare the modified feature data stream vectors (38) output by the feature extraction (11) with the distances to prototypical vectors, the aforementioned prototypes, within the framework of an emission calculation (12) by means of distance calculation, in particular Euclidean distance calculation. Introduction

[0003] The continuous recording of gestures and their parameterization and assignment parameters is an important basic function in many, especially mobile applications such as the control of computers, robots, mobile phones and smartphones.

[0004] A major problem with many gesture recognition systems is that they rely on processing image signals. This requires cameras and complex image processing devices. These require significant amounts of resources, such as computing power, memory, and electrical energy. However, in a mobile phone, these resources are very limited. This particularly applies to power consumption, storage space, and the lifespan of the processing units.

[0005] Alternative systems have lower sensitivity, which impairs recognition performance. Furthermore, these systems are not application-independent, requiring the time-consuming and therefore very expensive re-recording of gesture databases. Furthermore, most systems are sensitive to ambient light and other interference.

[0006] Finally, most application systems have additional sensors or sensory capabilities that should be combined with the operation of a gesture recognition system.

[0007] The following describes the state of the art: A number of disclosures are known from the prior art that implement gesture recognition for a touch pad. For example, US6677932B1 is based on a "Touch Data Structure" (designation 79, Fig. 3 in US6677932B1). It is therefore a gesture recognition device that requires touching the touchscreen. Consequently, only the x and y coordinates are used in this disclosure" (designation 80,

[0008] Fig.3 in US6677932B1). The determination of a movement in the Z direction cannot be performed. Crucially, however, this finding does not provide feedback to the physical sensor technology. The physical recognition process is not adapted to the already recognized sequence and to the two possible gestures to be distinguished (subsequent gesture hypothesis). As a result, the recognition result for the following gesture is not optimized.

[0009] The same applies to US20060026535A1. Here, too, a “touch screen” (designation 102, Fig. 2, US20060026535A1). As with US6677932B1, this is only a method for two-dimensional gesture recognition, which naturally allows for far fewer artistic design options. Here, too, the gesture recognizer configuration and thus the recognition result are not optimized for the subsequent gesture in a gesture sequence.

[0010] Furthermore, US20060026535A1 assumes a capacitive method, which is not claimed here.

[0011] Both US20060026535A1 and US6677932B1 are sensitive to interference, such as sunlight, because they lack a compensation signal that balances the measured signal at the sensor. This makes these methods susceptible to contamination and sensor drift.

[0012] Such a compensation signal is described in US20080158178A1 (claim 1, US20080158178A1). However, this document also describes a touchscreen. Therefore, only two-dimensional recognition is possible here. Here, too, the recognized gestures are not used to precondition the system for optimal recognition of the next gesture.

[0013] The article by Schuller, B.; Lang, M.; Rigoll, G.; Multimodal Emotion Recognition in Audiovisual Communication in ICME '02 2001, pp. 745-748 describes a DTW recognition system (Section 2 of the article) that is known to be speaker- and device-dependent, and therefore must be trained using test data sets. This requires device- and speaker-dependent gesture databases, which either prevent industrial production of such devices or require the user to undergo time-consuming training before using the device. This is generally a hurdle that the average consumer is unwilling to overcome. Furthermore, training by untrained individuals is usually inaccurate.

[0014] Section 2 of this article also refers to a Z component. However, since the sensor used is a SAW sensor, the Z component is detected by the contact pressure. Therefore, detaching the operator's hand from the touch pad and thus true 3D recognition are not possible.

[0015] Although the article evaluates the dialog history (section 3.4 of the article), there is no feedback to the physical level of the recognition system through reconfiguration of the sensor system. Thus, the sensor system is only optimized to an average value for all of these scripts. An extrapolation of the optimal setting of the physical parameters for distinguishing the presumably most likely next gestures does not occur.

[0016] What all the documents specifically cited above have in common is that they require direct contact between the operating hand and the sensor interface and are therefore not 3D-capable and only allow the detection of movement in the Z direction.

[0017] DE10300223B3 discloses an optoelectronic measuring arrangement comprising at least two first light sources that emit light in a time-sequential, time-synchronized manner, in phases, and comprising at least one receiver for receiving at least the clock-synchronous alternating light component originating from the first light sources, as well as a device for compensating for extraneous light by regulating the light intensity radiated into the measuring arrangement by at least one light source, so that the clock-synchronous component occurring between different phases becomes zero. The special feature of this device is that the at least one further light source used for the control is a light source independent of the first light sources, which is assigned to the receiver and whose light intensity can be regulated in terms of amplitude and sign.DE10300223B3 thus discloses a technical teaching that, in contrast to the previously mentioned documents, no longer requires contact between the operating hand and the device surface. However, the performance of this method is still limited in that recognized gestures do not influence the measurement itself. Thus, the false acceptance rate—that is, the rate of incorrectly accepted gestures—and the false rejection rate—that is, the rate of incorrectly rejected gestures—are unnecessarily increased. DE10300223B3 is only capable of determining a distance. DE10300223B3 discloses neither a method nor a device for associating the measured distances with gestures and identifying them. Task

[0018] The proposal aims to provide a speaker- and largely device-independent gesture interface for human-machine interfaces while simultaneously achieving low resource consumption and high robustness against ambient light.

[0019] This object is achieved by a device according to claim 1. Description of the basic idea of the proposal

[0020] The invention relates to a measuring system for detecting object parameters of an object. The object can perform a z-movement, an x-movement, and a y-movement during detection. Furthermore, the measuring system has a physical interface. According to the invention, the physical interface has at least one controller, at least one transmitter for light, and at least one receiver for the signal emitted by the at least one transmitter. According to the invention, the measuring system has at least a first unit in the form of an HMM gesture sequence recognition engine. According to the invention, the first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction and emission calculation.According to the invention, the at least one controller of the physical interface is configured to generate the signal emitted by at least one transmitter and to compare it with at least one signal received by the at least one receiver, and to determine the comparison result as a multidimensional signal. The first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction and emission calculation based on this multidimensional signal. According to the invention, the feature extraction step comprises at least one of the following substeps: • Dividing the multidimensional signal into individual frames of defined length, • Filtering of the multidimensional signal, • Normalization of the multidimensional signal, • Orthogonalization of the multidimensional signal, • non-linear mapping of the multidimensional signal, • Formation of derivatives of the thus generated values of the multidimensional signal.

[0021] According to the invention, the first unit in the form of the HMM gesture sequence recognition engine comprises a prototype database, wherein prototypes from example data streams are stored in this prototype database. According to the invention, the first unit in the form of the HMM gesture sequence recognition engine is configured to compare the pattern vectors output by the feature extraction with the distances to prototypical vectors, the aforementioned prototypes, within the framework of an emission calculation using distance calculation, in particular Euclidean distance calculation.

[0022] The following text explains the invention and its context in more detail.

[0023] The proposal is divided into the following parts: - The actual detection system, in particular a multidimensional feedback optical measuring system - A smartphone or robot suitable for use in the proposed detection system - A processing method and system for determining a gesture recognition result, which can be carried out fully automatically by the recognition system. - A procedure for adapting the proposed system and method and the associated databases to a specific hardware platform - A method for recording device and speaker independent gesture databases - The basic structure of an associated gesture language for controlling a smartphone and similar applications in which a system is to be influenced by gestures.

[0024] These are described in the following sections. Example recognition system

[0025] A first exemplary proposed measurement system primarily consists of a triangulation-capable optical measurement system, typically with a transmitter system and a receiver system, or a combined transmitter / receiver system that can switch very quickly between transmit and receive phases. Each transmitter system can consist of multiple transmitters. Likewise, each receiver system can consist of multiple receivers. The same applies to combined receiver / transmitter systems.

[0026] To clarify the operation, the device is first described using a system with only one transmitter and one receiver using the schematic Fig. 1. It should be noted that all figures in this document contain only those details that a person skilled in the art needs to understand and comprehend the proposed idea.

[0027] The exemplary device consists of at least one first transmitter (3) that transmits into a first transmission path consisting of transmission sections (4) and (6). This transmission path (4, 6) can be divided, for example, by an object into the said first transmission section (4) and the second transmission section (6). Further serial and parallel divisions into further serial and parallel transmission sections are conceivable. (This is equivalent to the possibility of a unidirectional transmission network with inputs and outputs.)

[0028] At the end of the first transmission section (4) there is an interaction zone (2) in which the flow of physical variables (1), for example the location, orientation and surface condition of an object, in particular a hand, interacts (5) with the signal current of the transmitter (3) emerging from the first transmission section (4). This interaction (5) changes the properties of the transmission signal of the transmitter (3). These properties can be, for example, amplitude, phase, spectrum, etc. The object thus imprints traces of its properties on the transmitter signal, which are later used to enable conclusions to be drawn about the object.

[0029] After this modification (5), the signal from the transmitter (3) typically passes through the second transmission section (6) and is picked up by a receiver or sensor (7).

[0030] The problem of independence from environmental influences remains. To achieve this, the received signal at the receiver (7) is typically supplemented by a compensation signal while still in the physical medium, resulting in a virtually constant overall signal. A compensation transmitter (9) sends a corresponding compensation signal over a second transmission path, which is superimposed in the receiver (7) on the signal from the transmitter (3) after it has passed through the transmission paths (4, 6). For the sake of completeness, it should be mentioned that the compensation signal can be generated both in the medium, here as light, for example, or as a signal, here as an electrical current signal emulating the photocurrent of an LED, fed directly into the amplifier line.

[0031] For this purpose, a controller (8) (labeled "controller") controls a compensation transmitter (9), which also feeds a signal into the receiver (7) via a predefined and typically separate transmission path (10) that is essentially unaffected by the flow of physical influences (1). There, the two signals overlap, as mentioned above, either linearly or non-linearly. Ideally, however, the receiver (7) is a linear receiver in which the signals from the transmitter (3) (after modification (5)) and the signal from the compensator (9) overlap linearly. If the receiver (7) is non-linear, a component proportional to the product of the two overlapping signals is added. Even if this case can be taken into account in terms of control technology, it will not be explained in detail here. However, it is easily possible for a person skilled in the art to modify the controller described here so that such non-linearities can be taken into account at any time.

[0032] The controller (8) also generates the transmission signal S5 of the transmitter (3). It is important when generating the transmission signal S5 that this transmission signal S5 is orthogonal to the transmission signals of other controllers with respect to the scalar product performed in this controller (8). This applies if there are additional controllers in the system.

[0033] The controller (8) provides a vector output signal (24), which is then analyzed. The dimension and content are selected to optimize gesture recognition.

[0034] For the analysis of the vectorial signal (24) output by the controller (8) as well as for the synthesis of the test signal, various methods and the devices or system components suitable for these methods can be used.

[0035] The result of this analysis when controlled by the controller (8) is a set of parameters, which can be supplemented by further parameters, which can be obtained, for example, from the signals of the optional additional sensors (37), touchpads, switches, buttons and other measuring systems (37). The result is a so-called quantization vector, the components of which, the measured parameters, will generally not be completely independent of each other. Each parameter on its own usually has too low a selectivity for the precise differentiation of gestures in complex contexts. By creating one or more such quantization vectors from the continuous stream of analog physical parameter values (24) at typically regular or regular time intervals by the physical interface (23) (see Fig. 1) by means of suitable means, a time- and value-quantized multidimensional parameter data stream (24) is created.

[0036] The resulting multidimensional signal, in the form of one or more streams (24) of quantization vectors, is first divided into individual frames of defined length, filtered, normalized, then orthogonalized, and if necessary, suitably distorted by a nonlinear mapping—e.g., logarithmization and cepstrum analysis. Furthermore, derivatives of the thus generated values are also formed here. This method has long been known as feature extraction (11) in pattern recognition, specifically speech recognition and material recognition. The possibilities are manifold and can be found in the literature. Finally, the multidimensional quantization sector is multiplied by a so-called LDA matrix.

[0037] The next step of detection can be carried out using two different methods: a) by a neural network or b) by an HMM recognizer c) by a Petri net First, the HMM recognizer (Fig. 1) is described:

[0038] Using the aforementioned predefined LDA matrix (14), the modified feature data stream vectors (38) are mapped from the multidimensional input parameter space to a new parameter space, thereby maximizing their selectivity. The components of the resulting new transformed feature vectors are selected not according to real physical or other parameters, but rather according to maximum significance, resulting in said maximum selectivity.

[0039] The LDA matrix was usually calculated offline based on sample data streams from the database (18) with known gesture datasets, previously through training (17). In the case of device-independent gesture recognition, the procedure deviates from the standard procedure in one crucial step: ( Fig. 2)

[0040] If care is taken that all elements of the feature extraction (11) consist of at least locally reversible functions, deviations in the geometry etc. can be taken into account in the form of an approximately linear transformation function.

[0041] Therefore, if the LDA matrix (14) has been determined for a geometric arrangement of transmitters and receivers in a well-defined application, it can be adapted to the respective geometry of the specific device by simple matrix multiplication with a device / application-specific LDA matrix (26). ( Fig.2) This results in a device-independent and speaker-independent gesture recognizer that can be easily transported from device to device - for example, from a game console to a smartphone.

[0042] This matrix (26) can be determined, for example, by having a number of test users perform predefined and standardized test gestures for the specific gesture recognition application to be recognized. Since the gestures are known, a simple linear transformation between the new vectors and the vectors previously recorded on another device in the database (18) can be calculated using a corresponding application-specific training tool (25).

[0043] In order to make the test gestures reproducible worldwide, it is advisable to have them performed in a standardized manner by suitable devices (e.g. motorized dolls or robots). ( Fig.9) This allows these artificial gestures to be compared precisely with the stored prototypes for these robot-based reference gestures. This is done using Fig. 11. It shows an exemplary mobile phone (33) above which a model hand (32) is located. This model hand can, for example, be the hand of a mannequin. The hand is moved by a segmented robot arm with joints and three rotational degrees of freedom (35). Between the segments (36) are joints (35). Fig.Figure 9 does not show exactly how the hand is moved, i.e., by which motor, but only how the recording of the standard gestures / standardization gestures works. If the size of the hand (32) and its shape, color, etc., as well as its orientation and distance from the mobile phone (33), the lighting, the reflective properties of the container and the hand itself, in which the entire device and the mobile phone (33) are located, and the movement, indicated here by a swiping movement as an elementary gesture (34), are defined, the gesture recognition result of such a defined movement (34) can be predicted and thus compared with the obtained result.

[0044] The result of this calculation is an application-specific LDA matrix (26). (see also Fig. 2) The advantage is that the feature data stream vectors (38) leaving the feature extraction are speaker-independent and largely device-independent.

[0045] The key economic advantage of this method for transferring a gesture database from one device type to another is that the expensive recording of the gesture databases only needs to be performed once with real gesture speakers. This cost can be as much as €1 million per application.

[0046] The other corresponding methods are known from speech and material recognition. At the same time, the prototypes are calculated from these sample data streams in the database (18) in the coordinates of the new parameter space and stored in a prototype database (15). In addition to these statistical data, this database can also contain instructions for a computer system regarding what should happen in the event of successful or failed recognition of the respective prototype. Typically, the computer system will be the computer system (e.g., a mobile phone) whose human-machine interface (hereinafter referred to as HMI) is intended to represent the proposed gesture recognition.

[0047] The resulting feature data stream vectors (38), which are output by feature extraction (11), are then compared with these previously stored, i.e., learned gesture prototypes of the prototype book (15), for example, by calculating the Euclidean distance between a quantization vector in the coordinates of the new parameter space and all of these previously stored prototypes of the prototype book (15) in the emission calculation (12). At least two recognitions are performed: 1. Does the detected quantization vector of the feature data stream vectors (38) correspond to one of the pre-stored quantization vector prototypes (= previously known gestures) or not and with what probability and reliability? 2. If it is one of the gestures already stored, which one is it and with what probability and reliability?

[0048] To facilitate initial detection, dummy prototypes are usually stored in the prototype database (15), which should cover all parasitic parameter combinations encountered during operation as far as possible. These prototypes are stored in a database (15), which is also referred to as the codebook (15) below.

[0049] If the minimum half of the known Euclidean distance between two quantization vectors of two different prototypes of elementary gestures in new coordinates, also stored in the database / codebook (15), is exceeded, the prototype is considered recognized. From this point on, it can be ruled out that further distances to other prototypes of the codebook (15) calculated during a further search could yield even smaller distances.

[0050] The minimum Euclidean distance is calculated using the following formula: DistFV_CbE=MinCb_cnt=Cb_anz1[∑dim_cnt=dim1(FVdim_cnt−CbCb_cnt,din_cnt)2]

[0051] Here, dim_cnt stands for the dimension index that runs up to the maximum dimension of the feature vector (38) dim.

[0052] FV dim_cnt stands for the component of the feature vector (38) corresponding to the index dim_cnt.

[0053] Cb_cnt stands for the number of the entry in the code book (15), i.e. for the elementary gesture corresponding to Cb_cnt.

[0054] Cb CB_cnt,dim_cnt accordingly stands for the component corresponding to dim_cnt of the prototypical codebook feature vector entry associated with the elementary gesture corresponding to Cb_cnt.

[0055] Dist FV_CbEtherefore represents the obtained minimum Euclidean distance. When searching for the smallest Euclidean distance, the number Cb_cnt that produces the smallest distance is noted.

[0056] For clarification, an example assembly code is provided: Beginning of the code Mov Cb_cnt,#Cb_anz initialize code book vector counter Mov C, #0 initialize register C with 0 Mov dist, maxvalue initialize the distance with maximum value Mov num, not_valid_num initialize the nearest neighbor number with an invalid value Mov Cb_adr, Cb_badr initialize codebook address with base address Label_A: / / next vector Mov SP, #0 initialize cache Mov dim_cnt, #dim initialize dimension counter with feature vector dimension Label B: / / next dimension MovA, $Cb_adr load value absolutely from code book address SubA, $FV_adr, dim_cnt subtract value absolute relative from feature vector value Mov BA load B register with result MulA B multiply A and B (=A 2 ) AddA, SP add result to intermediate result Mov SP, A and remember Dec dim_cnt next vector component Inc Cb_adr increase codebook pointer by one jnz dim_cnt, Label B but only if it wasn't the last Cmp SP, dist rate code book entry jmpgt Label C Mov dist, SP if better entry than previous optimum Mov num, Cb_cnt Remember entry number and distance Label C: dec_Cb_cnt next code book entry jnz Cb_cnt,Label A but only if it wasn't the last End of code

[0057] The confidence measure for a correct recognition is derived from the dispersion of the underlying base data streams for a prototype and the distance of the quantization vector from their center of gravity.

[0058] Fig. Figure 10 demonstrates various detection cases. For simplicity, the representation is chosen for a two-dimensional feature vector so that the methodology can be represented on a two-dimensional sheet of paper. In reality, the feature vectors (38) are typically always multidimensional.

[0059] The centers of gravity of various prototypes (41, 42, 43, 44) are shown. As described above, half the minimum distance between these prototypes can be stored in the codebook (15). This would then be a global parameter that would be equally valid for all prototypes. However, this decision based on the minimum distance assumes that the dispersion of the elementary gesture prototypes (41, 42, 43, 44) is more or less identical. If feature extraction is optimal, this is also the case. This would correspond to a circle around each of the prototypes with a gesture database-specific radius.

[0060] However, this is rarely achievable in reality. An improvement in recognition performance could be achieved by storing the scatter for the prototype of each elementary gesture. This would correspond to a circle around each prototype with a prototype-specific radius. The disadvantage is an increase in computing power.

[0061] A further improvement in recognition performance can be achieved by modeling the scatter width for the prototype of each elementary gesture using an ellipse. Instead of the radius as before, the major axis diameters of the scatter ellipse and their tilt relative to the coordinate system must now be stored. The disadvantage is a further, massive increase in computing power and memory requirements.

[0062] Of course, the calculation can be made even more complicated, but this usually only increases the effort massively and does not significantly improve the recognition performance.

[0063] Compact recognizers for mobile devices will therefore typically resort to the simplest of the options described.

[0064] The position of the gesture vectors determined by the emission calculation can vary considerably. It is conceivable that such a gesture vector of the unregistered gesture (46) is too far away from any other gesture. This distance threshold can, for example, be the minimum half of the prototype distance. It is also possible that the scatter ranges of the prototypes (43, 42) overlap, and a determined gesture vector of the imprecisely identified gesture (45) lies within the overlap range. In this case, a hypothesis list would contain both prototypes with different probabilities, due to different distances.

[0065] In the best case, the determined gesture vector of the recognized gesture (48) lies in the scatter range (threshold ellipsoid) (47) of a single prototype (41) which is thus reliably recognized.

[0066] To improve modeling of the ranges of variation of a single gesture, it is conceivable to model it using multiple prototypes with associated ranges of variation. Thus, multiple prototypes can represent the same gesture. The risk here is that, due to the distribution of the probability of a gesture among several such subprototypes, the probability of the individual subprototype gesture may become smaller than that of another gesture whose probability was lower than that of the original gesture. Thus, this other gesture may falsely prevail.

[0067] A key problem, therefore, is the computing power required to reliably recognize the elementary gestures. This will be discussed briefly: A crucial point is that the computational effort increases with Cb_anz * dim.

[0068] For a non-optimized HMM recognizer, the number of assembly instructions that must be executed to calculate a vector component is approximately 8 steps.

[0069] The number A_Abst of assembly steps required to calculate the distance of a single codebook entry (CbE) to a single feature vector (FV) is calculated approximately as follows: A_Abst=FV_Dimension*8+8

[0070] This results in the number A_CB of assembly steps for determining the code book entry with the smallest distance: A_CB=Cb_anz*(A_Abst)+4=Cb_anz*(FV_Dimension*8+8)+4

[0071] Using the example of a medium HMM recognizer with 50000 CbE (code book entries = CB_anz) and 24 FV_dimensions (feature vector dimensions = FV_dimension), the number of steps is: 50000*(24*8+8)+4~10 million operations per feature vector in the form of streams(24)

[0072] At a relatively low sampling rate of 8kHz = 8000 FV per second (feature vectors in the form of currents (24) per second), a computing power of 8GIpS (8 billion instructions per second) is required.

[0073] Especially for energy-autonomous and / or mobile systems, this simple example already leads to a resource consumption that is unsustainable.

[0074] In the case of an optimized HMM recognizer, as already mentioned, the smallest codebook distance between two prototypes is pre-calculated and stored in the codebook (15). This has the advantage that the search can be aborted if a distance is found that is less than half of this minimum codebook distance. This halves the average search time. Further optimizations can be achieved by sorting the codebook according to the statistical occurrence of the elementary gestures. This ensures that the most frequent elementary gestures are found immediately, further reducing the computation time.

[0075] For such an optimized HMM recognizer, the computing power requirements are as follows:

[0076] Again, there are 8 steps to calculate a vector component. The steps to calculate the distance A_Distance of a codebook entry (CbE) to the feature vector (FV) are: A_Abst=FV_Dimesnion *8+8

[0077] The number of steps for determining the codebook entry with the smallest distance A_CB with optimization is slightly higher: A_CB=Cb_anz*(A_Abst)+4=Cb_anz*(FV_Dimesnion*8+10)+4

[0078] The two additional assembly instructions are necessary to check whether the distance is less than half the smallest code book entry distance.

[0079] Furthermore, the number CB_Anz of code book entries for mobile and energy-autonomous applications is limited to 4000 code book entries or even less.

[0080] In addition, the number of feature vectors per second is reduced by lowering the sampling rate.

[0081] This is explained using a simple example:

[0082] The aforementioned medium HMM recognizer is now operated with a codebook with only about a tenth of the entries, e.g., with 4000 CbE and still with 24 FV dimensions.

[0083] The number of steps is now 4000*(24*8+10)+4~808004_ operations per feature vector

[0084] If the feature vector rate is reduced to 100 FV per second (feature vectors per second) extracted from 10ms parameter windows over, for example, 80 samples each and the search is aborted if the distance is less than half the smallest codebook entry distance, the effort is at least halved with suitable codebook sorting.

[0085] The required computing power then drops to <33 DSP MIPS (33 million operations per second). In reality, codebook sorting leads to even lower computing power requirements, for example, 30 MIPS. This makes the system real-time capable and integrable into a single IC.

[0086] The search space can be narrowed down by pre-selection. A prerequisite for this is an even distribution of the data, i.e., the centroids of the quadrants are in the geometric quadrant center.

[0087] The necessary reduction of the code book (15) has advantages and disadvantages:

[0088] Reducing the codebook (15) increases both the false acceptance rate (FAR), i.e. the incorrect gestures that are accepted as gestures, and the false rejection rate (FRR), i.e. the correctly executed gestures that are not recognized.

[0089] On the other hand, this reduces the resource requirements (computing power, chip area, memory, power consumption, etc.).

[0090] In addition, the history, i.e., previously recognized gestures, can be used to form hypotheses. A suitable model for this is the so-called Hidden Markov Model (HMM).

[0091] For each gesture prototype, a confidence measure and a distance to the measured quantization vector can be derived, which can be used directly by the system or further processed. It is also useful to output a hypothesis list for the recognized prototypes of each frame, for example, containing the ten most probable gesture prototypes for that frame with the respective probability and reliability of recognition.

[0092] With these data of the recognized gesture parameters (21), instructions or executable code correlated with the respective prototype can also be output to the said user computer system (e.g. a mobile phone), which were stored, for example, in the prototype database (18) or otherwise previously.

[0093] Since not every temporal and spatial gesture sequence makes sense, it is possible to evaluate the temporal sequence of the hypothesis lists of the consecutive frames.

[0094] The elementary gesture sequence path through the successive hypothesis lists that has the highest probability must be found.

[0095] Here, too, at least two detections are made: 1. Is the most probable elementary gesture sequence one of the already stored sequence prototypes or not, and with what probability and reliability? 2. If it is one of the sequences already stored, which one is it and with what probability and reliability?

[0096] For this purpose, a sequence prototype can be entered fully automatically into a sequence dictionary (16) using a learning program of the online training tool (20), or manually using a type-in tool (19) that allows these sequence prototypes to be entered via the keyboard. It is also conceivable, with suitable standardization, to download individual gesture sequences, for example, from the Internet.

[0097] Using a Viterbi decoder (13), the most probable of the predefined sequences for a sequence of elementary gestures, or more precisely, the recognized gesture prototypes, can be determined from the sequence of hypothesis lists. This is especially true if individual elementary gestures were incorrectly recognized due to measurement errors. Therefore, it is very useful to transfer sequences of gesture hypothesis lists (39) from the emission calculation (12), as described above, to the Viterbi sequence recognition of the Viterbi decoder (13). The result is the elementary gesture sequence (22) recognized as the most probable or, analogous to the emission calculation, a sequence hypothesis list. Here, too, instructions to the said computer unit, for example a mobile phone, as to what is to be done can be linked to the gesture sequence recognized as the most probable in the sequence database of the sequence lexicon (16) or otherwise. If necessary, a warning message is sent to the user.

[0098] Finally, we consider the functional software components of the gesture sequence recognition engine. This is described in Fig. 1 as a Viterbi search (13). This search (13) accesses the sequence dictionary (16). The sequence dictionary (16) is fed by a learning tool of the online training tool (20) and a tool of the type-in tool (19), where these elementary gesture sequences can be specified through textual input. The possibility of downloading has already been mentioned.

[0099] The basis of sequence recognition (22) is the Hidden Markov Model. The model is built up from different states. In the Fig. In the example given in Figure 23, these states are symbolized by numbered circles. In the example in Fig. 23, the circles are numbered from 1 to 6. There are transitions between the states. These transitions are described in the Fig.23 is denoted by the letter a and two indices i, j. The first index i denotes the number of the source node, the second index j the number of the destination node. In addition to the transitions between two different nodes, there are also transitions a ii or a jj which lead back to the starting node. In addition, there are transitions that allow to skip nodes from the sequence, thus resulting in a probability of a k-th observable b k actually observed. This results in sequences of observables that can be calculated with pre-calculable probabilities b k can be observed.

[0100] It is important to note that every Hidden Markov Model consists of unobservable states q i exists. Between two states q i and q j the transition probability a ij .

[0101] Thus, the probability p for the transition from qi on q j write as: p(qnj|qn−1i)≡aij

[0102] Here, n represents a discrete point in time. The transition takes place between step n and state q i and the step n+1 with the state q j instead of.

[0103] The emission distribution b i Ge) depends on the state q i As already explained, this is the probability of observing the elementary gesture Ge (the observable) when the system (Hidden Markov Model) is in state q i located: p(Ge|qi)≡bi(Ge)

[0104] To start the system, the initial states must be determined. This is done by a probability vector π i . It can then be stated that a state q i with probability π i an initial state is: p(qi1)≡πi

[0105] It is important that a new model must be created for each gesture sequence. In a model M, the observation probability for an observation sequence of elementary gestures Ge⇀=(Ge1,Ge2,…GeN) be determined

[0106] This corresponds to a non-directly observable state sequence that corresponds to the following sequence: Q⇀=(q1,q2,…qN)

[0107] The probability p of observing this state sequence Q, which depends on the model M, the state sequence Q and the observation sequence Ge, is: p(Ge→|Q→,M)=p(Ge1,Ge2,...GeN|q1,q2,.....qN)=p(Ge1|q1)⋅p(Ge2|q2)⋅........p(GeN|qN)=∏n=1Np(Gen|qn)=∏n=1Nbn(Gen)

[0108] This results in the probability of a state sequence Q→=(q1,q2,.....qN) in model M: p(Q→|M)=p(q1,q2,...qN|M)=p(q1)⋅p(q2|q1)⋅p(q3|q1,q2)⋅......p(qN|q1,q2,....qN−1)=p(q1)∏n=1Np(qn|qn−1)=π1∏n=2Na(n−1)n

[0109] This results in the probability of recognizing a gesture word equal to a gesture sequence: p(Ge→|Mj)=∑all Qkp(Ge→|Qk,Mj)p(Q→k|Mj)=∑all Qk(∏n=1Nbn(Gen))(π1∏n=2Na(n−1)n)

[0110] The most probable word model (gesture sequence) for the observed emission Ge is determined by summing the individual probabilities over all possible paths Q k that lead to this observed gesture sequence Ge. p(Ge→|Mj)=∑all Qkp(Ge→|Qk,Mj)p(Q→k|Mj)=∑all Qk(∏n=1Nbn(Gen))(π1∏n=2Na(n−1)n)

[0111] The summation over all possible paths Q is not without its problems due to the potential computational effort. Therefore, the process is usually aborted very early. Some recognizers only use the most probable path Qk . This will be discussed below.

[0112] The calculation is done by recursive calculation. The probability a n (i) at time n the system is in state q i to observe can be calculated as follows: αn(i)=(Ge1,Ge2,.....Gen;qn=qi)≡p(Gein,qni)αn+1(j)=[∑i=1Sαn(i)⋅aij]bj(Gen+1)

[0113] Here, all S possible paths leading to the state q are summed up. i+1 lead in

[0114] It is assumed that the total probability in the state qi n+1 to reach the best path is dominated. Then the sum can be simplified with little error. αn+1*(j)=[maxi(αn*(i)⋅aij)]bj(cn+1)

[0115] By tracing back from the last state, you now get the best path.

[0116] The probability of this path is a product. Therefore, a logarithmic calculation reduces the problem to a pure summation problem. The probability of recognizing a word corresponds to the probability of recognizing a model M j corresponds to the determination of the most probable word model for the observed emission X. This is now done exclusively via the best possible path Q best p(Ge→|Mj)=∑allQk(∏n=1Nbn(Gen))(π1∏n=2Na(n−1)n)

[0117] This will then become p(Ge→|Mj)=p(Ge→|Qbest,Mj)p(Q→best|Mj)=exp(ln(π1)+ln(b1(Ge1))+∑n=2Nln(bn(Gen))+ln(a(n−1)n))

[0118] It is now of particular importance that the Prototype Book (15) contains only elementary gestures.

[0119] The Fig. Figures 12 to 22 show some of the possible exemplary elementary gestures over a sketched smartphone as an exemplary device.

[0120] The figures show gestures performed with one hand. These can, of course, also be combined into two-handed gestures. It is certainly also conceivable to perform these gestures with other body parts. For example, the circular movement of the hand over the smartphone could correspond to a circular movement of the foot in front of a trunk. The recognition system would then be directed downward and, upon detecting this foot gesture in combination with other factors, such as the proximity of the correct car key, would open the trunk.

[0121] For simplicity, gestures can be broken down into different hierarchies. First, there are the orientations of the hand or object. Then, there are the transitions between these orientations. These are typically rotations. This is followed by the spatial positioning and its transitions, which correspond to the hand or object's trajectories in space above the device.

[0122] Let's first turn to the elementary object orientations. These, like the movements, are stored in the Prototype Book (15). As a first approximation, we can specify the following orientations of a hand, for example. For an object, an analogous basic orientation must first be determined in order to define an analogous orientation. a. A palm or object underside to the recognizer b. A hand from the side with little finger towards the recognizer (object from the right side) c. The back of the hand or the top of the object to the recognizer d. A hand from the side with thumb towards the recognizer (object from the left side) e. A hand or object oriented in one of the following directions: i. North, ii. North East iii. East iv. South East from South vi. South West vii. West viii. North West f. A hand tilted with the tip down (An object tilted in front downwards) g. A hand tilted with the tip upwards (An object tilted upwards in front) h. One hand vertically with tip upwards (object vertically upwards) i. One hand vertically with tip downwards (object vertically downwards) j. One hand bent (object rotated) k. To the right l. To the left m. The fingers of one hand spread n. The number of spread fingers on a hand is 1, 2, 3, 4, 5 o. One finger (especially the index finger) of one hand points forward, and so on. Regarding the last points j, k, l, and m, their applicability to another object depends heavily on its capabilities. Completely different properties may be important for a given object. If multiple objects / hands are present, combinations of these basic orientations are also possible. The above list, of course, also applies to objects other than hands. For example, feet or signaling aids. This also applies to the following sections.

[0123] The distinguishability of the orientations ultimately depends only on the number of transmitters T_Anz and receivers S_Anz. The number of parameters P_Anz that can be extracted without derivation is then proportional to P_Anz=T_Anz*S_Anz

[0124] However, this requires correspondingly extensive regulation.

[0125] The elementary gestures performed by a hand or an object can be divided into • those without an object (hand), • those with an object (hand) but without movement (e.g. no gesture, remain still), • Change in the number of objects (hands) in front of the recognizer (none, one, two) typically with a spatial orientation (entry from left, right, top, bottom and corresponding intermediate values such as diagonally above etc.), • those in which the object (the hand) itself is rotated around one of three spatial axes to the left or right ( Fig. 12, Fig. 13, Fig. 14), especially reorientations of the object / hand (e.g. edge of the hand, palm of the hand, top of the hand to the detector), • those in which the object (the hand) performs a circular, elliptical, or hyperbolic movement to the left or right in a coordinate direction of a polar coordinate system or a spherical coordinate system or another orthogonal coordinate system. ( Fig. 21 and Fig. 22 show circular movements. The third axis orientation corresponding Fig. 21 is not shown, as only the smartphone would be rotated. Of course, a restriction to curve segments is conceivable.) • those in which the object (the hand) is moved translationally forwards or backwards in one of the three spatial directions (x, y, z, and diagonal intermediate values) ( Fig. 18, Fig. 19, Fig. 20) • those in which the structure of the hand (the object) is changed (extending one, two or more fingers, e.g. counting, Fig. 15, where Fig.15 for simplicity the numbers 4 and 5 are not shown; or finger spreading or closing, Fig. 17; or e.g. open or close hand Fig. 16) • Combinations of these gestures or movements (e.g. circular movement with opening and closing of one hand, whereby the hand is always closed at one pole of the circle and open at the other.)

[0126] For every gesture, there is always an inverse gesture. This cancels out the result of a previously performed movement in a geometric sense. For example, a gesture that is inverse to a clockwise circular movement is a counterclockwise circular movement, and the inverse to a finger closing is a finger opening ( Fig. 17) etc.

[0127] It is important that combinations of these gestures are conceivable. For example, the opening and closing of the fingers can occur in six spatial directions (x, y, z, -x, -y, -z) relative to the line of movement of the fingers. The hand can be oriented in six directions (x, y, z, -x, -y, -z). (Diagonal intermediate orientations are of course also possible.) In total, the basic elementary gesture can be performed in 36 configurations. These can be transformed into one another by rotations. However, not all gestures are equally recognizable. In the case of a finger opening in the Z direction, the lower finger shadows the upper one. An expert will therefore typically not use such gestures.

[0128] An elementary gesture sequence “Swipe in X direction” ( Fig. 19) can therefore consist, for example, of the following actions: a) Elementary gesture “Idle” = no object b) Beginning of entry (= hand object appears at the edge of the measuring area) c) Entry process (= object does not change size and distance essentially, but performs a translational movement in the X-direction, thus appears at different times at different locations) d) Remain stationary (= object does not change size and distance essentially, does not perform any translational movement in X and Y directions, thus appears at the same locations at different times) (= end of entry process) e) Wipe (= object does not change size and distance essentially, but performs a translational movement in the X-direction, thus appearing at different times in different places) f) Remain stationary (= object does not change size and distance essentially, does not perform any translational movement in X and Y directions, thus appears at the same locations at different times) (= end of wipe) g) Exit process (= object does not change size and distance essentially, but performs a translational movement in the X-direction, thus appears at different times at different locations) h) End of entry (= hand object disappears at the opposite edge of the measuring area) i) Elementary gesture “Idle” = no object

[0129] A “gesture word” defined in this way, which is composed of the elementary gestures described above, can be recognized with high accuracy by the combined HMM-Viterbi recognizer (13) described above.

[0130] There are now basically two recognition scenarios possible: Single gesture word recognition ( Fig. 24) and continuous gesture word recognition ( Fig. 25).

[0131] With single-word gesture recognition, another event triggers the control system to initiate gesture recognition. This could be, for example, a button press. The advantage of this design is that it saves computing power and thus energy. This can be advantageous for mobile devices.

[0132] In continuous gesture word recognition ( Fig. 25), the gesture recognizer is permanently active. It is in at least one idle state, which corresponds to the permanent recognition of the gesture "no gesture." Experience has shown that this is represented by a large number of different prototypes. Furthermore, it is possible to put the gesture recognizer into a special power-saving mode. For example, it is conceivable that a self-powered device is powered by a solar cell, with the actual gesture recognizer switched off. The energy yield of the solar cell can be recorded as a sensor signal (37, Fig.1) can be used. The recognizer then jumps from the "sleep" state to the "wake up" state. This means that the transmitters and the rest of the system are started, which massively increases power consumption. If no permitted elementary gesture is found, the system returns to "sleep" mode. However, if a permitted elementary gesture is found (In Fig. 25, these are the gestures "no gesture" and "bcd," the system transitions to the respective states. If the system remains in the "no gesture" state for too long, it returns to the "sleep" state to conserve energy. As long as gestures are executed consecutively within a specified period of time, the system always returns to the "no gesture" state. In many cases, it is useful to always end the gesture words with a special "end gesture" state, which concludes the gesture words in a system-specific manner.

[0133] The gesture words, the recognition of which was described above, can now be combined to form more complex “gesture sentences”.

[0134] It is useful to group gesture words into gesture word classes. Examples of classes could be classes that group gesture words that, for example, • Subjects • Objects (e.g. items and people) or • Properties of these objects or • Changes in properties or • Properties of the changes themselves • Cause of these changes or • Describe the relationships between these objects.

[0135] The sequence of gesture words can be defined depending on the gesture word class. This allows subsequent HMM-Viterbi recognizers to correctly recognize the correct content of such a defined gesture message, a gesture sentence, despite potentially incorrect gesture word recognition from the previously described combined HMM-Viterbi recognizer. Since the process is analogous to the recognition of elementary gesture sequences (gesture words) based on a hypothesis list, it will not be explained further and is also omitted from the figures for simplicity.

[0136] A gesture grammar defined in this way allows for a significantly improved transmission of compact command sequences similar to a gesture language.

[0137] In order to make such a gesture language efficient, it is useful to design the transitions between the gesture words in such a way that the end of the preceding and the beginning of the following gesture can be coordinated.

[0138] Gesture words can contain elementary gestures and elementary gesture sequences that express dependencies on other gesture words or express other information such as temporal references (such as past and future, etc.). The elementary gesture sequences can also consist of a single elementary gesture.

[0139] At this point, reference should be made to the diverse possibilities of linguistics.

[0140] It is now particularly important that the gestures performed have consequences and achieve something. Such gestures are usually used to control a computer. This could also be the computer of a robot. Since the applications are generally not created by the device manufacturer, it is sensible for the gestures or gesture sequences to be assigned, or to be able to be assigned, to text strings or other symbol chains. This creates a link between the gestures or gesture sequences on the one hand and actions of the actuators of the computer and / or robot system on the other. Such a transfer of a link is expediently carried out using a data transfer protocol, for example an HTTP or XML protocol.

[0141] When using a neural network detector ( Fig.3) A neural network (27) is fed with the determined quantization vectors of the feature data stream vectors (38) of the feature extraction (11). However, a neural network (27) must first have been conditioned using sample data from the database (18) and a training method of the training tool (28). Here, too, feedback of the recognition result of the neural network, recognized gesture parameters (29), or the feature data stream vectors (38) to the controller (8) is conceivable. Fig. 8)

[0142] Neural network recognizers are significantly more compact and therefore easier to integrate than HMM recognizers. However, their ability to distinguish gestures is limited due to internal quantization noise. They are also more sensitive to interference. Therefore, attempting to distinguish more than 20 gestures, including parasitic gestures, with a single neural network is not recommended. Furthermore, neural network recognizers typically produce speaker- and device-dependent results. They therefore require training by the user, which severely limits their usability.

[0143] This type of recognition system is state of the art and has long been known for pattern recognition.

[0144] However, in order to provide the system with a particularly efficient set of feature vectors, it is advisable to control the transmitter (3) and the compensation transmitter (9) with an optimal set of signals. This idea of feeding back a recognition result to the physical interface (23) is new.

[0145] For this purpose, the Viterbi recognizer (13) can output a list of the expected next gestures (Gn) based on the sequence lexicon (16). This information can then be used to control the transmitters (3) so that the obtained vectors in the form of one or more currents (24) enable optimal gesture discrimination. Fig. 7)

[0146] This procedure will be explained using a specific example: Fig.Figure 4 shows an example schematic of a mobile device (mobile phone) with eight combinations of transmitter diode (H1-H8), receiver diode (D1-D8) and compensation diode (K1-K8).

[0147] After creating the forecast of the most likely subsequent gestures, the transmitter sends a temporal and, if necessary, spatial pattern that is most optimally suited to distinguishing the likely alternatives.

[0148] This pattern can be calculated in advance using the database (18).

[0149] The Fig. 27 to 58 show all possible activity patterns of the different transmitters of the exemplary mobile phone from Fig. 4.

[0150] It is clear to those skilled in the art that, assuming linear addition of the activity patterns is possible, the patterns can form a Hilbert space. Thus, by orthogonalizing the patterns, simple base patterns can be generated from which all other patterns are composed.

[0151] In the simplest case, these elementary activity patterns can be represented by the activity of a single diode.

[0152] By irradiating with an optimized transmission pattern, the significance of the quantization vectors is maximized.

[0153] The last problem is the derivation of the optimal transmission sequence.

[0154] For this purpose, the optimal stimulation pattern for this gesture is extracted from the database (18) based on the next elementary gesture (Gn) expected by the Viterbi recognizer (13) for verification by the controller (8). Therefore, when creating this database (18), a recognition performance must already be stored for each conceivable pattern. This can be achieved, for example, using a matrix (LDA_B).

[0155] Finally, the control algorithm of the physical interface (23) should be described using Fig. 5, with the transmitter (the LED H i ), compensation transmitter (the light-emitting diode K i ) and the sensor (the photodiode D i ) are linked together: Each of the transmitters H1 to H8 sends a signal into a transmission channel (I1 i) to the object to be measured (O), here typically a hand. There, the light is reflected and transmitted via a further transmission channel (I3 j ) to a receiver diode K1 to K8.

[0156] For each pairing of receiving diode D i and transmitting diode H j a controller C ij Since this would require a large amount of hardware, it is sensible to use only eight such controllers C ij and always only one transmitting diode H j This means that between the transmitting diodes H i is switched using time-division multiplexing. The corresponding feature vectors are therefore typically buffered and combined to form a feature vector, in this example eight times wider, in the form of one or more streams (24). In this example, they are passed to the feature extraction at one-eighth of the original feature vector frequency.

[0157] Each of the receiving diodes D i Each signal generator G is assigned to the input signal. This generates a signal that is orthogonal to the other generators. The orthogonality is explained in more detail below.

[0158] The signal from the generator G, which is applied to the receiver diode D i is converted by an orthogonalization unit into a short initial pulse S5 oi and a trailing signal S5 di split in such a way that their sum again produces the signal S5 i and that these signals are orthogonal with respect to the scalar product explained later.

[0159] The signal of the receiving diode D i is amplified by a preamplifier and connected to the signal S5 di to signal S10 dij in the multiplier M 1ij The signal is then multiplied by the filter F 1ij filtered. The previous multiplication in the multiplier M1ij and this filtering in filter F1 ij define the scalar product of the output signal of the receiver D 1j with signal S5 di . When orthogonality of further signals was mentioned in this description, it was meant that these two signals, when scalar multiplied together, result in zero.

[0160] The filter output signal is then passed through the amplifier V1 ij to signal S4 dij amplified. The inverse transformation is carried out by multiplying the signal S4 again dij with signal S5 di to signal S6 dij .

[0161] The same operation is performed with the other signal S5 oi This produces the signal S6 in an analog manner oij .

[0162] To determine the compensation signal of the compensation diode K 1j To obtain the signals S6 dij and S6 oijfor all generator signals i are added together and after adding a suitable offset B1 j the compensation diode K 1j supplied.

[0163] Each controller C ij In this example, outputs two values: A ij and D eij . In this example, these form a 16-dimensional analog vector with the seven other pairs of the other seven active controllers. Since the time-division multiplex allows only one transmitter H i active, eight of these vectors are generated within a measurement cycle in which all transmitters H i are active once. These are buffered as described and combined to form a 128-dimensional feature vector in the form of one or more streams (24). This 128-dimensional feature vector (24) forms the input signal for feature extraction (11). Since usually several transmitters H i and several sensors D1 jare present, the number of signals to a compensation signal S6 j contributing controller C ij be greater than one. This is in Fig. 5 is indicated by the fact that the signal from sensor D1 j is led out of the figure downwards and on the other side an analog signal from below into the summation point with the signal S6 oij This means that the next controller C kj where k≠i. This controller C kj will then of course be signaled with S5 k analogous to Controller C ij operated. So that the two controllers C ij and C kj do not interfere with each other, the signals S5 i and S5 k with respect to the scalar product by multiplication in the multiplier M1 ij (or M1 kj ) and subsequent filtering in filter F1 ij (or F1 kj) be orthogonal.

[0164] The compensation by the compensation transmitters K j takes place in the medium and is therefore largely independent of problems such as sensor contamination and aging. However, the compensation transmitter is always required. In the case of an optical system and an LED as the compensation transmitter, K j However, this is not absolutely necessary. It is also conceivable to configure the system as in Fig. 6 shown. Here is what is on Fig. 5 was not shown, a base current supply of the photodiode D1 i This is ensured by the resistance R. The compensation transmitter is the Fig. 6 drawn current source K j This emulates the compensation luminous flux of the transmission channel I2 j out of Fig. 5 by forcing a current through the resistor R which is equal to the current flowing through the light flux of the transmission channel I2 j in Fig.5 in the photo sensor D1 j generated electrical current.

[0165] Since the input of the sensors D j should always be at the same operating point, experience has shown that it is useful to use a signal S5 i for an i which is a constant signal. In this case, the value of the compensation current K j an equivalent value that only changes when the equivalent value of the signal received by Dj changes. Typically, this results in constant light compensation in optical systems. Therefore, the corresponding transmitter H i for this equivalent value is typically not operated and can therefore be omitted.

[0166] In terms of compensation, both systems are the same. Fig.5, however, has the advantage that the compensation signal of the transmission channel I2j and the measurement signal of the transmission channels I1i, I3j are subject to the same influences and are therefore typically affected in the same way. The system of Fig. 5 is therefore technically better, the Fig. 6 is the cheaper one because it has one less transmitter.

[0167] Considering what has been said so far, it becomes clear that a proposed measurement system can be constructed vectorially. This is Fig. 9 drawn.

[0168] The generator bank G generates a vector signal S5, whose components are the already known S5 i -signals. From these, the vector signal S5o is formed by a logic L and a vector delay unit Δt. Its components are the already known S5 di and S5o i .

[0169] The vector generator signal S5 is provided with a vector offset of the constant vector signal B1 if necessary and fed to the transmitter bank H. In this respect, the eight transmitters from Fig. 4 can be understood as a physical representation of a vector. The elements of this transmitter vector H are the already known transmitters H i . Each of the transmitters H i of the transmitter vector H sends a signal into a transmission channel I1 i in. This I1 i in turn form a vectorial transmission channel I1. This leads to at least one object O. (The discussion of multiple objects is not carried out here, as this is easily possible for a specialist.) From this object O, a transmission channel I3 leads j to one sensor D1 each j . The I3 jform the vectorial transmission channel I3. The D1j form the vectorial sensor bank D1. Their vectorial output signal is amplified and multiplied in the multiplier M by each signal of the vectorial S5o signal. It is easy to see that at this point the system is greatly expanded in terms of the necessary effort. A reduction to a subset of multiplications and / or a time-division multiplex therefore seems appropriate here. In this respect, this is not a mathematically ideal multiplication. The result is the signal vectorial signal S10, whose components are the already known signals S10o. ij and S10d ij These are each filtered individually in the vector filter bank F, where this filter consists of the known filters F1 ij and F2 ij This is followed by each of the filtered signals in a vector amplifier V, which in turn consists of the known amplifiers V1ij and V2 ij out of Fig. 5. The vectorial output signal S4 of the vectorial amplifier bank V consists of the signals S4o ij and S4 dij out of Fig. 5 together. Again, the multiplication of each element S4o ij and S4 dij of the signal vector S4 with the associated signals S5 di and S5o i . The resulting signals S6d are then calculated over all i ij and S6 oij to the respective component S6 jof the vector signal S6 is summed. If necessary, the vector bias signal B2 is added to the vector compensation signal S3. This drives the vector compensation transmitter bank K. This radiates back into the vector sensor bank D via the vector transmission path I2. As already discussed, the system is selected so that the fluctuations in the signal over the transmission path I3 are compensated by the compensation.

[0170] The recognition result is passed as a data stream to the feature extraction FE. This generates a feature vector pFV whose selectivity is not yet maximized. This vector is maximized with respect to selectivity by multiplying it with the LDA matrix (LDA_F) (14). The result is the feature vector FV (38). From this, the emission calculation EC (12) generates the hypothesis list HL (39). Using the sequence lexicon SL (16), the Viterbi search VS (13) generates the list of recognized gesture sequences G (22). At the same time, a prediction for the next expected gesture Gn is created. This vector is modified by multiplication, for example, using the illumination matrix (LDA_B) (40), and used to modify the generator signals S5. Fig. 9 this is done by signal-wise multiplication of the vector of generator signals with the vector thus obtained. Advantages of the proposal

[0171] The proposal enables the immediate discrimination of gesture sequences while providing improved immunity to ambient light and improved suppression of recognition errors due to redundancies.

[0172] It is particularly suitable for mobile systems and internet use. Figures Fig. 1: Example functional process of speaker-independent gesture or gesture sequence recognition with HMM recognition Fig. 2: Example functional process of speaker- and device-independent gesture or gesture sequence recognition with HMM recognition Fig. 3: Example functional process of speaker-dependent gesture or gesture sequence recognition with a neural network recognizer Fig. 4: Positioning of the transmitters (especially transmitter diodes) H i , the receiver (especially receiver diodes) D jand the compensation transmitter (especially compensation diodes) D j using the example of a mobile phone with eight exemplary transmitters, compensation transmitters and receivers. Fig. 5: Example control loop for the compensation control of the compensation transmitter K j Fig. 6: Example control loop for the compensation control of the compensation transmitter K j where the compensation transmitter is a current source. Fig. 7: Example functional sequence of speaker and device independent gesture or gesture sequence recognition with HMM recognition with influencing the transmission pattern of the transmitters depending on the recognition result Fig. 8: Example functional sequence of speaker-dependent gesture or gesture sequence recognition with a neural network recognizer with influencing the transmission pattern of the transmitters depending on the recognition result Fig. 9: Example vectorial overall control loop for the compensation control of the compensation transmitters and the adaptation of the transmission pattern of the transmitters to the detection result. Fig. 10: Diagram to illustrate the process of calculating emissions Fig. 11: Schematic representation of the principle of a mechanically automated calibration device for the reproducible execution of standardized gestures Fig. 12: Rotating the hand around the longitudinal axis Fig. 13: Rotating the hand around the axis perpendicular to the palm Fig. 14: Rotating the hand around the axis transverse to the longitudinal axis in the palm Fig. 15: Spreading a different number of fingers (0, 1, 2, 3) 4 and 5 fingers are not drawn for better clarity. Fig. 16: Changing the hand shape from a fist to a flat hand Fig. 17: Spread your fingers Fig. 18: Moving the hand up and down (Z-movement) Fig. 19: Translational movement across the device (x-movement) Fig. 20: Translational movement parallel to the device (y-movement) Fig. 21: Circular movement around the y-axis above the device (x-axis would be analogous) Fig. 22: Circular movement around the z-axis above the device Fig. 23: HMM model Fig. 24: Single word recognition (individual models) Fig. 25: Continuous Gesture Recognition Fig. 26-57: exemplary possible activity patterns List of names 1 Physical Value Stream / Stream of physical quantities 2 Interaction Region / Area of Interaction 3 First transmitter 4 First transmission path 5 Interaction or modification 6 Second transmission path 7 1 Sensor or receiver 8 controllers 9 Compensation Transmitter 10 Transmission path from the compensation transmitter (9) to the sensor (7) 11 Feature Extraction 12 Emission Computation 13 Viterbi Search / Viterbi decoder 14 LDA matrix 15 Prototype Book (Code Book or CB) 16 Sequence Lexicon / Sequence Lexicon 17 Training tool for creating the LDA matrix (14) and the prototype e-book (15) 18 Database of previously recorded gesture feature vectors of previously known elementary gestures in previously known elementary gesture sequences of a statistical selection of gesture speakers in a statistical selection of gesture situations, preferably including standard gestures and standard gesture sequences, for recalibrating the gesture recognizer when the physical interface changes (23) 19 Type-In tool for manually entering elementary gesture sequences (gestures) to be recognized 20 Online training tool for entering new elementary gesture sequences (gestures via the physical interface (23) (decomposition into previously known elementary gestures) 21 Recognized gesture parameters 22 Recognized gesture sequences 23 Physical Interface 24 Stream of physical values from the physical interface (23) This is transformed into a stream of gesture feature vectors in the feature extraction (11). 25 Application-specific training 26 Application-specific LDA matrix (This is always changed when the physical interface (23) is changed.) 27 Neural Network 28 Training tool for the neural network 29 Gesture parameters detected by the neural network 30 Example mobile phone 31 screen 32 Example Artificial Robot Hand for defining and displaying a standard gesture 33 Example device 34 Example elementary gestures 35 Example joints with three rotational degrees of freedom 36 Example arms with one translational degree of freedom each (length) 37 Signals from other sensors and measuring systems 38 Modified Feature Vector Data Stream 39 Elementary gesture hypothesis list 40 Illumination Matrix (LDA B) 41 Prototype 1 42 Prototype 2 43 Prototype 3 44 Prototype 4 45 Not precisely identified gesture 46 Unregistered gesture 47 Threshold Ellipsoid 48 Recognized gesture A ij Amplitude level of the radiation from the transmitter H i into the receiver D j B1: Constant vector signal that is added to vector signal S5. The resulting addition feeds transmitter bank H. B2 Constant vector signal that is added to the vector signal S6. The resulting addition, S3, feeds the compensation transmitter bank K. CB Code Book (Prototype database (15)) CBE Code Book entry (= entry in the prototype database (15)) Cij Controller for regulating the signal of the compensation diode K j depending on the transmitter H i . Each transmitter must have its own S5 signal to all other signals i orthogonal S5 i Send a signal. K1 compensation transmitter 1 (compensation LED 1) j=1 K2 compensation transmitter 2 (compensation LED 2) j=2 K3 compensation transmitter 3 (compensation LED 3) j=3 K4 compensation transmitter 4 (compensation LED 4) j=4 K5 compensation transmitter 5 (compensation LED 5) j=5 K6 compensation transmitter 6 (compensation LED 6) j=6 K7 compensation transmitter 7 (compensation LED 7) j=7 K8 compensation transmitter 8 (compensation LED 8) j=8 D1 Sensor 1 (photodiode 1) j=1 / Receiving diode 1 / Receiver 1 D2 Sensor 2 (photodiode 2) j=2 / Receiving diode 2 / Receiver 2 D3 Sensor 3 (photodiode 3) j=3 / Receiving diode 3 / Receiver 3 D4 Sensor 4 (photodiode 4) j=4 / receiving diode 4 / receiver 4 D5 Sensor 5 (photodiode 5) j=5 / Receiving diode 5 / Receiver 5 D6 Sensor 6 (photodiode 6) j=6 / receiving diode 6 / receiver 6 D7 Sensor 7 (photodiode 7) j=7 / Receiving diode 7 / Receiver 7 D8 Sensor 8 (photodiode 8) j=8 / receiving diode 8 / receiver 8 Dj Sensor j (photodiode j) j=j / receiving diode / receiver De ij Delay level of the transmitter's radiation H iinto the receiver D j DSP Digital Signal Processor Δt Vector delay unit. This is used by the logic L to generate the vector signal S5o from the vector signal S5 F Vector filter bank consisting of all filters F1 ij and F2 ij F1 ij Filter 1 of Controller C ij F2 ij Filter 2 of Controller C ij FV feature vector (38) G Vector Generator Bank / Signal Generator GS Recognized gesture sequence (22) Gn Expected next gesture (vector) H1 Transmitter 1 / Transmitter 1 / Transmitting diode 1 i=1 H2 transmitter 2 / transmitter 2 / transmitting diode 2 i=2 H3 transmitter 3 / transmitter 3 / transmitting diode 3 i=3 H4 transmitter 4 / transmitter 4 / transmitting diode 4 i=4 H5 transmitter 5 / transmitter 5 / transmitting diode 5 i=5 H6 transmitter 6 / transmitter 6 / transmitting diode 6 i=6 H7 transmitter 7 / transmitter 7 / transmitting diode 7 i=7 H8 transmitter 8 / transmitter 8 / transmitting diode 8 i=8 Hi Transmitter i / Transmitter i / Transmit diode ii=i HL Elementary Gesture Hypothesis List (39) I1 Vectorial first transmission section (transmission channel) from the vectorial transmitter bank H to the object O. I2 Vectorial second transmission section (transmission channel) from the vectorial compensation transmitter bank K to the object O. I3 Vectorial third transmission section (transmission channel) from object O to vectorial sensor bank D. I1 i First transmission section (transmission channel) from the Transmitte H i to object O. 12 j Second transmission section (transmission channel) from the compensation transmitter K j to object O. I3 j Third transmission section (transmission channel) from object O to sensor Dj . K1 compensation transmitter 1 (compensation LED 1) j=1 K2 compensation transmitter 2 (compensation LED 2) j=2 K3 compensation transmitter 3 (compensation LED 3) j=3 K4 compensation transmitter 4 (compensation LED 4) j=4 K5 compensation transmitter 5 (compensation LED 5) j=5 K6 compensation transmitter 6 (compensation LED 6) j=6 K7 compensation transmitter 7 (compensation LED 7) j=7 K8 compensation transmitter 8 (compensation LED 8) j=8 K j Compensation transmitter j Compensation LED j) j=j L The logic L generates the vector signal S5o from the vector signal S5 using the delay unit Δt. LDA_F LDA matrix (14) LDA_B Illumination Matrix (40) M1 ij Multiplexer 1 of controller C ij M2 ij Multiplexer 2 of controller C ij O Object pFV feature vector (e.g. after logarithmization, framing and filtering etc.) before multiplication with the LDA matrix (14) S3 Vectorial transmission signal for the vectorial compensation transmitter bank K. The signal is generated by vectorial addition of the vector B2 with S6. S3 j Transmission signal for the compensation transmitter K j . The signal is created by adding the signal B2 j with S6 j . S4d ij Amplitude level of the radiation / signal of the transmitter H i into the receiver D j S4o ij Delay level of the transmitter's radiation / signal H i into the receiver D j S5 Vectorial transmission signal consisting of the signals S5 i S5o Vector signal consisting of all signals S5o i and S5d i S5o iGate signal for reception delay measurement of signal S5 i of the transmitter H i S5d i Gate signal for receiving amplitude measurement of signal S5 i of the transmitter H i S5 i Transmission signal of transmitter H i . Here this signal S5 i to all other signals S5 i be orthogonal. S6 Vector signal consisting of all signals S6o i and S6d i S6oi Inverse transform of the receive delay measurement of the signal S5 i of the transmitter H i S6d i Inverse transform of the received amplitude measurement of the signal S5 i of the transmitter H i S6 j Compensation pre-signal for controlling the compensation transmitter K j If necessary, a constant B1 j added up. S10 Vectorial output signal of the multiplexer M consisting of all signals S10d ij and S10o ij S10d ij Output signal of the multiplexer M1 ij of the controller C ij S10o ij Output signal of the multiplexer M2 ij of the controller C ij SL Sequence Lexicon (16) V amplifier bank consisting of all amplifiers V1 ij and V2 ij V1 ij Amplifier 1 of Controller C ij V2 ij Amplifier 2 of Controller C ij

Claims

[1] Measuring system for recording object parameters of an object (O), where the object can perform a z-movement and an x-movement and a y-movement during the capture, wherein the measuring system has a physical interface (23) and where the physical interface (23) - at least one controller (C ij ) and - at least one transmitter (H i ) for light and - at least one recipient (D j ) for the signal from the at least one transmitter (H i ) emitted signal and wherein the measuring system comprises at least a first unit in the form of an HMM gesture sequence recognition engine and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to perform the steps of feature extraction (11) and emission calculation (12), and wherein the at least one controller (Cij ) the physical interface (23) is arranged to - that of the at least one transmitter (H i ) emitted signal and - with at least one by the at least one recipient (D j ) received signal and - to determine the comparison result as a multidimensional signal (24) and that the first unit in the form of the HMM gesture sequence recognition engine is configured to carry out the steps of feature extraction (11) and emission calculation (12) on the basis of this multidimensional signal (24), and that the feature extraction (11) step comprises at least one of the following sub-steps: - dividing the multidimensional signal (24) into individual frames of defined length, - filtering of the multidimensional signal (24), - Normalization of the multidimensional signal (24), - Orthogonalization of the multidimensional signal (24), - non-linear mapping of the multidimensional signal (24), - Formation of derivatives of the thus generated values of the multidimensional signal (24) and wherein the first unit in the form of the HMM gesture sequence recognition engine comprises a prototype database (15) and wherein prototypes from example data streams are stored in this prototype database (15) and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to - to compare pattern vectors (38) output by the feature extraction (11) with the distances to prototypical vectors, the said prototypes, by means of distance calculation, in particular Euclidean distance calculation, within the framework of the emission calculation (12). [2] Measuring system according to claim 1, where there is at least one compensation transmitter (K j ) and wherein the at least one compensation transmitter (K j ) is set up to be connected to a transmission channel (I2 j ) between this compensation transmitter (K j ) and at least one recipient (D j ) to feed in a signal, - that it aligns with the signal of at least one transmitter (H i ) in such a way that a cancellation of at least a predefined part of the signal of the transmitter at the receiver (D j ) and where the controller (C ij ) and the compensation transmitter (K j ) are set up in such a way - the signal of the compensation transmitter (K j ) by means of the controller (C ij ) and regulate it in such a way that the said extinction occurs. [3] Measuring system according to one of the preceding claims, wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to - the controllers (C ij ) supplied data stream in the form of the multidimensional signal (24) in the further course of processing with one of the speaker- and object-independent and device-dependent LDA matrix (26). [4] Measuring system according to one of the preceding claims, wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to - the controllers (C ij ) supplied data stream in the form of the multidimensional signal (24) with an application-independent matrix in the further course of processing. [5] Measuring system according to one of the preceding claims, wherein the first unit in the form of the HMM gesture sequence recognition engine comprises the prototype database (15) and wherein prototypes from example data streams are stored in this prototype database (15) and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to output at least one hypothesis list (39) for the recognized prototypes of each frame as a result of the emission calculation (12), and where the hypothesis list (39) contains the most probable of these prototypes of a frame with the respective probability of detection and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to form the hypothesis list (39), and where the hypothesis list (39) can contain only one element. [6] Measuring system according to claim 5, wherein the first unit in the form of the HMM gesture sequence recognition engine has a sequence dictionary (16) and wherein the sequence dictionary (16) contains sequence prototypes and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to evaluate the temporal sequence of the hypothesis lists (39) of the successive frames and wherein the first unit in the form of the HMM gesture sequence recognition engine is configured to determine from this sequence of hypothesis lists (39) the most probable of the predefined sequence prototypes for a sequence of recognized prototypes. [7] Measuring system according to one of the preceding claims and claim 5, wherein the measuring system is arranged so that the hypothesis list (39) contains at least one of at least one of the transmitters (H i ) transmitted signal is affected. [8] Measuring system according to one of the preceding claims, that the first unit in the form of the HMM gesture sequence recognition engine comprises a device-independent prototype database. [9] Device, in particular a robot or computer or smartphone, wherein the device has a measuring system in the form of a measuring system according to one or more of the preceding claims and wherein the device is configured to be able to control at least one actuator by means of at least one elementary gesture which is stored as a prototype in a prototype database (15) of the first unit in the form of the HMM gesture sequence recognition engine.

Citation Information

Patent Citations

  • Optoelectronic measuring device with stray light compensation provided by intensity regulation of additional light source

    DE10300223B3