Gesture Collection and Recognition System with Machine Learning Accelerator
Through the gesture recognition system of machine learning accelerator, the problem of inaccurate gesture recognition in the prior art is solved, and higher-precision gesture input is achieved, which is suitable for a variety of electronic devices and game applications.
Patent Information
- Application Number
- CN202011504228.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-12-18
AI Technical Summary
The prior art is difficult to effectively identify and personalize gesture input, resulting in limited input methods in complex control instructions and game applications.
The gesture collection and recognition system with machine learning accelerator is adopted, and the accuracy and correctness of gesture recognition are improved through transmission units, reception chains, gesture storage engines and machine learning accelerators, combined with a self-interference cancellation engine and machine learning model.
It realizes higher-precision gesture recognition, can avoid human interference, provide a better user experience, and supports gesture setting and training, suitable for a variety of electronic devices and gaming applications.
Smart Images

Figure CN114647302B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a gesture collection and recognition system, and particularly to a gesture collection and recognition system with a machine learning accelerator. Background Art
[0002] With the rapid development of electronic products in technology, the communication and interaction between users and electronic devices have become an important technical issue.
[0003] General input methods include touch screens, voice control, input methods using pen tips, etc. Although the above methods can all be used, there are still many limitations.
[0004] For example, when a user needs to input an instruction into an electronic device, the user still needs to touch the electronic device or make a sound as the input instruction. However, for the input instruction, the application has limitations in terms of distance. In addition, the above methods are difficult to implement for applications such as games or more complex control instructions.
[0005] In view of this, although technical means for controlling by detecting gestures have been disclosed, the technology for detecting gestures has problems of difficult gesture recognition and correct gesture recognition. Further, since the gestures generated by each user are not the same, currently, the technical means for detecting gestures are still difficult to achieve the result of customized recognition, and thus cannot be further developed into a more efficient input method. Summary of the Invention
[0006] In view of the above problems, the present invention provides a gesture collection and recognition system with a machine learning accelerator, comprising:
[0007] A transmission unit, having a self-interference cancellation engine, transmits a transmission signal to detect a gesture;
[0008] A first receiving chain, receives a first signal to generate first feature image data corresponding to the first signal, wherein the first signal is generated by the gesture reflecting the transmission signal;
[0009] A second receiving chain, receives a second signal to generate second feature image data corresponding to the second signal, wherein the second signal is generated by the gesture reflecting the transmission signal;
[0010] A third receiving chain, receives a third signal to generate third feature image data corresponding to the third signal, wherein the third signal is generated by the gesture reflecting the transmission signal;
[0011] A gesture storage engine, generates gesture data according to at least the first feature image data, the second feature image data and the third feature image data, comprising:
[0012] A first end, coupled to the first receiving chain and receiving the first feature image data;
[0013] A second end, coupled to the second receiving chain and configured to receive the second feature image data;
[0014] A third end, coupled to the third receiving chain and configured to receive the third feature image data; and an output end, configured to output the gesture data, the gesture data corresponding to at least the first feature image data, at least the second feature image data, and at least the third feature image data; and
[0015] A machine learning accelerator, configured to execute a machine learning model with the gesture data, including an input end, coupled to the output end of the gesture storage engine, for receiving the gesture data;
[0016] Wherein the first signal, the second signal, and the third signal are input into the self-interference cancellation engine.
[0017] As described above, through the above embodiments provided by the gesture recognition system, an anti-artificial interference / collision avoidance system can be realized. The accuracy and correctness of gesture recognition can also be improved through machine learning. The present invention also allows users to set gestures, and the gestures can also be trained by a server to obtain a better user experience. Brief Description of the Drawings
[0018] Figure 1 It is a block diagram of a gesture collection and recognition system with a machine learning accelerator according to the present invention. Detailed Description of the Embodiments
[0019] Please refer to Figure 1 , which is a block diagram of a gesture collection and recognition system with a machine learning accelerator according to the present invention. A predetermined gesture collection and recognition system 100 with a machine learning accelerator includes a transmission unit TX, a first receiving chain RX1, a second receiving chain RX2, a third receiving chain RX3, a gesture storage engine 130, and a machine learning accelerator 150.
[0020] The transmission unit TX is used to transmit a transmission signal St, and the transmission signal St is used to detect changes in gesture 199. The first receiving chain RX1 is used to receive a first signal Sr1 and generate first characteristic image data Dfm1 corresponding to the first signal Sr1. The transmission signal St detects changes in gesture 199 and generates a reflected signal, and the first signal Sr1 is generated from the reflected signal generated by the change in gesture 199 (reflecting). The second receiving chain R2 is used to receive a second signal Sr2 and generate second characteristic image data Dfm2 corresponding to the second signal Sr2. The transmission signal St detects changes in gesture 199 and generates a reflected signal, and the second signal Sr2 is generated from the reflected signal generated by the change in gesture 199. The third receiving chain R3 is used to receive a third signal Sr3 and generate third characteristic image data Dfm3 corresponding to the third signal Sr3. The transmission signal St detects changes in gesture 199 and generates a reflected signal, and the third signal Sr3 is generated from the reflected signal generated by the change in gesture 199.
[0021] The transmission unit TX includes a self-interference cancellation (SIC) engine, which is implemented based on an analog vector modulator. The output signal of the self-interference cancellation engine can automatically track the time-varying self-interference signal through the Least Mean Square algorithm. The first signal Sr1, the second signal Sr2, and the third signal Sr3 received by the above-mentioned first receiving chain RX1, second receiving chain RX2, and third receiving chain RX3 are input to the self-interference cancellation engine in the transmission unit TX. The self-interference cancellation engine can process the transmission signal St, the first signal Sr1, the second signal Sr2, and the third signal Sr3 with large power variations to achieve more accurate and extensive gesture detection results for the transmission signal St transmitted by the transmission unit TX and the first signal Sr1, the second signal Sr2, and the third signal Sr3 received by the first receiving chain R1, second receiving chain R2, and third receiving chain R3.
[0022] As Figure 1 shown, according to an embodiment of the present invention, the transmission unit TX is coupled to the antenna ANT TX, for transmitting the transmission signal St. The first receiving chain RX1 includes a first antenna ANT1, a first receiver RX1, a first signal processing engine SP1, and a first feature image generator FMG1, but is not limited to this in the embodiments of the present invention. The first antenna ANT1 is used to receive the first signal Sr1. The first receiver RX1 includes a first end, a second end, and an output end. The first end is coupled to the first antenna ANT1 and is used to receive the first signal Sr1. The second end is coupled to the self-interference cancellation engine of the transmission unit TX and is used to receive the transmission signal St. The output end is used to output the first signal Sr1. The first signal processing engine SP1 is used to generate first processed data Dp1 based on the first signal Sr1. The first signal processing engine SP1 includes an input end and an output end. The input end is coupled to the output end of the first receiver RX1 and is used to receive the first signal Sr1. The output end is used to output the first processed data Dp1. The first feature image generator FMG1 is used to generate first feature image data Dfm1 based on the first processed data Dp1. The first feature image generator FMG1 includes an input end and an output end. The input end is coupled to the output end of the first signal processing engine SP1 and is used to receive the first processed data Dp1. The output end is used to output the first feature image data Dfm1.
[0023] As Figure 1 shown, according to an embodiment of the present invention, the second receiving chain R2 includes a second antenna ANT2, a second receiver RX2, a second signal processing engine SP2, and a second feature image generator FMG2, but is not limited to this in the embodiments of the present invention. The second antenna ANT2 is used to receive the first signal Sr2. The second receiver RX2 includes a first end, a second end, and an output end. The first end is coupled to the second antenna ANT2 and is used to receive the second signal Sr2. The second end is coupled to the self-interference cancellation engine of the transmission unit TX and is used to receive the transmission signal St. The output end is used to output the second signal Sr2. The second signal processing engine SP2 is used to generate second processed data Dp2 based on the second signal Sr2. The second signal processing engine SP2 includes an input end and an output end. The input end is coupled to the output end of the second receiver RX2 and is used to receive the second signal Sr2. The output end is used to output the second processed data Dp2. The second feature image generator FMG2 is used to generate second feature image data Dfm2 based on the second processed data Dp2. The second feature image generator FMG2 includes an input end and an output end. The input end is coupled to the output end of the second signal processing engine SP2 and is used to receive the second processed data Dp2. The output end is used to output the second feature image data Dfm2.
[0024] As Figure 1As shown, according to an embodiment of the present invention, the third receiving chain R3 includes a third antenna ANT3, a third receiver RX3, a third signal processing engine SP3, and a third feature image generator FMG3, but is not limited thereto in the embodiments of the present invention. The third antenna ANT3 is used to receive the third signal Sr3. The third receiver RX3 includes a first end, a second end, and an output end. The first end is coupled to the third antenna ANT3 and is used to receive the third signal Sr3. The second end is coupled to the self-interference cancellation engine of the transmission unit TX and is used to receive the transmission signal St. The output end is used to output the third signal Sr3. The third signal processing engine SP3 is used to generate third processed data Dp3 according to the third signal Sr3. The third signal processing engine SP3 includes an input end and an output end. The input end is coupled to the output end of the third receiver RX3 and is used to receive the third signal Sr3. The output end is used to output the third processed data Dp3. The third feature image generator FMG3 is used to generate third feature image data Dfm3 according to the third processed data Dp3. The third feature image generator FMG3 includes an input end and an output end. The input end is coupled to the output end of the third signal processing engine SP3 and is used to receive the third processed data Dp3. The output end is used to output the third feature image data Dfm3.
[0025] The gesture storage engine 130 includes a first end, a second end, a third end, and an output end. The first end is coupled to the first receiving chain RX1 and is used to receive the first feature image data Dfm1. The second end is coupled to the second receiving chain R2 and is used to receive the second feature image data Dfm2. The third end is coupled to the third receiving chain R3 and is used to receive the third feature image data Dfm3. The output end is used to output gesture data Dg, and the gesture data Dg corresponds to at least the first feature image data Dfm1, the second feature image data Dfm2, and the third feature image data Dfm3. Further, the gesture storage engine 130 is respectively connected to the first feature image generator FMG1 of the first receiving chain RX1, the second feature image generator FMG2 of the second receiving chain R2, and the third feature image generator FMG3 of the third receiving chain R3 to receive the first feature image data Dfm1, the second feature image data Dfm2, and the third feature image data Dfm3 generated by the first feature image generator FMG1, the second feature image generator FMG2, and the third feature image generator FMG3. The gesture storage engine 130 stores predetermined gesture data Dg according to at least the first feature image data Dfm1, at least the second feature image data Dfm2, and at least the third feature image data Dfm3. For example, the first feature image data Dfm1 and at least the second feature image data Dfm2 include a person making a victory gesture. The victory gesture is to extend the index finger and middle finger of the hand. The gesture data Dg generates corresponding data according to the victory gesture image and stores it in the gesture storage engine 130.
[0026] The predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a three-dimensional coordinate tracking engine 160, connected to the output end of the gesture storage engine 130 to receive the first feature image data Dfm1, the second feature image data Dfm2, and the third feature image data Dfm3. The three-dimensional coordinate tracking engine 160 can calculate the spatial coordinates for the first feature image data Dfm1, the second feature image data Dfm2, and the third feature image data Dfm3 through the function of coordinate conversion, and transmit the calculation results and the gesture data Dg stored in the gesture storage engine 130 to the microcontroller 180 for coordinate comparison.
[0027] The predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a first fast Fourier transform channel FFT CH1 (FFT channel 1), a second fast Fourier transform channel FFT CH2 (FFT channel 2), and a third fast Fourier transform channel FFT CH3 (FFT channel 3), connected between the gesture storage engine 130 and the three-dimensional coordinate tracking engine 160, so as to convert the first feature image data Dfm1, the second feature image data Dfm2, and the third feature image data Dfm3 received by the gesture storage engine 130 to the three-dimensional coordinate tracking engine 160.
[0028] The machine learning accelerator 150 is used to perform machine learning with the gesture data Dg. The machine learning accelerator 150 includes an input end, coupled to the output end of the gesture storage engine 130, for receiving the gesture data Dg.
[0029] As Figure 1 shown, according to an embodiment of the present invention, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a frequency synthesizer FS, used to provide a reference vibration signal S LO . In this embodiment, the transmission unit TX includes an input end, the first receiver RX1 further includes a second end, the second receiver RX2 further includes a second end, and the third receiver RX3 further includes a second end. The frequency synthesizer FS includes a first end, a second end, a third end, and a fourth end. The first end is coupled to the second end of the first receiver RX1 for outputting the reference vibration signal S LO to the first receiver RX1. The second end is coupled to the second end of the second receiver RX2 for outputting the reference vibration signal S LO to the second receiver RX2. The third end is coupled to the input end of the transmission unit TX for outputting the reference vibration signal S LO to the transmission unit TX. The fourth end is coupled to the second end of the third receiver RX3 for outputting the reference vibration signal S LOto the third receiver RX3. According to an embodiment of the present invention, the transmission unit TX can modulate the transmission signal St according to the reference vibration signal S LO and the first receiving chain RX1 can modulate the first signal Sr1 according to the reference vibration signal S LO . The second receiving chain RX2 can modulate the second signal Sr2 according to the reference vibration signal S LO . The third receiving chain RX3 can modulate the third signal Sr3 according to the reference vibration signal S LO .
[0030] According to an embodiment of the present invention, as Figure 1 shown, the frequency synthesizer FS further includes an integrator-differentiator modulation waveform generator (Sigma-Delta modulator WG) SDM WG and an event synthesizer ES. The integrator-differentiator modulation waveform generator is used to modulate the generated waveform and the reference vibration signal S LO . The event synthesizer ES can increase the resolution of the waveform by increasing the number of antennas. In an embodiment of the present invention, the event synthesizer ES includes a main event synthesizer ES1 and a subsidiary event synthesizer ES2.
[0031] The predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a crystal oscillator XTALOSC, which is used to generate a stable reference frequency for the frequency synthesizer FS.
[0032] According to an embodiment of the present invention, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a feature acquisition engine 120, which is used to analyze the features of the received first signal Sr1, second signal Sr2, and third signal Sr3, and input the analyzed features into the machine learning hardware acceleration program machine 154 of the machine learning accelerator 150 for recognition learning. In an embodiment of the present invention, the feature acquisition engine 120 includes a range feature acquisition engine, a Doppler feature acquisition engine, and a phase difference feature acquisition engine. The feature acquisition engine 120 is used to detect the up, down, left, and right position changes of the gesture 199 in space.
[0033] According to an embodiment of the present invention, as Figure 1 shown, the machine learning accelerator 150 includes a weight regulation engine 158 and an array processor 1510. The weight regulation engine 158 is used to store a weight value Wc. The array processor 1510 is connected to the weight regulation engine 158 and is used to receive the weight value Wd and the gesture data Dg, and use the recognition algorithm to recognize the gesture 199 according to the weight value Wd and the gesture data Dg. Since the weight value Wc has been stored in the weight regulation engine 158, the required memory storage space can thus be reduced.
[0034] According to an embodiment of the present invention, asFigure 1 As shown, the machine learning accelerator 150 further includes a machine learning hardware acceleration program machine 154, which is connected to the feature acquisition engine 120, the microcontroller 180, and the array processor 1510, and receives the features after the feature acquisition engine 120 analyzes the first signal Sr1, the second signal Sr2, and the third signal Sr3, and is used as the connection interface between the array processor 1510 and the microcontroller 180. The machine learning hardware acceleration program machine 154 includes a microcontroller unit controller 1541, a direct memory access (DMA) controller 1542, a memory 1543, and a normalization activation function module (Softmax activation function) 1544. The microcontroller unit controller 1541 controls the signal input and output of the weight regulation engine 158, the array processor 1510, and the memory 1543 through the direct memory access controller 1542. In an embodiment of the present invention, the microcontroller unit controller 1541 is a neural network, which performs operations according to the parameters received by the microcontroller 180 and generates a sequence of control signals to the direct memory access controller 1542 to further control the signal input and output of the weight regulation engine 158 and the memory 1543. The array processor 1510 performs operations using the array data of the weight regulation engine 158 and the memory 1543, and stores the operation results in the memory 1543. The normalization activation function module 1544 is used for input and output signals, and the operation results finally generated by the array processor 1510 (the array data of the array processor 1510) are transmitted and output to the microcontroller 180 and the application program AP through the normalization activation function module 1544 for operations. The microcontroller (MCU) 180 is used for operations related programs, such as a mobile application (APP) for gesture recognition. The microcontroller 180 can also be used to transmit data to the cloud server for weight value training.
[0035] According to an embodiment of the present invention, as Figure 1 shown, the gesture data Dg can be transmitted to the cloud server 388 for training through the cloud server 388 to generate updated weight values Wu. The updated weight values Wu can be transmitted to the weight regulation engine 158 to update the weight values Wc stored in the weight regulation engine 158. Then, the machine learning accelerator 150 can perform gesture recognition calculations using the gesture data Dg and the updated weight values Wu stored in the weight regulation engine 158. By performing training on the cloud server 388, the weight values used by the machine learning accelerator 150 can be updated and adjusted in real time, and the accuracy and correctness of gesture recognition can also be improved. Furthermore, the setting and training of gestures can also be realized.
[0036] According to an embodiment of the present invention, as Figure 1 shown, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes an external host 170, which has a wireless connection function and is set between the cloud server 388 and the machine learning accelerator 150. The gesture data Dg can be transmitted to the cloud server 388 through the external host 170 and trained by the cloud server 388 to generate updated weight values Wu. The external host 170 can also be connected to the gesture storage engine 130 to store and transmit the first feature image data Dfm1, the second feature image data Dfm2, the third feature image data Dfm3, and the predetermined gesture data Dg, and transmit them to the cloud server 388 for training.
[0037] As Figure 1 shown, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes an application program AP, which is connected to the microcontroller 180. The application program AP can perform position tracking of the gesture 199 according to the result of the coordinate comparison calculation of the above microcontroller 180. Further, the present invention combines the technical features of the self-interference cancellation engine to achieve more precise gesture tracking. For example, when the user strokes or writes a single character in the air, the application program AP can use optical recognition technology to identify the character written by the user, such as identifying Chinese characters. In addition, the present invention can also integrate the gesture recognition functions of the entire system, fingers, hands, and palms into wearable devices, smart devices, laptops, smart home appliances, home appliance products, electrical products, humanized interface devices, or human-machine interface devices to use the gesture recognition function as a control command for inputting these products. Furthermore, by combining the functions of gesture recognition and object tracking, it can also be further applied to somatosensory games and Chinese font input devices. Further, when applying the functions of gesture recognition and object tracking to somatosensory games, by recognizing and tracking the gesture actions of the operator, control commands for inputting into the game can be further generated according to the results of recognizing and tracking the gesture actions. When applying the functions of gesture recognition and object tracking to a Chinese font input device, by recognizing and tracking which Chinese character the operator's gesture action is drawing, the output of Chinese characters can be further generated according to the results of recognizing and tracking the gesture actions.
[0038] According to an embodiment of the present invention, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a power management unit PMU for receiving the voltage V. The above Figure 1 Each functional block can be implemented by hardware, software, and / or firmware. Each functional block can be formed independently or combined with each other into a single functional block. The terminals of each functional block are used for signal and data transmission, which is only an example and not used to limit the content disclosed in the present invention, and it can be adjusted and changed according to various embodiments.
[0039] According to an embodiment of the present invention, the above-mentioned predetermined gesture collection and recognition system 100 with a machine learning accelerator can be implemented through an anti-artificial interference / collision avoidance system. According to an embodiment of the present invention, the predetermined gesture collection and recognition system 100 with a machine learning accelerator further includes a Frequency Modulated Continuous Waveform (FMCW) radar system, and the frequency modulated continuous waveform radar system is an application program for hand / finger gesture recognition using a hardware deep neural network accelerator (such as the machine learning accelerator 150) and a gesture training platform. The predetermined gesture collection and recognition system 100 with a machine learning accelerator can process high-frequency signals such as 60 GHz. The predetermined gesture collection and recognition system 100 with a machine learning accelerator can be implemented through a System on Chip (SoC), a chipset, or an integrated device having at least one chip, and other components that can be connected to a circuit board.
[0040] For example, the anti-artificial interference / collision avoidance system can be implemented by turning on the receivers RX1 and RX2 and scanning the spectrum. For example, the spectrum scan can be the entire 57 - 67 GHz spectrum. The predetermined gesture collection and recognition system 100 with a machine learning accelerator can skip the part of the spectrum occupied by other users or devices to avoid collisions. The anti-artificial interference / collision avoidance algorithm can be completed on a basic framework. The entire algorithm for gesture recognition can be implemented based on machine learning and a Deep Neural Network (DNN). Regarding the circuits for machine learning and the deep neural network, such as Figure 1The machine learning accelerator 150 can receive gesture recognition output from the feature image generators FMG1, FMG2 and frames. Due to the computer workload, real-time processing and low latency, the recognition algorithm is implemented with a specific hardware array processor (such as the array processor 1510). A specific scheduler (such as the machine learning hardware acceleration scheduler 154) can serve as an interface between the array processor 1510 and the microcontroller 180. In addition, specific algorithms can be applied to reduce the memory requirements for storing weight values. Therefore, before re-authorizing the weight values to the accelerator 150, a specific engine (such as the weight regulation engine 158) can be used to process the weight values (such as the weight value Wc). According to an embodiment of the present invention, in a predetermined gesture collection and recognition system 100 having a machine learning accelerator, the machine learning accelerator 150 can be dedicated to gesture detection and recognition and is disposed in a local system. The predetermined gesture collection and recognition system 100 having a machine learning accelerator can be a stand-alone system and can be used for independent gesture recognition. Therefore, the predetermined gesture collection and recognition system 100 having a machine learning accelerator can be more conveniently integrated into other devices (such as smart phones, tablets, and desktop computers, etc.) and can effectively improve the computing efficiency. For example, the time and / or power consumption required for gesture recognition can also be reduced. The machine learning accelerator 150 is used to reduce the gesture processing time required by the predetermined gesture collection and recognition system 100 having a machine learning accelerator, and the weight values used by the machine learning accelerator 150 can also be obtained from gesture training. Gesture training can be performed by a remote machine learning server, such as the cloud server 388.
[0041] In a typical application scenario, a fixed number of gestures can be collected and used for training. Gesture recognition using a plurality of weight values can be improved and the accuracy can be enhanced by performing training, and the training process utilizes a set of collected gestures. For example, a single gesture can be performed by one thousand people to generate one thousand samples, and the one thousand samples can be processed by a cloud machine learning server (such as the cloud server 388). The cloud machine learning server can perform gesture training, and the gesture training uses these samples to obtain corresponding results. The results can be a set of weight values used in the gesture inference process. Therefore, when the user performs or makes a gesture, this set of weight values can be applied to the calculation process to enhance the execution of recognition.
[0042] A set of basic gestures can be implemented by using the weight values for training the set. In addition, a predetermined gesture collection and recognition system 100 with a machine learning accelerator can allow users to have customized gestures. The personal gestures of the user can be recorded and transmitted to a cloud machine learning server (such as cloud server 388) through an external host processor (such as microcontroller 180) or an external device or system with network transmission capabilities for gesture training. The external host processor (such as microcontroller 180) and the external device or system with network transmission capabilities can execute a gesture collection application and can be connected by wired or wireless means. The results of the training (such as updated weight value Wu) can be downloaded, so that the user can use the gestures they own.
[0043] As described above, the signals for gesture sensing can have a frequency range of 60 GHz. Since the signals correspond to wavelengths in the millimeter range, the processing system can detect the movement of tiny hands / fingers with millimeter-range accuracy. Specific processing of the phase information of the radar signals is also a prerequisite. Figure 1 The specific phase processing engine (such as feature acquisition engine 120) in can be used to achieve this purpose.
[0044] In summary, through the above-described embodiments provided by the gesture recognition system, an anti-artificial interference / collision avoidance system can be implemented. The accuracy and correctness of gesture recognition can also be improved through machine learning. The present invention also allows users to set gestures, and the gestures can also be trained by the server to obtain a better user experience.
Claims
1. A gesture collection and recognition system with a machine learning accelerator, characterized in that, Comprising: A transmission unit having a self-interference cancellation engine for transmitting a transmission signal to detect a gesture; A first receiving chain for receiving a first signal to generate first feature image data corresponding to the first signal, wherein the first signal is generated by the gesture reflecting the transmission signal; A second receiving chain for receiving a second signal to generate second feature image data corresponding to the second signal, wherein the second signal is generated by the gesture reflecting the transmission signal; A third receiving chain for receiving a third signal to generate third feature image data corresponding to the third signal, wherein the third signal is generated by the gesture reflecting the transmission signal; A gesture storage engine for generating gesture data based on at least the first feature image data, the second feature image data, and the third feature image data, comprising: A first end coupled to the first receiving chain and receiving the first feature image data; A second end coupled to the second receiving chain and for receiving the second feature image data; A third end coupled to the third receiving chain and for receiving the third feature image data; And An output end for outputting the gesture data, the gesture data corresponding to at least the first feature image data, at least the second feature image data, and at least the third feature image data; And A machine learning accelerator for executing a machine learning model with the gesture data, comprising an input end coupled to the output end of the gesture storage engine for receiving the gesture data; Wherein the first signal, the second signal, and the third signal are input to the self-interference cancellation engine.
2. The gesture collection and recognition system with a machine learning accelerator as described in claim 1, characterized in that, The first receiving chain comprises: A first antenna for receiving the first signal; A first receiver comprising: A first end coupled to the first antenna for receiving the first signal; and An output end for outputting the first signal; A first signal processing engine for generating first processed data based on the first signal, comprising: An input end coupled to the output end of the first receiver for receiving the first signal; and An output end for outputting the first processed data; and A first feature image generator for generating the first feature image data based on the first processed data, comprising: An input end coupled to the output end of the first signal processing engine for receiving the first processed data; and An output end for outputting the first feature image data; The second receiving chain comprises: A second antenna for receiving the second signal; A second receiver comprising: A second end coupled to the second antenna for receiving the second signal; and An output end for outputting the second signal; A second signal processing engine for generating second processed data based on the second signal, comprising: An input end coupled to the output end of the second receiver for receiving the second signal; and An output end for outputting the second processed data; and A second feature image generator for generating the second feature image data based on the second processed data, comprising: An input end coupled to the output end of the second signal processing engine for receiving the second processed data; and An output terminal for outputting the second feature image data; And The third receiving chain includes: A third antenna for receiving the third signal; A third receiver, including: A third terminal coupled to the third antenna for receiving the third signal; and An output terminal for outputting the third signal; A third signal processing engine for generating a third processed data according to the third signal, including: An input terminal coupled to the output terminal of the third receiver for receiving the third signal; and An output terminal for outputting the third processed data; and A third feature image generator for generating the third feature image data according to the third processed data, including: An input terminal coupled to the output terminal of the third signal processing engine for receiving the third processed data; and An output terminal for outputting the third feature image data.
3. The gesture collection and recognition system with a machine learning accelerator according to claim 2, characterized in that, Further includes: a frequency synthesizer for providing a reference vibration signal; Wherein The transmission unit includes an input terminal; The first receiver further includes a second terminal; The second receiver further includes a second terminal; The frequency synthesizer includes: A first terminal coupled to the second terminal of the first receiver, outputting the reference vibration signal to the first receiver; A second terminal coupled to the second terminal of the second receiver, outputting the reference vibration signal to the second receiver; A third terminal coupled to the input terminal of the transmission unit, outputting the reference vibration signal to the transmission unit; A fourth terminal coupled to the third terminal of the third receiver, outputting the reference vibration signal to the third receiver; Wherein the transmission unit modulates the transmission signal according to the reference vibration signal, the first receiving chain modulates the first signal according to the reference vibration signal, the second receiving chain modulates the second signal according to the reference vibration signal, and the third receiving chain modulates the third signal according to the reference vibration signal.
4. The gesture collection and recognition system with a machine learning accelerator as described in claim 3, wherein, The frequency synthesizer includes: An integral-differential modulation waveform generator for modulating a generated waveform and modulating the reference vibration signal; and An event synthesizer for increasing the resolution of the waveform.
5. The gesture collection and recognition system with a machine learning accelerator according to claim 2, characterized in that Further includes: A feature acquisition engine for analyzing a phase of the first signal, the second signal, and the third signal according to the first feature image data, the second feature image data, and the third feature image data. The feature acquisition engine includes: A first terminal coupled to the output terminal of the first feature image generator A second terminal coupled to the output terminal of the second feature image generator; and A third terminal coupled to the output terminal of the third feature image generator.
6. The gesture collection and recognition system with a machine learning accelerator as described in claim 1, characterized in that The machine learning accelerator further includes: A weight regulation engine for storing a weight value; An array processor connected to the weight regulation engine for receiving the weight value and the gesture data, and identifying the gesture according to the weight value and the gesture data using an identification algorithm; And A machine learning hardware acceleration program machine connected to the weight regulation engine and storing the weight value.
7. The gesture collection and recognition system with a machine learning accelerator according to claim 6, characterized in that, The machine learning hardware acceleration program machine further includes: A direct memory access controller; A micro control unit controller for controlling the weight regulation engine and the array processor through the direct memory access controller; A memory stores an array of data for the array processor; and a normalization activation function module outputs the array of data for the array processor.
8. The gesture collection and recognition system with a machine learning accelerator according to claim 6, characterized in that, The gesture data is transmitted to a cloud server for training, and an updated weight value is generated by the cloud server and transmitted to the learning hardware acceleration program machine to update the weight value.
9. The gesture collection and recognition system with a machine learning accelerator as claimed in claim 8, wherein It further includes: An external host is connected to the learning hardware acceleration program machine, the gesture storage engine, and the cloud server for transmitting the gesture data to the cloud server and receiving the updated weight value by the cloud server.
10. The gesture collection and recognition system with a machine learning accelerator as claimed in claim 1, characterized in that, It further includes a three-dimensional coordinate tracking engine connected to the output end of the gesture storage engine to convert the spatial coordinates of the first feature image data, the second feature image data, and the third feature image data.
11. The gesture collection and recognition system with a machine learning accelerator according to claim 1, characterized in that, The gesture collection and recognition system with a machine learning accelerator is provided in a human-machine interface device, a smart device, a wearable device, a motion sensing game, or a Chinese font input device.
Citation Information
Patent Citations
Smartphone-based radar system for facilitating awareness of user presence and orientation
TW202009654A
Gesture recognition system having machine-learning accelerator
US20190383903A1