Interactive Gesture Recognition Method and System Based on Multi-View Structured Graph Convolutional Network
By constructing hand skeleton point sequence and multi-level graph convolution algorithm, combined with the principle of space-time consistency, the problem of missing feature information in interactive gesture recognition under multiple perspectives is solved, and the recognition accuracy is improved, especially fine-grained feature extraction of fingers and hands.
Patent Information
- Application Number
- CN202510735915.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The prior art under multi-view conditions, the lack of feature information during interactive gesture recognition leads to insufficient recognition accuracy, especially at the front and back view angles, the hand joint connection relationship and detailed features are difficult to accurately obtain, affecting the recognition of hand and finger morphological features.
The method based on multi-view structured graph convolution network is adopted to construct a hand skeleton point sequence, obtain the basic features of the hand, and use the multi-level graph convolution algorithm to perform multi-level enhancement processing, combine the principle of space-time consistency to perform space-time fusion, enhance the detailed characteristics of the fingers, and balance local and overall attention.
It improves the accuracy of interactive gesture recognition, avoids the lack or inadequate hand information from multiple perspectives, and enhances the ability to recognize finger details and overall hand structure.
Smart Images

Figure CN120260137B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and particularly to an interactive gesture recognition method and system based on a multi-view structured graph convolutional network. Background Art
[0002] With the rapid development of computer vision technology, gesture recognition has gradually become an important recognition task. Especially for interactive gesture recognition, due to its wide application scenarios and complex recognition requirements, interactive gesture recognition has gradually become a crucial part of the gesture recognition task.
[0003] In the prior art, interactive gestures are often recognized from a single fixed view, which makes the available feature information for recognition relatively limited and the problem of misrecognition occurs frequently. Therefore, recognition under multi-view conditions has become an important direction. Under multi-view conditions, there are significant differences in the spatial structure and detailed features of the hand between the front view and the back view, resulting in a lack of information on the connection relationships between hand joints and the information of the joints themselves, and insufficient exploration of the hierarchical information of the hand. At the same time, the fine-grainedness of the similar features between individual fingers and the fine-grainedness of the global features of the entire hand are poor in the local view and the global view of the hand, resulting in the inability to recognize the morphological features and movement trajectories of the hand and fingers in a detailed and sufficient manner, further affecting the accuracy of recognition.
[0004] Therefore, how to design an interactive gesture recognition method that can avoid the loss of feature information and enhance the fine-grainedness of the local and the whole under multi-view conditions to improve the recognition accuracy has become an urgent problem to be solved. Summary of the Invention
[0005] Based on this, the interactive gesture recognition method and system based on a multi-view structured graph convolutional network proposed by the present invention constructs a hand skeleton point sequence based on a topological graph, avoids the influence of the complex structure and rich actions of the hand, provides accurate hand joint and bone information, and accurately obtains the hand movement state. Then, by designing a multi-level graph convolutional algorithm for multi-level enhancement, the fingertips, finger bodies, and palms of the hand are decomposed into multiple levels in the order from outside to inside to avoid the problem of missing or insufficient hand information. It also combines the spatio-temporal consistency principle for spatio-temporal fusion, extracts and optimizes the gesture sequence information from the time and space perspectives respectively, further integrates information in different dimensions, thereby obtaining rich detailed information contained under multi-views, while avoiding the loss of the original gesture information. It also designs a fine-grained enhancement based on part and whole. According to the structural characteristics of the fingers, it fully excavates the finger detailed features, and effectively balances the attention degree to the finger part details and the hand overall structure. The present invention improves the accuracy of the interactive gesture recognition method.
[0006] An interactive gesture recognition method based on a multi-view structured graph convolutional network proposed by the present invention includes:
[0007] Obtain multi-view hand images and construct a hand skeleton point sequence to obtain basic hand features. The construction of the hand skeleton point sequence is based on a topological graph structure. The joints in the hand skeleton point sequence serve as the vertices of the topological graph, and the bones serve as the edges of the topological graph. The basic hand features include joint features, bone features, joint motion features, and bone motion features;
[0008] Perform multi-level enhancement processing on the basic hand features according to the multi-level graph convolutional algorithm to obtain multi-level enhanced features. The multi-level enhancement processing is based on hierarchical division of the hand skeleton point sequence, and the multi-level enhanced features are based on dimension expansion processing and differential enhancement processing;
[0009] Perform spatio-temporal fusion enhancement processing on the basic hand features and the multi-level enhanced features according to the spatio-temporal consistency principle to obtain multi-level spatio-temporal fusion enhanced features. The spatio-temporal fusion enhancement processing is based on a time difference matrix and a space difference matrix;
[0010] Perform fine-grained enhancement on the multi-level spatio-temporal fusion enhanced features to obtain the final recognition result. The fine-grained enhancement is based on partial enhancement processing and overall enhancement processing. The partial enhancement processing is based on a single finger region, and the overall enhancement processing is based on the overall hand region.
[0011] In summary, according to the above-mentioned interactive gesture recognition method based on a multi-view structured graph convolutional network, by constructing a hand skeleton point sequence based on a topological graph, the influence of the complex structure and rich movements of the hand is avoided, accurate hand joint and bone information is provided, and the hand movement state is accurately obtained. Then, by designing a multi-level graph convolutional algorithm for multi-level enhancement, the fingertips, finger bodies, and palms of the hand are decomposed into multiple levels in the order from outside to inside to avoid the problem of missing or insufficient hand information. Combining the principle of spatio-temporal consistency for spatio-temporal fusion, the gesture sequence information is extracted and optimized from the temporal and spatial perspectives respectively, further integrating information in different dimensions, thereby obtaining rich detailed information contained in multiple perspectives while avoiding the loss of original gesture information. A fine-grained enhancement based on part and whole is also designed. According to the structural characteristics of the fingers, the detailed finger features are fully excavated, and at the same time, the attention degree to the partial details of the fingers and the overall structure of the hand is effectively balanced. The present invention improves the accuracy of the interactive gesture recognition method. Specifically, multi-view hand images are obtained and a hand skeleton point sequence is constructed to obtain the basic hand features. The construction of the hand skeleton point sequence is based on a topological graph structure. The joints in the hand skeleton point sequence are used as the vertices of the topological graph, and the bones are used as the edges of the topological graph. The basic hand features include joint features, bone features, joint movement features, and bone movement features, avoiding the influence of the complex structure and rich movements of the hand, providing accurate hand joint and bone information, and accurately obtaining the hand movement state. According to the multi-level graph convolutional algorithm, the basic hand features are subjected to multi-level enhancement processing to obtain multi-level enhanced features. The multi-level enhancement processing is based on the hierarchical division of the hand skeleton point sequence. The multi-level enhanced features are based on dimension expansion processing and difference enhancement processing. The fingertips, finger bodies, and palms of the hand are decomposed into multiple levels in the order from outside to inside to avoid the problem of missing or insufficient hand information. According to the principle of spatio-temporal consistency, the basic hand features and the multi-level enhanced features are subjected to spatio-temporal fusion enhancement processing to obtain multi-level spatio-temporal fusion enhanced features. The spatio-temporal fusion enhancement processing is based on a temporal difference matrix and a spatial difference matrix, further integrating information in different dimensions, thereby obtaining rich detailed information contained in multiple perspectives while avoiding the loss of original gesture information. The multi-level spatio-temporal fusion enhanced features are subjected to fine-grained enhancement to obtain the final recognition result. The fine-grained enhancement is based on part enhancement processing and whole enhancement processing. The part enhancement processing is based on a single finger region, and the whole enhancement processing is based on the overall hand region. According to the structural characteristics of the fingers, the detailed finger features are fully excavated, and at the same time, the attention degree to the partial details of the fingers and the overall structure of the hand is effectively balanced. The present invention improves the accuracy of the interactive gesture recognition method.
[0012] Further, the step of obtaining multi-view hand images and constructing a hand skeleton point sequence to obtain the basic hand features specifically includes:
[0013] Obtain multi-view hand images and construct a hand skeleton point sequence. The hand skeleton point sequence is based on a topological graph structure, where joints are the vertices of the topological graph and bones are the edges of the topological graph. The edges connect adjacent vertices according to the hand skeleton structure. Obtain the point set of joints and the edge set of bones according to the data structure, obtain the basic feature matrix according to the point set of joints, and obtain the adjacency matrix according to the edge set of bones;
[0014] Obtain the hand basic features according to the basic feature matrix and the adjacency matrix. The hand basic features include joint features, bone features, joint motion features, and bone motion features.
[0015] Further, the step of performing multi-level enhancement processing on the hand basic features according to the multi-level graph convolution algorithm to obtain multi-level enhanced features specifically includes:
[0016] Perform hierarchical partitioning on the hand skeleton point sequence to obtain multi-level hand features;
[0017] Perform dimension expansion on the multi-level hand features according to multi-level global pooling. The specific algorithm for dimension expansion is as follows:
[0018] ,
[0019] ,
[0020] ,
[0021] Among them, represents the multi-level hand features after dimension expansion, represents the dimension expansion process, represents the multi-level hand features, represents global average pooling, represents expanding one dimension of the vector in different dimensions, represents the dimension, respectively represent the dimensions after and before expansion, represents the type of global average pooling, respectively represent the multi-level type, spatial type, and temporal type, represents the divided level; [[ID=,50]]
[0022] Then perform differential enhancement on the multi-level hand features after dimension expansion according to self-pairing subtraction. The specific algorithm for differential enhancement is as follows:
[0023] ,
[0024] Among them, represents the multi-level enhanced features, Represents the level of division, Represents the hyperbolic tangent activation function, Represents self-pairing subtraction for two dimensions of the feature matrix, Represents concatenating feature matrices of multiple levels.
[0025] Furthermore, the step of hierarchically dividing the opponent's skeleton point sequence to obtain multi-level hand features specifically includes:
[0026] Hierarchically divide the opponent's skeleton point sequence to obtain multi-level hand features. The hierarchical division divides the skeleton point set according to different hand regions. The hand regions include the palm region, the first finger joint region, the second finger joint region, the third finger joint region, and the fingertip region. Each hand region has a uniquely corresponding level, and each level includes all the skeleton points within the corresponding hand region;
[0027] The specific algorithm for the hierarchical division is as follows:
[0028] ,
[0029] Among them, Represents multi-level hand features, Represents the multi-level decomposition function, Represents the hand base feature, Represents the level of division.
[0030] Furthermore, the step of performing spatio-temporal fusion enhancement processing on the hand base feature and the multi-level enhancement feature according to the spatio-temporal consistency principle to obtain multi-level spatio-temporal fusion enhancement features specifically includes:
[0031] Perform spatio-temporal difference enhancement processing on the hand base feature according to the spatial dimension and the time dimension to obtain spatio-temporal difference features. The spatio-temporal difference features include spatial difference features and time difference features. The specific algorithm for the spatio-temporal difference enhancement processing is as follows:
[0032] ,
[0033] ,
[0034] Among them, and respectively represent the spatial difference feature and the time difference feature, Represents the hyperbolic tangent activation function, and respectively represent spatial global pooling and time global pooling, Represents self-pairing subtraction for two dimensions of the feature matrix, representation convolution;
[0035] Perform matrix multiplication on the spatio-temporal difference feature and the original hand basic feature to obtain an optimized spatio-temporal difference feature, and then perform weighted fusion on the multi-level enhanced feature and the optimized spatio-temporal difference feature to obtain a multi-level spatio-temporal fusion enhanced feature.
[0036] Further, the steps of performing matrix multiplication on the spatio-temporal difference feature and the original hand basic feature to obtain an optimized spatio-temporal difference feature, and then performing weighted fusion on the multi-level enhanced feature and the optimized spatio-temporal difference feature to obtain a multi-level spatio-temporal fusion enhanced feature specifically include:
[0037] Perform matrix multiplication on the spatio-temporal difference feature and the original hand basic feature to obtain an optimized spatio-temporal difference feature. The optimized spatio-temporal difference feature includes an optimized spatial difference feature and an optimized temporal difference feature. The specific algorithm of the matrix multiplication is as follows:
[0038] ,
[0039] ,
[0040] wherein, and respectively represent the optimized spatial difference feature and the optimized temporal difference feature, and represent the spatial difference feature and the temporal difference feature, represents the weight coefficient, represents the adjacency matrix, represents the hand basic feature, represents convolution;
[0041] Perform weighted fusion on the multi-level enhanced feature and the optimized spatio-temporal difference feature to obtain a multi-level spatio-temporal fusion enhanced feature. The specific algorithm of the weighted fusion is as follows:
[0042] ,
[0043] wherein, represents the multi-level spatio-temporal fusion enhanced feature, represents the multi-level enhanced feature, , , represent hyperparameters.
[0044] Further, the steps of performing fine-grained enhancement on the multi-level spatio-temporal fusion enhanced feature to obtain the final recognition result specifically include:
[0045] The multi-level spatio-temporal fusion enhanced features are finely grained split to obtain single-finger fine-grained features. The specific algorithm for the fine-grained split is as follows:
[0046] ,
[0047] wherein, represents the single-finger fine-grained feature, represents the multi-level spatio-temporal fusion enhanced feature, represents the feature matrix split by finger refinement, represents the finger ordinal number;
[0048] The single-finger fine-grained feature and the multi-level spatio-temporal fusion enhanced feature after overall information processing are fused by weight ratio to obtain the partial-global fusion feature. The overall information processing is used to enhance the global features of the overall hand region. The specific algorithm for the weight ratio fusion is as follows:
[0049] ,
[0050] wherein, represents the partial-global fusion feature, represents the average pooling operation, represents the concatenation of the single-finger fine-grained features of five fingers, represents the convolution processing of the single-finger fine-grained feature, and represent weight hyperparameters, represents the multi-level spatio-temporal fusion enhanced feature after overall information processing;
[0051] The final recognition result is obtained according to the partial-global fusion feature.
[0052] An interactive gesture recognition system based on a multi-view structured graph convolutional network proposed by the present invention includes:
[0053] A skeleton construction module, configured to obtain multi-view hand images and construct a hand skeleton point sequence to obtain hand basic features. The hand skeleton point sequence construction is based on a topological graph structure. The joints in the hand skeleton point sequence serve as the vertices of the topological graph, and the bones serve as the edges of the topological graph. The hand basic features include joint features, bone features, joint motion features, and bone motion features;
[0054] A multi-level graph convolution module, configured to perform multi-level enhancement processing on the hand basic features according to a multi-level graph convolution algorithm to obtain multi-level enhanced features. The multi-level enhancement processing is based on hierarchical partitioning of the hand skeleton point sequence. The multi-level enhanced features are based on dimension expansion processing and difference enhancement processing;
[0055] A spatio-temporal fusion module, configured to perform spatio-temporal fusion enhancement processing on the hand basic features and the multi-level enhancement features according to the spatio-temporal consistency principle, so as to obtain multi-level spatio-temporal fusion enhancement features, and the spatio-temporal fusion enhancement processing is based on a time difference matrix and a space difference matrix;
[0056] A local-global enhancement module, configured to perform fine-grained enhancement on the multi-level spatio-temporal fusion enhancement features to obtain a final recognition result, and the fine-grained enhancement is based on partial enhancement processing and global enhancement processing, the partial enhancement processing is based on a single finger region, and the global enhancement processing is based on an overall hand region.
[0057] The present invention also provides a storage medium, which stores one or more programs, and when the programs are executed by a processor, the interactive gesture recognition method based on a multi-view structured graph convolutional network as described above is implemented.
[0058] The present invention also provides a computer device, which includes a memory and a processor, wherein:
[0059] The memory is used for storing a computer program;
[0060] When the processor executes the computer program stored in the memory, the interactive gesture recognition method based on a multi-view structured graph convolutional network as described above is implemented. Description of the Drawings
[0061] Figure 1 It is a flowchart of an interactive gesture recognition method based on a multi-view structured graph convolutional network proposed in the first embodiment of the present invention;
[0062] Figure 2 It is a schematic structural diagram of an interactive gesture recognition system based on a multi-view structured graph convolutional network proposed in the second embodiment of the present invention;
[0063] Figure 3 It is a logical flowchart of a part-and-whole module proposed in the first embodiment of the present invention.
[0064] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0065] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0066] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of this invention are only for the purpose of describing specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0068] Please refer to Figure 1 , which shows a flowchart of an interactive gesture recognition method based on a multi-view structured graph convolutional network proposed in the first embodiment of the present invention. This interactive gesture recognition method based on a multi-view structured graph convolutional network includes steps S01 to S04, where:
[0069] Step S01: Obtain multi-view hand images and construct a hand skeleton point sequence to obtain basic hand features;
[0070] It should be noted that in this embodiment, the construction of the hand skeleton point sequence is based on a topological graph structure. The joints in the hand skeleton point sequence serve as the vertices of the topological graph, and the bones serve as the edges of the topological graph. The basic hand features include joint features, bone features, joint motion features, and bone motion features;
[0071] Obtain multi-view hand images and construct a hand skeleton point sequence. The hand skeleton point sequence is based on a topological graph structure. The joints serve as the vertices of the topological graph, and the bones serve as the edges of the topological graph. The edges connect adjacent vertices according to the hand skeleton structure. Obtain the point set of joints and the edge set of bones according to the data structure. Obtain the basic feature matrix according to the point set of joints, and obtain the adjacency matrix according to the edge set of bones;
[0072] Obtain the basic hand features according to the basic feature matrix and the adjacency matrix. The basic hand features include joint features, bone features, joint motion features, and bone motion features.
[0073] Step S02: Perform multi-level enhancement processing on the basic hand features according to the multi-level graph convolutional algorithm to obtain multi-level enhanced features;
[0074] It should be noted that in this embodiment, the multi-level enhancement processing is based on the hand skeleton point sequence for hierarchical division, and the multi-level enhancement features are based on dimension expansion processing and differential enhancement processing;
[0075] Perform hierarchical division on the hand skeleton point sequence to obtain multi-level hand features;
[0076] According to multi-level global pooling, perform dimension expansion on the multi-level hand features. The specific algorithm for dimension expansion is as follows:
[0077] ,
[0078] ,
[0079] ,
[0080] Among them, represents the multi-level hand features after dimension expansion, represents dimension expansion processing, represents the multi-level hand features, represents global average pooling, represents expanding one dimension of the vector in different dimensions, represents the dimension, respectively represent the dimension after expansion and the dimension before expansion, represents the type of global average pooling, respectively represent the multi-level type, spatial type, and temporal type, represents the divided level;
[0081] Then, perform differential enhancement on the multi-level hand features after dimension expansion according to self-pairing subtraction. The specific algorithm for differential enhancement is as follows:
[0082] ,
[0083] Among them, represents the multi-level enhancement features, represents the divided level,[[ID= 56]] represents the hyperbolic tangent activation function, represents performing self-pairing subtraction on two dimensions of the feature matrix, represents splicing the feature matrices of multiple levels;
[0084] Hierarchically partition the opponent's skeleton point sequence to obtain multi-level hand features. The hierarchical partition divides the skeleton point set according to different hand regions, and the hand regions include the palm region, the first finger joint region, the second finger joint region, the third finger joint region, and the fingertip region. Each hand region has a unique corresponding level, and each level includes all the skeleton points within the corresponding hand region;
[0085] The specific algorithm for the hierarchical partition is as follows:
[0086] ,
[0087] where, represents the multi-level hand features, represents the multi-level decomposition function, represents the basic hand features, represents the divided level.
[0088] Step S03: Perform spatio-temporal fusion enhancement processing on the basic hand features and the multi-level enhanced features according to the spatio-temporal consistency principle;
[0089] It should be noted that in this embodiment, in order to obtain the multi-level spatio-temporal fusion enhanced features, the spatio-temporal fusion enhancement processing is based on the time difference matrix and the space difference matrix;
[0090] Perform spatio-temporal difference enhancement processing on the basic hand features according to the spatial dimension and the time dimension to obtain spatio-temporal difference features. The spatio-temporal difference features include spatial difference features and time difference features. The specific algorithm for the spatio-temporal difference enhancement processing is as follows:
[0091] ,
[0092] ,
[0093] where, and respectively represent the spatial difference feature and the time difference feature, represents the hyperbolic tangent activation function, and respectively represent spatial global pooling and temporal global pooling, represents self-pairing subtraction of the two dimensions of the feature matrix, represents convolution;
[0094] Perform matrix multiplication on the spatio-temporal difference feature and the original basic hand features to obtain the spatio-temporal difference optimized feature, and then perform weighted fusion on the multi-level enhanced feature and the spatio-temporal difference optimized feature to obtain the multi-level spatio-temporal fusion enhanced feature;
[0095] Multiply the spatio-temporal difference features and the original hand base features in matrix form to obtain spatio-temporal difference optimized features. The spatio-temporal difference optimized features include spatial difference optimized features and temporal difference optimized features. The specific algorithm for the matrix multiplication is as follows:
[0096] ,
[0097] ,
[0098] where and represent the spatial difference optimized feature and the temporal difference optimized feature respectively, and represent the spatial difference feature and the temporal difference feature, represents the weight coefficient, represents the adjacency matrix, represents the hand base feature, represents convolution;
[0099] Perform weighted fusion on the multi-level enhanced features and the spatio-temporal difference optimized features to obtain multi-level spatio-temporal fusion enhanced features. The specific algorithm for the weighted fusion is as follows:
[0100] ,
[0101] where represents the multi-level spatio-temporal fusion enhanced feature, represents the multi-level enhanced feature, , , represent hyperparameters.
[0102] Step S04: Perform fine-grained enhancement on the multi-level spatio-temporal fusion enhanced features to obtain the final recognition result;
[0103] It should be noted that in this embodiment, the fine-grained enhancement is based on partial enhancement processing and overall enhancement processing. The partial enhancement processing is based on a single finger region, and the overall enhancement processing is based on the overall hand region;
[0104] Perform fine-grained splitting on the multi-level spatio-temporal fusion enhanced features to obtain single finger fine-grained features. The specific algorithm for the fine-grained splitting is as follows:
[0105] ,
[0106] where represents the single finger fine-grained feature, Represents the multi-level spatio-temporal fusion enhanced feature, Represents splitting the feature matrix by finger refinement, Represents the finger ordinal number;
[0107] Perform weighted ratio fusion on the single-finger fine-grained feature and the multi-level spatio-temporal fusion enhanced feature after overall information processing to obtain the part-whole fusion feature. The overall information processing is used to enhance the global feature of the overall hand region. The specific algorithm of the weighted ratio fusion is as follows:
[0108] ,
[0109] Among them, Represents the part-whole fusion feature, Represents the average pooling operation, Represents concatenating the single-finger fine-grained features of five fingers, Represents performing convolution processing on the single-finger fine-grained feature, and Represents the weight hyperparameter, Represents the multi-level spatio-temporal fusion enhanced feature after overall information processing;
[0110] Obtain the final recognition result according to the part-whole fusion feature;
[0111] The fine-grained enhancement in this embodiment is based on the part-whole module. For the specific logic flow of the part-whole module, please refer to Figure 3 .
[0112] In summary, according to the above-mentioned interactive gesture recognition method based on a multi-view structured graph convolutional network, by constructing a hand skeleton point sequence based on a topological graph, the influence of the complex structure and rich movements of the hand is avoided, accurate hand joint and bone information is provided, and the hand movement state is accurately obtained. Then, by designing a multi-level graph convolutional algorithm for multi-level enhancement, the fingertips, finger bodies, and palms of the hand are decomposed into multiple levels in the order from outside to inside to avoid the problem of missing or insufficient hand information. Also, the principle of spatio-temporal consistency is combined for spatio-temporal fusion, and gesture sequence information is extracted and optimized from the time and space perspectives respectively, further integrating information in different dimensions, thereby obtaining rich detailed information contained in multiple views while avoiding losing the original gesture information. Additionally, a fine-grained enhancement based on part and whole is designed. According to the structural characteristics of the fingers, the detailed features of the fingers are fully excavated, and at the same time, the attention degree to the partial details of the fingers and the overall structure of the hand is effectively balanced. The present invention improves the accuracy of the interactive gesture recognition method. Specifically, multi-view hand images are acquired and a hand skeleton point sequence is constructed to obtain the basic hand features. The construction of the hand skeleton point sequence is based on a topological graph structure. The joints in the hand skeleton point sequence serve as the vertices of the topological graph, and the bones serve as the edges of the topological graph. The basic hand features include joint features, bone features, joint movement features, and bone movement features, avoiding the influence of the complex structure and rich movements of the hand, providing accurate hand joint and bone information, and accurately obtaining the hand movement state. According to the multi-level graph convolutional algorithm, multi-level enhancement processing is performed on the basic hand features to obtain multi-level enhanced features. The multi-level enhancement processing is based on hierarchical division of the hand skeleton point sequence. The multi-level enhanced features are based on dimension expansion processing and differential enhancement processing. The fingertips, finger bodies, and palms of the hand are decomposed into multiple levels in the order from outside to inside to avoid the problem of missing or insufficient hand information. According to the principle of spatio-temporal consistency, spatio-temporal fusion enhancement processing is performed on the basic hand features and the multi-level enhanced features to obtain multi-level spatio-temporal fusion enhanced features. The spatio-temporal fusion enhancement processing is based on a time difference matrix and a space difference matrix, further integrating information in different dimensions, thereby obtaining rich detailed information contained in multiple views while avoiding losing the original gesture information. Fine-grained enhancement is performed on the multi-level spatio-temporal fusion enhanced features to obtain the final recognition result. The fine-grained enhancement is based on part enhancement processing and whole enhancement processing. The part enhancement processing is based on a single finger region, and the whole enhancement processing is based on the overall hand region. According to the structural characteristics of the fingers, the detailed features of the fingers are fully excavated, and at the same time, the attention degree to the partial details of the fingers and the overall structure of the hand is effectively balanced. The present invention improves the accuracy of the interactive gesture recognition method.
[0113] Please refer to Figure 2, which shows a schematic structural diagram of an interactive gesture recognition system based on a multi-view structured graph convolutional network proposed in the second embodiment of the present invention. The system includes:
[0114] A skeleton construction module 10, configured to obtain multi-view hand images and construct a hand skeleton point sequence to obtain basic hand features. The construction of the hand skeleton point sequence is based on a topological graph structure, where the joints in the hand skeleton point sequence serve as the vertices of the topological graph and the bones serve as the edges of the topological graph. The basic hand features include joint features, bone features, joint motion features, and bone motion features;
[0115] A multi-level graph convolution module 20, configured to perform multi-level enhancement processing on the basic hand features according to a multi-level graph convolution algorithm to obtain multi-level enhanced features. The multi-level enhancement processing is based on hierarchical partitioning of the hand skeleton point sequence, and the multi-level enhanced features are based on dimension expansion processing and differential enhancement processing;
[0116] A spatio-temporal fusion module 30, configured to perform spatio-temporal fusion enhancement processing on the basic hand features and the multi-level enhanced features according to the principle of spatio-temporal consistency to obtain multi-level spatio-temporal fusion enhanced features. The spatio-temporal fusion enhancement processing is based on a time difference matrix and a space difference matrix;
[0117] A local-global enhancement module 40, configured to perform fine-grained enhancement on the multi-level spatio-temporal fusion enhanced features to obtain a final recognition result. The fine-grained enhancement is based on partial enhancement processing and global enhancement processing. The partial enhancement processing is based on a single finger region, and the global enhancement processing is based on the overall hand region.
[0118] The present invention also proposes a computer storage medium, on which one or more programs are stored. When the program is executed by a processor, the above-mentioned interactive gesture recognition method based on a multi-view structured graph convolutional network is implemented.
[0119] The present invention also proposes a computer device, including a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned interactive gesture recognition method based on a multi-view structured graph convolutional network.
[0120] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0121] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0122] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0123] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0124] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. An interactive gesture recognition method based on a multi-view structured graph convolutional network, characterized in that, Including: Obtain multi-view hand images and construct a hand skeleton point sequence to obtain basic hand features. The construction of the hand skeleton point sequence is based on a topological graph structure, where the joints in the hand skeleton point sequence are used as the vertices of the topological graph and the bones are used as the edges of the topological graph. The basic hand features include joint features, bone features, joint motion features, and bone motion features. Perform multi-level enhancement processing on the basic hand features according to the multi-level graph convolution algorithm to obtain multi-level enhanced features. The multi-level enhancement processing is based on the hierarchical division of the hand skeleton point sequence, and the multi-level enhanced features are based on dimension expansion processing and differential enhancement processing. Perform spatio-temporal fusion enhancement processing on the basic hand features and the multi-level enhanced features according to the spatio-temporal consistency principle to obtain multi-level spatio-temporal fusion enhanced features. The spatio-temporal fusion enhancement processing is based on the time difference matrix and the space difference matrix. Perform fine-grained enhancement on the multi-level spatio-temporal fusion enhanced features to obtain the final recognition result. The fine-grained enhancement is based on partial enhancement processing and overall enhancement processing. The partial enhancement processing is based on the single finger region, and the overall enhancement processing is based on the overall hand region.
2. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 1, wherein The step of obtaining multi-view hand images and constructing a hand skeleton point sequence to obtain basic hand features specifically includes: Obtain multi-view hand images and construct a hand skeleton point sequence. The hand skeleton point sequence is based on a topological graph structure, with joints as the vertices of the topological graph and bones as the edges of the topological graph. The edges connect adjacent vertices according to the hand skeleton structure. Obtain the point set of joints and the edge set of bones according to the data structure, obtain the basic feature matrix according to the point set of joints, and obtain the adjacency matrix according to the edge set of bones. Obtain the basic hand features according to the basic feature matrix and the adjacency matrix. The basic hand features include joint features, bone features, joint motion features, and bone motion features.
3. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 1, wherein The step of performing multi-level enhancement processing on the basic hand features according to the multi-level graph convolution algorithm to obtain multi-level enhanced features specifically includes: Perform hierarchical division on the hand skeleton point sequence to obtain multi-level hand features. Perform dimension expansion on the multi-level hand features according to multi-layer global pooling. The specific algorithm for dimension expansion is as follows: , , , Among them, represents the multi-level hand features after dimension expansion, represents the dimension expansion process, represents the multi-level hand features, represents global average pooling, represents a dimension for expanding vectors in different dimensions, represents the dimension, respectively represent the dimensions after and before expansion, represents the type of global average pooling, respectively represent the multi-level type, spatial type, and temporal type, represents the divided level; Then perform differential enhancement on the multi-level hand features after dimension expansion according to self-pairing subtraction. The specific algorithm for differential enhancement is as follows: , Among them, represents multi-level enhanced features, represents the divided levels, represents the hyperbolic tangent activation function, represents self-pairing subtraction for two dimensions of the feature matrix, represents concatenating the feature matrices of multiple levels.
4. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 3, wherein The step of performing hierarchical division on the hand skeleton point sequence to obtain multi-level hand features specifically includes: Perform hierarchical division on the hand skeleton point sequence to obtain multi-level hand features. The hierarchical division divides the skeleton point set according to different hand regions. The hand regions include the palm region, the first finger joint region, the second finger joint region, the third finger joint region, and the fingertip region. Each hand region has a unique corresponding level, and each level includes all the skeleton points within the corresponding hand region. The specific algorithm for the hierarchical division is as follows: , Among them, represents multi-level hand features, represents the multi-level decomposition function, represents the basic hand features, represents the divided levels.
5. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 1, wherein The step of performing spatio-temporal fusion enhancement processing on the hand basic features and the multi-level enhancement features according to the spatio-temporal consistency principle to obtain multi-level spatio-temporal fusion enhancement features specifically includes: Performing spatio-temporal difference enhancement processing on the hand basic features according to the spatial dimension and the time dimension to obtain spatio-temporal difference features, where the spatio-temporal difference features include spatial difference features and time difference features. The specific algorithm of the spatio-temporal difference enhancement processing is as follows: , , Among them, and represent spatial difference feature and temporal difference feature respectively, represents the hyperbolic tangent activation function, and represent spatial global pooling and temporal global pooling respectively, represents self-paired subtraction on two dimensions of the feature matrix, represents convolution; Performing matrix multiplication processing on the spatio-temporal difference features and the original hand basic features to obtain spatio-temporal difference optimized features, and then performing weighted fusion processing on the multi-level enhancement features and the spatio-temporal difference optimized features to obtain multi-level spatio-temporal fusion enhancement features.
6. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 5, characterized in that The step of performing matrix multiplication processing on the spatio-temporal difference features and the original hand basic features to obtain spatio-temporal difference optimized features, and then performing weighted fusion processing on the multi-level enhancement features and the spatio-temporal difference optimized features to obtain multi-level spatio-temporal fusion enhancement features specifically includes: Performing matrix multiplication processing on the spatio-temporal difference features and the original hand basic features to obtain spatio-temporal difference optimized features, where the spatio-temporal difference optimized features include spatial difference optimized features and time difference optimized features. The specific algorithm of the matrix multiplication processing is as follows: , , Among them, and represent the spatial difference optimization feature and the temporal difference optimization feature respectively, and represent the spatial difference feature and the temporal difference feature, represents the weight coefficient, represents the adjacency matrix, represents the hand basic feature, represents convolution; Performing weighted fusion processing on the multi-level enhancement features and the spatio-temporal difference optimized features to obtain multi-level spatio-temporal fusion enhancement features. The specific algorithm of the weighted fusion processing is as follows: , Among them, represents the multi-level spatio-temporal fusion enhanced feature, represents the multi-level enhanced feature, , , represent hyperparameters.
7. The interactive gesture recognition method based on a multi-view structured graph convolutional network according to claim 1, wherein The step of performing fine-grained enhancement on the multi-level spatio-temporal fusion enhancement features to obtain the final recognition result specifically includes: Performing fine-grained splitting on the multi-level spatio-temporal fusion enhancement features to obtain single-finger fine-grained features. The specific algorithm of the fine-grained splitting is as follows: , Among them, represents the fine-grained feature of a single finger, represents the multi-level spatio-temporal fusion enhanced feature, represents the feature matrix split and refined by finger, represents the finger ordinal number; Performing weight ratio fusion on the single-finger fine-grained features and the multi-level spatio-temporal fusion enhancement features after overall information processing to obtain part-whole fusion features. The overall information processing is used to enhance the global features of the overall hand region. The specific algorithm of the weight ratio fusion is as follows: , Among them, represents the partial-global fusion feature, represents the average pooling operation, represents the single-finger fine-grained feature obtained by splicing five fingers, represents the convolution processing of the single-finger fine-grained feature, and represents the weight hyperparameter, represents the multi-level spatio-temporal fusion enhanced feature after the overall information processing; Obtaining the final recognition result according to the part-whole fusion features.
8. An interactive gesture recognition system based on a multi-view structured graph convolutional network, characterized in that, Including: A skeleton construction module, which is used to obtain multi-view hand images and construct a hand skeleton point sequence to obtain hand basic features. The hand skeleton point sequence construction is based on a topological graph structure, where the joints in the hand skeleton point sequence are used as the vertices of the topological graph and the bones are used as the edges of the topological graph. The hand basic features include joint features, bone features, joint motion features, and bone motion features; A multi-level graph convolution module, which is used to perform multi-level enhancement processing on the hand basic features according to the multi-level graph convolution algorithm to obtain multi-level enhancement features. The multi-level enhancement processing is based on the hierarchical division of the hand skeleton point sequence, and the multi-level enhancement features are based on dimension expansion processing and difference enhancement processing; A spatio-temporal fusion module, which is used to perform spatio-temporal fusion enhancement processing on the hand basic features and the multi-level enhancement features according to the spatio-temporal consistency principle to obtain multi-level spatio-temporal fusion enhancement features. The spatio-temporal fusion enhancement processing is based on a time difference matrix and a spatial difference matrix; A local-global enhancement module is used to perform fine-grained enhancement on the multi-level spatio-temporal fusion enhancement features to obtain a final recognition result. The fine-grained enhancement is based on partial enhancement processing and global enhancement processing. The partial enhancement processing is based on a single finger region, and the global enhancement processing is based on the overall hand region.
9. A storage medium, characterized in that, The storage medium stores one or more programs, which when executed by a processor implement the interactive gesture recognition method based on a multi-view structured graph convolutional network according to any one of claims 1-7.
10. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store a computer program; When the processor is used to execute the computer program stored on the memory, it implements the interactive gesture recognition method based on a multi-view structured graph convolutional network according to any one of claims 1-7.
Citation Information
Patent Citations
Human body behavior recognition method and system based on graph convolution network
CN110796110A
Traffic police gesture recognition method based on double-branch space-time diagram convolutional network
CN111881802A