Multi-level gesture recognition method and system for large screen interaction

CN122261394BActive Publication Date: 2026-09-08SHENZHEN MAZHE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610375871.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-25
Publication Date
2026-09-08
Estimated Expiration
2046-03-25

AI Technical Summary

Technical Problem

[0005]为了解决现有技术中因大屏多用户手部互相遮挡,导致关键点归属无法区分,难以准确补全和识别用户手势的技术问题,本发明的目的在于提供一种面向大屏交互的多层级手势识别方法及系统,所采用的技术方案具体如下:

Benefits of technology

通过为每个用户建立专属的、表征其发出不同操作指令时手部运动特征的手势行为模型,在识别到目标用户后调取其专属模型,且在目标用户实时手势数据因遮挡出现手部关键点缺失时,结合遮挡前预设时段内目标用户未被遮挡的手势数据与该专属模型预测缺失关键点位置,有效解决了多用户手部互相遮挡时现有技术难以区分手部关键点归属、无法准确补全缺失关键点的技术问题,实现了对目标用户被遮挡手部关键点的精准预测,让大屏交互中的手势识别更具针对性,大幅提升了多用户交互场景下大屏手势识别的准确性与有效性,适配多用户同时操作大屏的实际交互需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122261394B_ABST
    Figure CN122261394B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of gesture interaction, in particular to a multi-level gesture recognition method and system for large-screen interaction, which solves the technical problem that in the prior art, due to mutual shielding of hands of multiple users in large-screen multi-user interaction, key points cannot be distinguished, and it is difficult to accurately complete and recognize user gestures. The method comprises the following steps: in a multi-user interaction scene, historical gesture data corresponding to each user and historical system operation instructions associated with the historical gesture data are acquired, and a gesture behavior model of each user is established; in an interaction process, when a target user is recognized, real-time gesture data of the target user is extracted, and the gesture behavior model of the target user is called; if missing of hand key points caused by shielding is found in the real-time gesture data of the target user, according to gesture data of the target user that is not shielded in a preset period before shielding occurs and the gesture behavior model of the target user, positions of the hand key points of the shielded part are predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gesture interaction technology, specifically to a multi-level gesture recognition method and system for large-screen interaction. Background Technology

[0002] With the development of technology, interactive large screens have been widely used in education, medical care, commercial display, smart home and other fields. Traditional mouse and keyboard interaction methods have shown obvious shortcomings on large-screen devices with large size and multiple users. Gesture interaction, with its high degree of freedom, has become the ideal choice for interaction on large screen devices. The operating range of large screens also greatly increases the possibility of multiple users interacting on large screens at the same time. Users can directly operate the digital content on the screen through various gestures, and the system needs to recognize each user's unique gestures to make corresponding responses.

[0003] When multiple users perform complex operations on the large screen simultaneously, the complexity and range of motion of the gestures increase accordingly. This can easily lead to situations where some of the user's gestures are obscured. In such cases, the existing gesture recognition algorithm struggles to accurately recognize the user's gestures, ultimately causing the large screen to be unable to interact with the user or to produce operational errors, thus affecting the user experience and efficiency of the large screen interaction.

[0004] When recognizing user gestures, existing technologies mostly first identify the joints of the hand and then match them with models in the gesture library to obtain the command information corresponding to the gesture. When faced with the problem of missing hand information caused by hand occlusion, existing technologies usually combine general hand structure and identified hand key points to complete the occluded hand information before performing gesture recognition. However, in scenarios where multiple users interact and their hands overlap, the system has difficulty accurately distinguishing the users to which the overlapping hand feature points belong, and cannot effectively separate the gesture information of each user. This ultimately leads to the failure of hand key point completion and errors in gesture recognition. Summary of the Invention

[0005] To address the technical problem in existing technologies where multiple users' hands occlude each other on large screens, making it difficult to distinguish key points and accurately complete and recognize user gestures, the present invention aims to provide a multi-level gesture recognition method and system for large-screen interaction. The specific technical solution adopted is as follows: Firstly, a multi-level gesture recognition method for large-screen interaction is provided, including: in a multi-user interaction scenario, acquiring historical gesture data corresponding to each user and historical system operation commands associated with the historical gesture data, and establishing a gesture behavior model for each user; the gesture behavior model is used to characterize the hand movement features of the user when issuing different operation commands; when a target user is identified during the interaction, extracting the target user's real-time gesture data and calling the target user's gesture behavior model; if the real-time gesture data of the target user is found to contain missing hand key points due to occlusion, predicting the position of the occluded hand key points based on the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and the target user's gesture behavior model.

[0006] Based on the above technical solution, in the multi-level gesture recognition method for large-screen interaction provided by this invention, a unique gesture behavior model is established for each user, representing the hand movement characteristics when issuing different operation commands. After identifying the target user, the unique model is retrieved. When the target user's real-time gesture data is missing key points due to occlusion, the missing key point positions are predicted by combining the target user's unoccluded gesture data within a preset time period before occlusion with the unique model. This effectively solves the technical problem that existing technologies cannot distinguish the ownership of key points and accurately fill in missing key points when multiple users' hands are mutually occluded. It achieves accurate prediction of key points of the target user's occluded hands, making gesture recognition in large-screen interaction more targeted, significantly improving the accuracy and effectiveness of large-screen gesture recognition in multi-user interaction scenarios, and adapting to the actual interaction needs of multiple users operating the large screen simultaneously.

[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the method for obtaining historical gesture data corresponding to each user specifically includes: acquiring historical image frames containing user interaction scenarios; identifying the person frame where each user is located from the historical image frames using a preset target detection algorithm; matching pre-created user feature information with user feature information within the person frame using a preset face recognition algorithm to determine the user affiliation of the person frame; and extracting hand key point coordinate information from the person frame as the historical gesture data of the user.

[0008] In conjunction with the first aspect mentioned above, in one possible implementation, the method for establishing a gesture behavior model for each user specifically includes: determining the gesture recognition weight of each hand key point based on the distance from each hand key point to the hand reference center in historical gesture data; the hand reference center is the geometric center of multiple hand key points or a wrist key point; the gesture recognition weight is used to characterize the distinguishing importance of hand key points in different gesture actions; and performing similarity analysis on gesture actions at different times in historical gesture data based on the gesture recognition weight, classifying gesture actions with similarity greater than a preset similarity threshold into the same gesture action type.

[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the method for establishing a gesture behavior model for each user further includes: assigning a corresponding action code to each type of gesture action, and dividing the historical gesture data of the same user into multiple gesture behavior sequences in chronological order; each gesture behavior sequence consists of action codes arranged in chronological order; determining the segmentation point between gesture behavior sequences based on the time interval between adjacent gesture actions and the similarity between adjacent gesture behavior sequences, and dividing the continuous gesture behavior sequence into multiple gesture behavior units with independent operational intentions; clustering the gesture behavior units in the historical gesture data of the same user to classify multiple different types of gesture behaviors; statistically analyzing the historical operation instructions generated after each type of gesture behavior occurs, and weighting the frequency of each historical operation instruction in the current type of gesture behavior based on the similarity between gesture behaviors within the same type, and determining the historical operation instruction with the highest weighted frequency as the operation instruction corresponding to each type of gesture behavior.

[0010] In conjunction with the first aspect above, in one possible implementation, the method for establishing a gesture behavior model for each user further includes: determining the gesture validity of the current gesture behavior based on the similarity between each gesture behavior and other gesture behaviors of the same type; and marking gesture behaviors whose gesture validity meets preset validity conditions as valid gesture operations corresponding to operation instructions.

[0011] In conjunction with the first aspect above, in one possible implementation, the method for predicting the location of the hand key points in the occluded portion specifically includes: determining the sequence similarity between a first gesture sequence consisting of unoccluded gesture data of the target user within a preset time period before the occlusion occurs, and a second gesture sequence corresponding to each valid gesture operation in the target user's gesture behavior model; assigning weights to each valid gesture operation in the target user's gesture behavior model based on the sequence similarity, and performing a weighted summation of the complete hand key point coordinates corresponding to future moments in the historical data for each valid gesture operation, and using the result of the weighted summation as the predicted location of the hand key points in the occluded portion.

[0012] In conjunction with the first aspect above, in one possible implementation, the method for determining the sequence similarity between the first gesture sequence and the second gesture sequence specifically includes: determining the matching distance between the first gesture sequence and the second gesture sequence through a preset sequence similarity analysis algorithm; and determining the sequence similarity between the first gesture sequence and the second gesture sequence based on the negative correlation of the matching distance.

[0013] In conjunction with the first aspect above, in one possible implementation, the method for extracting real-time gesture data of the target user specifically includes: acquiring interactive image frames containing the target user in real time using a visual sensor; identifying the person frame where the target user is located from the interactive image frames using a preset target detection algorithm; detecting the key points of the target user's hand from the image region corresponding to the person frame where the target user is located using a preset key point detection algorithm, and obtaining the coordinate information of the key points of the hand.

[0014] In conjunction with the first aspect mentioned above, in one possible implementation, after predicting the position of the hand key points of the occluded part, the method further includes: merging the predicted position of the hand key points of the occluded part with the identified unoccluded hand key points to obtain the complete hand key point information of the target user; matching the complete hand key point information with a preset gesture operation instruction set to obtain the gesture instructions of the target user, and controlling the large screen to execute the operation corresponding to the gesture instructions.

[0015] Secondly, a multi-level gesture recognition system for large-screen interaction is provided, including: a historical data analysis module, used to acquire historical gesture data corresponding to each user and historical system operation commands associated with the historical gesture data in multi-user interaction scenarios, and establish a gesture behavior model for each user; the gesture behavior model is used to characterize the hand movement features of the user when issuing different operation commands; a real-time data acquisition module, used to extract the real-time gesture data of the target user when the target user is identified during the interaction process, and call the gesture behavior model of the target user; and an occlusion prediction module, used to predict the position of the occluded hand key points if the real-time gesture data of the target user is found to be missing due to occlusion, based on the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and the gesture behavior model of the target user.

[0016] The present invention has the following beneficial effects: By establishing a unique gesture behavior model for each user, representing the hand movement characteristics when issuing different operation commands, the system retrieves the unique model after identifying the target user. When the target user's real-time gesture data is missing key hand points due to occlusion, the system combines the target user's unoccluded gesture data within a preset time period before occlusion with the location of the missing key points predicted by the unique model. This effectively solves the technical problem of existing technologies being unable to distinguish the ownership of key hand points and accurately fill in missing key points when multiple users' hands are mutually occluded. It achieves accurate prediction of key hand points of the target user's occluded hands, making gesture recognition in large-screen interaction more targeted, significantly improving the accuracy and effectiveness of large-screen gesture recognition in multi-user interaction scenarios, and adapting to the actual interaction needs of multiple users operating large screens simultaneously. Attached Figure Description

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A system structure diagram of a multi-level gesture recognition system for large-screen interaction provided in one embodiment of the present invention; Figure 2 A flowchart illustrating a multi-level gesture recognition method for large-screen interaction, provided as an embodiment of the present invention; Figure 3 This is a schematic diagram of the hardware structure of a multi-level gesture recognition device for large-screen interaction, provided as an embodiment of the present invention. Detailed Implementation

[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a multi-level gesture recognition method and system for large-screen interaction proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0021] The following description, in conjunction with the accompanying drawings, details the specific solution of a multi-level gesture recognition method and system for large-screen interaction provided by this invention.

[0022] Please see Figure 1 The diagram illustrates a system structure of a multi-level gesture recognition system for large-screen interaction provided by an embodiment of the present invention. The multi-level gesture recognition system for large-screen interaction includes: a historical data analysis module 1, a real-time data acquisition module 2, and an occlusion prediction module 3.

[0023] This multi-level gesture recognition system for large-screen interaction is an intelligent interactive system adapted to multi-user interaction scenarios on large screens and solving the problem of inaccurate gesture recognition caused by hand occlusion. The system's core functional modules are historical data analysis module 1, real-time data acquisition module 2, and occlusion prediction module 3. A gesture command matching and execution module 4 is added to complete the full large-screen interaction loop. Each module can be implemented collaboratively by hardware devices and software algorithms. The hardware primarily relies on visual sensing devices, image acquisition terminals, edge computing processors, data storage servers, high-performance artificial intelligence (AI) processing servers, and the large-screen main control terminal deployed around the large screen. The software incorporates algorithms such as object detection, face recognition, key point detection, and sequence similarity analysis. User gesture feature data and dedicated model data output from preceding modules serve as the core input for subsequent modules. Each module is interconnected and shares data, jointly achieving accurate gesture recognition in multi-user occlusion scenarios. The following provides a detailed introduction to each module and its sub-modules: Historical data analysis module 1 is the system's foundational data processing and model building module. It is primarily implemented by a data storage server and a high-performance data processing server on the large screen, equipped with algorithm programs. Its core function is to acquire historical gesture data and associated historical system operation commands from multiple users, establishing a unique gesture behavior model for each user. This model is stored on the data storage server, providing core personalized gesture features for subsequent model retrieval in real-time data acquisition module 2 and key point prediction in occlusion prediction module 3. To achieve refined model building, this module comprises four sub-modules: historical data acquisition sub-module 11, user feature matching sub-module 12, gesture behavior modeling sub-module 13, and effective gesture filtering sub-module 14.

[0024] Among them, the historical data acquisition submodule 11 continuously acquires historical image frames containing multi-user large screen interaction scenarios through high-definition visual sensors and industrial cameras deployed in the large screen interaction area. It is equipped with a preset target detection algorithm to identify the person frame of each user from the historical image frames, and then extracts the coordinate information of the user's hand key points from the person frame through a preset key point detection algorithm, thus initially forming the original historical gesture data.

[0025] The user feature matching submodule 12 is equipped with a preset face recognition algorithm. It matches the user feature information pre-created in the system with the user feature information in the person frame identified by the historical data collection submodule 11, accurately determines the user affiliation of each person frame, associates and binds the key hand coordinate information with the corresponding user, forms the exclusive historical gesture data of each user, and synchronously associates and stores the historical system operation instructions generated after each user's operation.

[0026] The gesture behavior modeling submodule 13 performs feature analysis on the associated user-specific historical gesture data. Based on the distance from each hand key point to the hand reference center in the historical gesture data, it determines the gesture recognition weight of each hand key point. Then, it combines the weight to perform similarity analysis on gesture actions at different times, and classifies gesture actions that meet the preset similarity requirements into the same gesture action type. It assigns a unique action code to each gesture action type, divides the user's historical gesture data into gesture behavior sequences composed of action codes in chronological order, determines the segmentation point based on the time interval between adjacent gesture behavior sequences and the similarity of gesture actions, and divides the continuous gesture behavior sequence into multiple gesture behavior units with independent operation intentions. Then, it clusters all gesture behavior units to classify multiple gesture behavior types, counts the historical system operation commands generated after each gesture behavior type occurs, and determines the operation command corresponding to the type of gesture behavior as the historical system operation command with the highest frequency.

[0027] The effective gesture filtering submodule 14 calculates the similarity between each gesture behavior and other gesture behaviors of the same type, determines the validity of the current gesture behavior based on the similarity, marks the gesture behaviors whose validity meets the preset valid conditions as valid gesture operations for the corresponding operation instructions, and finally integrates all information such as the user's gesture action type, gesture behavior sequence, valid gesture operations and associated system operation instructions to establish a unique gesture behavior model for each user that represents the hand movement characteristics when issuing different operation instructions.

[0028] The real-time data acquisition module 2 is the system's real-time perception and data extraction module. It is mainly implemented by high-definition visual sensors, real-time image acquisition terminals, and edge computing processors deployed around the large screen. Its core function is to identify the target user during real-time interaction on the large screen, extract the target user's real-time gesture data, and retrieve the exclusive gesture behavior model established for the target user by the historical data analysis module 1 from the data storage server. The real-time gesture data and the exclusive model are then synchronously transmitted to the occlusion prediction module 3 to provide real-time data and model support for occlusion judgment and key point prediction. This module has three sub-modules: real-time image acquisition sub-module 21, target user identification sub-module 22, and real-time gesture extraction sub-module 23.

[0029] Among them, the real-time image acquisition submodule 21 continuously acquires real-time image frames containing the large screen interaction scene through the high-definition visual sensor in the large screen interaction area at a high frame rate. After performing preliminary preprocessing such as noise reduction and deblurring on the acquired image frames, the preprocessed image frames are transmitted to the edge computing processor in real time for subsequent processing.

[0030] The target user identification submodule 22 is equipped with a preset target detection algorithm in the edge computing processor. It identifies the person frames of all users from the real-time image frames transmitted by the real-time image acquisition submodule 21, and then retrieves the user feature information stored in the historical data analysis module 1. Through feature matching, it determines the target user in the current interaction scenario and completes the identification of the target user.

[0031] Real-time gesture extraction submodule 23: Based on the target user person frame determined by the target user recognition submodule 22, the submodule detects the key points of the target user's hand from the corresponding image area using a preset key point detection algorithm, obtains the coordinate information of the key points of the hand in real time, and integrates them to generate real-time gesture data of the target user. At the same time, it retrieves the exclusive gesture behavior model established for the target user by the historical data analysis module 1 from the data storage server, and transmits the real-time gesture data and the exclusive gesture behavior model together to the occlusion prediction module 3.

[0032] The occlusion prediction module 3 is the core intelligent processing module of the system. It is mainly implemented by a high-performance AI processing server and an AI inference chip equipped with a dedicated algorithm model. Its core function is to determine whether there are missing hand key points due to occlusion in the real-time gesture data of the target user. If so, it combines the unoccluded gesture data within a preset time period before occlusion with the target user's exclusive gesture behavior model to predict the position of the hand key points in the occluded part. The predicted key point position information will provide core data support for the complete gesture generation of the subsequent gesture command matching and execution module 4. This module has three sub-modules: occlusion state judgment sub-module 31, sequence similarity analysis sub-module 32, and key point prediction calculation sub-module 33.

[0033] The occlusion state judgment submodule 31 receives the target user's real-time gesture data transmitted by the real-time data acquisition module 2. By statistically analyzing the actual number of hand key points detected in real time and judging whether the spatial distribution of hand key points conforms to the inherent structural characteristics of the human hand, it accurately identifies whether there are missing hand key points in the real-time gesture data due to mutual occlusion of multiple users' hands. If it is determined that there is no occlusion, the real-time gesture data is directly transmitted to the gesture command matching and execution module 4. If it is determined that there is occlusion, the workflow of subsequent submodules is triggered.

[0034] After determining that occlusion exists, the sequence similarity analysis submodule 32 extracts the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and integrates them to form a first gesture sequence. At the same time, it retrieves the second gesture sequence corresponding to the effective gesture operation in the target user gesture behavior model in the historical data analysis module 1, calculates the matching distance between the first gesture sequence and each second gesture sequence through a preset sequence similarity analysis algorithm, and then determines the sequence similarity between the two based on the negative correlation of the matching distance.

[0035] The key point prediction calculation submodule 33 assigns a corresponding weight to each valid gesture operation in the target user gesture behavior model based on the sequence similarity obtained by the sequence similarity analysis submodule 32. It performs a weighted summation of the complete hand key point coordinates corresponding to the future time in the historical data for each valid gesture operation, and uses the weighted summation result as the predicted position of the hand key point of the occluded part. The predicted position information is then transmitted to the gesture command matching and execution module 4.

[0036] The gesture command matching and execution module 4 is the final execution module of the system. It is mainly implemented by the large screen main control module and the large screen interactive execution terminal. Its core function is to integrate hand key point information to generate complete gesture data, match the corresponding gesture commands, and control the large screen to perform related operations. It is a key module connecting gesture recognition and actual interaction with the large screen. The operation of this module relies on the model built by the historical data analysis module 1 and the predicted key point positions output by the occlusion prediction module 3 to complete the closed loop of the entire large screen gesture interaction. This module has three sub-modules: complete gesture generation sub-module 41, gesture command matching sub-module 42, and large screen operation execution sub-module 43.

[0037] The complete gesture generation submodule 41 receives the predicted position of the occluded part of the hand key points output by the occlusion prediction module 3, and merges the predicted position with the coordinates of the unoccluded hand key points identified by the real-time data acquisition module 2 to generate complete hand key point data containing all key point information of the target user's hand.

[0038] The gesture command matching submodule 42 retrieves the preset gesture operation command set in the system, performs feature matching between the complete hand key point data obtained by the complete gesture generation submodule 41 and the gesture operation command set, calculates the similarity between the complete hand key point data and each standard gesture feature in the command set, and selects the system operation command corresponding to the standard gesture feature with the highest similarity as the gesture command of the target user.

[0039] The large screen operation execution submodule 43 transmits the gesture command determined by the gesture command matching submodule 42 to the large screen main control module. The large screen main control module then sends an operation signal to the large screen interaction execution terminal based on the gesture command, controlling the large screen to complete the corresponding interactive operation, such as content zooming, page switching, command confirmation, etc., to achieve precise linkage between the target user's gesture operation and the large screen interaction.

[0040] Please see Figure 2 The diagram illustrates a flowchart of a multi-level gesture recognition method for large-screen interaction provided by an embodiment of the present invention. This multi-level gesture recognition method for large-screen interaction includes: S1. In multi-user interaction scenarios, obtain the historical gesture data corresponding to each user and the historical system operation instructions associated with the historical gesture data, and establish a gesture behavior model for each user.

[0041] In practical applications of large-screen multi-user interaction, to achieve accurate gesture recognition under occlusion, it is necessary to first acquire historical gesture data and associated historical system operation commands for each user, and then build a unique gesture behavior model for each user based on this data. This model can accurately represent the hand movement characteristics of the user when issuing different large-screen system operation commands, providing personalized model support for gesture recognition and occlusion key point prediction in subsequent real-time interaction. The overall process involves first extracting user-level historical gesture data through visual acquisition and intelligent recognition methods, and then combining multi-dimensional feature analysis and algorithm calculation to complete the construction of the gesture behavior model. At the same time, the historical system operation commands fed back by the large-screen system after the generation of each historical gesture data are retrieved simultaneously to realize the association and binding of gesture data and operation commands.

[0042] It should be noted that the core of the multi-level gesture recognition method for large-screen interaction proposed in this invention lies in constructing a personalized gesture behavior model using the user's historical gesture data to address key point prediction in occluded scenarios. Therefore, this method is applicable to users who have performed at least one or more interactive operations and have historical gesture data. For users interacting for the first time and without historical data, the system can first use a general hand model for gesture recognition while simultaneously collecting the user's real-time gesture data; after accumulating a certain amount of historical data, a personalized gesture behavior model is then built for them, and the occlusion prediction function is enabled in subsequent interactions. In specific implementation, the initial model building strategy can be flexibly configured according to the actual application scenario.

[0043] In some implementations, the method for obtaining historical gesture data for each user may specifically include: First, visual sensors (such as high-definition industrial cameras and depth cameras) deployed on the same side of the interactive screen continuously collect continuous image frames containing multi-user operations within the interactive area of ​​the screen at a preset frame rate (such as 60 frames / second). These frames are stored in the data server as historical image frames containing user interaction scenarios, providing a raw and continuous visual data foundation for subsequent user recognition and gesture data extraction, ensuring the integrity and real-time nature of data collection.

[0044] Subsequently, a preset object detection algorithm is used to identify the human body bounding box of each user from historical image frames. The preset object detection algorithm can be YOLO (you only look once) v5, faster region-based convolutional neural networks (Faster R-CNN), etc. This type of algorithm can accurately select the human body region of each user from complex interactive scene images, achieve effective segmentation of user regions, avoid confusion of visual features of different users, and lay the foundation for subsequent accurate extraction of single user gesture data.

[0045] Next, the pre-created user feature information is matched with the user feature information within the person frame using a preset face recognition algorithm. The preset face recognition algorithm can be ArcFace, FaceNet, etc. The system will pre-collect and store the facial feature vectors of each user as user feature information. After extracting the user facial feature information within the person frame, the similarity is calculated with the pre-stored feature information. The preset similarity matching threshold is 0.7. When the matching similarity is greater than or equal to this threshold, the user belonging to the person frame can be determined, accurately binding the person frame with the specific user, realizing user-level differentiation of historical gesture data, and avoiding the mixing of gesture data from multiple users from the source.

[0046] Finally, the coordinate information of key hand points is extracted from the user's assigned character frame as the user's historical gesture data. Pre-set key point detection algorithms such as MediaPipeHands and OpenPose are used for extraction. First, the absolute coordinates of each key hand point are obtained, and then they are converted into three-dimensional relative coordinates (X, Y, Z) relative to the wrist key points through coordinate system transformation. The core data representing the gesture action are accurately extracted from the user's body area. The coordinate normalization process ensures the uniformity and comparability of gesture data under different acquisition angles and distances, so that the subsequent feature analysis has a unified data foundation.

[0047] In some implementations, after acquiring each user's historical gesture data and associated historical system operation commands, the method for establishing a gesture behavior model for each user may specifically include: First, based on the distance from each hand keypoint to the hand reference center in the historical gesture data, the gesture recognition weight of each hand keypoint is determined. The hand reference center can be the geometric center of multiple hand keypoints on the palm or a wrist keypoint. This gesture recognition weight is used to characterize the distinguishing importance of the hand keypoint in different gesture actions. The larger the weight value, the higher the contribution of the keypoint to the distinction of different gesture actions. In implementation, for each historical gesture data of the target user, the Euclidean distance from each hand keypoint to the hand reference center in the data is calculated. Then, based on this distance data, the gesture recognition weight of each hand keypoint is calculated using a preset formula, expressed as: In the formula, The gesture recognition weight for the j-th hand key point of the target user; The total number of historical gesture data of the target user, which is a positive integer; Let Euclidean distance be the distance from the j-th hand keypoint to the wrist keypoint in the t-th historical gesture data of the target user. It is a positive real number and the unit of measurement is length (such as pixel, centimeter). Let Euclidean distance be the distance from the i-th hand key point to the hand reference center in the t-th historical gesture data of the target user. It is a positive real number and the unit of measurement is length (such as pixel, centimeter). Let Euclidean distance be the maximum Euclidean distance from all hand key points to the hand reference center in the t-th historical gesture data of the target user. It is a positive real number and the unit of measurement is length (such as pixel, centimeter). The weight represents the ratio of the distance from the j-th keypoint to the wrist during the t-th operation of the target user to the maximum distance from the keypoint to the hand reference center during this operation. A larger ratio indicates a higher contribution of that keypoint to gesture recognition during this operation. The average contribution value of that keypoint is obtained by eliminating the influence of the number of operations through arithmetic averaging, which is the gesture recognition weight. This highlights the role of key points such as fingertips in gesture recognition, which have high distinguishability, making subsequent gesture similarity analysis more targeted and improving the accuracy of the analysis results.

[0048] Subsequently, based on this gesture recognition weight, a similarity analysis was performed on the gesture actions at different times in the historical gesture data, represented as follows: In the formula, The similarity of gestures between the target user's α-th and δ-th historical gestures; The total number of key hand points on a single hand of the target user, which is a positive integer; Let be the cosine similarity of the coordinates of the j-th hand key point in the target user's α-th and δ-th historical gestures, with a range of [-1, 1] and dimensionless; norm is the positive correlation normalization function, and the min-max normalization function can be selected, which maps the input value to the (0, 1) interval to achieve data normalization. The gesture recognition weight of the j-th hand key point of the target user is multiplied by the normalization result to obtain the weighted similarity term of the key point. The weighted similarity terms of all key points are accumulated to obtain the overall weighted similarity sum of the two gesture operations. The sum of gesture recognition weights for all hand key points is used as a normalization factor to map the overall weighted similarity to a reasonable range. The resulting ratio is the gesture similarity. .

[0049] Gestures with a similarity score greater than a preset similarity threshold are categorized into the same gesture type. The preset similarity threshold can be set to 0.7. During analysis, the coordinate similarity of different key points is weighted according to gesture recognition weights before calculating the overall similarity of the gestures. This effectively aggregates similar gestures from the target user, achieving categorized classification of gestures, simplifying the complexity of subsequent model construction, and enabling the model to accurately capture the characteristic patterns of similar gestures from users. Based on this, each type of gesture is assigned a corresponding action code. The action code can be a unique character code such as a, b, c, d, etc. Then, the historical gesture data of the same user is arranged in chronological order. Based on the action code, continuous gestures are converted into a chronologically arranged code sequence. After generating the initial code sequence in chronological order, chronological deduplication is performed, and continuous and repeated action codes are merged. Only the state switching nodes where the gesture changes are retained, and then it is divided into multiple gesture behavior sequences. Each gesture behavior sequence consists of action codes arranged in chronological order. The user's continuous gestures are chronologically and structurally transformed to accurately correspond to the user's hand movement process.

[0050] Subsequently, a sliding window with a preset length equal to the preset step size (e.g., 1 second, set according to the typical duration of gesture actions so that a window can completely contain a gesture action in most cases) is set. The window is divided into multiple consecutive, non-overlapping windows with the start time of the interaction (e.g., the timestamp of the first frame when the system starts collecting historical image frames) as the zero point. If no new gesture is performed within a certain period after a user completes a gesture, or if the characteristics of the preceding and following gestures differ significantly, this indicates that the moment may be the dividing point between different operational intentions. To quantify the probability of each gesture acting as a dividing point, the action codes of multiple consecutive windows before (and including) the end time of the k-th gesture are extracted, using the sliding window containing the k-th gesture as a benchmark, and these codes are constructed in chronological order to form a sequence of previous gesture behaviors. Extract the action codes of multiple consecutive windows after the end time of the current window (excluding the current window), and construct a sequence of subsequent gesture behaviors in chronological order. The length of the preceding and following gesture sequences can be determined based on a preset number of windows, for example, encoding three consecutive windows each. Then, it is calculated using dynamic time warping (DTW). and The similarity between them is such that the smaller the distance, the more similar the gesture features are.

[0051] Using the time of the k-th gesture as a reference, calculate the time difference between that time and the time of the next gesture, and use this time as the time interval after the gesture. When the actions are continuous and unobstructed, the k-th gesture is continuous with the (k+1)-th gesture, and the time interval is zero. The larger the time interval, the more likely that moment is a segmentation moment.

[0052] Based on this, taking into account both the time elapsed after a gesture and the similarity between preceding and following gestures, a pre-defined formula is used to determine the segmentation points between gesture sequences. This divides a continuous sequence of gestures into multiple gesture units with independent operational intentions, represented as follows: In the formula, The segmentation probability corresponding to the k-th historical gesture action of the target user; norm is the positive correlation normalization function, and the minimum-maximum (min-max) normalization function can be selected. Its function is to map the input value to the (0,1) interval to achieve data normalization. The time interval between the k-th historical gesture action of the target user and the time corresponding to the next historical gesture action is a positive real number, and its dimension is time unit (such as second). Let be the dynamic time-normalized distance between two gesture sequences before and after the k-th historical gesture action of the target user. It is a positive real number, dimensionless. A smaller value indicates a higher similarity between the two time sequences, suggesting that the user's actions are essentially the same and no gesture control operation was performed. The corresponding negative correlation normalization term is also included. The larger the value, the higher the probability of splitting; The time interval normalization term represents the contribution of the time interval after the k-th gesture to the segmentation probability. The larger the time interval, the closer this value is to 1, and the higher the segmentation probability. For the complement term, when the time interval normalization term is close to 1, the complement term is close to 0, and the influence of temporal similarity is negligible. When the time interval normalization term is close to 0, the complement term is close to 1, and temporal similarity becomes the core factor affecting the segmentation probability. Multiplying the complement term by the time distance negative correlation normalization term yields the temporal similarity weighted term, achieving complementary analysis of the two dimensions of time interval and temporal similarity. Finally, the time interval normalization term and the temporal similarity weighted term are summed to obtain the final segmentation probability. .

[0053] When the segmentation probability is greater than or equal to a preset segmentation probability threshold, the end position of the window corresponding to the gesture is marked as the segmentation point. This effectively segments the gesture behaviors corresponding to different user operation intentions, avoiding confusion between gesture behaviors corresponding to different large-screen operation commands, and ensuring a precise correspondence between gesture behaviors and user operation intentions. The preset segmentation probability threshold is typically determined by combining massive amounts of historical gesture behavior data collected in multi-user interaction scenarios on large screens. It is determined through statistical analysis of the time intervals between gesture behaviors with different user operation intentions and the distribution characteristics of gesture similarity, using percentile methods or experimental calibration methods. For example, after statistical analysis of 1000 sets of effective gesture behavior sequence data from user large-screen interactions, the segmentation probability value corresponding to the 90th percentile is set as the preset segmentation probability threshold.

[0054] Next, by clustering the gesture behavior units in the historical gesture data of the same user, various types of gesture behaviors are divided. During clustering, a preset clustering algorithm is selected, such as the density-based spatial clustering of applications with noise (DBSCAN) density clustering algorithm. The edit distance between gesture behavior sequences is used as the distance metric. The preset minimum number of points in the cluster is 10. The cluster radius is determined by the k-distance graph method. Complete operation gestures with similar user characteristics are aggregated to realize the typological summary of user gesture behaviors.

[0055] After clustering, the historical operation commands generated after each type of gesture behavior are counted. Based on the similarity between gesture behaviors within the same type, the frequency of each historical operation command in the current type of gesture behavior is weighted. That is, the command accuracy of each historical system operation command and the command accuracy of that type of gesture behavior are calculated using the command accuracy calculation formula, expressed as: In the formula, The operation instruction γ is the instruction correctness corresponding to a certain type of gesture behavior, and represents the weighted frequency of operation instruction γ in a certain type of gesture behavior; The total number of gesture operations that generate operation instructions γ in a certain type of gesture behavior is a positive integer; Let be the edit distance between the corresponding gesture action units of the w-th gesture operation and the s-th gesture operation in a certain type of gesture behavior. It is a positive integer, dimensionless, and the smaller the value, the higher the similarity between the two sequences. For negative correlation normalization function, inverse min-max normalization can be used to map the input value to the (0, 1) interval. The larger the input value, the smaller the output value, thus achieving negative correlation transformation. The average function represents the arithmetic mean of the similarity between a certain gesture operation and other gesture operations of the same type. Let be the total number of gesture operations within a certain type of gesture behavior, and be a positive integer. ≥ First, the average similarity of all gesture operations that generate operation command γ is accumulated to obtain the total matching degree of this type of gesture operation. This total matching degree is then divided by the total number of operations of a certain type of gesture behavior to map the total matching degree to a reasonable range. The final ratio is the command accuracy. The closer the value is to 1, the higher the degree of matching between the operation instruction and the type of gesture behavior, and the more likely it is to be used as the system operation instruction corresponding to the type of gesture behavior.

[0056] The historical system operation command with the highest accuracy value (i.e. the highest weighted frequency) is selected and identified as the system operation command corresponding to this type of gesture behavior. A one-to-one correspondence between user gesture behavior and large screen system operation command is established, so that the gesture behavior model can accurately map the user's operation intention and system command.

[0057] Furthermore, the validity of each gesture is determined based on its similarity to other gestures of the same type. The calculation is based on the edit distance between units of the same type of gesture, converting the distance into a similarity using a preset formula, and then calculating the average similarity. This average similarity is used as the gesture validity, expressed as: In the formula, The validity of the v-th gesture operation in a certain type of gesture behavior; The total number of gesture operations that may generate the same system operation command within a certain type of gesture behavior; denoted as the edit distance between the corresponding gesture behavior units of the v-th gesture operation and the s-th gesture operation, is a positive integer with no dimension. The smaller the value, the higher the similarity between the two sequences. For negative correlation normalization, inverse min-max normalization can be used to map the input value to the (0, 1) interval. The larger the input value, the smaller the output value, thus achieving negative correlation transformation. By arithmetically averaging the similarity between the v-th gesture operation and all other gesture operations of the same type and with the same command, the influence of the number of operations on the result is eliminated, and the average similarity of the gesture operation is obtained, which is the gesture effectiveness. The closer the value is to 1, the higher the similarity between the gesture and the standard gesture, the stronger the effectiveness of the operation, and the more accurately it can trigger the corresponding system operation command.

[0058] Then, gesture behaviors that meet the preset validity conditions are marked as valid gesture operations corresponding to operation commands. The preset validity conditions are gesture validity greater than or equal to 0.6 (which can be set according to the actual application scenario and the requirements for recognition accuracy, for example, it can be set to 0.7 in scenarios requiring high accuracy). Invalid gesture behaviors in the user's historical operations are removed, such as gesture data without operation intent generated by accidental touch or redundant actions. This ensures that the gesture data used in the gesture behavior model for subsequent prediction of occluded key points are all valid data that can accurately trigger system operation commands, which greatly improves the accuracy of subsequent predictions by the model.

[0059] S2. When a target user is identified during the interaction, extract the target user's real-time gesture data and call the target user's gesture behavior model.

[0060] In the process of real-time multi-user interaction on a large screen, this step, as the core link connecting the historical gesture behavior model and the real-time occlusion key point prediction, will first complete the accurate identification of the target user in the interaction when the system detects that a user has initiated a large-screen operation in the interaction area, and then specifically extract the user's real-time gesture data. At the same time, based on the user's identity, it will retrieve the gesture behavior model built specifically for them, and transmit the real-time gesture data and the personalized model synchronously to the occlusion prediction module for subsequent occlusion judgment and key point prediction. This provides real-time data support and personalized model basis for subsequent accurate analysis of the user's real-time gesture state and prediction of the key point position in the occlusion area. The whole process revolves around the identity of the target user to achieve accurate matching of data and model, avoiding confusion between the gesture data and model of different users.

[0061] In some implementations, during the user identification process, visual sensors deployed around the large-screen interaction area continuously collect real-time image information of the interaction scene at a preset high frame rate. These visual sensors are consistent with the type of device used to collect historical image frames; high-definition industrial cameras or depth cameras can be used. With a preset frame rate of 60 frames per second, they can completely capture the user's real-time hand gestures, providing continuous and clear raw visual data for user identification, ensuring the real-time nature and accuracy of identity recognition. The system first uses a preset object detection algorithm to identify the bounding boxes of all users in the interactive state from the real-time image information. Preset object detection algorithms can include YOLOv5, Faster R-CNN, etc., which can quickly and accurately select the human body area from the complex large-screen interaction background, achieving effective segmentation of the user and background within the interaction area, thus defining the scope for subsequent single-user identity recognition.

[0062] The system then calls up the pre-stored feature information of each user and uses a preset face recognition algorithm to match the facial features of the user in the real-time identified person frame with the pre-stored feature information. The preset face recognition algorithm can be ArcFace, FaceNet, etc., and the preset face feature similarity threshold is 0.7. When the similarity value obtained by the matching is greater than or equal to this threshold, it can be determined that the user corresponding to the person frame is the target user of this interaction, thus completing the accurate identification of the target user in the large screen interaction. It locks the core object that needs to be analyzed for gestures from the multi-user interaction scenario, laying the identity foundation for subsequent extraction of exclusive real-time gesture data and retrieval of exclusive models.

[0063] After identifying the target user, the system extracts the target user's real-time gesture data according to a preset method. First, it continues to acquire interactive image frames containing only the target user in real time using the aforementioned visual sensors. Simultaneously, it performs real-time preprocessing on the acquired image frames, including noise reduction and deblurring, to eliminate image interference caused by ambient light and device vibration. This provides high-quality visual data focused on the target user for real-time gesture data extraction, avoiding interference from affecting the detection accuracy of hand key points. Next, the preset target detection algorithm accurately identifies the target user's bounding box from the preprocessed interactive image frames. Compared to full-area recognition of the interactive scene, this process focuses on the target user's body area for secondary bounding box selection, further narrowing the detection range of hand key points, accurately locking the target user's hand operation area, and eliminating interference from other users, large screen backgrounds, and other irrelevant areas, thus improving the efficiency and accuracy of hand key point detection. Then, a preset keypoint detection algorithm is used to detect the key points of the target user's hand in the image region corresponding to the person frame where the target user is located, and obtain the coordinate information of the key points of the hand. The preset keypoint detection algorithm can be MediaPipeHands, OpenPose, etc., which is consistent with the algorithm for extracting historical gesture data. During detection, the absolute coordinates of each key point of the hand in the image are first obtained, and then they are converted into three-dimensional relative coordinates (X, Y, Z) relative to the wrist key points through coordinate system transformation. The core feature data of the target user's real-time hand operation are accurately extracted. Furthermore, through coordinate normalization processing consistent with historical gesture data, the consistency of the real-time gesture data and the constructed gesture behavior model in terms of data format is ensured, so that subsequent model matching and feature analysis can be carried out smoothly.

[0064] After successfully extracting the target user's real-time gesture data, the system will accurately retrieve a gesture behavior model specifically built for that target user from the data storage server based on the identified target user identity. This retrieval process is executed in real time. At the same time as the real-time gesture data extraction is completed, the gesture behavior model will form a matching dataset with the real-time gesture data, associating the target user's real-time hand operation features with the user's historical gesture operation habits. This allows subsequent occlusion judgment and key point prediction to be carried out based on the user's personalized hand movement features, completely abandoning the general hand structure completion method in existing technologies. This ensures the targeting and accuracy of subsequent key point prediction from the data source. At the same time, the synchronous transmission of the model and real-time data also ensures the real-time performance of the large screen interaction, meeting the user's need for immediate response to large screen operations.

[0065] S3. If it is detected that there are missing hand key points in the real-time gesture data of the target user due to occlusion, predict the position of the occluded hand key points based on the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and the gesture behavior model of the target user.

[0066] This step, as the core execution link of the multi-level gesture recognition method for large-screen interaction, first determines the integrity of real-time gesture data, identifying whether there is a problem of missing key hand points caused by multiple users' hands occluding each other. Then, for occlusion scenarios, it combines the effective real-time gesture data before occlusion with the user's exclusive gesture behavior model to accurately predict the position of the occluded key points. Subsequently, the real-time gesture data will be completed and matched with the corresponding large-screen operation commands, ultimately achieving accurate interactive response on the large screen. The entire process revolves around the personalized gesture movement characteristics of the target user, completely solving the problem of confusion in the attribution of key points in the general structure completion of existing technologies, ensuring the accuracy of gesture recognition and the smoothness of large-screen interaction in occlusion scenarios.

[0067] In some implementations, when identifying whether hand key points are missing due to occlusion in the target user's real-time gesture data, the system first retrieves a preset number of complete hand key points. This number is set based on the hand skeletal features and the calibration rules of the key point detection algorithm. For example, when using a mainstream hand key point detection algorithm, the preset number of complete key points is 21. This number is used to represent the baseline value of all joints and feature points that need to be detected in the human hand. Subsequently, the system counts the number of hand key points actually detected in the extracted real-time gesture data of the target user. At the same time, it uses preset hand structure judgment rules to determine whether the spatial distribution of the actually detected key points conforms to the inherent structural features of the human hand. These judgment rules are set based on the skeletal connection relationship of the palm and fingers. For example, the relative coordinate offset between the fingertip key point and the corresponding palm and finger key point must be within a reasonable range, and the key points at the base of the palm must form a spatial distribution that conforms to the outline of the human hand. If the actual number of detected key points is lower than the preset total number, or if the spatial distribution of the actual key points does not match the inherent structural characteristics of the human hand, the system determines that there are missing hand key points in the target user's real-time gesture data due to occlusion. It accurately identifies the hand occlusion state in large-screen interaction, avoids misjudging the user's natural hand posture as occlusion, or judges the actual occlusion state as a normal gesture, and provides accurate triggering conditions for the subsequent key point prediction stage, ensuring the targeted execution of the solution.

[0068] If the system determines that there are missing hand key points due to occlusion, it will first extract the unoccluded gesture data of the target user within a preset time period before the occlusion occurred and construct a first gesture sequence. This preset time period can be experimentally calibrated based on the frequency of user gesture operations and the speed of hand movements in large-screen interaction scenarios, for example, it can be preset to 10 seconds. The system will filter out the gesture data that has been determined to be unoccluded through integrity detection within this time period from the real-time gesture data storage cache, and then integrate them into a continuous gesture feature sequence, i.e., the first gesture sequence, in chronological order. This sequence is used to characterize the real-time hand movement characteristics of the target user before the occlusion occurred. At the same time, the system will retrieve all gesture sequences marked as valid gesture operations from the gesture behavior model and use them as the second gesture sequence. This sequence is used to characterize the hand movement characteristics of the target user in historical operations that can accurately trigger system commands, providing timely real-time samples and valid historical samples for subsequent sequence similarity analysis, ensuring the rationality of the samples for similarity analysis, and laying a data foundation for accurate prediction of key points.

[0069] The system then determines the sequence similarity between the first gesture sequence and each second gesture sequence. First, it performs feature matching on the first and second gesture sequences using a preset sequence similarity analysis algorithm to determine the matching distance between them. The preset sequence similarity analysis algorithm can be dynamic time warping algorithm, sequence edit distance algorithm, etc. This type of algorithm can effectively solve the problem of length difference and movement speed difference of gesture sequences in the time dimension, accurately quantify the degree of feature difference between the two sequences. The matching distance is used to characterize the size of the feature difference between the first and second gesture sequences. The larger the value, the more obvious the difference in hand movement features between the two. The sequence similarity between the first and second gesture sequences is then determined based on the negative correlation of the matching distance. Specifically, this can be achieved through a negative correlation normalization function, such as using an inverse min-max normalization function to map the matching distance to the interval between 0 and 1, thus converting the matching distance into sequence similarity. After conversion, the smaller the matching distance, the larger the sequence similarity value, indicating that the hand movement features of the first and second gesture sequences are more similar. The feature differences between the real-time gesture sequence and the historical valid gesture sequence are quantified into a similarity index, providing an objective and accurate quantitative basis for subsequent weight allocation, so that subsequent key point prediction can fit the real-time hand movement trend of the target user.

[0070] It should be noted that the various gesture behavior types and their corresponding operation commands in the gesture behavior model are mainly used to filter valid gesture operations (i.e., gestures that can accurately trigger commands) and provide a benchmark for subsequent gesture command matching. In occlusion prediction, the similarity weighting based directly on the historical data of all valid gesture operations and the real-time gesture sequence can make full use of the user's overall movement habits, avoid prediction bias caused by pre-classification errors, and the similarity weight itself can adaptively highlight the historical operations most relevant to the current gesture, so there is no need to explicitly identify the type of real-time gesture.

[0071] After calculating the sequence similarity, the system assigns weights to each valid gesture operation in the target user's gesture behavior model based on the sequence similarity. It then performs a weighted summation of the complete hand keypoint coordinates for each valid gesture operation in historical data corresponding to future moments. The weighted summation result is normalized and used as the predicted location of the hand keypoints in the occluded portion, expressed as: In the formula, The predicted coordinates of the g-th occluded hand key point of the target user at the current time are three-dimensional spatial coordinates (X, Y, Z) and the units of measurement are length units (such as pixels, centimeters). The total number of valid gesture operations in the target user's gesture behavior model is a positive integer. For the o-th valid gesture operation in historical data, the complete coordinate position of the g-th hand key point at a future time corresponding to the time node of the first gesture sequence is given, which is a three-dimensional spatial coordinate (X, Y, Z) and the unit of measurement is length (such as pixel, centimeter). Let be the sequence similarity between the second gesture sequence corresponding to the o-th valid gesture operation and the first gesture sequence. This is a dimensionless parameter with a value range of (0, 1). The sequence similarity is used as a weight value; the higher the sequence similarity, the greater the contribution of the corresponding valid gesture operation in the weighted calculation. Then, the complete hand keypoint coordinates of each valid gesture operation in the historical data corresponding to the time node of the first gesture sequence are retrieved for future moments. By weighted summation and then dividing by the sum of the sequence similarities corresponding to all valid gesture operations, the influence of the total similarity value on the coordinate position is eliminated. The predicted position of each hand keypoint in the occluded part is calculated. Combined with the target user's real-time hand movement characteristics and historical operation habits, the spatial position of the occluded keypoint is accurately predicted, avoiding the attribution error problem caused by general hand structure completion and ensuring the consistency between the predicted position and the target user's actual hand posture.

[0072] Specifically, the hand key point coordinates corresponding to the time node of the first gesture sequence in the future time refer to: taking the start time of the first gesture sequence as the alignment reference, extracting the hand key point coordinates after the time point corresponding to the start time of the first gesture sequence in the historical data of each valid gesture operation (determined by a sequence alignment algorithm, such as a dynamic time warping algorithm), and taking the part of the historical data that exceeds the length of the first gesture sequence after alignment as the complete coordinate position of the future time.

[0073] After predicting the location of the key points of the occluded part of the hand, the system merges the predicted location of the key points of the occluded part of the hand with the coordinates of the identified key points of the uncrowded hand, and integrates them into coordinate information containing all key points of the target user's hand, i.e., complete hand key point information. This completes the real-time gesture data of the target user that is missing due to occlusion, restores the complete three-dimensional spatial features of the hand, and allows the gesture data to accurately represent the actual hand operation posture of the target user, providing complete and accurate feature data for subsequent gesture command matching.

[0074] The system then matches the complete hand keypoint information with a preset gesture operation command set to obtain the target user's gesture command and controls the large screen to execute the corresponding operation. The preset gesture operation command set is a pre-configured dataset containing a one-to-one correspondence between various standard hand keypoint feature templates and large screen system operation commands. The standard hand keypoint feature templates serve as the baseline hand posture features that trigger the corresponding large screen operation command, defined by the functional requirements of the large screen interaction. The system concatenates the 3D coordinates of all keypoints in the complete hand keypoint information into a feature vector in a preset order. Similarly, each standard feature template in the gesture operation command set is represented as a feature vector of the same dimension. Then, a preset feature matching algorithm is used to calculate the matching degree between the two feature vectors. The preset feature matching algorithm can use cosine similarity matching or Euclidean distance matching, with a preset matching degree threshold, such as 0.7. The system selects the standard feature template with the highest matching degree that is greater than or equal to the preset threshold and determines its corresponding large screen system operation command as the target user's gesture command. If all matching degrees are below the preset threshold, the system determines it as an invalid gesture operation and does not trigger any response from the large screen. After determining the target user's gesture command, the system transmits the command to the main control module of the large screen. The main control module then sends operation signals to the interactive execution terminal of the large screen to control the large screen to complete the corresponding interactive operation, such as content zooming, page switching, command confirmation, and content dragging. The effect of this action is to realize a complete closed loop from gesture feature recognition to actual operation on the large screen, accurately converting the target user's hand operation intention into the interactive response of the large screen, and completing accurate gesture interaction in multi-user occlusion scenarios on the large screen.

[0075] Based on the above technical solution, by establishing a unique gesture behavior model for each user that represents the hand movement characteristics when issuing different operation commands, the unique model is retrieved after the target user is identified. When the target user's real-time gesture data is missing key hand points due to occlusion, the model combines the target user's unoccluded gesture data within a preset time period before occlusion with the location of the missing key points predicted by the unique model. This effectively solves the technical problem that existing technologies cannot distinguish the ownership of key hand points and accurately fill in missing key points when multiple users' hands are mutually occluded. It achieves accurate prediction of key hand points of the target user's occluded hands, making gesture recognition in large-screen interaction more targeted, significantly improving the accuracy and effectiveness of large-screen gesture recognition in multi-user interaction scenarios, and adapting to the actual interaction needs of multiple users operating large screens simultaneously.

[0076] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0077] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0078] In this embodiment of the invention, the multi-level gesture recognition device for large-screen interaction can be divided into functional units according to the above method example. For example, each function can be divided into its own functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0079] This invention also provides a hardware structure diagram of a multi-level gesture recognition device for large-screen interaction, see [link / reference]. Figure 3 The multi-level gesture recognition device 300 for large-screen interaction includes a processor 301, and optionally, a memory 302 connected to the processor 301.

[0080] In the first possible implementation, see Figure 3The multi-level gesture recognition device 300 for large-screen interaction also includes a transceiver 303. The processor 301, memory 302, and transceiver 303 are connected via a bus. The transceiver 303 is used to communicate with other devices or communication networks. Optionally, the transceiver 303 may include a transmitter and a receiver. The device in the transceiver 303 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of the present invention. The device in the transceiver 303 that implements the sending function can be considered as a transmitter, which is used to perform the sending steps in the embodiments of the present invention.

[0081] Based on the first possible implementation method Figure 3 The structural diagram shown can be used to illustrate the structure of the multi-level gesture recognition device for large-screen interaction involved in the above embodiments.

[0082] in, Figure 3 This can also be illustrated by the system chip in a multi-level gesture recognition device for large-screen interaction. In this case, the actions performed by the aforementioned multi-level gesture recognition device for large-screen interaction can be implemented by this system chip. The specific actions performed can be found above and will not be repeated here.

[0083] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In this invention, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several of the functions listed in this invention.

[0084] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications of the invention fall within the scope of the invention and its equivalents, the invention is also intended to include such modifications and modifications.

Claims

1. A multi-level gesture recognition method for large-screen interaction, characterized in that, include: In multi-user interaction scenarios, acquire historical gesture data for each user and historical system operation commands associated with the historical gesture data to establish a gesture behavior model for each user. The gesture behavior model is used to characterize the hand movement features of the user when issuing different operation commands; The process of acquiring historical gesture data for each user includes: collecting historical image frames containing user interaction scenarios; identifying the person frame of each user from the historical image frames using a preset target detection algorithm; matching pre-created user feature information with user feature information within the person frame using a preset face recognition algorithm to determine the user affiliation of the person frame; and extracting hand key point coordinate information from the person frame as the historical gesture data of the user. Establishing a gesture behavior model for each user includes: determining the gesture recognition weight of each hand keypoint based on the distance from each hand keypoint to the hand reference center in the historical gesture data; the hand reference center is the geometric center of multiple hand keypoints or a wrist keypoint; the gesture recognition weight is used to characterize the distinguishing importance of hand keypoints in different gesture actions; performing similarity analysis on gesture actions at different times in the historical gesture data based on the gesture recognition weight, classifying gesture actions with similarity greater than a preset similarity threshold into the same gesture action type; assigning a corresponding action code to each gesture action type, and dividing the historical gesture data of the same user into multiple gesture behavior sequences according to time sequence; each gesture behavior sequence consists of action codes arranged in time sequence; and determining the gesture recognition weight based on the time interval between adjacent gesture actions and adjacent gesture actions. The similarity between sequences is used to determine the segmentation points between gesture behavior sequences, dividing continuous gesture behavior sequences into multiple gesture behavior units with independent operational intentions. By clustering gesture behavior units in the historical gesture data of the same user, various types of gesture behaviors are identified. Historical operation commands generated after each type of gesture behavior are statistically analyzed, and the frequency of each historical operation command in the current type of gesture behavior is weighted based on the similarity between gesture behaviors within the same type. The historical operation command with the highest weighted frequency is determined as the operation command corresponding to each type of gesture behavior. The gesture validity of the current gesture behavior is determined based on the similarity between each gesture behavior and other gesture behaviors within the same type. Gesture behaviors whose gesture validity meets preset validity conditions are marked as valid gesture operations corresponding to the operation commands. When a target user is identified during the interaction, the target user's real-time gesture data is extracted, and the target user's gesture behavior model is invoked. If it is detected that there are missing key hand points due to occlusion in the real-time gesture data of the target user, the position of the key hand points of the occluded part is predicted based on the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and the gesture behavior model of the target user.

2. The multi-level gesture recognition method for large-screen interaction according to claim 1, characterized in that, Predict the location of key hand points in the occluded area, including: Determine the sequence similarity between the first gesture sequence, which consists of the unobstructed gesture data of the target user within a preset time period before the occlusion occurs, and the second gesture sequence corresponding to each valid gesture operation in the target user's gesture behavior model; Based on the sequence similarity, weights are assigned to each valid gesture operation in the target user's gesture behavior model, and the coordinates of the complete hand key points corresponding to the future time in the historical data for each valid gesture operation are weighted and summed. The result of the weighted sum is used as the predicted position of the hand key points of the occluded part.

3. The multi-level gesture recognition method for large-screen interaction according to claim 2, characterized in that, Determining the sequence similarity between the first gesture sequence and the second gesture sequence includes: The matching distance between the first gesture sequence and the second gesture sequence is determined by a preset sequence similarity analysis algorithm; The sequence similarity between the first gesture sequence and the second gesture sequence is determined based on the negative correlation of the matching distance.

4. The multi-level gesture recognition method for large-screen interaction according to claim 1, characterized in that, Extract real-time gesture data of the target user, including: Real-time acquisition of interactive image frames containing the target user's information using a visual sensor; The target user's frame is identified from the interactive image frame using a preset target detection algorithm; The key points of the target user's hand are detected from the image region corresponding to the person frame where the target user is located using a preset key point detection algorithm, and the coordinate information of the key points of the hand is obtained.

5. The multi-level gesture recognition method for large-screen interaction according to claim 1, characterized in that, After predicting the location of key hand points in the obscured area, the process also includes: The predicted hand keypoints of the occluded part are merged with the identified unoccluded hand keypoints to obtain the complete hand keypoint information of the target user. The complete hand key point information is matched with a preset gesture operation instruction set to obtain the target user's gesture instruction, and the large screen is controlled to execute the operation corresponding to the gesture instruction.

6. A multi-level gesture recognition system for large-screen interaction, characterized in that, include: The historical data analysis module is used to obtain historical gesture data for each user and historical system operation commands associated with the historical gesture data in multi-user interaction scenarios, and to establish a gesture behavior model for each user. The gesture behavior model is used to characterize the hand movement features of the user when issuing different operation commands; The process of acquiring historical gesture data for each user includes: collecting historical image frames containing user interaction scenarios; identifying the person frame of each user from the historical image frames using a preset target detection algorithm; matching pre-created user feature information with user feature information within the person frame using a preset face recognition algorithm to determine the user affiliation of the person frame; and extracting hand key point coordinate information from the person frame as the historical gesture data of the user. Establishing a gesture behavior model for each user includes: determining the gesture recognition weight of each hand keypoint based on the distance from each hand keypoint to the hand reference center in the historical gesture data; the hand reference center is the geometric center of multiple hand keypoints or a wrist keypoint; the gesture recognition weight is used to characterize the distinguishing importance of hand keypoints in different gesture actions; performing similarity analysis on gesture actions at different times in the historical gesture data based on the gesture recognition weight, classifying gesture actions with similarity greater than a preset similarity threshold into the same gesture action type; assigning a corresponding action code to each gesture action type, and dividing the historical gesture data of the same user into multiple gesture behavior sequences according to time sequence; each gesture behavior sequence consists of action codes arranged in time sequence; and determining the gesture recognition weight based on the time interval between adjacent gesture actions and adjacent gesture actions. The similarity between sequences is used to determine the segmentation points between gesture behavior sequences, dividing continuous gesture behavior sequences into multiple gesture behavior units with independent operational intentions. By clustering gesture behavior units in the historical gesture data of the same user, various types of gesture behaviors are identified. Historical operation commands generated after each type of gesture behavior are statistically analyzed, and the frequency of each historical operation command in the current type of gesture behavior is weighted based on the similarity between gesture behaviors within the same type. The historical operation command with the highest weighted frequency is determined as the operation command corresponding to each type of gesture behavior. The gesture validity of the current gesture behavior is determined based on the similarity between each gesture behavior and other gesture behaviors within the same type. Gesture behaviors whose gesture validity meets preset validity conditions are marked as valid gesture operations corresponding to the operation commands. The real-time data acquisition module is used to extract the real-time gesture data of the target user when the target user is identified during the interaction, and to call the target user's gesture behavior model. The occlusion prediction module is used to predict the position of the occluded hand key points if the real-time gesture data of the target user is found to be missing due to occlusion. This prediction is based on the unoccluded gesture data of the target user within a preset time period before the occlusion occurs and the gesture behavior model of the target user.

Citation Information

Patent Citations

  • Hand information identification method and device, hand information control method and device, electronic equipment and medium

    CN115797963A

  • Multi-user holographic sand table collaborative resolving system and method based on 360-degree omni-directional perception gesture recognition

    CN120723072A