Gesture recognition method and system for screen flying interaction

Through video data processing and hand skeleton modeling technology, combined with distance measurement and collision detection algorithm, the misjudgment problem of gesture overlap and crossover in multi-user gesture recognition is solved, and the accuracy and stability of gesture recognition are improved.

CN120103966AActive Publication Date: 2025-06-06GUANGDONG POWER GRID CO LTD

Patent Information

Application Number
CN202510010773.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

In multi-user interaction scenarios, conventional gesture recognition methods lack stability when identifying whether gestures have crossover and overlap, which can easily lead to misjudgment and reduce the accuracy of gesture recognition.

Method used

By obtaining video data of user gestures, using background subtraction algorithm and human posture estimation technology, a hand skeleton model is constructed, hand movement is simulated, and the distance measurement and collision detection algorithm are combined to determine whether the gestures of different users have spatial overlap and crossing.

Benefits of technology

Effectively distinguish between independent gestures and interactive gestures of each user, avoid misjudgment problems in overlapping gestures and crossing gestures, and improve the accuracy and stability of gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103966A_ABST
    Figure CN120103966A_ABST
Patent Text Reader

Abstract

The invention provides a gesture recognition method and system for screen flying interaction, and relates to the technical field of gesture recognition, and the method comprises the steps: processing video data through employing a background subtraction algorithm, and obtaining a foreground region in the video data; extracting a foreground region in the video data by adopting human body posture estimation to obtain a hand key point set of each human body; constructing a hand skeleton model by adopting forward kinematics according to the hand key point set of each human body; utilizing the hand skeleton model to simulate the movement of a hand, and extracting a gesture area and an actual space position of each human body; according to the gesture area and the actual space position of each human body, by combining distance measurement and a collision detection algorithm, whether the gestures of different users are spatially overlapped and crossed or not is judged, the independent gestures and the interactive gestures of the users are distinguished, the problem of misjudgment of a conventional gesture recognition method in the gesture overlapping and crossing state is avoided, and the user experience is improved. Therefore, the accuracy of gesture recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gesture recognition technology, and in particular to a gesture recognition method and system for flying screen interaction. Background Art

[0002] In modern human-computer interaction, users can achieve cross-device operations and interactions through gesture recognition. However, in actual applications, in multi-user interaction scenarios, if multiple users use gestures to operate at the same time, the spatial positions of user gestures may overlap and intersect. Conventional gesture recognition methods focus on single-user scenarios, and have insufficient stability when identifying whether gestures intersect and overlap. Due to the lack of stability, conventional gesture recognition methods are prone to misjudge gestures in overlapping and intersecting states, thereby reducing the accuracy of gesture recognition. Summary of the invention

[0003] The present application provides a gesture recognition method and system for flying screen interaction, aiming to solve the problem of insufficient stability in the related art when identifying whether gestures are crossed and overlapped, thereby avoiding the problem of misjudging gestures in overlapping and crossed states.

[0004] In order to achieve the above object, the present application provides a gesture recognition method for flying screen interaction, including the following:

[0005] Get video data of user gestures.

[0006] The video data is processed using a background subtraction algorithm to obtain the foreground area in the video data.

[0007] Human body posture estimation is used to extract the foreground area in the video data to obtain a set of hand key points of each human body.

[0008] According to the set of key points of each human hand, the hand skeleton model is constructed using forward kinematics.

[0009] The hand skeleton model is used to simulate the movement of the hand and extract the gesture area and actual spatial position of each human body.

[0010] According to the gesture area and actual spatial position of each human body, distance measurement combined with collision detection is used to determine whether the gestures of different users overlap or intersect in space.

[0011] Trigger interactive feedback based on the judgment results.

[0012] As a preferred solution of the present invention, a background subtraction algorithm is used to process the video data to obtain the foreground area in the video data, specifically including:

[0013] Initialize Gaussian mixture model parameters; the Gaussian mixture model parameters include: Gaussian distribution number, learning rate and foreground determination threshold.

[0014] Construct a Gaussian mixture model based on the initialized Gaussian mixture model parameters.

[0015] A plurality of video frames containing a human portrait are extracted from the video data, each of the video frames containing a plurality of pixel points.

[0016] A background model corresponding to each pixel is established using a Gaussian mixture model; the background model includes at least one Gaussian distribution.

[0017] The total matching degree corresponding to each pixel is calculated based on each Gaussian distribution in the background model corresponding to each pixel, and the current pixel is judged whether it is a foreground area or a background area according to the total matching degree.

[0018] If the current pixel is considered to be the foreground area, the background model is updated to adapt to the current scene changes; if the current pixel is considered to be the background, the background model is not updated and the current background model is still used.

[0019] The pixels in the foreground area of ​​each video frame are marked as 1, and the pixels in the background area of ​​each video frame are marked as 0, so as to generate a foreground mask map of each video frame.

[0020] The foreground mask images in each video frame are merged to track and extract the foreground area in the video data.

[0021] As a preferred solution of the present invention, human posture estimation is used to extract the foreground area in the video data to obtain a set of hand key points of each human body, specifically including:

[0022] A human posture estimation model is used to identify the foreground area in the video data to obtain multiple limb key points related to the human body in the video data; the multiple limb key points include the head, torso, limbs and hands.

[0023] According to the fixed index of each hand corresponding to the human body in the key point sequence, the key point information related to the hand is output; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions.

[0024] A set of key points of the hands corresponding to each human body is generated according to key point information related to the hands corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

[0025] As a preferred solution of the present invention, a hand skeleton model is constructed using forward kinematics based on a set of hand key points of each human body, specifically including:

[0026] The connection relationship between joint nodes and joints is defined according to the set of key points of each human hand, and the spatial transformation relationship between each joint is calculated.

[0027] According to the spatial transformation relationship between each joint, a hand skeleton model is constructed, which includes the three-dimensional positions and relative motion relationships of all joints.

[0028] As a preferred solution of the present invention, a hand skeleton model is used to simulate the movement of the hand, and the gesture area and actual spatial position of each human body are extracted, which specifically includes:

[0029] A three-dimensional bounding box is constructed based on all key three-dimensional positions in the hand skeleton model; within the three-dimensional bounding box, the gesture area is calculated based on the relative motion relationship in the hand skeleton model.

[0030] Based on the continuous frames in the video data, the changes of the hand skeleton model corresponding to each human body are tracked, and the motion trajectory of each joint corresponding to each human body in three-dimensional space is calculated.

[0031] The actual position of each human hand in space is extracted according to the motion trajectory of each joint corresponding to each human body in three-dimensional space.

[0032] As a preferred solution of the present invention, according to the gesture area and actual spatial position of each human body, whether the gestures of different users overlap and intersect in space is determined by distance measurement combined with collision detection, specifically including:

[0033] The center point distance between each human gesture is calculated based on the gesture area and actual spatial position of each human gesture.

[0034] It is determined whether the center point distance is less than the overlap threshold; when the center point distance is less than the overlap threshold, it is determined that the two gesture areas have preliminary overlap, otherwise there is no preliminary overlap.

[0035] When there is a preliminary overlap between two gesture areas, the axis-aligned bounding box collision detection algorithm is used to determine whether the two bounding boxes have an intersection based on the spatial overlap condition; if there is an intersection, it is determined that the two gesture areas have a spatial overlap; if there is no intersection, it is determined that the two gesture areas do not have a spatial overlap.

[0036] When two gesture areas overlap in space, the intersection volume of the two gesture areas is calculated to determine whether the intersection volume is greater than 0; if the intersection volume is greater than 0, it is determined that the two gesture areas overlap in space, otherwise it is considered that there is no spatial intersection.

[0037] As a preferred solution of the present invention, triggering interactive feedback according to the judgment result specifically includes:

[0038] When there is no spatial overlap between gestures, the distance between the center points is greater than the judgment threshold, or the bounding boxes have no intersection, and when there is no spatial overlap between gesture areas, each gesture area is considered to be independent, and the instructions of each gesture are processed independently without considering the interference between users. Otherwise, the current state of the gesture is specifically judged and feedback is triggered.

[0039] The present application further provides a gesture recognition system for flying screen interaction, comprising the following modules:

[0040] The data acquisition module is used to acquire video data of user gestures.

[0041] The background subtraction module is used to process the video data using a background subtraction algorithm to obtain a foreground area in the video data.

[0042] The data extraction module is used to extract the foreground area in the video data by using human posture estimation to obtain a set of hand key points of each human body.

[0043] The hand modeling module is used to construct a hand skeleton model using forward kinematics based on the set of hand key points of each human body.

[0044] The simulation recognition module is used to simulate the movement of the hand using the hand skeleton model and extract the gesture area and actual spatial position of each human body.

[0045] The comprehensive judgment module is used to judge whether the gestures of different users overlap and intersect in space based on the gesture area and actual spatial position of each human body through distance measurement combined with collision detection.

[0046] The feedback module is used to trigger interactive feedback according to the judgment results.

[0047] As a preferred solution of the present invention, the background subtraction module specifically includes:

[0048] The initialization unit is used to initialize the parameters of the Gaussian mixture model; the Gaussian mixture model parameters include: the number of Gaussian distributions, the learning rate and the foreground determination threshold.

[0049] The Gaussian mixture model construction unit is used to construct a Gaussian mixture model according to the initialized Gaussian mixture model parameters.

[0050] The multiple video frame extraction units are used to extract multiple video frames containing human portraits from the video data, each of the video frames containing multiple pixel points.

[0051] The background model building unit is used to build a background model corresponding to each pixel point by using a Gaussian mixture model; the background model includes at least one Gaussian distribution.

[0052] The judgment unit is used to calculate the total matching degree corresponding to each pixel point based on each Gaussian distribution in the background model corresponding to each pixel point, and judge whether the current pixel point is a foreground area or a background area according to the total matching degree.

[0053] The background model updating unit is used to update the background model to adapt to the current scene changes if the current pixel point is considered to be the foreground area; if the current pixel point is considered to be the background, the background model is not updated and the current background model is still used.

[0054] The foreground mask image generating unit is used to mark the pixel points in the foreground area of ​​each video frame as 1, and mark the pixel points in the background area of ​​each video frame as 0, so as to generate a foreground mask image of each video frame.

[0055] The foreground mask image merging unit is used to merge the foreground mask images in each video frame, and track and extract the foreground area in the video data.

[0056] As a preferred solution of the present invention, the data extraction module specifically includes:

[0057] The limb key point recognition unit is used to identify the foreground area in the video data using a human posture estimation model, and obtain multiple limb key points related to the human body in the video data; the multiple limb key points include the head, torso, limbs and hands.

[0058] The key point information determination unit is used to output the key point information related to the hand according to the fixed index of the hand corresponding to each human body in the key point sequence; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions.

[0059] The hand key point set generation unit is used to generate a set of hand key points corresponding to each human body according to the key point information related to the hand corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

[0060] The beneficial effects of this application are:

[0061] 1. By combining distance measurement and collision detection algorithm, it is determined whether different users' gestures overlap or intersect in space, and the independent gestures and interactive gestures of each user are effectively distinguished, avoiding the misjudgment problem of conventional gesture recognition methods in the state of gesture overlap and intersection, thereby improving the accuracy of gesture recognition.

[0062] 2. Human body posture estimation and hand skeleton modeling technology are used to not only dynamically track the gesture area, but also simulate hand movements through forward kinematics. Combined with the motion trajectory analysis of continuous frames, it can alleviate the interference problem when multiple user gestures overlap and cross, and improve the stability of gesture recognition.

[0063] 3. Dynamic background update and processing are implemented based on Gaussian mixture model to adapt to changing environmental conditions, reduce the impact of background noise, and improve processing efficiency to ensure that gesture recognition remains stable under complex backgrounds.

[0064] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0066] Figure 1 A flowchart of a gesture recognition method for flying screen interaction provided in an embodiment of the present application.

[0067] Figure 2 An overall architecture diagram of a gesture recognition system for flying screen interaction provided in an embodiment of the present application.

[0068] Description of the accompanying drawings: 21. Data acquisition module; 22. Background subtraction module; 23. Data extraction module; 24. Hand modeling module; 25. Simulation recognition module; 26. Comprehensive judgment module; 27. Feedback module. DETAILED DESCRIPTION

[0069] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0070] See also Figure 1 , Figure 1 A flowchart of a gesture recognition method for flying screen interaction is provided for an embodiment of the present application.

[0071] In this embodiment, a gesture recognition method for flying screen interaction includes steps S100, S200, S300, S400, S500, S600 and S700.

[0072] Step S100, obtaining video data of user gestures, specifically includes:

[0073] Configure the video acquisition parameters of the video capture device, which include resolution, frame rate, and color mode. After the configuration is complete, start the video capture device and obtain video data from the video capture device. The video data includes continuous video frames to reflect the dynamic changes of user gestures.

[0074] Among them, the video capture device includes a depth camera and a camera.

[0075] Step S200, using a background subtraction algorithm to process the video data to obtain a foreground area in the video data, specifically including:

[0076] Initialize Gaussian mixture model parameters; the Gaussian mixture model parameters include: Gaussian distribution number, learning rate and foreground determination threshold.

[0077] It should be noted that the number of Gaussian distributions is determined through 5-fold cross validation, and the initial value of the learning rate is limited to between 0.01 and 0.1; the foreground judgment threshold is determined based on the matching degree of each Gaussian distribution in the Gaussian mixture model, and the initial value is set to 0.5, which is dynamically adjusted according to the color distribution of the foreground and background in the subsequent judgment process.

[0078] Construct a Gaussian mixture model based on the initialized Gaussian mixture model parameters.

[0079] A plurality of video frames containing a human portrait are extracted from the video data, each of the video frames containing a plurality of pixel points.

[0080] A background model corresponding to each pixel is established using a Gaussian mixture model; the background model includes at least one Gaussian distribution.

[0081] The total matching degree corresponding to each pixel is calculated based on each Gaussian distribution in the background model corresponding to each pixel, and the current pixel is judged whether it is a foreground area or a background area according to the total matching degree.

[0082] It should be noted that the total matching degree corresponding to each pixel is calculated based on each Gaussian distribution in the background model corresponding to each pixel, and the calculation formula is:

[0083]

[0084] In the formula, ω k is the weight of the kth Gaussian distribution, which is determined according to the contribution of the kth Gaussian distribution to the overall model; is a Gaussian distribution function, which is used to indicate the matching degree of the kth Gaussian distribution to the color value I(x, y), where μ k and σ k are the mean and standard deviation of the Gaussian distribution respectively; P(I(x, y)) is the total matching degree corresponding to the current pixel color value I(x, y), p(x, y) represents the current pixel, and I(x, y) represents the color value of the current pixel.

[0085] It should be further explained that it is determined whether the total matching degree is less than the foreground determination threshold; if the matching degree of the current pixel point is less than the foreground determination threshold, the pixel point is considered to belong to the foreground area; otherwise, it is considered to belong to the background area. The judgment process is:

[0086] if P(I(x,y))<T foreground , then p(x, y) = foreground

[0087] else if P(I(x,y))≥T foreground , then p(x, y) = background

[0088] Where P(I(x, y)) represents the matching degree of the pixel point (x, y) in the current video frame I(x, y); T foreground Represents the foreground judgment threshold, and p(x, y) indicates whether the pixel point (x, y) belongs to the foreground area.

[0089] If the current pixel is considered to be the foreground area, the background model is updated to adapt to the current scene changes; if the current pixel is considered to be the background, the background model is not updated and the current background model is still used.

[0090] The pixels in the foreground area of ​​each video frame are marked as 1, and the pixels in the background area of ​​each video frame are marked as 0, so as to generate a foreground mask map of each video frame.

[0091] The foreground mask images in each video frame are merged to track and extract the foreground area in the video data.

[0092] It should be noted that the foreground mask images obtained in each video frame are merged to generate a global foreground area map, and Kalman filtering is used as a tracking algorithm to extract and track the foreground area in the video frame. According to the global foreground area map and the tracking results, the foreground area in the video data is extracted.

[0093] Step S300, extracting the foreground area in the video data by using human posture estimation to obtain a set of hand key points of each human body, specifically including:

[0094] A human posture estimation model is used to identify the foreground area in the video data to obtain multiple limb key points related to the human body in the video data; the multiple limb key points include the head, torso, limbs and hands.

[0095] It should be noted that OpenPose in the deep learning human posture estimation model is used for limb recognition, multiple limb key points of the human body are detected in the foreground area, and the information of multiple limb key points is extracted through a convolutional neural network.

[0096] According to the fixed index of each hand corresponding to the human body in the key point sequence, the key point information related to the hand is output; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions.

[0097] It should be noted that in the key point sequence, the index of the hand key point is fixed, and the hand-related key points are directly located through the index.

[0098] A set of key points of the hands corresponding to each human body is generated according to key point information related to the hands corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

[0099] It should be noted that the specific set is:

[0100] P hand ={(x i ,y i )|i=1,2,...,h}

[0101] Where P hand is the set of hand key points corresponding to each human body, (x i ,y i ) represents the two-dimensional coordinates of each key point corresponding to each human body, and h is the number of key points on the hand.

[0102] Step S400: constructing a hand skeleton model using forward kinematics based on the hand key point set of each human body, specifically including:

[0103] The connection relationship between joint nodes and joints is defined according to the set of key points of each human hand, and the spatial transformation relationship between each joint is calculated.

[0104] It should be noted that, based on the set of key points of the hand, the various joint nodes of the hand are defined, the connection relationship between the joints is defined through anatomical knowledge, and then the spatial transformation relationship between the joints is calculated.

[0105] It should be further explained that the forward kinematics method is used to calculate the rotation angles and displacements between adjacent joints to obtain the spatial transformation relationship between the joints.

[0106] According to the spatial transformation relationship between each joint, a hand skeleton model is constructed, which includes the three-dimensional positions and relative motion relationships of all joints.

[0107] Step S500: simulating the movement of the hand using the hand skeleton model to extract the gesture area and actual spatial position of each human body, specifically including:

[0108] A three-dimensional bounding box is constructed based on all key three-dimensional positions in the hand skeleton model; within the three-dimensional bounding box, the gesture area is calculated based on the relative motion relationship in the hand skeleton model.

[0109] It should be noted that the size and shape of the 3D bounding box are updated with the dynamic changes of the hand. As the hand moves, the 3D bounding box will be adjusted in real time to enclose all key points of the hand model to ensure that the entire gesture area is covered.

[0110] Based on the continuous frames in the video data, the changes of the hand skeleton model corresponding to each human body are tracked, and the motion trajectory of each joint corresponding to each human body in three-dimensional space is calculated.

[0111] The actual position of each human hand in space is extracted according to the motion trajectory of each joint corresponding to each human body in three-dimensional space.

[0112] Step S600: judging whether the gestures of different users overlap or intersect in space by distance measurement combined with collision detection according to the gesture areas and actual spatial positions of the human bodies, specifically including:

[0113] The center point distance between each human gesture is calculated based on the gesture area and actual spatial position of each human gesture.

[0114] It is determined whether the center point distance is less than the overlap threshold; when the center point distance is less than the overlap threshold, it is determined that the two gesture areas have preliminary overlap, otherwise there is no preliminary overlap.

[0115] When there is a preliminary overlap between two gesture areas, the axis-aligned bounding box collision detection algorithm is used to determine whether the two bounding boxes have an intersection based on the spatial overlap condition; if there is an intersection, it is determined that the two gesture areas have a spatial overlap; if there is no intersection, it is determined that the two gesture areas do not have a spatial overlap.

[0116] It should be noted that the spatial overlap condition is:

[0117]

[0118]

[0119] In the formula, Indicates the maximum coordinate of gesture area i on the x-axis; Indicates the minimum coordinate of gesture area i on the x-axis; Indicates the maximum coordinate of gesture area j on the x-axis; Indicates the minimum coordinate of gesture area j on the x-axis; Indicates the maximum coordinate of gesture area i on the y-axis; Indicates the minimum coordinate of gesture area i on the y-axis; Indicates the maximum coordinate of gesture area j on the y-axis; Indicates the minimum coordinate of gesture area j on the y-axis; Indicates the maximum coordinate of gesture area i on the z-axis; Indicates the minimum coordinate of gesture area i on the z-axis; Indicates the maximum coordinate of gesture area j on the z-axis; Indicates the minimum coordinate of gesture area j on the z-axis.

[0120] The purpose of the spatial overlap formula is to determine whether two 3D gesture areas intersect. The judgment logic is:

[0121] For the overlap condition in the x-axis direction: For the x-axis, Indicates that the right side of gesture area i exceeds the left side of gesture area j; It means that the left side of gesture area i is located before the right side of gesture area j. If these two conditions are met, it means that there is overlap in the x-axis direction.

[0122] For the overlap condition in the y-axis direction: For the y-axis, and Represents the overlap in the y-axis direction.

[0123] For the overlap condition in the z-axis direction: For the z-axis, and Indicates the overlap in the z-axis direction.

[0124] If the overlapping conditions of the three axes are met, that is, the ranges of the x, y and z axes intersect, the two gesture areas are considered to overlap in the three-dimensional space, and further collision detection is performed.

[0125] When two gesture areas overlap in space, the intersection volume of the two gesture areas is calculated to determine whether the intersection volume is greater than 0; if the intersection volume is greater than 0, it is determined that the two gesture areas overlap in space, otherwise it is considered that there is no spatial intersection.

[0126] It should be noted that the calculation formula for the intersection volume of two gesture areas is:

[0127]

[0128] Where V intersection represents the intersection volume of the two gesture area bounding boxes, Indicates the maximum and minimum values ​​of the intersection area in the x-axis direction. Indicates the maximum and minimum values ​​of the intersection area in the y-axis direction. Indicates the maximum and minimum values ​​of the intersection area in the z-axis direction.

[0129] Step S700, triggering interactive feedback according to the judgment result, specifically includes:

[0130] When there is no spatial overlap between gestures, the distance between the center points is greater than the judgment threshold, or the bounding boxes have no intersection, and when there is no spatial overlap between gesture areas, each gesture area is considered to be independent, and the instructions of each gesture are processed independently without considering the interference between users. Otherwise, the current state of the gesture is specifically judged and feedback is triggered.

[0131] It should be noted that when the gestures overlap in space but do not intersect, the intersection volume V intersection =0, the user is prompted that the current gesture has a potential conflict, and the bounding box is highlighted to prompt the user to adjust the gesture position;

[0132] When gestures overlap and intersect in space, V intersection >0, it is considered that there is a conflict in the overlapping and intersecting areas, and the processing of gestures in the conflicting area is suspended. The user is reminded by voice that there is a conflict in the gestures and that the user should make adjustments in time.

[0133] At this point, a gesture recognition method for flying screen interaction is completed.

[0134] See also Figure 2 , Figure 2 The figure is an overall architecture diagram of a gesture recognition system for flying screen interaction.

[0135] In this embodiment, a gesture recognition system for flying screen interaction includes the following modules:

[0136] The data acquisition module 21 is used to acquire video data of user gestures.

[0137] The background subtraction module 22 is used to process the video data using a background subtraction algorithm to obtain a foreground area in the video data.

[0138] The data extraction module 23 is used to extract the foreground area in the video data by using human posture estimation to obtain a set of hand key points of each human body.

[0139] The hand modeling module 24 is used to construct a hand skeleton model using forward kinematics based on a set of hand key points of each human body.

[0140] The simulation recognition module 25 is used to simulate the movement of the hand using the hand skeleton model and extract the gesture area and actual spatial position of each human body.

[0141] The comprehensive determination module 26 is used to determine whether the gestures of different users overlap or intersect in space based on the gesture area and actual spatial position of each human body by combining distance measurement with collision detection.

[0142] The feedback module 27 is used to trigger interactive feedback according to the judgment result.

[0143] The background subtraction module 22 mentioned in the above content specifically includes:

[0144] The initialization unit is used to initialize the parameters of the Gaussian mixture model; the Gaussian mixture model parameters include: the number of Gaussian distributions, the learning rate and the foreground determination threshold.

[0145] The Gaussian mixture model construction unit is used to construct a Gaussian mixture model according to the initialized Gaussian mixture model parameters.

[0146] The multiple video frame extraction units are used to extract multiple video frames containing human portraits from the video data, each of the video frames containing multiple pixel points.

[0147] The background model building unit is used to build a background model corresponding to each pixel point by using a Gaussian mixture model; the background model includes at least one Gaussian distribution.

[0148] The judgment unit is used to calculate the total matching degree corresponding to each pixel point based on each Gaussian distribution in the background model corresponding to each pixel point, and judge whether the current pixel point is a foreground area or a background area according to the total matching degree.

[0149] The background model updating unit is used to update the background model to adapt to the current scene changes if the current pixel point is considered to be the foreground area; if the current pixel point is considered to be the background, the background model is not updated and the current background model is still used.

[0150] The foreground mask image generating unit is used to mark the pixel points in the foreground area of ​​each video frame as 1, and mark the pixel points in the background area of ​​each video frame as 0, so as to generate a foreground mask image of each video frame.

[0151] The foreground mask image merging unit is used to merge the foreground mask images in each video frame, and track and extract the foreground area in the video data.

[0152] The data extraction module 23 mentioned in the above content specifically includes:

[0153] The limb key point recognition unit is used to identify the foreground area in the video data using a human posture estimation model, and obtain multiple limb key points related to the human body in the video data; the multiple limb key points include the head, torso, limbs and hands.

[0154] The key point information determination unit is used to output the key point information related to the hand according to the fixed index of the hand corresponding to each human body in the key point sequence; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions.

[0155] The hand key point set generation unit is used to generate a set of hand key points corresponding to each human body according to the key point information related to the hand corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

[0156] At this point, a gesture recognition system for flying screen interaction is completed.

[0157] The above is only a specific implementation of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the embodiments of the present invention, which should be included in the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention shall be based on the protection scope of the claims.

Claims

1. A gesture recognition method for flying screen interaction, characterized in that: These include: Get video data of user gestures; The video data is processed using a background subtraction algorithm to obtain a foreground area in the video data; Human body posture estimation is used to extract the foreground area in the video data to obtain the set of key points of each human hand; According to the set of key points of each human hand, the hand skeleton model is constructed using forward kinematics; Use the hand skeleton model to simulate the movement of the hand and extract the gesture area and actual spatial position of each human body; Based on the gesture area and actual spatial position of each human body, distance measurement combined with collision detection is used to determine whether the gestures of different users overlap or intersect in space; Trigger interactive feedback based on the judgment results.

2. A gesture recognition method for flying screen interaction as claimed in claim 1, characterized in that: The background subtraction algorithm is used to process the video data to obtain the foreground area in the video data, including: Initialize Gaussian mixture model parameters; the Gaussian mixture model parameters include: Gaussian distribution number, learning rate and foreground determination threshold; Construct a Gaussian mixture model according to the initialized Gaussian mixture model parameters; Extracting multiple video frames containing human portraits from the video data, each of the video frames containing multiple pixel points; A background model corresponding to each pixel is established using a Gaussian mixture model; the background model includes at least one Gaussian distribution; The total matching degree of each pixel is calculated based on the Gaussian distribution in the background model corresponding to each pixel, and the current pixel is judged as the foreground area or the background area according to the total matching degree; If the current pixel is considered to be the foreground area, the background model is updated to adapt to the current scene changes; if the current pixel is considered to be the background, the background model is not updated and the current background model is still used; Mark the pixels in the foreground area of ​​each video frame as 1, and mark the pixels in the background area of ​​each video frame as 0, to generate a foreground mask image of each video frame; The foreground mask images in each video frame are merged to track and extract the foreground area in the video data.

3. A gesture recognition method for flying screen interaction as claimed in claim 1, characterized in that: Human body posture estimation is used to extract the foreground area in the video data to obtain a set of key points of each human hand, including: A human posture estimation model is used to identify the foreground area in the video data, and multiple limb key points related to the human body in the video data are obtained; the multiple limb key points include the head, torso, limbs and hands; According to the fixed index of each hand corresponding to the human body in the key point sequence, the key point information related to the hand is output; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions; A set of key points of the hands corresponding to each human body is generated according to key point information related to the hands corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

4. A gesture recognition method for flying screen interaction as claimed in claim 1, characterized in that: According to the set of key points of each human hand, forward kinematics is used to construct the hand skeleton model, including: Define the connection relationship between joint nodes and joints according to the key point set of each human hand, and calculate the spatial transformation relationship between each joint; According to the spatial transformation relationship between each joint, a hand skeleton model is constructed, which includes the three-dimensional positions and relative motion relationships of all joints.

5. A gesture recognition method for flying screen interaction as claimed in claim 1, characterized in that: The hand skeleton model is used to simulate the movement of the hand and extract the gesture area and actual spatial position of each human body, including: A 3D bounding box is constructed based on all key 3D positions in the hand skeleton model; within the 3D bounding box, a gesture area is calculated based on the relative motion relationship in the hand skeleton model; Based on the continuous frames in the video data, the changes of the hand skeleton model corresponding to each human body are tracked, and the motion trajectory of each joint corresponding to each human body in three-dimensional space is calculated; The actual position of each human hand in space is extracted according to the motion trajectory of each joint corresponding to each human body in three-dimensional space.

6. A gesture recognition method for flying screen interaction as claimed in claim 5, characterized in that: According to the gesture area and actual spatial position of each human body, distance measurement combined with collision detection is used to determine whether the gestures of different users overlap or intersect in space. Specifically, it includes: Calculate the center point distance between each human gesture according to the gesture area and actual spatial position of each human gesture; Determine whether the center point distance is less than the overlap threshold; when the center point distance is less than the overlap threshold, it is determined that the two gesture areas have preliminary overlap, otherwise there is no preliminary overlap; When there is a preliminary overlap between two gesture areas, the axis-aligned bounding box collision detection algorithm is used to determine whether the two bounding boxes have an intersection based on the spatial overlap condition; if there is an intersection, it is determined that the two gesture areas have a spatial overlap; if there is no intersection, it is determined that the two gesture areas do not have a spatial overlap; When two gesture areas overlap in space, the intersection volume of the two gesture areas is calculated to determine whether the intersection volume is greater than 0; if the intersection volume is greater than 0, it is determined that the two gesture areas overlap in space, otherwise it is considered that there is no spatial intersection.

7. A gesture recognition method for flying screen interaction as claimed in claim 6, characterized in that: Trigger interactive feedback based on the judgment results, including: When there is no spatial overlap between gestures, the center point distance is greater than the judgment threshold, or the bounding boxes have no intersection, and when there is no spatial overlap between gesture areas, each gesture area is considered to be independent, and the instructions of each gesture are processed independently without considering the interference between users. Otherwise, the current state of the gesture is specifically judged and feedback is triggered.

8. A gesture recognition system for flying screen interaction, characterized in that: Includes the following modules: A data acquisition module, used to acquire video data of user gestures; A background subtraction module is used to process the video data using a background subtraction algorithm to obtain a foreground area in the video data; A data extraction module is used to extract the foreground area in the video data by using human posture estimation to obtain a set of hand key points of each human body; The hand modeling module is used to construct a hand skeleton model using forward kinematics based on the key point set of each human hand; The simulation recognition module is used to simulate the hand movement using the hand skeleton model and extract the gesture area and actual spatial position of each human body; The comprehensive judgment module is used to judge whether the gestures of different users overlap or intersect in space based on the gesture area and actual spatial position of each human body through distance measurement combined with collision detection; The feedback module is used to trigger interactive feedback according to the judgment results.

9. A gesture recognition system for flying screen interaction according to claim 8, characterized in that: The background subtraction module specifically includes: An initialization unit, used to initialize Gaussian mixture model parameters; the Gaussian mixture model parameters include: Gaussian distribution quantity, learning rate and foreground determination threshold; A Gaussian mixture model construction unit, used for constructing a Gaussian mixture model according to the initialized Gaussian mixture model parameters; A plurality of video frame extraction units, used for extracting a plurality of video frames containing a human image from the video data, each of the video frames containing a plurality of pixel points; A background model building unit, used to build a background model corresponding to each pixel point using a Gaussian mixture model; the background model includes at least one Gaussian distribution; A judgment unit, used to calculate the total matching degree corresponding to each pixel point based on each Gaussian distribution in the background model corresponding to each pixel point, and judge whether the current pixel point is a foreground area or a background area according to the total matching degree; A background model updating unit, used to update the background model to adapt to the current scene change if the current pixel is considered to be the foreground area; if the current pixel is considered to be the background, the background model is not updated and the current background model is still used; A foreground mask image generating unit is used to mark the pixel points in the foreground area of ​​each video frame as 1, and mark the pixel points in the background area of ​​each video frame as 0, so as to generate a foreground mask image of each video frame; The foreground mask image merging unit is used to merge the foreground mask images in each video frame, and track and extract the foreground area in the video data.

10. The gesture recognition system for flying screen interaction according to claim 8, characterized in that: The data extraction module specifically includes: A limb key point recognition unit is used to recognize the foreground area in the video data by using a human posture estimation model, and obtain multiple limb key points related to the human body in the video data; the multiple limb key points include the head, torso, limbs and hands; A key point information determination unit, used to output key point information related to the hand according to the fixed index of the hand corresponding to each human body in the key point sequence; the key point information related to the hand includes the two-dimensional coordinates of the wrist and finger joint positions; The hand key point set generation unit is used to generate a set of hand key points corresponding to each human body according to the key point information related to the hand corresponding to each human body, and the set includes two-dimensional coordinate data of each key point of the hand.

Citation Information

Patent Citations

  • Medical image display control method and device, storage medium and electronic equipment

    CN118097792A

  • Multi-user information recommendation method, system and equipment based on mirror screen and medium

    CN118134607A

  • Three-dimensional gesture estimation method and head-mounted display device

    CN118212683A

Cited By

  • Human-computer interaction method and system for network game

    CN120748038A

  • Large-screen interaction-oriented multi-level gesture recognition method and system

    CN122261394A