Screen projection cooperative control method based on gesture recognition

By collecting and analyzing user gesture video data in multi-person interaction scenarios and generating a deterministic operation instruction set, the difficulty of traditional screen projection control mechanisms to identify the sender and intention of gestures in multi-person interactions is solved, and the coordinated control and stability of screen projection devices are achieved.

CN120428867AInactive Publication Date: 2025-08-05SUZHOU YUNGAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510866848.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional gesture recognition and screen projection control mechanisms are difficult to adapt to natural interaction scenarios where multiple people exist at the same time and gestures occur freely. They cannot accurately identify the sender of the gesture and their control intentions, resulting in problems such as false triggering, control conflicts or loss of commands, affecting the stability and response accuracy of the interactive system.

Method used

By collecting real-time video data of all users in the preset interactive space, performing spatial feature segmentation of multi-user gesture candidate areas, analyzing gesture spatial conflict characteristics, and generating a deterministic operation instruction set through time series similarity and user identity association to realize collaborative control of screen projection devices.

Benefits of technology

Accurately capture user control intentions in multi-person interaction environments, effectively identify and locate conflicting areas, reduce the risk of error identification, improve the accuracy of belonging judgment, and ensure the stability and real-time response in complex interactive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428867A_ABST
    Figure CN120428867A_ABST
Patent Text Reader

Abstract

The invention discloses a screen projection cooperative control method based on gesture recognition, and particularly relates to the technical field of cooperative control. The method comprises the following steps: acquiring real-time video data of all users in a preset interaction space in real time, and performing spatial feature segmentation to obtain an initial user gesture spatial feature data set; performing spatial position proximity and overlapping degree analysis on the initial user gesture spatial feature data set to obtain a gesture spatial conflict feature set, and performing time sequence similarity analysis to obtain a gesture validity undetermined data set; through user identity association analysis, a gesture control intention affiliation probability matrix is output, an effective control user and an effective gesture instruction are determined by executing gesture ownership dynamic judgment, a deterministic operation instruction set is generated, and efficient cooperation of projection screen control affiliation judgment and operation execution under multi-user natural gesture input is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of collaborative control technology, and more specifically, to a screen projection collaborative control method based on gesture recognition. Background Art

[0002] Traditional gesture recognition and screen projection control mechanisms typically rely on single-user identity binding or explicit control handoff processes, making them difficult to adapt to natural interaction scenarios where multiple people are present and gestures can flow freely. This is especially true in conferencing, teaching, or collaborative environments, where multiple users may simultaneously initiate control attempts without explicit permission switching instructions.

[0003] Existing technologies generally lack the ability to dynamically determine the ownership of gestures and are unable to accurately identify the sender of the gesture and its control intention, which can lead to problems such as false triggering, control conflicts or command loss, seriously affecting the stability and response accuracy of the interactive system.

[0004] In order to solve the above problems, a technical solution is now provided. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a screen projection collaborative control method based on gesture recognition to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A screen projection collaborative control method based on gesture recognition includes the following steps:

[0008] Collect real-time video data of all users in the preset interaction space, perform spatial feature segmentation of multi-user gesture candidate areas on the real-time video data, and obtain an initial user gesture spatial feature dataset;

[0009] Perform spatial proximity and overlap analysis on all gesture candidate regions in the initial user gesture spatial feature dataset to obtain a gesture spatial conflict feature set;

[0010] Based on the gesture spatial conflict feature set, a gesture validity dataset was obtained through time series similarity analysis.

[0011] Perform user identity association analysis on the gesture validity dataset. Based on the association records between historical gesture actions and user control rights, output the attribution probability matrix of the current gesture control intention.

[0012] Perform dynamic judgment of gesture ownership based on the ownership probability matrix, and determine the valid user and valid gesture command of the current screen projection control according to the maximum ownership probability of the gesture control intention in the ownership probability matrix;

[0013] Based on the valid users and valid gesture instructions of the current screen projection control, a deterministic operation instruction set for the current screen projection collaborative control is generated to realize the collaborative control operation of the screen projection device.

[0014] In a preferred embodiment, real-time video data of all users in a preset interactive space is collected, and spatial feature segmentation of multiple user gesture candidate areas is performed on the real-time video data to obtain an initial user gesture spatial feature dataset, specifically:

[0015] Collect real-time video data of all users in the preset interactive space;

[0016] Determine preliminary position coordinates of multiple user hand contour areas based on pixel color distribution information of the user's limb area in the real-time video data;

[0017] Perform edge detection on the preliminary position coordinates of multiple user hand contour areas to obtain accurate spatial contour information of multiple user gesture candidate areas;

[0018] According to the precise spatial contour information of multiple user gesture candidate areas, spatial feature quantization processing is performed on the multiple user gesture candidate areas respectively to obtain an initial user gesture spatial feature dataset.

[0019] In a preferred embodiment, spatial proximity and overlap analysis is performed on all gesture candidate regions in the initial user gesture spatial feature dataset to obtain a gesture spatial conflict feature set, and a gesture validity pending dataset is obtained through time series similarity analysis, specifically:

[0020] Based on all the gesture candidate regions in the initial user gesture spatial feature dataset, the Euclidean distance between the center coordinates of any two gesture candidate regions and the overlap ratio of the contour pixels between the gesture candidate regions are calculated;

[0021] According to the preset spatial proximity distance threshold and contour pixel overlap ratio threshold, a group of gesture candidate regions that meet the conditions of spatial proximity and contour overlap are determined and marked as the gesture spatial conflict feature set.

[0022] In a preferred embodiment, based on the gesture spatial conflict feature set, a gesture validity pending dataset is obtained through time series similarity analysis, specifically:

[0023] For each group of gesture candidate regions marked in the gesture spatial conflict feature set, the spatial coordinate position data in the real-time video data is extracted, and the similarity of the spatial motion trajectories of each group of gesture candidate regions is calculated;

[0024] Based on the preset similarity threshold, the gesture candidate region groups with similarity higher than the similarity threshold are screened and marked as gesture validity pending datasets.

[0025] In a preferred embodiment, user identity association analysis is performed on the gesture validity pending dataset, and based on the association records between historical gesture actions and user control rights, an attribution probability matrix of the current gesture control intention is output, specifically:

[0026] Extract the spatial location data of each gesture candidate region in the gesture validity pending dataset, combine it with the facial feature information of each user in the real-time video data, and establish the spatial correspondence between the gesture candidate region and the user identity;

[0027] Based on the preset user historical control behavior records, the frequency of each user gaining control after completing a specific gesture action in the historical session is counted. Combined with the spatial correspondence, the control attribution probability values of the gesture candidate area for multiple users are calculated to construct the attribution probability matrix of the current gesture control intention.

[0028] In a preferred embodiment, dynamic judgment of gesture ownership is performed based on the attribution probability matrix, and the valid user and valid gesture instruction of the current screen projection control are determined according to the maximum attribution probability of the gesture control intention in the attribution probability matrix, specifically:

[0029] Extract the belonging probability value of each user corresponding to all gesture candidate areas in the belonging probability matrix and record the maximum belonging probability value;

[0030] Compare the recorded maximum attribution probability value with the decision-making probability threshold;

[0031] When the maximum attribution probability value is higher than or equal to the judgment decision probability threshold, the single user corresponding to the maximum attribution probability value is determined as the valid user of the current projection screen control, and the spatial feature data of the gesture candidate area corresponding to the maximum attribution probability value is determined as the valid gesture instruction of the current projection screen control according to the preset control instruction mapping rules.

[0032] In a preferred embodiment, based on the valid user and valid gesture instructions of the current screen projection control, a deterministic operation instruction set for the current screen projection collaborative control is generated to implement the collaborative control operation of the screen projection device, specifically:

[0033] Obtain user identity feature information of the valid user of the current screen projection control, and identify the screen projection control semantic information corresponding to the valid gesture instruction based on the spatial coordinate position data and shape contour data of the valid gesture instruction of the current screen projection control;

[0034] According to the mapping relationship between the projection control semantic information and the deterministic operation instructions, the projection control semantic information is mapped and converted into the corresponding deterministic operation instructions;

[0035] Based on the deterministic operation instructions, a deterministic operation instruction set for the current screen projection collaborative control is generated in accordance with the operation execution order associated with the deterministic operation instructions, and sent to the screen projection device to implement the collaborative control operation of the screen projection device.

[0036] The technical effects and advantages of the screen projection collaborative control method based on gesture recognition of the present invention are as follows:

[0037] By collecting user gesture video data in real time and accurately segmenting and extracting gesture spatial features, it is possible to accurately capture the control intentions of users in a multi-person interactive environment. By analyzing the proximity and overlap of spatial positions, it is possible to effectively identify and locate conflicting areas in the multi-person gesture interaction process. By using time series similarity analysis, it is possible to identify gestures that may lead to ambiguous judgments and effectively reduce the risk of misidentification. Based on the correlation analysis between user identity and historical control records, it quantifies the probability of attribution of gesture control to different users and improves the accuracy of attribution judgments. It performs dynamic judgment of gesture ownership through the attribution probability matrix to determine the effective user with the most control intention and their gesture instructions. Finally, it generates a deterministic set of operation instructions to achieve precise collaborative control of projection devices. It effectively improves the intelligent judgment capability in multi-person collaborative control scenarios, reduces the problems of gesture misattribution, control conflict and response delay, and ensures stability, real-time performance and user experience in complex interactive scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a schematic diagram of a screen projection collaborative control method based on gesture recognition in the present invention. DETAILED DESCRIPTION

[0039] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0040] Example 1

[0041] Figure 1 The present invention provides a screen projection collaborative control method based on gesture recognition, which includes the following steps:

[0042] Collect real-time video data of all users in the preset interaction space, perform spatial feature segmentation of multi-user gesture candidate areas on the real-time video data, and obtain an initial user gesture spatial feature dataset;

[0043] Perform spatial proximity and overlap analysis on all gesture candidate regions in the initial user gesture spatial feature dataset to obtain a gesture spatial conflict feature set;

[0044] Based on the gesture spatial conflict feature set, a gesture validity dataset was obtained through time series similarity analysis.

[0045] Perform user identity association analysis on the gesture validity dataset. Based on the association records between historical gesture actions and user control rights, output the attribution probability matrix of the current gesture control intention.

[0046] Perform dynamic judgment of gesture ownership based on the ownership probability matrix, and determine the valid user and valid gesture command of the current screen projection control according to the maximum ownership probability of the gesture control intention in the ownership probability matrix;

[0047] Based on the valid users and valid gesture instructions of the current screen projection control, a deterministic operation instruction set for the current screen projection collaborative control is generated to realize the collaborative control operation of the screen projection device.

[0048] Specifically, real-time video data of all users in a preset interaction space is collected, and spatial feature segmentation of multi-user gesture candidate areas is performed on the real-time video data to obtain an initial user gesture spatial feature dataset, including:

[0049] Collect real-time video data of all users in the preset interactive space;

[0050] Specifically, multiple wide-angle camera devices are installed within the preset interaction space. The field of view of each wide-angle camera device covers the entire preset interaction space and overlaps with each other to ensure that there are no blind spots in the real-time video data collection for all users. The video data resolution of each wide-angle camera device is uniformly set to a high-resolution format to ensure the accuracy of feature extraction of the user's body area. The video data acquisition frame rate of the wide-angle camera device is configured to the real-time video data stream standard to ensure the temporal continuity of gesture movements.

[0051] Determine preliminary position coordinates of multiple user hand contour areas based on pixel color distribution information of the user's limb area in the real-time video data;

[0052] Specifically, an image segmentation method based on color histogram statistical analysis is selected. By analyzing the user's skin color distribution information in real-time video data, the user's hand contour area is determined: each frame of the real-time video data is divided into multiple image units, and the proportion of pixels in each image unit whose color value falls within the typical skin color range is counted. When the color value of an image unit satisfies the pixel ratio greater than a preset skin color pixel ratio threshold, the image unit is marked as the preliminary location area where the user's hand is located. By calculating the center position of the boundary of the marked image unit area, the preliminary position coordinates of multiple user hand contour areas are obtained.

[0053] For example, if the user raises a hand gesture in the preset interactive space, the proportion of skin-colored pixels in the corresponding image unit in the real-time video data increases significantly, and the corresponding image unit is identified as a preliminary position area, and the center position coordinates of the area are calculated.

[0054] Perform edge detection on the preliminary position coordinates of multiple user hand contour areas to obtain accurate spatial contour information of multiple user gesture candidate areas;

[0055] Specifically, a gradient-based edge detection algorithm is used to process the determined preliminary location region. The preliminary location region is first converted into a grayscale image, and the grayscale image is subjected to image noise smoothing filtering to reduce local noise interference. The processed grayscale image is then subjected to horizontal and vertical pixel gradient changes to determine candidate pixel locations at the image edge. When the degree of pixel gradient change exceeds a preset gradient change threshold, the edge pixel locations of multiple user gesture candidate regions are obtained. The collection of edge pixel locations constitutes the precise spatial contour information of the gesture candidate region, defining the external contour morphology of the gesture.

[0056] Based on the precise spatial contour information of multiple user gesture candidate areas, the spatial features of the multiple user gesture candidate areas are quantified to obtain an initial user gesture spatial feature dataset;

[0057] Specifically, a multi-dimensional spatial feature quantization method is selected, which includes the following three specific contents of spatial feature quantization: spatial coordinate position data, which is obtained by calculating the geometric center position of the precise spatial contour of each user gesture candidate area, so as to clarify the specific position of each gesture in space; shape contour data, which is obtained by performing Fourier contour descriptor transformation on the spatial contour boundary of each user gesture candidate area to obtain a shape feature vector, which is used to characterize the difference between the gesture shapes of different users; geometric size data, which is obtained by calculating the average distance between the two farthest points on the contour boundary of the gesture candidate area to obtain gesture size features, which is used to distinguish the size scales of gestures of different users.

[0058] For example, if the spatial outline of a gesture candidate region is a closed ellipse, a set of feature vectors is obtained through Fourier transform of the contour descriptor to represent the ellipse shape. Simultaneously, the average distance between the two farthest points on the ellipse contour boundary is calculated to obtain the geometric size characteristics of the gesture region.

[0059] Specifically, we perform spatial proximity and overlap analysis on all gesture candidate regions in the initial user gesture spatial feature dataset to obtain a gesture spatial conflict feature set. We also perform time series similarity analysis to obtain a gesture validity dataset, including:

[0060] Based on all the gesture candidate regions in the initial user gesture spatial feature dataset, the Euclidean distance between the center coordinates of any two gesture candidate regions and the overlap ratio of the contour pixels between the gesture candidate regions are calculated;

[0061] Specifically, the center position coordinates of each gesture candidate area are obtained from the initial user gesture spatial feature dataset. Each center position coordinate is determined by the coordinate value in the horizontal axis direction and the coordinate value in the vertical axis direction of the image. The center position coordinates of each two gesture candidate areas in the initial user gesture spatial feature dataset are paired and the Euclidean distance is calculated separately: the square value of the coordinate difference of the two center position coordinates in the horizontal axis direction is added to the square value of the coordinate difference in the vertical axis direction, and then the arithmetic square root of the sum is taken to obtain the Euclidean distance between the two gesture candidate areas.

[0062] The Euclidean distance value of each pair of gesture candidate areas in the initial user gesture spatial feature dataset is calculated in sequence as a basis for judging the proximity of gesture space.

[0063] The contour boundary of each gesture candidate region is determined based on its shape contour data, and the pixel set occupied by each gesture candidate region is obtained. For each pixel set occupied by two gesture candidate regions in the initial user gesture spatial feature dataset, the ratio between the number of pixels contained in the overlapping portion between the two pixel sets and the total number of pixels contained in one pixel set is calculated. That is, the value obtained by dividing the number of overlapping pixels by the total number of pixels in the pixel set itself is the contour pixel overlap ratio between the two gesture candidate regions.

[0064] The contour pixel overlap ratio values between each two gesture candidate regions are determined in sequence to determine the degree of overlap of the gesture contours.

[0065] According to the preset spatial proximity distance threshold and contour pixel overlap ratio threshold, a group of gesture candidate regions that meet the conditions of spatial proximity and contour overlap are determined and marked as a gesture spatial conflict feature set;

[0066] Specifically, the spatial proximity threshold is set by calculating the minimum reasonable distance between gestures that are adjacent but not interfering with each other during user interaction. The threshold for the outline pixel overlap ratio is determined based on the acceptable pixel overlap threshold between different gestures in actual interaction scenarios.

[0067] The Euclidean distance between each two candidate gesture regions is compared with the spatial proximity distance threshold. If the Euclidean distance is less than or equal to the spatial proximity distance threshold, the two candidate gesture regions are determined to meet the spatial proximity condition. The outline pixel overlap ratio between each two candidate gesture regions is compared with the outline pixel overlap ratio threshold. If the outline pixel overlap ratio is greater than or equal to the outline pixel overlap ratio threshold, the two candidate gesture regions are determined to meet the outline overlap condition.

[0068] Any two gesture candidate regions that meet the spatial proximity and contour overlap conditions are recorded as a set of gesture space conflict feature combinations, and all gesture space conflict features that meet the spatial proximity and contour overlap conditions are summarized and uniformly marked as the gesture space conflict feature set.

[0069] Specifically, based on the gesture spatial conflict feature set and through time series similarity analysis, a gesture validity dataset is obtained, including:

[0070] For each group of gesture candidate regions marked in the gesture spatial conflict feature set, the spatial coordinate position data in the real-time video data is extracted, and the similarity of the spatial motion trajectories of each group of gesture candidate regions is calculated;

[0071] Specifically, the spatial center coordinates of each gesture candidate region in multiple consecutive real-time video frames are obtained from the gesture spatial conflict feature set. The spatial center coordinates are determined by taking the arithmetic mean of the horizontal coordinates of all pixels along the boundary of the region in each frame as the horizontal coordinate, and the arithmetic mean of the vertical coordinates of all pixels as the vertical coordinate. This method generates the spatial center coordinates of each gesture candidate region in each consecutive video frame.

[0072] For example, in a specific interaction scenario, if a set of gesture spatial conflict feature sets includes gestures performed simultaneously by two users, the spatial center position coordinate data of each of the two users in multiple consecutive real-time video frames are extracted separately to clearly record the spatial motion trajectory of their gestures over time.

[0073] The dynamic time warping algorithm is used to calculate the similarity of the spatial motion trajectories between each group of gesture candidate areas. The dynamic time warping algorithm obtains the similarity value between the two trajectory sequences by finding the best alignment path between the spatial coordinate sequences of the two motion trajectories. Specifically, the spatial position coordinate points in the two spatial motion trajectory sequences are compared one by one, the Euclidean distance between the corresponding coordinate points is calculated, the Euclidean distance values are accumulated one by one, and the corresponding relationship that minimizes the total accumulated distance is found, thereby obtaining the best matching path of the two spatial motion trajectories and the similarity value corresponding to the path. Through the application of the dynamic time warping algorithm, the differences in motion rhythm and speed between different gesture candidate areas can be effectively handled, and accurate and stable similarity evaluation results can be obtained.

[0074] Based on the preset similarity threshold, the gesture candidate region groups with similarity higher than the similarity threshold are screened and marked as gesture validity pending datasets;

[0075] Specifically, a threshold value of similarity between user gestures within a reasonable range of differences is calculated as the similarity threshold. The motion trajectory similarity values between each group of gesture candidate regions, calculated by the dynamic time warping algorithm, are compared with a preset similarity threshold. If the calculated similarity value is greater than or equal to the similarity threshold, the group of gesture candidate regions is marked as a gesture validity pending dataset.

[0076] Specifically, we conduct user identity association analysis on the gesture validity dataset. Based on the association records between historical gesture actions and user control rights, we output the attribution probability matrix of the current gesture control intention, including:

[0077] Extract the spatial location data of each gesture candidate region in the gesture validity pending dataset, combine it with the facial feature information of each user in the real-time video data, and establish the spatial correspondence between the gesture candidate region and the user identity;

[0078] Specifically, the spatial coordinate position data of each gesture candidate region is extracted one by one from the gesture validity dataset. The spatial coordinate position data of each gesture candidate region is determined by calculating the arithmetic mean of the horizontal and vertical coordinate values of all pixel positions on the region's outline boundary, ensuring that the position data accurately represents the exact position of the gesture candidate region in the image space.

[0079] Based on the preset user control behavior history records, the frequency of each user gaining control after completing a specific gesture action in the historical session is counted. Combined with the spatial correspondence, the control attribution probability values of the gesture candidate area for multiple users are calculated to construct the attribution probability matrix of the current gesture control intention;

[0080] Specifically, a facial feature recognition algorithm is used to analyze the facial area of each user in the real-time video data, extract the key facial feature points of each user, including the positions of the corners of the eyes, the tip of the nose, and the corners of the mouth, and use the spatial coordinate positions of the key facial feature points as a clear spatial identifier of each user's identity. The distance between the spatial position data of each gesture candidate area and the spatial position data of the key facial feature points of each user is calculated one by one: the sum of the distance value between the spatial position of the gesture candidate area and the center position of the user's facial area in the horizontal coordinate direction and the square of the distance value in the vertical coordinate direction is calculated, and the arithmetic square root of the sum is taken to determine the spatial distance value between the gesture candidate area and the user identity. The user identity with the smallest spatial distance value is clearly associated with the gesture candidate area, thereby establishing a spatial correspondence.

[0081] For example, when there are gesture candidate areas raised by two users in a specific video scene, the position data of the two gesture candidate areas are first extracted, and then the spatial distance between each gesture candidate area and the facial areas of all users is calculated respectively. Then, each gesture candidate area is clearly associated with the identity of the user closest to it, so as to establish a clear spatial correspondence.

[0082] A pre-built database of user control history records the number of times each user successfully gained control of the projection screen after completing a specific gesture during their interactions, as well as the total number of times a specific gesture was performed by the user. Frequency is calculated by dividing the total number of times a user successfully gained control after completing a specific gesture by the total number of times the user completed that gesture. This determines the probability of each user gaining control after performing a specific gesture. Statistical analysis reveals the historical patterns of each user's gesture control behavior, providing accurate historical data support for probability calculations.

[0083] For example, a specific user has historically performed a right-swiping gesture a total of ten times, of which eight times he successfully gained control. The probability of the user gaining control after performing the right-swiping gesture is 80%, which accurately represents the actual situation of the historical control behavior.

[0084] When there's a clear spatial correspondence between a gesture candidate region and a specific user, the probability of a user historically gaining control by performing a specific gesture is used as the probability of the current gesture candidate region being controlled by that user. If the spatial correspondence shows that multiple users are relatively close to the same gesture candidate region, making it unclear where they belong, a normalized weight adjustment is performed based on each user's historical probability of gaining control by performing that gesture, combined with the spatial distance. Specifically, each user's probability is multiplied by the inverse of the user's spatial distance from the gesture candidate region. The adjusted probability values are then normalized across all users to determine the probability of control belonging to the gesture candidate region for multiple users.

[0085] For example, if the spatial position of a gesture candidate area is close to the spatial distance between two users, and there are differences in the historical probability values of the two users performing specific gesture actions to obtain control, the normalized weight adjustment calculation method is used to clearly determine the control probability value of the gesture candidate area for each user, ensuring that the analysis results fully combine historical data and current spatial correspondence, and are objective and accurate.

[0086] By calculating the control attribution probability values for each user across all the gesture candidate areas, we construct an attribution probability matrix for the current gesture control intent. Specifically, the rows of the matrix are the gesture candidate areas, and the columns are the user identities. The control attribution probability values for each gesture candidate area are filled into the corresponding matrix positions to form the attribution probability matrix.

[0087] Specifically, dynamic gesture ownership judgment is performed based on the ownership probability matrix, and the valid user and valid gesture command of the current screen projection control are determined according to the maximum ownership probability of the gesture control intention in the ownership probability matrix, including:

[0088] Extract the belonging probability value of each user corresponding to all gesture candidate areas in the belonging probability matrix and record the maximum belonging probability value;

[0089] Specifically, the attribution probability values for each gesture candidate region for each user are obtained from the attribution probability matrix. The attribution probability values explicitly reflect the probability that each gesture candidate region is controlled by each user. The attribution probability values are explicitly organized in the form of a matrix, where each row represents a gesture candidate region, each column represents a user, and each matrix cell contains a value that explicitly represents the probability that the gesture candidate region is controlled by the corresponding user.

[0090] The probability values within each cell of the attribution probability matrix are analyzed and compared one by one. The highest probability value is selected, and the single user identity and specific gesture candidate area that this value clearly corresponds to are recorded. The specific probability value comparison method is as follows: starting with the probability value corresponding to the first gesture candidate area and the first user, the current value is compared with the previously analyzed values one by one, and the larger probability values and their corresponding user identities and gesture candidate areas are continuously recorded until all matrix cells are traversed.

[0091] For example, in a specific video scene, the attribution probability matrix contains two gesture candidate regions and the control probability data of two users. Through a one-by-one comparison method, the single user with the highest control probability and its corresponding gesture candidate region are determined, and the corresponding highest probability value is recorded.

[0092] Compare the recorded maximum probability value with the decision-making probability threshold;

[0093] Specifically, the method for setting the judgment decision probability threshold is: through statistical analysis of a large amount of real gesture data in actual scenarios, the probability critical value between the user gesture control behavior being misjudged and correctly judged is used as the basis for determining the judgment decision probability threshold, ensuring that the set judgment decision probability threshold can effectively distinguish between reliable control behaviors and behaviors with uncertain attribution.

[0094] When the maximum attribution probability value is higher than or equal to the judgment decision probability threshold, the single user corresponding to the maximum attribution probability value is determined as the valid user for the current screen projection control, and the spatial feature data of the gesture candidate area corresponding to the maximum attribution probability value is determined as the valid gesture instruction for the current screen projection control according to the preset control instruction mapping rule;

[0095] Specifically, when the maximum probability value recorded is higher than or equal to the judgment decision probability threshold, it means that the corresponding user has a clear control intention to perform the corresponding gesture, and the recorded user is determined to be the actual effective screen projection control user.

[0096] A control instruction mapping rule database is established to store the one-to-one mapping relationship between each gesture spatial feature and the specific control instruction. The spatial feature data of the gesture candidate area corresponding to the maximum probability value, including shape contour data, geometric dimension data and spatial position data, are feature matched one by one with the mapping rules stored in the control instruction mapping rule database. The matching method is specifically as follows: the similarity between the spatial feature data of the gesture candidate area and various types of spatial feature data recorded in the control instruction mapping rule database is compared one by one. When the degree of matching between the spatial feature data of the gesture candidate area and the feature data corresponding to a control instruction in the database exceeds the preset matching threshold, the matched control instruction is clearly determined as a valid gesture instruction for the current screen projection control.

[0097] For example, in a real-world scenario, if the spatial feature data of a user's gesture is determined to be a straight line swiping from left to right, and by matching it with the control instruction mapping rule database, it can be clearly determined that the control instruction corresponding to the spatial feature data is the "next page" operation instruction on the projection screen interface. Therefore, the operation corresponding to the spatial feature data is clearly determined to be a valid gesture instruction for the current projection screen control, effectively implementing the projection screen control operation.

[0098] Specifically, based on the valid user and valid gesture instructions of the current screen projection control, a deterministic operation instruction set for the current screen projection collaborative control is generated to implement the collaborative control operation of the screen projection device, including:

[0099] Obtain user identity feature information of the valid user of the current screen projection control, and identify the screen projection control semantic information corresponding to the valid gesture instruction based on the spatial coordinate position data and shape contour data of the valid gesture instruction of the current screen projection control;

[0100] Specifically, based on the valid user currently controlling the screen projection, the identity characteristic information of the valid user stored in the user identity information database is called. User identity characteristic information includes the user's unique user number, user name, user role, and user permission level. By obtaining clear user identity characteristic information, the identity of the current operator can be accurately identified, and the scope of operations and permissions allowed to be performed can be determined based on the user's identity, thereby ensuring the effectiveness of collaborative control operations.

[0101] The spatial coordinate position data and shape contour data of effective gesture instructions are fused, specifically by combining the motion direction features of the spatial coordinate position data and the geometric features of the shape contour data into a unified gesture feature vector. A pre-trained gesture semantic classification and recognition algorithm is used to perform screen projection control semantic recognition on the fused gesture feature vector. Specifically, a classification algorithm in a machine learning algorithm is used to train a large amount of labeled gesture training data, establish a classification mapping rule between gesture feature vectors and screen projection control semantic information, and apply the classification mapping rule to classify and recognize the actual gesture feature vector to obtain specific screen projection control semantic information.

[0102] According to the mapping relationship between the projection control semantic information and the deterministic operation instructions, the projection control semantic information is mapped and converted into the corresponding deterministic operation instructions;

[0103] Specifically, a mapping relationship database of control semantics to deterministic operation instructions is constructed. The mapping relationship database records the clear mapping relationship between each type of screen projection control semantic information and the corresponding deterministic operation instruction. The screen projection control semantic information is matched and compared with the mapping relationship database. The specific matching method is: the identified screen projection control semantic information is accurately text-matched with each record stored in the mapping relationship database. When the screen projection control semantic information in the mapping relationship database is exactly the same as the identified semantic information text, the corresponding deterministic operation instruction in the mapping relationship database is recorded as the current deterministic operation instruction.

[0104] Based on the deterministic operation instructions, a deterministic operation instruction set for the current screen projection collaborative control is generated in accordance with the operation execution order associated with the deterministic operation instructions, and sent to the screen projection device to implement the collaborative control operation of the screen projection device;

[0105] Specifically, an operation sequence execution rule table is constructed to record the specific execution order of each deterministic operation instruction in actual operation and its associated projection device control command. Specifically, the deterministic operation instruction is input into the operation sequence execution rule table for matching. By matching the projection device control command associated with the deterministic operation instruction, the corresponding commands are arranged one by one according to the order of records in the operation sequence execution rule table to form a set of deterministic operation instructions.

[0106] By recording the operation sequence, we can ensure the sequence and stability of the screen projection control operation, and avoid control errors or equipment misoperation caused by confusion in the operation sequence.

[0107] For example, if the deterministic operation instruction is "start video playback", the deterministic operation instruction is mapped into a clear control command according to the operation sequence execution rule table, such as: open the video player program, load the video resource, execute the playback instruction, and arrange them in sequence into a deterministic operation instruction set.

[0108] Use a network communication interface to send a deterministic set of operation instructions to the target projection device via a standardized data transmission protocol, and then execute the corresponding operation commands on the projection device. The specific sending method is as follows: first, clearly establish a stable network communication connection with the target projection device, then send each operation command in the deterministic operation instruction set, and wait for the projection device to confirm the successful execution of each command before continuing to send the next command to ensure accurate and reliable command implementation.

[0109] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0110] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0111] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0112] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0113] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0114] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0115] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0116] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0117] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0118] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A screen projection collaborative control method based on gesture recognition, characterized in that: The steps include: S1: Collect real-time video data of all users in the preset interaction space, perform spatial feature segmentation of multi-user gesture candidate areas on the real-time video data, and obtain the initial user gesture spatial feature dataset; S2: Perform spatial proximity and overlap analysis on all gesture candidate regions in the initial user gesture spatial feature dataset to obtain a gesture spatial conflict feature set; S3: Based on the gesture spatial conflict feature set, the gesture validity dataset is obtained through time series similarity analysis; S4: Perform user identity association analysis on the gesture validity dataset to be determined. Based on the association records between historical gesture actions and user control rights, output the attribution probability matrix of the current gesture control intention. S5: Perform dynamic judgment on gesture ownership based on the ownership probability matrix, and determine the valid user and valid gesture command of the current screen projection control according to the maximum ownership probability of the gesture control intention in the ownership probability matrix; S6: Based on the valid users and valid gesture instructions of the current screen projection control, a deterministic operation instruction set for the current screen projection collaborative control is generated to realize the collaborative control operation of the screen projection device.

2. The method for collaborative screen projection control based on gesture recognition according to claim 1, characterized in that: Step S1 is specifically as follows: Collect real-time video data of all users in the preset interactive space; Determine preliminary position coordinates of multiple user hand contour areas based on pixel color distribution information of the user's limb area in the real-time video data; Perform edge detection on the preliminary position coordinates of multiple user hand contour areas to obtain accurate spatial contour information of multiple user gesture candidate areas; According to the precise spatial contour information of multiple user gesture candidate areas, spatial feature quantization processing is performed on the multiple user gesture candidate areas respectively to obtain an initial user gesture spatial feature dataset.

3. The method for collaborative screen projection control based on gesture recognition according to claim 2, characterized in that: Step S2 is specifically as follows: Based on all the gesture candidate regions in the initial user gesture spatial feature dataset, the Euclidean distance between the center coordinates of any two gesture candidate regions and the overlap ratio of the contour pixels between the gesture candidate regions are calculated; According to the preset spatial proximity distance threshold and contour pixel overlap ratio threshold, a group of gesture candidate regions that meet the conditions of spatial proximity and contour overlap are determined and marked as the gesture spatial conflict feature set.

4. The method for collaborative screen projection control based on gesture recognition according to claim 3, characterized in that: Step S3 is specifically as follows: For each group of gesture candidate regions marked in the gesture spatial conflict feature set, the spatial coordinate position data in the real-time video data is extracted, and the similarity of the spatial motion trajectories of each group of gesture candidate regions is calculated; Based on the preset similarity threshold, the gesture candidate region groups with similarity higher than the similarity threshold are screened and marked as gesture validity pending datasets.

5. The method for collaborative screen projection control based on gesture recognition according to claim 4, characterized in that: Step S4 is specifically as follows: Extract the spatial location data of each gesture candidate region in the gesture validity pending dataset, combine it with the facial feature information of each user in the real-time video data, and establish the spatial correspondence between the gesture candidate region and the user identity; Based on the preset user historical control behavior records, the frequency of each user gaining control after completing a specific gesture action in the historical session is counted. Combined with the spatial correspondence, the control attribution probability values of the gesture candidate area for multiple users are calculated to construct the attribution probability matrix of the current gesture control intention.

6. The method for collaborative screen projection control based on gesture recognition according to claim 5, characterized in that: Step S5 is specifically as follows: Extract the belonging probability value of each user corresponding to all gesture candidate areas in the belonging probability matrix and record the maximum belonging probability value; Compare the recorded maximum probability value with the decision-making probability threshold; When the maximum attribution probability value is higher than or equal to the judgment decision probability threshold, the single user corresponding to the maximum attribution probability value is determined as the valid user of the current projection screen control, and the spatial feature data of the gesture candidate area corresponding to the maximum attribution probability value is determined as the valid gesture instruction of the current projection screen control according to the preset control instruction mapping rules.

7. The method for collaborative screen projection control based on gesture recognition according to claim 6, characterized in that: Step S6 is specifically as follows: Obtain user identity feature information of the valid user of the current screen projection control, and identify the screen projection control semantic information corresponding to the valid gesture instruction based on the spatial coordinate position data and shape contour data of the valid gesture instruction of the current screen projection control; According to the mapping relationship between the projection control semantic information and the deterministic operation instructions, the projection control semantic information is mapped and converted into the corresponding deterministic operation instructions; Based on the deterministic operation instructions, a deterministic operation instruction set for the current screen projection collaborative control is generated in accordance with the operation execution order associated with the deterministic operation instructions, and sent to the screen projection device to implement the collaborative control operation of the screen projection device.

Citation Information

Cited By

  • In-vehicle multi-screen linkage control method and system based on gesture recognition

    CN120821420A