PPT control method based on human posture estimation
Through the PPT control method based on human posture estimation, the camera is used to collect image data for target detection and posture estimation, and PPT control instructions are generated. This solves the space limitations and equipment dependence problems of traditional PPT control methods and realizes contactless and multifunctional PPT operation.
Patent Information
- Application Number
- CN202310859550.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-07-13
AI Technical Summary
The existing PPT control method relies on the mouse, keyboard and laser pen, which leads to problems such as user space limitations and device loss and damage, making it difficult to achieve contactless, convenient and fast multi-functional control.
A method based on human posture estimation is adopted. Through the client and server architecture, the camera is used to collect image data, perform target detection, tracking and posture estimation, generate PPT control instructions, and realize contactless operation.
It realizes contactless PPT control without carrying additional equipment. It can complete functions such as page turning, laser pointer movement, page zooming and volume adjustment through simple human movements. It is stable to use and has high accuracy.
Smart Images

Figure CN117032448B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human posture estimation in computer vision, and in particular to a PPT control method based on human posture estimation. Background Art
[0002] The use of PPT is becoming more and more widespread, and PPT is used in many occasions such as teaching, office work, conference reports, etc. The user needs to perform a series of operations on PPT, such as switching PPT pages, zooming in and out of PPT pages, and moving the laser pointer. Since the mouse, keyboard, and laser pen are essentially extensions of computer peripherals, many unfavorable factors will arise. For example, when using the keyboard and mouse to control PPT, the user needs to stand next to the computer, which affects communication with the audience. Although the laser pen can free the user from spatial limitations, the user may forget to bring it, forget to charge it, or it may be damaged during use. Taking into account the many inconveniences brought about by the above traditional ways of controlling PPT, in order to achieve a more efficient way of using PPT, the present invention applies the human posture estimation technology in computer vision to PPT control.
[0003] In the field of computer vision, human posture estimation is an extremely important research direction. The task goal is to locate and identify the key points of the human body. These key point information are often connected in the order of joints. The key point information can be used to further estimate the human posture, and then send corresponding instructions to the computer through the corresponding human posture information to achieve contactless PPT control operations.
[0004] Therefore, the present invention discloses a contactless PPT control method, which can get rid of the spatial limitations of mouse and keyboard control of PPT, avoid many disadvantages brought by traditional hardware laser pointers, and realize a contactless and multifunctional PPT control method. Summary of the Invention
[0005] (1) Technical issues to be resolved
[0006] The technical problems to be solved by the present invention are:
[0007] When using PPT, users need to control PPT through traditional computer hardware extension devices such as keyboard, mouse, laser pen remote control, etc., and it is difficult to achieve a convenient, fast, contactless, and multi-functional control method that does not require carrying any additional equipment.
[0008] (2) Technical solution
[0009] To solve the above problems, the present invention discloses a PPT control method based on human posture estimation. The method of the present invention can avoid the various disadvantages of traditional control methods and provide a convenient, fast, efficient, contactless and multifunctional PPT control method, which is characterized by comprising the following steps:
[0010] Step S1: Collect clear human body image data through the client data acquisition module, perform image encoding, and send the encoding result to the server;
[0011] Furthermore, the step S1 specifically includes:
[0012] Step S1.1: This method adopts a client / server architecture, where the client's data acquisition module collects human image data in the environment through traditional cameras such as laptop cameras and classroom surveillance cameras;
[0013] Step S1.2: The client data acquisition module encodes the captured human body image data;
[0014] Step S1.3: The client data acquisition module sends the encoded human body image data to the server.
[0015] Step S2: The server-side posture calculation module performs target detection, target tracking, and posture estimation on the human image data, and outputs the target prediction result;
[0016] Furthermore, the step S2 specifically includes:
[0017] Step S2.1: The server-side posture calculation module receives the client request and transmits the received human body image to the target detection unit for processing;
[0018] Step S2.2: The target detection unit uses a target detection algorithm to detect the user object in the scene and send the detection results to the target tracking unit and the pose estimation unit. The algorithm uses a pre-collected and annotated human body dataset for model training.
[0019] Step S2.3: The target tracking unit uses a target tracking algorithm to complete target tracking of the user object in the scene and achieve global single tracking of the user;
[0020] Step S2.4: The posture estimation unit uses a human posture estimation algorithm to estimate the human posture of the input frame. The algorithm model is trained using the COCO human key point dataset.
[0021] Step S2.5: The target prediction result is transferred to the server feedback and control module for action classification and instruction generation.
[0022] Step S3: The server-side feedback and control module classifies and identifies the prediction results of the server-side posture calculation module, generates corresponding PPT control instructions, and sends the instructions to the client. The client controls the PPT program to perform corresponding operations according to the instructions.
[0023] (3) Beneficial effects
[0024] When using PPT, users do not need to carry any additional hardware equipment or stand in a fixed position. They can use simple human movements to achieve PPT functions such as turning pages, moving the laser pointer, zooming and moving pages, adjusting video volume, etc. The use process is stable and the accuracy is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a schematic diagram of the overall module of the method of the present invention
[0026] Figure 2 This is a module workflow diagram of the method of the present invention
[0027] Figure 3 Schematic diagram of the left and right virtual frames of the method of the present invention
[0028] Figure 4 This is a schematic diagram of the switching PPT posture of the method of the present invention
[0029] Figure 5 This is a schematic diagram of the zoomed PPT page of the method of the present invention
[0030] Figure 6 Schematic diagram of controlling the volume of PPT media according to the method of the present invention DETAILED DESCRIPTION
[0031] The following embodiments are used to illustrate the present invention. The present invention is based on a PPT control method for human posture estimation, which comprises the following steps:
[0032] The overall module schematic diagram of the present invention is as follows Figure 1 As shown, the user needs to install the software on the computer and collect human posture image information through the data acquisition module of the client; if the computer cannot access the server, the posture solution module built into the program is called to perform human posture solution on the collected image locally on the computer; if the server can be accessed, the data acquisition module collects the image data and sends it to the server for more accurate and efficient human posture solution, and returns the corresponding instructions; the control and feedback module of the client executes the corresponding control instructions. The specific module workflow diagram is shown in the figure below. Figure 2 As shown;
[0033] Step S1: The client data acquisition module collects human body image data, performs image encoding, and sends the encoding result to the server;
[0034] Furthermore, the step S1 specifically includes:
[0035] Step S1.1: This method adopts a client / server architecture, where the client's data acquisition module collects human image data in the environment through traditional cameras such as laptop cameras and classroom surveillance cameras;
[0036] Step S1.2: The client data acquisition module encodes the collected human body image data;
[0037] Step S1.3: The client data acquisition module sends the encoded human body image data to the server.
[0038] Step S2: The server-side posture calculation module performs target detection, target tracking, and posture estimation on the human image data, thereby outputting the human key point prediction results;
[0039] Furthermore, the step S2 specifically includes:
[0040] Step S2.1: The server-side posture calculation module receives the client request and transmits the received human body image to the target detection unit and the target tracking unit for processing;
[0041] Step S2.2: The target detection unit uses a target detection algorithm to detect the user object in the scene and input the detection results into the target tracking unit and the pose estimation unit. The algorithm uses a pre-collected and annotated human body dataset for model training;
[0042] Step S2.3: The target tracking unit uses a target tracking algorithm to globally track the user object in the scene. In real-world scenarios, there are often many human objects, which can significantly interfere with user experience. Therefore, the present invention not only locates the target user but also tracks them, ensuring that the program is always operated by the designated user during execution.
[0043] Specifically, in step S2.3, the target tracking algorithm predicts the trajectory of the previous and next frames through the Kalman filter, continuously searches for the augmented path through the Hungarian algorithm, and finds the maximum match to achieve the tracking and matching of the target result;
[0044] Step S2.4: The posture estimation unit uses a human posture estimation algorithm to estimate the human posture of the user object in the scene, and the algorithm uses the COCO human key point dataset to pre-train the posture estimation model; the final result of the human posture estimation in the posture estimation unit is represented by a prediction vector of the human posture, where the prediction vector is defined as:
[0045]
[0046] Where p is the prediction vector, is the x-axis coordinate of the nth key point in the two-dimensional plane, is the y-axis coordinate of the nth key point in the two-dimensional plane, is the confidence of the nth key point, n = 1, 2, …, 17.
[0047] Step S3: The server-side feedback and control module classifies and identifies the human body key point prediction results, generates corresponding PPT control instructions, and sends the instructions to the client. The client controls the PPT program to perform corresponding operations according to the instructions.
[0048] Step S3.1: Filter the key points of the prediction result p of the posture solution module and filter out the confidence The key point of the key point and All are set to 0 to prevent the accuracy of the prediction results from affecting user usage;
[0049] Step S3.2: In order to achieve more functions through human body posture and avoid crosstalk between commands generated by different human body postures, the present invention uses a virtual frame strategy to solve this problem;
[0050] Specifically, the virtual frame in step S3.2 is as follows: Figure 3 As shown, it is divided into two virtual frames, left and right. The virtual frame is based on the left shoulder key point. To the right shoulder key point Euclidean distance For the length, For width;
[0051] Step S3.3: by classifying the predicted result p of the human body posture, different PPT control instructions are issued for different human body postures;
[0052] Specifically, the different postures and control instructions in step S3.3 include:
[0053] like Figure 4As shown, this diagram is a diagram of the posture of switching PPT pages. The left half of the figure is to issue a forward page turning instruction by waving the arm to the left, and the right half of the figure is to issue a backward page turning instruction by waving the arm to the right;
[0054] By moving the right arm wrist keypoint Place the laser pointer in the right virtual frame to control the movement of the laser pointer. The user's right wrist key points The position in the right virtual frame will be projected proportionally to the corresponding position on the PPT page;
[0055] like Figure 5 As shown, this diagram is a schematic diagram of zooming in on the PPT page, through the left arm wrist key points and right arm wrist key points Euclidean distance in the left virtual box Multiply by the corresponding zoom factor to control the PPT page zoom;
[0056] like Figure 6 As shown, this diagram is a diagram for controlling the volume of PPT media, through the left arm wrist key points and right arm wrist key points Euclidean distance in the right virtual box Multiply by the corresponding zoom factor to control the volume of PPT media.
Claims
1. The PPT control method based on human posture estimation is characterized by: Use human posture estimation technology to analyze the user's human posture, generate corresponding PPT control instructions based on the analysis results, and realize the control of PPT, including the following steps: Step S1: The client data acquisition module collects human body image data, performs image encoding, and sends the encoding result to the server; Step S2: The server-side posture calculation module performs target detection, target tracking, and posture estimation on the human image data, thereby outputting the human key point prediction results; Step S3: The control module classifies and identifies the predicted results of the human body key points, generates corresponding PPT control instructions, and sends the PPT control instructions to the client. The client controls the PPT program to perform corresponding operations according to the PPT control instructions. The step S2 specifically includes: Step S2.1: The server-side posture calculation module receives the client request and transmits the received human body image to the target detection unit and the target tracking unit for processing; Step S2.2: The target detection unit uses a target detection algorithm to complete target detection of the user object in the scene and input the detection results into the target tracking unit and the posture estimation unit; Step S2.3: The target tracking unit uses a target tracking algorithm to complete target tracking of the user object in the scene, thereby achieving global tracking of the user; Step S2.4: The posture estimation unit uses a human posture estimation algorithm to estimate the human posture of the user object in the scene, and the human posture estimation algorithm uses the COCO human key point dataset to pre-train the posture estimation model; the final result of the human posture estimation in the posture estimation unit is represented by a prediction vector of the human posture, where the prediction vector is defined as: in is the prediction vector, It is The key points in the two-dimensional plane Axis coordinates, It is The key points in the two-dimensional plane Axis coordinates, It is The confidence of the key points, ; The step S3 specifically includes: Step S3.1: Prediction vector of attitude solution module Perform key point filtering and filter out confidence The key point of the key point and All are set to 0 to prevent the accuracy of the prediction results from affecting user usage; Step S3.2: In order to achieve more functions through human body posture and avoid crosstalk between commands generated by different human body postures, a virtual frame strategy is used to solve the problem of crosstalk between commands generated by different human body postures. Specifically, the virtual frame in step S3.2 is divided into two virtual frames, left and right, with the virtual frame having twice the left shoulder key points. To the right shoulder key point Euclidean distance For the length, The virtual frame ratio conforms to the standard PPT aspect ratio; Step S3.3: Predicting the human body posture vector Classify and issue different PPT control instructions for different human postures; Specifically, the different postures and control instructions in step S3.3 include: Swing your arm to the left to turn the page forward, and to the right to turn the page backward. By moving the right arm wrist keypoint Place the laser pointer in the right virtual frame to control the movement of the laser pointer. The user's right wrist key points The position in the right virtual frame will be projected proportionally to the corresponding position on the PPT page; Pass the left arm wrist key point and right arm wrist key points Euclidean distance in the left virtual box Multiply by the corresponding zoom factor to control the PPT page zoom; Pass the left arm wrist key point and right arm wrist key points Euclidean distance in the right virtual box Multiply by the corresponding zoom factor to control the volume of PPT media.
Citation Information
Patent Citations
Posture recognition method and device based on artificial intelligence, terminal and storage medium
CN111931701A
Integrated gesture control PPT system and method based on cloud edge cooperation technology, and platform
CN113918009A