Action recognition method, head-mounted display device, and storage medium

By identifying the feature information and limb key points of human image data, calculating affinity distance and angle, and combining the generation of an adversarial network model, the problem of low accuracy in motion recognition of head-mounted display devices is solved, and the accuracy of user motion recognition is improved.

CN115393962BActive Publication Date: 2025-07-08GEER TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211049678.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-08
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The head-mounted display device has low accuracy when recognizing user actions.

Method used

By determining the characteristic information of human image data, identifying the limb key points, and calculating the affinity distance and angle between the limb key points, the generated adversarial network model matches the action, and improving the accuracy of action recognition.

Benefits of technology

Improves the accuracy of motion recognition of head-mounted display devices during user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393962B_ABST
    Figure CN115393962B_ABST
Patent Text Reader

Abstract

The present invention discloses an action recognition method, a head-mounted display device, and a storage medium. The method includes: determining feature information corresponding to the collected human body image data, and identifying limb key points corresponding to the human body image data according to the feature information; determining the affinity distance and limb angle between the limb key points; and identifying an action that matches the limb key points, the affinity distance, and the limb angle as a control action corresponding to the human body image data, which solves the problem of low recognition accuracy of user actions by the head-mounted display device, and improves the recognition accuracy of user actions through the technical solution of the present application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality technology, and in particular, to a method for action recognition, a head-mounted display device, and a storage medium. Background Art

[0002] In recent years, due to its light weight and convenient portability, head-mounted display devices have been used by more and more users. A head-mounted display device has a graphical interface that can interact with a three-dimensional environment, and it has become an important medium for human-computer interaction. In the process of human-computer interaction, the process of recognizing user actions is particularly important. However, when capturing and recognizing user actions through a head-mounted display device, the accuracy is relatively low. Summary of the Invention

[0003] Embodiments of the present application aim to solve the problem of low accuracy in recognizing user actions by a head-mounted display device by providing a method for action recognition, a head-mounted display device, and a storage medium.

[0004] The present application provides a method for action recognition, which includes:

[0005] Determine the feature information corresponding to the collected human body image data, and identify the limb key points corresponding to the human body image data according to the feature information;

[0006] Determine the affinity distance and limb angle between the limb key points;

[0007] Identify the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

[0008] Optionally, the step of determining the affinity distance and limb angle between the limb key points includes:

[0009] Obtain the position information of the first limb key point, the position information of the second limb key point, and the position information of the third limb key point;

[0010] Determine the first affinity distance between the first limb key point and the second limb key point according to the position information of the first limb key point and the position information of the second limb key point;

[0011] Determine the second affinity distance between the second limb key point and the third limb key point according to the position information of the second limb key point and the position information of the third limb key point;

[0012] Determine the radian value between the first limb key point and the third limb key point according to the first affinity distance and the second affinity distance;

[0013] Determine the limb angle according to the radian value.

[0014] Optionally, the step of identifying the action that matches the limb key point, the affinity distance, and the limb angle as the control action corresponding to the human body image data includes:

[0015] Input the limb key point, the affinity distance, the limb angle, and the standard limb action with labels into a generative adversarial network model;

[0016] Use the discriminator of the generative adversarial network model to determine the similarity between the limb action corresponding to the limb key point, the affinity distance, and the limb angle and the standard limb action with labels;

[0017] When the similarity reaches a preset threshold, determine the action that matches the limb key point, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

[0018] Optionally, after the step of using the discriminator to determine the similarity between the limb action corresponding to the limb key point, the affinity distance, and the limb angle and the standard limb action with labels, it further includes:

[0019] When the similarity does not reach the preset threshold, return to execute the step of inputting the limb key point, the affinity distance, the limb angle, and the standard limb action with labels into the discriminator of the generative adversarial network model.

[0020] Optionally, the step of determining the feature information corresponding to the acquired human body image data includes:

[0021] Segment the currently acquired human body image data to obtain target image data;

[0022] Input the target image data into a first neural network model, and obtain the feature information according to the output result of each layer of the first neural network model. The target image data sequentially passes through the input layer, pooling layer, convolutional layer, fully connected layer, and softmax layer of the first neural network model.

[0023] Optionally, the feature information includes finger feature information and arm feature information; the step of identifying the limb key points corresponding to the human body image data according to the feature information includes:

[0024] Input the finger feature information and the arm feature information into a second neural network model, and identify the finger key points corresponding to the finger feature information and the arm key points corresponding to the arm feature information;

[0025] Generate the limb key points based on the positions of the finger key points and the positions of the arm key points.

[0026] Optionally, after the step of identifying the limb key points corresponding to the human body image data according to the feature information, the method further includes:

[0027] Obtain the positions of the limb end key points;

[0028] Determine the positions of other key points except the limb end key points based on inverse kinematics and the positions of the limb end key points;

[0029] Use the positions of the other key points to correct the corresponding limb key points, obtain the corrected limb key points, and perform determining the affinity distances and limb angles between the corrected limb key points.

[0030] Optionally, after the step of identifying the control action corresponding to the human body image data by matching the limb key points, the affinity distances, and the limb angles, the method further includes:

[0031] Generate an operation signal corresponding to the control action;

[0032] Respond to the operation corresponding to the operation signal on the interaction interface of the head-mounted display device.

[0033] Optionally, before the step of determining the feature information corresponding to the collected human body image data and identifying the limb key points corresponding to the human body image data according to the feature information, the method further includes:

[0034] Control a hand-held camera to start to collect lower limb image data through the hand-held camera, and the hand-held camera is communicatively connected to the head-mounted display device.

[0035] Optionally, the step of determining the feature information corresponding to the collected human body image data and identifying the limb key points corresponding to the human body image data according to the feature information includes:

[0036] Determine the feature information corresponding to the upper limb image data collected by an external camera of the head-mounted display device and the feature information corresponding to the lower limb image data collected by the hand-held camera;

[0037] Identify the limb key points corresponding to the human body image data according to the feature information corresponding to the upper limb image data and the feature information corresponding to the lower limb image data.

[0038] In addition, to achieve the above object, the present invention further provides a head-mounted display device, which includes a storage unit, a control unit, and an action recognition program stored on the storage unit and executable on the control unit. When the action recognition program is executed by the control unit, the steps of the above-mentioned action recognition method are implemented.

[0039] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which an action recognition program is stored. When the action recognition program is executed by a control unit, the steps of the above-mentioned action recognition method are implemented.

[0040] The technical solutions of an action recognition method, a head-mounted display device, and a storage medium provided by this application first determine the feature information corresponding to the collected human body image data, and then identify the limb key points in the human body image data according to the feature information. Then, after obtaining the limb key points, determine the affinity distance between the limb key points and determine the limb angle according to the limb key points. Finally, identify the action that matches the limb key points, affinity distance, and limb angle as the action corresponding to the human body image data. Since the collected human body image data is subjected to feature recognition and limb key points are extracted, and then the affinity distance and limb angle are used to constrain the human body key points, the problem of low accuracy of action recognition of the head-mounted display device for users is solved, and the recognition accuracy of user actions is improved through the technical solutions proposed in this application. Description of the Drawings

[0041] Figure 1 It is a schematic structural diagram of the head-mounted display device related to the embodiment solution of the present invention;

[0042] Figure 2 It is a schematic flowchart of the first embodiment of the action recognition method of the present invention;

[0043] Figure 3 It is a schematic diagram of the limb key points of the present invention.

[0044] The realization of the object of this application, functional features, and advantages will be further described in conjunction with the embodiments with reference to the drawings. The above drawings are only drawings of one embodiment and not all of the invention. Detailed Embodiments

[0045] This application aims to solve the problem of low accuracy in recognizing user actions by a head-mounted display device. This application proposes an action recognition method. The action recognition method first determines the feature information corresponding to the collected human body image data, and then identifies the limb key points in the human body image data based on the feature information. Next, after obtaining the limb key points, it determines the affinity distance between the limb key points and determines the limb angles based on the limb key points. Finally, it identifies the action that matches the limb key points, affinity distance, and limb angles as the action corresponding to the human body image data. By performing feature recognition on the collected human body image data, extracting limb key points, and using the affinity distance and limb angles to constrain the human body key points, the accuracy of user action recognition by the head-mounted display device is improved during the interaction with the user.

[0046] To better understand the above technical solution, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0047] As Figure 1 shown, Figure 1 is a schematic structural diagram of the hardware operating environment of the head-mounted display device involved in the embodiment of the present invention. Some embodiments of the present invention provide a head-mounted display device that can be an external head-mounted display device or an integrated head-mounted display device, where the external head-mounted display device needs to be used in cooperation with an external processing system (such as a computer processing system). Optionally, the head-mounted display device can also be a virtual reality head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, etc. This application takes a virtual reality head-mounted device as an example.

[0048] Figure 1 shows a schematic internal configuration diagram of the head-mounted display device 500 in some embodiments. The display unit 501 may include a display panel, and the display panel is disposed inside the head-mounted display device 500 and can be a single panel or composed of multiple small panels separately arranged. The display panel can be an electroluminescent (EL) element, a liquid crystal display, or a micro display with a similar structure, or a retina-direct display or a similar laser scanning display.

[0049] The virtual image optical unit 502 captures the image displayed on the display unit 501 in a magnified manner and allows the user to view the displayed image as a magnified virtual image. As the display image output to the display unit 501, it can be an image of a virtual scene provided by a content reproduction device (Blu-ray Disc or DVD player) or a streaming media server, or an image of a real scene captured using an external camera 510.

[0050] In some embodiments, the virtual image optical unit 502 may include a lens unit, such as a spherical lens, an aspherical lens, a Fresnel lens, etc. The input operation unit 503 includes at least one operation component for performing input operations, such as a key, a button, a switch, or other components with similar functions, receives user instructions through the operation component, and outputs the instructions to the control unit 507.

[0051] The status information acquisition unit 504 can be used to acquire the status information of the user of the wearable head-mounted display device 500. The status information acquisition unit 504 may include various types of sensors for detecting status information by itself, and can also acquire status information from external devices (such as smartphones, wristwatches, and other multifunctional terminals worn by the user) through the communication unit 505. The status information acquisition unit 504 can also acquire the position information and / or attitude information of the user's head. The status information acquisition unit 504 may include one or more of a gyroscope sensor, an acceleration sensor, a Global Positioning System (GPS) sensor, a geomagnetic sensor, a Doppler effect sensor, an infrared sensor, and a radio frequency field strength sensor. In addition, the status information acquisition unit 504 acquires the status information of the user of the wearable head-mounted display device 500, such as acquiring the operation status of the user (whether the user is wearing the head-mounted display device 500), the action status of the user (such as stationary, walking, running, and other moving states, the posture of the hand or fingertips, the open or closed state of the eyes, the line of sight direction, the pupil size, the limb movement), the mental state (whether the user is immersed in observing the displayed image, etc.), and even the physiological state, etc.

[0052] The communication unit 505 performs communication processing, modulation and demodulation processing, and encoding and decoding processing of communication signals with external devices. In addition, the control unit 507 can send transmission data to external devices from the communication unit 505. The communication method can be in a wired or wireless form, such as Mobile High-Definition Link (MHL) or Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Wi-Fi, Bluetooth communication or Low Energy Bluetooth communication, and a mesh network of the IEEE 802.11s standard, etc. In addition, the communication unit 505 can be a cellular radio transceiver operating according to Wideband Code Division Multiple Access (W-CDMA), Long-Term Evolution (LTE), and similar standards.

[0053] In some embodiments, the head-mounted display device 500 may further include a storage unit. The storage unit 506 is a large-capacity storage device configured to have a solid-state drive (SSD) or the like. In some embodiments, the storage unit 506 may store an action recognition program or various types of data. For example, the content viewed by the user using the head-mounted display device 100 may be stored in the storage unit 506.

[0054] The image processing unit 508 is used to perform signal processing, such as image quality correction related to the image signal output from the control unit 507, and convert its resolution to the resolution according to the screen of the display unit 501. Then, the display driving unit 509 sequentially selects each row of pixels of the display unit 501 and scans each row of pixels of the display unit 501 row by row, thereby providing a pixel signal based on the signal-processed image signal.

[0055] In some embodiments, the head-mounted display device 500 may further include an external camera. The external camera 510 may be disposed on the front surface of the main body of the head-mounted display device 500, and the external camera 510 may be one or more. The external camera 510 can acquire three-dimensional information and can also be used as a distance sensor. In addition, a position-sensitive detector (PSD) that detects the reflected signal from an object or other types of distance sensors can be used together with the external camera 510. The external camera 510 and the distance sensor can be used to detect the body position, posture, and shape of the user wearing the head-mounted display device 500. In addition, under certain conditions, the user can directly view or preview the real scene through the external camera 510. Optionally, the external camera 510 may also be a handheld camera disposed in the user's scene and communicatively connected to the head-mounted display device 500. In this scenario, the handheld camera can be used to acquire lower limb images.

[0056] In some embodiments, the head-mounted display device 500 may further include a sound processing unit. The sound processing unit 511 can perform sound quality correction or sound amplification of the sound signal output from the control unit 507, as well as signal processing of the input sound signal, etc. Then, the sound input / output unit 512 outputs the sound to the outside after sound processing and inputs the sound from the microphone.

[0057] It should be noted that Figure 1 The structures or components shown in the dashed boxes may be independent of the head-mounted display device 500. For example, they may be disposed in an external processing system (such as a computer system) and used in cooperation with the head-mounted display device 500; or, the structures or components shown in the dashed boxes may be disposed inside or on the surface of the head-mounted display device 500.

[0058] Those skilled in the art can understand that Figure 1The head-mounted display device structure shown does not constitute a limitation on the head-mounted display device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0059] In Figure 1 In the head-mounted display device shown, the control unit 507 can be used to call the action recognition program stored in the storage unit 506. In this embodiment, the head-mounted display device includes: a storage unit 506, a control unit 507, and an action recognition program stored on the storage unit 506 and executable on the control unit 507, where:

[0060] When the control unit 507 calls the action recognition program stored in the storage unit 506, it performs the following operations:

[0061] Determine the feature information corresponding to the collected human body image data, and identify the limb key points corresponding to the human body image data according to the feature information;

[0062] Determine the affinity distance and limb angle between the limb key points;

[0063] Identify the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

[0064] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0065] Obtain the position information of the first limb key point, the position information of the second limb key point, and the position information of the third limb key point;

[0066] Determine the first affinity distance between the first limb key point and the second limb key point according to the position information of the first limb key point and the position information of the second limb key point;

[0067] Determine the second affinity distance between the second limb key point and the third limb key point according to the position information of the second limb key point and the position information of the third limb key point;

[0068] Determine the radian value between the first limb key point and the third limb key point according to the first affinity distance and the second affinity distance;

[0069] Determine the limb angle according to the radian value.

[0070] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0071] Input the limb key points, the affinity distance, the limb angles, and the standard limb actions with labels into a generative adversarial network model;

[0072] Use the discriminator of the generative adversarial network model to determine the similarity between the limb actions corresponding to the limb key points, the affinity distance, and the limb angles and the standard limb actions with labels;

[0073] When the similarity reaches a preset threshold, determine the action matching the limb key points, the affinity distance, and the limb angles as the control action corresponding to the human body image data.

[0074] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0075] When the similarity does not reach the preset threshold, return to execute the step of inputting the limb key points, the affinity distance, the limb angles, and the standard limb actions with labels into the discriminator of the generative adversarial network model.

[0076] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0077] Segment the currently acquired human body image data to obtain target image data;

[0078] Input the target image data into a first neural network model, and obtain the feature information according to the output results of each layer of the first neural network model. The target image data sequentially passes through the input layer, pooling layer, convolutional layer, fully connected layer, and softmax layer of the first neural network model.

[0079] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0080] Input the finger feature information and the arm feature information into a second neural network model to identify the finger key points corresponding to the finger feature information and the arm key points corresponding to the arm feature information;

[0081] Generate the limb key points based on the positions of the finger key points and the positions of the arm key points.

[0082] When the control unit 507 calls the action recognition program stored in the storage unit 506, it also performs the following operations:

[0083] Obtain the positions of the key points at the ends of the limbs;

[0084] Determine the positions of other key points except the key points at the end of the limb based on inverse kinematics and the positions of the key points at the end of the limb;

[0085] Use the positions of the other key points to correct the corresponding limb key points, obtain the corrected limb key points, and perform determining the affinity distances and limb angles between the corrected limb key points.

[0086] When the control unit 507 calls the action recognition program stored in the storage unit 506, the following operations are also performed:

[0087] Generate an operation signal corresponding to the control action;

[0088] Respond to the operation corresponding to the operation signal on the interaction interface of the head-mounted display device.

[0089] When the control unit 507 calls the action recognition program stored in the storage unit 506, the following operations are also performed:

[0090] Control the hand-held camera to start to collect lower limb image data through the hand-held camera, and the hand-held camera is communicatively connected to the head-mounted display device.

[0091] When the control unit 507 calls the action recognition program stored in the storage unit 506, the following operations are also performed:

[0092] Determine the feature information corresponding to the upper limb image data collected by the external camera of the head-mounted display device and the feature information corresponding to the lower limb image data collected by the hand-held camera;

[0093] Identify the limb key points corresponding to the human body image data according to the feature information corresponding to the upper limb image data and the feature information corresponding to the lower limb image data.

[0094] The technical solution of the present application will be described below by way of embodiments.

[0095] As Figure 2 shown, in the first embodiment of the present application, the action recognition method of the present application includes the following steps:

[0096] Step S110, determine the feature information corresponding to the collected human body image data, and identify the limb key points corresponding to the human body image data according to the feature information.

[0097] In this embodiment, the action recognition method of the present application is applied to a head-mounted display device, namely a VR all-in-one machine. Optionally, the action recognition method can also be applied to other terminal devices. Human body image data can be obtained through an external camera. The external camera can be arranged on the front surface of the head-mounted display device body, and there can be one or more such external cameras. The external camera can also be an image acquisition device arranged in the user's scene and communicatively connected to the head-mounted display device. In this scenario, the external camera can be used to collect human body image data.

[0098] Optionally, after the external camera collects the human body image data, the human body image data is sent to the head-mounted display device. Optionally, the human body image data collected by the external camera within a preset time period can be sent to the head-mounted display device at regular intervals. It is also possible to send the human body image data collected by the external camera in real time to the head-mounted display device. Optionally, the human body image data can be human body image frame data obtained by video decoding.

[0099] Optionally, after the human body image data is sent to the head-mounted display device, the control unit of the head-mounted display device can process the human body image data to determine a control action. Specifically, the feature information corresponding to the collected human body image data can be determined, and the limb key points corresponding to the human body image data are identified based on the feature information. Among them, the feature information is also image feature information. Image feature information refers to a set of a series of attributes that can characterize the characteristics or content of an image, mainly including image natural features such as brightness, color, texture, etc., and image artificial features such as image spectrum, image histogram, etc. Image feature information mainly includes color features, texture features, shape features, and spatial relationship features of the image. Image feature extraction can be divided into two categories: global feature extraction and local feature extraction according to its relative scale. Global feature extraction focuses on the overall representation of the image. Common global features include color features, texture features, shape features, spatial position relationship features, etc. Local feature extraction focuses on the special properties of a certain local area of the image. An image often contains several regions of interest, and several local features can be extracted from these regions in varying numbers.

[0100] Optionally, the present application obtains feature information by inputting the human body image data into a first neural network model. The first neural network model is used to convert the human body image data into feature information. The first neural network model can be a VGG19 network model, and the first neural network model can also be other models with image feature conversion functions. Among them, the first neural network model includes an input layer, a pooling layer, a convolutional layer, a fully connected layer, and a softmax layer. When the human body image data passes through different layers of the first neural network model, it will be converted into different feature information.

[0101] In some application scenarios, an all-in-one VR device can capture and recognize the user's limb movements, and use these limb movements as an index for the interaction interface in a three-dimensional environment. The user can use different movements to perform corresponding operations on the graphical interface. In this process, to improve the accuracy of movement recognition. After collecting the human body image data, it is necessary to segment the currently obtained human body image data to delete some unnecessary information, such as deleting the environmental information, to obtain the target image data including the target image data. Then, the target image data is input into the first neural network model. Among them, the target image data will sequentially pass through the input layer, pooling layer, convolutional layer, fully connected layer, and softmax layer of the first neural network model, and obtain feature information through the output results of each layer. Optionally, the target image data can be limb image data, such as upper limb image data, or head image data, or data of other parts of the user's body.

[0102] In this embodiment, after obtaining the feature information, the limb key points corresponding to the human body image data are recognized according to the feature information. Optionally, the feature information can be input into the second neural network model to obtain the limb key points. Among them, the fingers and arms in the upper limb image data can be respectively transformed to obtain finger feature information and arm feature information. After obtaining the finger feature information and arm feature information, it is necessary to input the finger feature information and arm feature information into the second neural network model to recognize and extract the finger key points corresponding to the finger feature information, and recognize and extract the arm key points corresponding to the arm feature information. After obtaining the finger key points and arm key points, the finger key points and arm key points can be sorted according to the positions of the finger key points and arm key points to obtain the limb key points with a connection sequence.

[0103] Optionally, the above-mentioned second neural network model can be a CNN or RNN neural network model. The finger feature information and arm feature information can be respectively input into the second neural network model to recognize and extract the finger feature information and arm feature information. Among them, the process of using the second neural network model to recognize and extract the feature information to obtain the limb key points belongs to conventional technical means and will not be elaborated here.

[0104] Optionally, after obtaining the limb key points, inverse kinematics technology can also be used to correct the limb key points, so as to obtain the corrected limb key points. Among them, inverse kinematics means that given the position and orientation of the end of a limb, the positions of the corresponding key points of the robot are calculated. The methods that can be used to solve the positions of the key points by inverse kinematics include, but are not limited to: analytical method, numerical method. Among them, the numerical method includes, but is not limited to: Jacobian inverse matrix method, Newton method, numerical drive method, hybrid method, biomechanical constraint, etc. Optionally, the positions of the key points at the end of the limb can be obtained, and based on the inverse kinematics and the positions of the key points at the end of the limb, the positions of other key points except the key points at the end are determined, and then the positions of other key points are used to correct the corresponding limb key points, so as to obtain the corrected limb key points. After obtaining the corrected limb key points, the affinity distance and limb angle of the corrected limb key points can be determined. By determining the positions of each key point through inverse kinematics and correcting the originally determined limb key points, the finally obtained limb key points are made more accurate.

[0105] Optionally, the positions of other key points can also be determined according to the inverse kinematics technology, and the correction coefficients of other key points are determined, and then the originally determined limb key points are corrected by using the correction coefficients, so as to obtain the corrected limb key points, and then the affinity distance and limb angle of the corrected limb key points are determined.

[0106] Step S120, determine the affinity distance and limb angle between the limb key points.

[0107] In this embodiment, after identifying the limb key points corresponding to the human body image data according to the feature information or after determining the corrected limb key points, for the marked limb key points, since the data collected by the external camera is a two-dimensional image, it is necessary to determine its specific position and behavior through feature constraints, including the affinity distance and angle of the key points. Therefore, the affinity distance and limb angle between the limb key points can be further determined. Among them, the number of the limb key points includes multiple, and the positions of each limb key point can be marked. The affinity distance is also the Euclidean distance, and the Euclidean distance refers to the distance between two points in space. Optionally, the affinity distance between two adjacent limb key points can be calculated, or the affinity distance between any two limb key points can be calculated. For example, there are three limb key points, which are the first limb key point located at the finger ( Figure 3 the right hand first in Figure 3 ), the second limb key point located at the elbow ( Figure 3If it is the right shoulder among them, the affinity distance between the first limb key point and the second limb key point can be calculated, and the affinity distance between the first limb key point and the third limb key point can also be calculated. Wherein, the limb angle is the limb radian, which can be the included angle between two line segments formed by the limb key points. For example, the included angle formed by the first line segment between the first limb key point and the second limb key point and the second line segment between the third limb key point and the second limb key point, that is, the included angle at the right elbow. It can also be when the limb has a radian, and the corresponding limb angle can be calculated for it.

[0108] Optionally, determining the affinity distance and limb angle between limb key points can be specifically: obtaining the position information of the first limb key point, the position information of the second limb key point, and the position information of the third limb key point, determining the first affinity distance between the first limb key point and the second limb key point according to the position information of the first limb key point and the position information of the second limb key point, and determining the second affinity distance between the second limb key point and the third limb key point according to the position information of the second limb key point and the position information of the third limb key point. Among them, the affinity distance of the key point can be calculated by the following Euclidean distance. For example, the Euclidean distance between the arm key point a and the wrist key point b is:

[0109]

[0110] Wherein, xa and ya respectively represent the abscissa and ordinate of the arm key point a, and xb and yb respectively represent the abscissa and ordinate of the wrist key point b.

[0111] After calculating the Euclidean distance between limb key points in the above manner, in order to judge the relative position movement of limb behaviors, an angle constraint is introduced. Therefore, the radian value between the first limb key point and the third limb key point can be further determined according to the first affinity distance and the second affinity distance; wherein, the radian value can be calculated by the following formula:

[0112]

[0113] Wherein, a, b, and c are respectively calculated according to the first affinity distance and the second affinity distance determined according to the positions of the first limb key point, the second limb key point, and the third limb key point in the space coordinate system.

[0114] After obtaining the radian value, the radian value can be further converted into a limb angle according to the conversion relationship between the radian and the angle.

[0115] Step S130, identify the action that matches the limb key point, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

[0116] In this embodiment, after obtaining the limb key points, the affinity distances between the limb key points, and the limb angles, the limb key points, the affinity distances, and the limb angles can be input into a generative adversarial network to identify the control actions corresponding to the human body image data. Optionally, the set of limb key points, the affinity distances between the key points, and the angle values are input into the generative adversarial network model. The attributes are input into the generator G for predicting limb control actions, and are input into the discriminator D together with the standard limb actions with labels for confrontation. After achieving the confrontation effect, that is, reaching the threshold or the result tending to be stable, the generator G can effectively identify the user's limb control actions.

[0117] Optionally, identifying the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data specifically includes the following steps:

[0118] Step S131, input the limb key points, the affinity distance, the limb angle, and the standard limb actions with labels into the generative adversarial network model;

[0119] Step S132, use the discriminator of the generative adversarial network model to determine the similarity between the limb action corresponding to the limb key points, the affinity distance, and the limb angle and the standard limb actions with labels.

[0120] Step S133, when the similarity reaches a preset threshold, determine the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

[0121] Among them, the standard limb actions with labels are standard actions, and each limb key point, limb angle, and the affinity distance between the key points in this standard action can be labeled. Then, compare each limb key point, limb angle, and the affinity distance between the key points labeled by this standard action with the limb key points, limb angles, and affinity distances input into the discriminator of the generative adversarial network model to observe the corresponding similarities. Optionally, the preset threshold can be set according to the actual situation. When its similarity reaches the preset threshold, the discriminant result can be output as "true". At this time, determine the standard action that matches the limb key points, affinity distance, and limb angle as the control action corresponding to the human body image data.

[0122] Optionally, when the similarity does not reach the preset threshold, the discrimination result can be output as "false". At this time, continue to predict the limb control action according to the generator of the generative adversarial network model. Input the limb key points, affinity distance, limb angle, and standard limb actions with labels into the generative adversarial network model for discrimination until the discrimination result is "true", then end the discrimination, and determine the matching standard action as the control action corresponding to the human body image data, so as to effectively recognize the user's limb actions.

[0123] Optionally, identifying the action that matches the limb key points, affinity distance, and limb angle as the control action corresponding to the human body image data can also be to determine multiple standard limb actions according to the limb angle. Determine the standard affinity distance corresponding to this affinity distance among the respective standard limb actions, calculate the difference between this affinity distance and each standard affinity distance, and select the target standard affinity distance corresponding to the smallest difference. Determine the target standard limb action corresponding to the target standard affinity distance with the smallest difference as the control action corresponding to the human body image data, thereby improving the recognition accuracy of the control action.

[0124] Optionally, after determining the control action, corresponding control can be performed. This control action can be used as the index of the interaction interface in the three-dimensional environment, and the user can use different control actions to perform corresponding operations on the graphical interface. Optionally, an operation signal corresponding to this control action can be generated, and the operation corresponding to this operation signal can be responded to on the interaction interface of the head-mounted display device. For example, waving the arm can perform a page-turning operation, and "making a victory sign" can be used to confirm, etc., thereby improving the usage experience of the VR all-in-one machine.

[0125] According to the above technical solution of this embodiment, first, the feature information corresponding to the collected human body image data is determined, and then the limb key points in the human body image data are recognized according to the feature information; then, after obtaining the limb key points, the affinity distance between the limb key points is determined, and the limb angle is determined according to the limb key points; finally, the action that matches the limb key points, affinity distance, and limb angle is recognized as the action corresponding to the human body image data. Since the feature recognition of the collected human body image data and the extraction of limb key points are performed, and the affinity distance and limb angle are used to constrain the human body key points, the accuracy of user action recognition is improved during the interaction between the head-mounted display device and the user.

[0126] Optionally, when it is necessary to recognize lower limb movements, if only the external camera of the head-mounted display device is turned on for recognition, since the external camera is generally worn on the user's head, it may cause inaccurate recognition of lower limb movements. Therefore, when it is necessary to recognize lower limb movements or overall movements, the handheld camera can be controlled to turn on to recognize lower limb movements. Among them, the handheld camera is communicatively connected to the head-mounted display device, and the data collected by the mobile phone camera can be transmitted to the head-mounted display device or the terminal device for processing. For example, when jumping on a dance mat, the external camera of the handheld camera and the head-mounted display device can be turned on simultaneously to collect limb image data, and then the limb movements of the user can be determined according to the limb image data, improving the recognition accuracy of the user's overall limb movements.

[0127] Optionally, after controlling the handheld camera to start, the external camera of the head-mounted display device is used to collect upper limb image data, and the handheld camera is used to collect lower limb image data. After obtaining the upper limb image data and the lower limb image data, the feature information corresponding to the upper limb image data collected by the external camera of the head-mounted display device and the feature information corresponding to the lower limb image data collected by the handheld camera can be determined. Then, since there may be overlapping image data in the collected upper limb image data and lower limb image data, the overlapping image data can be determined based on the feature information corresponding to the upper limb image data and the feature information corresponding to the lower limb image data, the overlapping image data is filtered, and then the upper limb image data and the lower limb image data after filtering the overlapping image data are fused, and the limb key points corresponding to the human body image data are recognized according to the fused upper limb image data and lower limb image data, and then the overall movement is determined according to the limb key points corresponding to the human body image data.

[0128] According to the above technical solution of this embodiment, when it is necessary to recognize lower limb movements or overall movements, the handheld camera and the external camera can be used in combination. The lower limb image data is obtained through the handheld camera, and the upper limb image data is obtained through the external camera, so as to obtain the user's overall movement, making the action recognition accuracy more accurate.

[0129] The embodiment of the present invention provides an embodiment of the action recognition method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0130] Based on the same inventive concept, the embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores an action recognition program. When the action recognition program is executed by a processor, it realizes the above-mentioned various steps of action recognition and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0131] Since the computer-readable storage medium provided by the embodiments of the present application is the computer-readable storage medium used to implement the methods of the embodiments of the present application, based on the methods introduced in the embodiments of the present application, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium used in the methods of the embodiments of the present application falls within the scope of protection of the present application.

[0132] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage units, CD-ROMs, optical storage units, etc.) that contain computer-usable program code.

[0133] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the control unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the control unit of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0134] These computer program instructions can also be stored in a computer-readable storage unit that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable storage unit generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0136] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

[0137] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0138] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for action recognition, characterized in that, The described action recognition method includes: Determine the feature information corresponding to the collected human body image data, and recognize the limb key points corresponding to the human body image data according to the feature information; Determine the affinity distance and limb angle between the limb key points, including: obtaining the position information of the first limb key point, the position information of the second limb key point, and the position information of the third limb key point; determining the first affinity distance between the first limb key point and the second limb key point according to the position information of the first limb key point and the position information of the second limb key point; determining the second affinity distance between the second limb key point and the third limb key point according to the position information of the second limb key point and the position information of the third limb key point; determining the radian value between the first limb key point and the third limb key point according to the first affinity distance and the second affinity distance; determining the limb angle according to the radian value; Identify the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

2. The action recognition method according to claim 1, wherein The step of identifying the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data includes: Input the limb key points, the affinity distance, the limb angle, and the standard limb action with labels into the generative adversarial network model; Use the discriminator of the generative adversarial network model to determine the similarity between the limb action corresponding to the limb key points, the affinity distance, and the limb angle and the standard limb action with labels; When the similarity reaches the preset threshold, determine the action that matches the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data.

3. The action recognition method according to claim 2, wherein After the step of using the discriminator to determine the similarity between the limb action corresponding to the limb key points, the affinity distance, and the limb angle and the standard limb action with labels, it further includes: When the similarity does not reach the preset threshold, return to execute the step of inputting the limb key points, the affinity distance, the limb angle, and the standard limb action with labels into the discriminator of the generative adversarial network model.

4. The action recognition method according to claim 1 or 3, characterized in that The feature information includes finger feature information and arm feature information; The step of recognizing the limb key points corresponding to the human body image data according to the feature information includes: Input the finger feature information and the arm feature information into the second neural network model to recognize the finger key points corresponding to the finger feature information and the arm key points corresponding to the arm feature information; Generate the limb key points based on the positions of the finger key points and the arm key points.

5. The action recognition method according to claim 1, characterized in that, After the step of recognizing the limb key points corresponding to the human body image data according to the feature information, it further includes: Obtain the position of the limb end key point; Determine the positions of other key points except the limb end key point based on inverse kinematics and the position of the limb end key point; Use the positions of the other key points to correct the corresponding limb key points, obtain the corrected limb key points, and perform the determination of the affinity distance and limb angle between the corrected limb key points.

6. The action recognition method according to claim 1, wherein After the step of identifying the action matching the limb key points, the affinity distance, and the limb angle as the control action corresponding to the human body image data, the method further includes: Generate an operation signal corresponding to the control action; Respond to the operation corresponding to the operation signal on the interaction interface of the head-mounted display device.

7. The action recognition method according to claim 1, characterized in that Before the step of determining the feature information corresponding to the collected human body image data and identifying the limb key points corresponding to the human body image data according to the feature information, the method further includes: Control the handheld camera to start to collect lower limb image data through the handheld camera, and the handheld camera is communicatively connected to the head-mounted display device.

8. The action recognition method according to claim 7, wherein The step of determining the feature information corresponding to the collected human body image data and identifying the limb key points corresponding to the human body image data according to the feature information includes: Determine the feature information corresponding to the upper limb image data collected by the external camera of the head-mounted display device and the feature information corresponding to the lower limb image data collected by the handheld camera; Identify the limb key points corresponding to the human body image data according to the feature information corresponding to the upper limb image data and the feature information corresponding to the lower limb image data.

9. A head-mounted display device, characterized in that, The head-mounted display device includes: a storage unit, a control unit, and an action recognition program stored on the storage unit and executable on the control unit. When the action recognition program is executed by the control unit, the steps of the action recognition method according to any one of claims 1-8 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an action recognition program. When the action recognition program is executed by a control unit, the steps of the action recognition method according to any one of claims 1-8 are implemented.

Citation Information

Patent Citations

  • Human-computer interaction method and system based on human body action recognition

    CN114721509A