A multi-source virtual reality human-computer interaction method and system
By analyzing the changing trends of key feature points in the historical video of user gesture actions, and predicting and adjusting the area of key feature points in the user gesture action image at the current moment, the problem of high-quality calculations in the existing VR system is solved, and high-precision gesture recognition and natural virtual environment control are achieved.
Patent Information
- Application Number
- CN202411355117.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-26
AI Technical Summary
When detecting key feature points in user gesture action videos, existing VR systems need to detect each frame of images, resulting in large amount of calculation and affecting system efficiency.
By analyzing the changing trend of key feature points in the historical video of user gesture actions, predict and adjust the area to which the key feature points belongs in the user gesture action image at the current moment, obtain the key feature prediction area, and use this area to identify the user's gesture actions.
It reduces the amount of computing on the computer, maintains the accuracy of key feature points, and realizes high-precision gesture recognition, allowing users to naturally control the virtual environment.
Smart Images

Figure CN119649442B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and particularly to a multi-source virtual reality human-computer interaction method and system. Background Art
[0002] In the context of the rapid development of virtual reality (VR) technology, multi-source virtual reality human-computer interaction (HCI) systems are becoming a research hotspot; existing VR systems usually rely on input sources such as gamepads or sensors to achieve user interaction with the virtual environment; however, image processing technology can provide a more intuitive and natural interaction experience by analyzing the user's gesture movements in real time, enabling the virtual environment to better respond to the user's dynamic behavior.
[0003] The prior art determines the user's gesture movements by extracting key feature points in the user's gesture movement video. To ensure the accuracy of the key feature points, it is necessary to detect the key feature points for each frame of the user's gesture movement video, which will greatly increase the computational load of the computer. Summary of the Invention
[0004] To solve the technical problem of how to reduce the computational load while ensuring the accuracy of key feature points, the purpose of the present invention is to provide a multi-source virtual reality human-computer interaction method and system, and the specific technical solutions adopted are as follows:
[0005] In a first aspect, a multi-source virtual reality human-computer interaction method, the method includes: obtaining user gesture movement images at a plurality of historical moments; analyzing the change trend of key feature points in the user gesture movement images at a plurality of historical moments, predicting the region to which the key feature points in the user gesture movement image at the current moment belong, and adjusting its area to obtain a key feature prediction region; identifying the user's gesture movement according to the feature points in the key feature prediction region, and mapping the identified user gesture movement into the virtual reality environment for interactive operations.
[0006] This method analyzes the change trend of key feature points in the historical video of the user's gesture movement, predicts and adjusts the region to which the key feature points in the user's gesture movement at the current moment belong. The user's gesture movement is identified by real-time obtaining the key feature points of the user's gesture movement, and the identified user gesture movement is mapped into the virtual reality environment for interactive operations.
[0007] Further, the method for obtaining user gesture action images at a plurality of historical moments is as follows: capturing the user's gestures by using a depth camera and generating a user gesture action video; decomposing the user gesture action video into a plurality of frames of user gesture action images; and forming a set of user gesture action images within a historical time period. After decomposing the user gesture action video into a plurality of frames of user gesture action images, it further includes: preprocessing the plurality of frames of user gesture action images, and the preprocessed plurality of frames of user gesture action images form a set of user gesture action images within a historical time period. The preprocessing of the plurality of frames of user gesture action images specifically includes: grayscaling each frame of user gesture action image and performing filtering by using a median filtering algorithm.
[0008] Further, the method for obtaining the key feature prediction region is as follows: obtaining the key feature points in the user gesture action image at each historical moment; analyzing the change trend of the key feature points within a historical time period to obtain the prediction accuracy of the prediction points of the key feature points; and determining all key feature prediction regions in the user gesture action image at the current moment according to the prediction accuracy of the prediction points.
[0009] Further, obtaining the key feature points in the user gesture action image at each historical moment specifically includes: for the user gesture action image at any historical moment, performing corner detection on the user gesture action image at this historical moment by using the SIFT corner detection algorithm to obtain all the corner points in the user gesture action image at this historical moment; for the t-th corner point in the user gesture action image at this historical moment, taking the t-th corner point as the window center, obtaining a window with a window size of 5*5, and denoting this window as the neighborhood range window of the t-th corner point; t is a positive integer; obtaining the possibility that the t-th corner point is a key feature point according to the relationship between the gray value of the t-th corner point and the gray values within the neighborhood range window of the t-th corner point; if the possibility that the t-th corner point is a key feature point is greater than or equal to the key threshold parameter, then the t-th corner point is denoted as the key feature point in the user gesture action image at the historical moment; traversing all historical moments within the historical time period to obtain all the key feature points in the user gesture action image at each historical moment.
[0010] Further, analyze the change trend of the key feature points within a historical time period to obtain the prediction accuracy of the prediction points of the key feature points, specifically including: determining the prediction point of the j-th key feature point in the user gesture action image at the N-th historical moment according to the coordinate position of the j-th key feature point in the user gesture action image at the (N - 1)-th historical moment, as well as the predicted change direction and predicted change distance of the j-th key feature point; determining the complexity of the change trend of the j-th key feature point according to the fitting error value of the fitting straight line of the j-th key feature point, the number of all clustering clusters in the clustering result of the j-th key feature point, and the absolute value of the difference between the number of data points in the m-th clustering cluster and the number of data points in the (m + 1)-th clustering cluster in the clustering result of the j-th key feature point; determining the prediction accuracy of the prediction point corresponding to the j-th key feature point according to the complexity of the change trend of the j-th key feature point, the Euclidean distance between the j-th key feature point and the prediction point corresponding to the j-th key feature point in the user gesture action image at the n-th historical moment, and the absolute value of the difference between the possibility that the j-th key feature point is a key feature point and the possibility that other corner points in the neighborhood corner point set of the j-th key feature point are key feature points; m, N, and j are all positive integers.
[0011] Further, according to the prediction accuracy of the prediction points, determine all key feature prediction regions in the user gesture action image at the current moment, specifically including: for the j-th key feature point, in the user gesture action image at the current moment, take the Euclidean distance between the j-th key feature point and the prediction point corresponding to the j-th key feature point in the user gesture action image at the N-th historical moment as the initial prediction window size of the j-th key feature point; round up the product of the prediction accuracy of the prediction point corresponding to the j-th key feature point and the initial prediction window size of the j-th key feature point to obtain the target prediction window size value U of the j-th key feature point. j ; Determine the prediction point of the j-th key feature point in the user gesture action image at the current moment according to the coordinate position of the j-th key feature point in the user gesture action image at the N-th historical moment, as well as the predicted change direction and predicted change distance of the j-th key feature point; take the prediction point of the j-th key feature point as the window center, and obtain a window with a size of U j ×U j and record this window as the key feature prediction region of the j-th key feature point; traverse all key feature points in the user gesture action image at the current moment to determine all key feature prediction regions in the user gesture action image at the current moment.
[0012] Further, the feature points in the key feature prediction region are used to identify the user's gesture actions, and the identified user gesture actions are mapped to the virtual reality environment for interactive operations, specifically including: for any key feature prediction region, the SIFT corner detection algorithm is used to detect the corners of the key feature prediction region, obtain the corners of the key feature prediction region, identify the user's gesture actions through the corners of all key feature prediction regions, and map the identified user gesture actions to the virtual reality environment for interactive operations.
[0013] In a second aspect, the present invention is a multi-source virtual reality human-computer interaction system, and the system includes: a data acquisition module: used to acquire user gesture action images at a plurality of historical moments; a key feature prediction region analysis module: used to analyze the change trend of key feature points in the user gesture action images at a plurality of historical moments, predict the region to which the key feature points in the user gesture action image at the current moment belong, and adjust its area to obtain a key feature prediction region; an interactive operation module: used to identify the user's gesture actions according to the feature points in the key feature prediction region, and map the identified user gesture actions to the virtual reality environment for interactive operations.
[0014] The present invention has the following beneficial effects:
[0015] By analyzing the change trend of key feature points in the historical video of the user's gesture actions, predicting and adjusting the region to which the key feature points in the user's gesture actions at the current moment belong, the present invention can not only greatly reduce the computing amount of the computer, but also maintain the accuracy of the key feature points, so as to achieve high-precision gesture recognition, enabling the user to control the interface in the virtual environment through natural hand movements. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of a multi-source virtual reality human-computer interaction method provided by an embodiment of the present invention;
[0018] Figure 2 It is a structural diagram of a multi-source virtual reality human-computer interaction system provided by an embodiment of the present invention. Detailed Embodiments
[0019] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a multi-source virtual reality human-computer interaction method and system proposed according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0021] The following specifically describes the specific solution of a multi-source virtual reality human-computer interaction method and system provided by the present invention in conjunction with the accompanying drawings.
[0022] Please refer to Figure 1 , which shows a flowchart of a multi-source virtual reality human-computer interaction method provided by an embodiment of the present invention.
[0023] Step S1: Collect user gesture action images at a plurality of historical moments.
[0024] In the present invention, there are N historical moments before the current moment, where n ∈ (1, N), and both n and N are positive integers. The nth moment represents any one of the plurality of historical moments, and the Nth moment represents the last one of the plurality of historical moments.
[0025] The specific method for obtaining user gesture action images at a plurality of historical moments is as follows:
[0026] 1. Precisely track the movement of the user's hand through the motion capture device in the multi-source virtual reality human-computer interaction system, and use a depth camera to capture the user's gestures and generate user gesture action videos.
[0027] 2. Use FFmpeg software to decompose the user gesture action video into a plurality of frames of user gesture action images.
[0028] 3. Since the collected images are RGB three-channel images and may contain a certain degree of noise, preprocessing is required. The preprocessing adopted in the present invention is as follows: First, grayscale each frame of user gesture action image to obtain the grayscale image corresponding to each user gesture action image, and then filter it using the median filter algorithm to obtain the filtered grayscale image; all the preprocessed plurality of frames of user gesture action images are recorded as user gesture action images at historical moments and form the set A of user gesture action images within the historical time period, specifically:
[0029] A = {A1, A2, …, A n , …, A N}
[0030] wherein, A n represents the user gesture action image at the nth historical moment; N represents the total number of user gesture action images at all historical moments in the set of user gesture action images within the historical time period.
[0031] Step S2: By analyzing the change trend of key feature points in the user gesture action images at several historical moments, predict the region to which the key feature points in the user gesture action image at the current moment belong, and perform area adjustment on it to obtain the key feature prediction region.
[0032] The process of obtaining the key feature prediction region is as follows:
[0033] a. Obtain the key feature points in the user gesture action image at each historical moment.
[0034] In an image, key feature points usually appear as prominent parts of the image, such as corner points, edges, or texture regions, which can help identify important structures or details in the image; furthermore, it also shows that key feature points are used to describe the detailed features of the image. For image details, they generally occur in regions of drastic changes. Therefore, by calculating the stability of the difference in gray values between a pixel point and other pixel points within the local region of a certain pixel point, it is determined whether the region to which the pixel point belongs is a detail region. Key feature points generally have obvious features for the neighborhood region, such as being local maxima or local minima, so that information is easy to retrieve and not easily lost during subsequent image capture.
[0035] The specific method for obtaining the key feature points in the user gesture action image at each historical moment is as follows:
[0036] 1. For the user gesture action image at any historical moment, use the SIFT corner detection algorithm to perform corner detection on the user gesture action image at this historical moment to obtain all the corner points in the user gesture action image at this historical moment.
[0037] 2. For the tth corner point in the user gesture action image at this historical moment, take the tth corner point as the center of the window, obtain a window with a size of 5 * 5, and denote this window as the neighborhood range window of the tth corner point.
[0038] 3. The calculation formula for the possibility that the tth corner point is a key feature point is:
[0039]
[0040] wherein, Pgt Indicates the possibility that the t-th corner point is a key feature point; K t Indicates the number of all pixel points in the neighborhood range window of the t-th corner point; h t Indicates the gray value of the t-th corner point; h t,i Indicates the gray value of the i-th pixel point in the neighborhood range window of the t-th corner point; σ t Indicates the variance of the gray values of all pixel points in the neighborhood range window of the t-th corner point; || represents taking the absolute value; norm() represents the linear normalization function; exp() represents the exponential function with the natural constant as the base; Indicates the sum of the absolute differences between the gray value of the t-th corner point and the gray values of all pixel points in the neighborhood range window of the t-th corner point; To ensure that the denominator is not zero, a tuning parameter ∈ is added to the denominator, ∈ = 0.01.
[0041] 4. Preset a key threshold parameter T1 = 0.75. If the possibility that the t-th corner point is a key feature point is greater than or equal to the key threshold parameter T1, then the t-th corner point is recorded as a key feature point in the user gesture action image at the historical moment; furthermore, all key feature points in the user gesture action image at each historical moment are obtained.
[0042] b. Analyze the change trend of key feature points in the historical time period to obtain the predicted points of key feature points and the corresponding prediction accuracies.
[0043] When the degree of change in the position distance and direction of key feature points in the historical time period is more different, it indicates that the amplitude of the user's changing gesture is greater. Therefore, the changes of key feature points in the historical time period can be clustered according to the amplitude of the user's changing gesture to obtain the key feature points at the last historical moment, and the predicted points corresponding to the key feature points are obtained therefrom; when the prediction accuracy of the predicted points corresponding to the key feature points is greater, the smaller the adjustment of the prediction area is required. Each key feature point corresponds to a prediction area, and the area of this prediction area is related to the prediction accuracy of the predicted points corresponding to the key feature points. The greater the accuracy, the smaller the area of the prediction area.
[0044] The calculation method for obtaining the predicted points of key feature points and the corresponding prediction accuracies is as follows:
[0045] 1. Since the more different the degree of change in the position distance and direction of key feature points in the historical time period, the greater the amplitude of the user's changing gesture, the greater the difficulty in prediction, and the greater the adjustment of the prediction area is required. Therefore, it is necessary to obtain the predicted points of key feature points and the complexity of the change trend. The specific obtaining method is as follows:
[0046] (1) For the j-th key feature point, obtain the position coordinates of the j-th key feature point in the user gesture action images at all historical moments, and use the least squares method for linear fitting to obtain the fitting error value of the fitting line of the j-th key feature point.
[0047] (2) Denote the sine value of the angle between the straight line connecting the coordinate positions of the j-th key feature point in the user gesture action image at the n-th historical moment and the coordinate position of the j-th key feature point in the user gesture action image at the n+1-th historical moment and the horizontal line as the change direction value of the j-th key feature point in the user gesture action image at the n-th historical moment; Denote the Euclidean distance between the coordinate position of the j-th key feature point in the user gesture action image at the n-th historical moment and the coordinate position of the j-th key feature point in the user gesture action image at the n+1-th historical moment as the change distance of the j-th key feature point in the user gesture action image at the n-th historical moment.
[0048] (3) Construct a two-dimensional space through the change distance and change direction value of the key feature point, input the change direction value and change distance of the j-th key feature point obtained from the user gesture action images at all historical moments into the two-dimensional space for clustering to obtain several clustering clusters; Obtain the clustering cluster to which the j-th key feature point in the user gesture action image at the (N-1)-th historical moment belongs, denoted as the target clustering cluster; Take the mean value of all change direction values in the target clustering cluster as the predicted change direction of the j-th key feature point; Take the mean value of all change distances in the target clustering cluster as the predicted change distance of the j-th key feature point; Determine the predicted point of the j-th key feature point in the user gesture action image at the N-th historical moment according to the coordinate position of the j-th key feature point in the user gesture action image at the (N-1)-th historical moment, and the predicted change direction and predicted change distance of the j-th key feature point.
[0049] (4) The calculation formula for obtaining the complexity of the change trend of the j-th key feature point is:
[0050]
[0051] In the formula, B j represents the complexity of the change trend of the j-th key feature point; RS j represents the fitting error value of the fitting line of the j-th key feature point; M j represents the number of all clustering clusters in the clustering result of the j-th key feature point; c j,m represents the number of data points in the m-th clustering cluster in the clustering result of the j-th key feature point; c j,m+1 represents the number of data points in the (m+1)-th clustering cluster in the clustering result of the j-th key feature point; || represents taking the absolute value.
[0052] 2. Since the greater the complexity of the change trend of the key feature points, the greater the prediction difficulty, and the lower the corresponding prediction accuracy; in the user gesture action image at the last historical moment, the smaller the distance between the key feature point and the corresponding prediction point of the key feature point, the greater the prediction accuracy of the corresponding prediction point; when the difference in the possibility of the key feature point among all the corner points in its neighborhood corner point set is greater, it indicates that the possibility of the corresponding prediction point of the key feature point being the key feature point is greater, and then the prediction accuracy is greater at this time. The calculation method for obtaining the prediction accuracy of the corresponding prediction point of the j-th key feature point is as follows:
[0053] (1) In the user gesture action image at the n-th historical moment, sort the other corner points in ascending order according to the Euclidean distance from the j-th key feature point, and take the first 10 corner points in the sorted corner point set as the neighborhood corner point set of the j-th key feature point.
[0054] (2) The calculation method for the prediction accuracy of the corresponding prediction point of the j-th key feature point is as follows:
[0055]
[0056] In the formula, Q j represents the prediction accuracy of the corresponding prediction point of the j-th key feature point; L j represents the number of all corner points in the neighborhood corner point set of the j-th key feature point; Pg j,l represents the possibility that the l-th corner point in the neighborhood corner point set of the j-th key feature point is the key feature point; Pg j represents the possibility that the j-th key feature point is the key feature point; B j represents the complexity of the change trend of the j-th key feature point; d N,j represents the Euclidean distance between the j-th key feature point and the corresponding prediction point of the j-th key feature point in the user gesture action image at the N-th historical moment; to ensure that the denominator is not zero, a tuning parameter ∈ is added to the denominator, ∈ = 0.01; norm() represents the normalization function.
[0057] c. Determine all the key feature prediction regions in the user gesture action image at the current moment according to the prediction accuracy of the prediction points.
[0058] When the prediction accuracy of the corresponding prediction point of the key feature point is greater, the area of the prediction region needs to be adjusted smaller.
[0059] The specific method for obtaining all the key feature prediction regions in the user gesture action image at the current moment is as follows:
[0060] (1) For the j-th key feature point, in the user gesture action image at the current moment, the Euclidean distance between the j-th key feature point and its corresponding predicted point in the user gesture action image at the N-th historical moment is used as the initial predicted window size R of the j-th key feature point. j .
[0061] (2) The specific formula for obtaining the target predicted window size value of the j-th key feature point is:
[0062]
[0063] In the formula, U j represents the target predicted window size value of the j-th key feature point; Q j represents the prediction accuracy of the predicted point corresponding to the j-th key feature point; R j represents the initial predicted window size of the j-th key feature point; represents rounding up.
[0064] (3) Through the above method, according to the coordinate position of the j-th key feature point in the user gesture action image at the N-th historical moment, as well as the predicted change direction and predicted change distance of the j-th key feature point, determine the predicted point of the j-th key feature point in the user gesture action image at the current moment; use the predicted point of the j-th key feature point as the window center, obtain a window with a size of U j ×U j , and record this window as the key feature prediction region of the j-th key feature point.
[0065] Step S3: Identify the user's gesture action based on the feature points in the key feature prediction region, and map the identified user gesture action to the virtual reality environment for interactive operations.
[0066] For any key feature prediction region, use the SIFT corner detection algorithm to detect the corners of this key feature prediction region, obtain the corners of this key feature prediction region, identify the user's gesture action through the corners of all key feature prediction regions, and map the identified user gesture action to the virtual reality environment for interactive operations.
[0067] Based on the same inventive concept as the above method embodiment, the embodiment of the present invention also provides a multi-source virtual reality human-computer interaction system, as Figure 2 shown, for implementing the above method embodiment, the system includes:
[0068] Data acquisition module: used to acquire user gesture action images at several historical moments.
[0069] Key feature prediction region analysis module: used to analyze the change trend of key feature points in user gesture action images at a number of historical moments, predict the region to which the key feature points in the user gesture action image at the current moment belong, and adjust its area to obtain the key feature prediction region.
[0070] Interaction operation module: used to identify the user's gesture action according to the feature points in the key feature prediction region, and map the identified user gesture action to the virtual reality environment for interaction operations.
[0071] It should be noted that: the above-mentioned sequence of embodiments of the present invention is only for description and does not represent the advantages or disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0072] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. The key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A multi-source virtual reality human-computer interaction method, characterized in that: The method comprises: Obtain user gesture action images at several historical moments; By analyzing the change trend of key feature points in the user gesture action image at several historical moments, predicting the area to which the key feature points in the user gesture action image at the current moment belong, and adjusting the area thereof, a key feature prediction area is obtained; Recognize the user's gestures based on the feature points in the key feature prediction area, and map the recognized user gestures to the virtual reality environment for interactive operations; The method to obtain the key feature prediction area is: Obtain key feature points in the user's gesture action image at each historical moment; Analyze the change trend of the key feature points in the historical time period to obtain the prediction accuracy of the prediction points of the key feature points; According to the prediction accuracy of the prediction point, all key feature prediction areas in the user's gesture action image at the current moment are determined; Obtain the key feature points in the user gesture action image at each historical moment, including: For a user gesture action image at any historical moment, use the SIFT corner point detection algorithm to perform corner point detection on the user gesture action image at the historical moment to obtain all corner points in the user gesture action image at the historical moment; For the tth corner point in the user gesture action image at the historical moment, take the tth corner point as the window center, obtain a window with a window size of 5*5, and record the window as the neighborhood range window of the tth corner point; t is a positive integer; According to the relationship between the gray value of the t-th corner point and the gray value in the neighborhood range window of the t-th corner point, the possibility of the t-th corner point being a key feature point is obtained; If the probability that the t-th corner point is a key feature point is greater than or equal to the key threshold parameter, the t-th corner point is recorded as the key feature point in the user gesture action image at the historical moment; Traverse all historical moments in the historical time period and obtain all key feature points in the user gesture action image at each historical moment; Analyze the change trend of the key feature points in the historical time period to obtain the prediction accuracy of the prediction points of the key feature points, specifically including: Determine the predicted point of the jth key feature point in the user gesture action image at the Nth historical moment according to the coordinate position of the jth key feature point in the user gesture action image at the N-1th historical moment, and the predicted change direction and predicted change distance of the jth key feature point; Determine the complexity of the change trend of the jth key feature point according to the fitting error value of the fitting straight line of the jth key feature point, the number of all clusters in the clustering result of the jth key feature point, the absolute value of the difference between the number of data points in the mth cluster in the clustering result of the jth key feature point and the number of data points in the m+1th cluster in the clustering result of the jth key feature point; Determine the prediction accuracy of the prediction point corresponding to the jth key feature point according to the complexity of the change trend of the jth key feature point, the Euclidean distance between the jth key feature point in the user gesture action image at the Nth historical moment and the prediction point corresponding to the jth key feature point in the user gesture action image at the Nth historical moment, and the absolute value of the difference between the possibility that the jth key feature point is a key feature point and the possibility that other corner points in the neighborhood corner point set of the jth key feature point are key feature points; m, N and j are all positive integers.
2. A multi-source virtual reality human-computer interaction method according to claim 1, characterized in that: Get user gesture action images at several historical moments. The specific steps are as follows: Using a depth camera to capture the user's gestures and generate a user gesture action video; Decompose the user gesture action video into a plurality of frames of user gesture action images; and form a set of user gesture action images within a historical time period.
3. The multi-source virtual reality human-computer interaction method according to claim 2, characterized in that: The user gesture action video is decomposed into a plurality of frames of user gesture action images, and then the method further includes: preprocessing the plurality of frames of user gesture action images, wherein the preprocessed plurality of frames of user gesture action images constitute a set of user gesture action images within a historical time period.
4. The multi-source virtual reality human-computer interaction method according to claim 3, characterized in that: Preprocessing the plurality of frames of user gesture action images specifically includes: Each frame of user gesture action image is grayed out and filtered using a median filtering algorithm.
5. The multi-source virtual reality human-computer interaction method according to claim 1, characterized in that: According to the prediction accuracy of the prediction point, all key feature prediction areas in the user gesture action image at the current moment are determined, including: For the jth key feature point, in the user gesture action image at the current moment, the Euclidean distance between the jth key feature point and the corresponding prediction point of the jth key feature point in the user gesture action image at the Nth historical moment is used as the initial prediction window size of the jth key feature point; The product of the prediction accuracy of the prediction point corresponding to the j-th key feature point and the initial prediction window size of the j-th key feature point is rounded up to obtain the target prediction window size value U of the j-th key feature point. j ; According to the coordinate position of the jth key feature point in the user gesture action image at the Nth historical moment, as well as the predicted change direction and predicted change distance of the jth key feature point, determine the predicted point of the jth key feature point in the user gesture action image at the current moment; take the predicted point of the jth key feature point as the window center, and obtain the window size U j ×U j The window is recorded as the key feature prediction area of the jth key feature point; All key feature points in the user gesture action image at the current moment are traversed to determine all key feature prediction areas in the user gesture action image at the current moment.
6. The multi-source virtual reality human-computer interaction method according to claim 1, characterized in that: The user's gestures are identified based on the feature points in the key feature prediction area, and the identified user gestures are mapped to the virtual reality environment for interactive operations, including: For any key feature prediction area, the SIFT corner detection algorithm is used to perform corner detection on the key feature prediction area to obtain the corner points of the key feature prediction area. The user's gestures are identified through the corner points of all key feature prediction areas, and the identified user gestures are mapped to the virtual reality environment for interactive operations.
7. A multi-source virtual reality human-computer interaction system, characterized in that: The system comprises: Data acquisition module: used to obtain user gesture action images at several historical moments; Key feature prediction region analysis module: used to analyze the change trend of key feature points in the user gesture action image at several historical moments, predict the region to which the key feature points in the user gesture action image at the current moment belong, and adjust the area to obtain the key feature prediction region; Interactive operation module: used to identify the user's gestures based on the feature points in the key feature prediction area, and map the identified user gestures to the virtual reality environment for interactive operations; The method to obtain the key feature prediction area is: Obtain key feature points in the user's gesture action image at each historical moment; Analyze the change trend of the key feature points in the historical time period to obtain the prediction accuracy of the prediction points of the key feature points; According to the prediction accuracy of the prediction point, all key feature prediction areas in the user's gesture action image at the current moment are determined; Obtain the key feature points in the user gesture action image at each historical moment, including: For a user gesture action image at any historical moment, use the SIFT corner point detection algorithm to perform corner point detection on the user gesture action image at the historical moment to obtain all corner points in the user gesture action image at the historical moment; For the tth corner point in the user gesture action image at the historical moment, take the tth corner point as the window center, obtain a window with a window size of 5*5, and record the window as the neighborhood range window of the tth corner point; t is a positive integer; According to the relationship between the gray value of the t-th corner point and the gray value in the neighborhood range window of the t-th corner point, the possibility of the t-th corner point being a key feature point is obtained; If the probability that the t-th corner point is a key feature point is greater than or equal to the key threshold parameter, the t-th corner point is recorded as the key feature point in the user gesture action image at the historical moment; Traverse all historical moments in the historical time period and obtain all key feature points in the user gesture action image at each historical moment; Analyze the change trend of the key feature points in the historical time period to obtain the prediction accuracy of the prediction points of the key feature points, specifically including: Determine the predicted point of the jth key feature point in the user gesture action image at the Nth historical moment according to the coordinate position of the jth key feature point in the user gesture action image at the N-1th historical moment, and the predicted change direction and predicted change distance of the jth key feature point; Determine the complexity of the change trend of the jth key feature point according to the fitting error value of the fitting straight line of the jth key feature point, the number of all clusters in the clustering result of the jth key feature point, the absolute value of the difference between the number of data points in the mth cluster in the clustering result of the jth key feature point and the number of data points in the m+1th cluster in the clustering result of the jth key feature point; Determine the prediction accuracy of the prediction point corresponding to the jth key feature point according to the complexity of the change trend of the jth key feature point, the Euclidean distance between the jth key feature point in the user gesture action image at the Nth historical moment and the prediction point corresponding to the jth key feature point in the user gesture action image at the Nth historical moment, and the absolute value of the difference between the possibility that the jth key feature point is a key feature point and the possibility that other corner points in the neighborhood corner point set of the jth key feature point are key feature points; m, N and j are all positive integers.
Citation Information
Patent Citations
Posture recognition method for VR electronic toy
CN117994856A
Devices and methods for single or multi-user gesture detection using computer vision
US20230419733A1