A response mapping method and device for an AR glasses collaborative gesture operation terminal
Through visual algorithms and deep learning technology, the spatial position of fingers in AR glasses is calculated and displayed in real time, which solves the problem that users cannot see the position of fingers when operating in AR glasses, and improves the accuracy and convenience of the operation.
Patent Information
- Application Number
- CN202411494545.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-24
AI Technical Summary
When wearing existing AR glasses, users cannot see the position of their fingers in real time, resulting in operation errors and inconvenience.
The vision-based terminal-gestor positioning tracking algorithm is adopted, combined with deep learning technology and visual geometry algorithms, and the spatial position and posture of mobile terminals and gestures are calculated in real time, and the spatial position of gestures is projected to the mobile terminal plane through the projection algorithm to realize the visual display of finger positions.
It solves the problem that users cannot see the finger position in real time when operating in AR glasses, improves the accuracy and convenience of the operation, and provides a better augmented reality experience.
Smart Images

Figure CN119472991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of augmented reality, and in particular to a response mapping method and device for an AR glasses to cooperate with a gesture operation terminal. Background Art
[0002] Augmented Reality (AR) technology is a technology that combines the real world with virtual information based on real-time computer calculations and multi-sensor fusion. This technology simulates and re-outputs human sensations such as vision, hearing, smell, and touch, and superimposes virtual information on real information to provide users with an experience that transcends the real world.
[0003] As a popular entertainment interaction medium, AR provides smooth and open audio-visual and gaming experiences for more and more users. Currently, the most popular method is that users wear AR glasses, and the screen of the mobile phone terminal is displayed inside the AR glasses through a data cable. Users view videos or game interfaces through the glasses, but the control operations still remain on the mobile phone terminal side. In this scenario, since users can only see the screen projected from the mobile phone through the AR glasses and cannot see the real-world scene, they cannot see the real positions of the mobile phone terminal and their hands. When controlling video playback or performing game operations, it will lead to the phenomenon of "blind operation", which is prone to misoperation and affects the real experience.
[0004] In response to this, researchers designed a processing device that can collect finger images and the screen of the mobile phone terminal. However, it was further found that during the process of projecting the spatial pose of the gesture onto the plane of the mobile phone terminal through a projection algorithm to obtain the position of the hand relative to the mobile phone terminal, it is impossible to track the real-time coordinate position of the thumb. At the same time, the determination of the pose of the projected thumb is relatively rigid, and the position of the thumb above the mobile phone screen will cause deviations in the projected coordinates of the thumb due to the size of the mobile phone screen during the projection calculation. Summary of the Invention
[0005] To address this pain point, the present invention proposes a vision-based terminal-gesture positioning and tracking algorithm and device. Through deep learning technology and visual geometry algorithms, the spatial positions and poses of the mobile phone terminal and the gesture are calculated in real time, and then the spatial pose of the gesture is projected onto the plane of the mobile phone terminal through a projection algorithm to obtain the position of the hand relative to the mobile phone terminal. These position information and the mobile phone screen image are uploaded to the AR device, so that the AR device can simultaneously display the position of the user's finger while displaying the mobile phone screen, providing visual reference and assistance for the user's operation.
[0006] The purpose of the present invention is to provide a response mapping method and device for an AR glasses to cooperate with a gesture operation terminal, which solves the above-mentioned technical problems pointed out in the prior art.
[0007] The present invention provides a response mapping method for an AR glasses to cooperate with a gesture operation terminal, including the following operation steps: obtaining a binocular view mobile phone image containing a complete mobile phone contour, obtaining the mobile phone corner coordinates after corner detection of the collected binocular view mobile phone image, and then obtaining the mobile phone corner spatial coordinates by means of triangulation and plane fitting using the mobile phone corner coordinates; the binocular view mobile phone image includes a left-eye view mobile phone image and a right-eye view mobile phone image; obtaining a thumb image of a binocular view containing a thumb picture, obtaining the thumb corner coordinates after corner detection of the collected thumb image, and then processing the thumb corner coordinates through triangulation, Kalman filter algorithm and ICP algorithm in sequence to obtain the thumb corner spatial coordinates and the thumb posture; the thumb image includes a left-eye view thumb image and a right-eye view thumb image; calculating the direction size ratio of the thumb projected onto the mobile phone screen based on the thumb corner spatial coordinates and the mobile phone corner spatial coordinates; the direction size ratio includes an X-direction size ratio r_X and a Y-direction size ratio r_Y; transmitting the binocular view mobile phone image to the AR device to obtain an AR display image; fusing the thumb image and the AR display image based on the direction size ratio and the thumb posture to obtain a fused image.
[0008] Preferably, the operation steps of obtaining the mobile phone corner coordinates after corner detection of the collected binocular view mobile phone image, and then obtaining the mobile phone corner spatial coordinates by means of triangulation and plane fitting using the mobile phone corner coordinates include the following operation steps: performing corner detection on the binocular view mobile phone image through a pre-trained key point detection network model to obtain the mobile phone corner coordinates corresponding to the binocular view mobile phone image; calculating the three-dimensional coordinates corresponding to the mobile phone corner coordinates by means of a triangulation algorithm using the internal and external parameters of the camera; performing plane fitting on the three-dimensional coordinates to obtain the mobile phone corner spatial coordinates.
[0009] Preferably, the pre-trained key point detection network model is trained based on the YOLO model; the obtaining of the pre-trained key point detection network model includes the following operation steps:
[0010] Obtaining a plurality of sample mobile phone contour images by means of web crawler method; performing preprocessing operations on the sample mobile phone contour images to obtain preprocessed sample mobile phone contour images; the preprocessing operations include picture size adjustment operations; performing annotation on the preprocessed sample mobile phone contour images to obtain a sample data set; establishing an initial key point detection network model; training the initial key point detection network model based on the sample data set, and outputting a trained key point detection network model after the initial key point detection network model converges.
[0011] Preferably, for the planar fitting of the three-dimensional coordinates to obtain the spatial coordinates of the mobile phone corner points, the general expression of the prior plane is first determined through the average point coordinates of each three-dimensional coordinate, and then a parameter-free prior plane is constructed through the average point coordinates. Next, the plane equation parameters of the general expression of the prior plane are solved, and finally, the general expression of the target prior plane is obtained. Furthermore, the three-dimensional coordinates are projected onto the target prior plane to obtain the spatial coordinates of the mobile phone corner points. The specific operation steps are as follows:
[0012] Calculate the average point coordinate \(P_m(x_m,y_m,z_m)\) based on each of the three-dimensional coordinates;
[0013] Construct a prior plane based on the average point coordinate \(P_m(x_m,y_m,z_m)\);
[0014] Construct an error matrix \(H\) based on the error terms of each of the three-dimensional coordinates to the prior plane;
[0015] The error matrix \(H\) is:
[0016]
[0017] In the formula, \((x1,y1,z1)\), \((x2,y2,z2)\), \((x3,y3,z3)\), and \((x4,y4,z4)\) respectively represent the three-dimensional coordinates corresponding to the mobile phone corner point coordinates; \((x m ,y m ,z m ) represents the average coordinate of the four three-dimensional coordinates;
[0018] Perform a singular value decomposition operation on the error matrix \(H\) to obtain a decomposition matrix \(H'\);
[0019] The decomposition matrix \(H'\) is:
[0020] \(H' = UDV\ T ;
[0021] In the formula, \(U\) is the left singular vector matrix, \(D\) is the singular value diagonal matrix, and \(V T is the right singular vector matrix;
[0022] Obtain the last column eigenvector of the right singular vector matrix \(V T as the plane parameters \(A\), plane parameter \(B\), and plane parameter \(C\);
[0023] Calculate the plane parameter \(D\) according to the three-dimensional coordinates and the prior plane;
[0024] Obtain the target prior plane according to the plane parameters \(A\), plane parameter \(B\), plane parameter \(C\), and the plane parameter \(D\);
[0025] The general expression of the target prior plane is as follows:
[0026] Ax_m + By_m + Cz_m + D = 0;
[0027] Wherein, A, B, C, and D are respectively plane equation parameters;
[0028] By vertically projecting each of the three-dimensional coordinates onto the target prior plane, the spatial coordinates of the mobile phone corner points corresponding to each of the three-dimensional coordinates are obtained.
[0029] Preferably, the method for processing the thumb corner point coordinates respectively through triangulation, Kalman filtering algorithm, and ICP algorithm to obtain the spatial coordinates and posture of the thumb includes the following operation steps: obtaining the internal and external parameters of the camera; calculating the initial three-dimensional coordinates of the thumb corner points through triangulation according to the internal and external parameters of the camera and the thumb corner point coordinates; obtaining the frame thumb corner point coordinates of consecutive frames; calculating the three-dimensional coordinates of the frame thumb corner points corresponding to each frame through triangulation; performing prediction and correction operations on the three-dimensional coordinates of the frame thumb corner points through the Kalman filtering algorithm to obtain the corrected three-dimensional coordinates of the frame thumb corner points; calculating the thumb postures corresponding to the corrected three-dimensional coordinates of each frame thumb corner point through the ICP algorithm based on the initial three-dimensional coordinates of the thumb corner points and the corrected three-dimensional coordinates of the frame thumb corner points;
[0030] The thumb posture includes a rotation matrix R and a translation vector T.
[0031] Preferably, the method for calculating the initial three-dimensional coordinates of the thumb corner points through triangulation according to the internal and external parameters of the camera and the thumb corner point coordinates includes the following operation steps:
[0032] Obtaining N frames of first thumb corner point coordinates of the thumb corner point coordinates in a static state on the mobile phone screen surface;
[0033] Calculating the first three-dimensional coordinates of the thumb corner points through triangulation according to the internal and external parameters of the camera for the first thumb corner point coordinates; calculating the average coordinates of the first three-dimensional coordinates of the thumb corner points corresponding to each of the thumb corner point coordinates; determining the average coordinates as the initial three-dimensional coordinates of the thumb corner points.
[0034] Preferably, the method for performing prediction and correction operations on the three-dimensional coordinates of the frame thumb corner points through the Kalman filtering algorithm to obtain the corrected three-dimensional coordinates of the frame thumb corner points includes the following operation steps:
[0035] Establishing a motion model for each of the three-dimensional coordinates of the frame thumb corner points;
[0036] Traverse the three-dimensional coordinates of each frame thumb corner point, and predict based on the three-dimensional coordinates of the frame thumb corner point corresponding to the thumb corner point coordinates in consecutive frames and the motion model to obtain the predicted three-dimensional coordinates corresponding to the three-dimensional coordinates of the frame thumb corner point in the t-th frame;
[0037] Correct the predicted three-dimensional coordinates based on the three-dimensional coordinates of the frame thumb corner point in the (t - 1)-th frame to obtain the corrected three-dimensional coordinates of the thumb corner point;
[0038] Repeat the above operations until all the three-dimensional coordinates of the frame thumb corner points are traversed.
[0039] Preferably, calculating the thumb posture corresponding to each corrected three-dimensional coordinate of the frame thumb corner point through the ICP algorithm based on the initial three-dimensional coordinates of the thumb corner point and the corrected three-dimensional coordinates of the frame thumb corner point includes the following operation steps:
[0040] Calculate the initial centroid coordinates of each of the initial three-dimensional coordinates of the thumb corner point and the frame centroid coordinates of each of the corrected three-dimensional coordinates of the frame thumb corner point;
[0041] Calculate the initial relative centroid coordinates based on the initial centroid coordinates and each of the initial three-dimensional coordinates of the thumb corner point; and calculate the frame relative centroid coordinates based on the corrected three-dimensional coordinates of the thumb corner point and the frame centroid coordinates;
[0042] The calculation method of the initial relative centroid coordinates is:
[0043] Initial relative centroid coordinates = Initial three-dimensional coordinates of the thumb corner point - Initial centroid coordinates;
[0044] The calculation method of the frame relative centroid coordinates is:
[0045] Frame relative centroid coordinates = Corrected three-dimensional coordinates of the thumb corner point - Frame centroid coordinates;
[0046] Perform a normalization operation on each of the initial relative centroid coordinates to obtain an initial normalized point set; perform a normalization operation on each of the frame relative centroid coordinates to obtain a frame normalized point set;
[0047] Calculate the thumb posture (R, T) corresponding to each corrected three-dimensional coordinate of the frame thumb corner point through the least squares method based on the initial normalized point set and the frame normalized point set;
[0048] The calculation method of the thumb posture (R, T) is:
[0049]
[0050] Wherein, R is a rotation matrix; T is a translation vector; argmin is the input value corresponding to the minimum value of the function; sp_i is the i-th initial relative centroid coordinate in the initial normalized point set; dp_i is the i-th frame relative centroid coordinate in the frame normalized point set.
[0051] Preferably, calculating the direction dimension ratio of the thumb projected onto the mobile phone screen based on the spatial coordinates of the thumb corner points and the spatial coordinates of the mobile phone corner points includes the following operation steps:
[0052] The spatial intersection points obtained by vertically projecting each of the spatial coordinates of the thumb corner points onto the target prior plane;
[0053] The spatial intersection points represent the intersection points of the projections of the spatial coordinates of each thumb corner point onto the target prior plane and the target prior plane;
[0054] Calculate the left Euclidean distance dist_x of each of the spatial intersection points perpendicular to the left edge of the target prior plane and the upper Euclidean distance dist_y of each of the spatial intersection points perpendicular to the upper edge of the target prior plane;
[0055] Calculate the length L of the target prior plane and the width W of the target prior plane based on the spatial coordinates of each of the mobile phone intersection points;
[0056] Calculate the ratio r_X in the X direction based on the left Euclidean distance dist_x and the length L of the target prior plane; calculate the ratio r_Y in the Y direction based on the upper Euclidean distance dist_y and the width W of the target prior plane;
[0057] The calculation method of the ratio r_X in the X direction is:
[0058] r_X = dist_x ÷ L;
[0059] The calculation method of the ratio r_Y in the Y direction is:
[0060] r_Y = dist_y ÷ W.
[0061] Correspondingly, the present invention also proposes a response mapping device for an AR glasses collaborative gesture operation terminal, including an acquisition and calculation unit, a calculation unit, and a fusion module;
[0062] Wherein, the acquisition and calculation unit includes a binocular vision module and a calculation module;
[0063] Wherein, the binocular vision module is used to obtain a binocular view mobile phone image containing a complete mobile phone contour;
[0064] The calculation module is used to obtain the mobile phone corner coordinates after corner detection of the collected binocular view mobile phone images, and then use the mobile phone corner coordinates to obtain the mobile phone corner space coordinates through triangulation and plane fitting;
[0065] The binocular view mobile phone images include the left-eye view mobile phone image and the right-eye view mobile phone image;
[0066] The binocular vision module is further used to obtain a thumb image containing a thumb picture;
[0067] The calculation module is further used to obtain the thumb corner coordinates after corner detection of the collected thumb image, and then use the thumb corner coordinates to process through triangulation, Kalman filter algorithm and ICP algorithm in sequence to obtain the thumb corner space coordinates and the thumb posture;
[0068] The thumb image includes the left-eye view thumb image and the right-eye view thumb image;
[0069] The calculation unit is used to calculate the direction size ratio of the thumb projected onto the mobile phone screen based on the thumb corner space coordinates and the mobile phone corner space coordinates;
[0070] The direction size ratio includes the X-direction size ratio r_X and the Y-direction size ratio r_Y;
[0071] The fusion module is used to transmit the binocular view mobile phone image to the AR device to obtain an AR display image; and fuse the thumb image and the AR display image based on the direction size ratio and the thumb posture to obtain a fused image.
[0072] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:
[0073] Analyzing the above response mapping method and device for an AR glasses collaborative gesture operation terminal provided by the present invention, it can be seen that in specific applications, first, the accurate position of the mobile phone screen in the three-dimensional space is determined through the corner space coordinates of the mobile phone, providing the position basis of the mobile phone screen for subsequent steps; by obtaining an image containing the thumb, the two-dimensional corner coordinates of the thumb are obtained through corner detection, and then these two-dimensional corner coordinates are processed through triangulation, Kalman filtering, and ICP algorithms to obtain the three-dimensional space coordinates and posture of the thumb, accurately tracking the position and posture of the thumb, and providing data support for fusing and displaying the actual position of the thumb; based on the corner space coordinates of the mobile phone and the corner space coordinates of the thumb, the projection size ratio of the thumb on the mobile phone screen (the size ratio in the X direction and the Y direction) is calculated, and the space coordinates of the thumb are converted into relative sizes on the mobile phone screen to ensure the accurate display of the relative position of the thumb on the AR device; the binocular view mobile phone image is transmitted to the AR device, and based on the direction size ratio and the thumb posture, the thumb image is fused with the AR display image to finally generate a fused image, and then the synthesized image is displayed on the AR device to achieve the synchronous mapping of the binocular view mobile phone image and the thumb image, providing an enhanced reality experience for user interaction;
[0074] Through deep learning technology and visual geometry algorithms, the spatial positions and postures of the mobile phone terminal and the gesture are calculated in real time, and then the spatial pose of the gesture is projected onto the mobile phone terminal plane through a projection algorithm to obtain the position of the hand relative to the mobile phone terminal. These position information and the mobile phone screen image are uploaded to the AR device, enabling the AR device to simultaneously display the position of the user's finger while displaying the mobile phone screen, providing visual reference and assistance for the user's operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0076] Figure 1 It is a schematic diagram of the overall operation steps of a response mapping method for an AR glasses collaborative gesture operation terminal provided in Embodiment 1 of the present invention;
[0077] Figure 2 It is a schematic diagram of thumb corner simulation of a response mapping method for an AR glasses collaborative gesture operation terminal provided in Embodiment 1 of the present invention;
[0078] Figure 3Schematic diagram of the operation steps for obtaining the spatial coordinates of the thumb corner point and the thumb gesture in a response mapping method of an AR glasses collaborative gesture operation terminal provided in Embodiment 1 of the present invention;
[0079] Figure 4 Schematic diagram of simulating the Euclidean distance from the spatial intersection point perpendicular to the left edge of the mobile phone plane to the upper edge of the Euclidean distance in a response mapping method of an AR glasses collaborative gesture operation terminal provided in Embodiment 1 of the present invention;
[0080] Figure 5 Overall architecture diagram of a response mapping device of an AR glasses collaborative gesture operation terminal provided in Embodiment 2 of the present invention.
[0081] Reference numerals: acquisition and calculation unit 10; calculation unit 20; fusion module 30; binocular vision module 11; calculation module 12. Detailed implementation manners
[0082] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0083] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.
[0084] Embodiment 1
[0085] As Figure 1 shown, the present invention proposes a response mapping method for an AR glasses collaborative gesture operation terminal, including the following operation steps:
[0086] Step S10: (By shooting with a binocular vision module) Obtain a binocular view mobile phone image containing a complete mobile phone contour, obtain the mobile phone corner coordinates after corner detection of the collected binocular view mobile phone image, and then use the mobile phone corner coordinates to obtain the mobile phone corner spatial coordinates through triangulation and plane fitting;
[0087] The binocular view mobile phone image includes a left-eye view mobile phone image and a right-eye view mobile phone image;
[0088] It should be noted that the present invention only considers mobile phone terminals with a rectangular contour and does not consider special-shaped mobile phone terminals.
[0089] Step S20: (Captured by a binocular vision module) Obtain a thumb image containing a thumb picture. After performing corner detection on the captured thumb image, obtain the thumb corner coordinates, and then use the thumb corner coordinates to sequentially process through triangulation, Kalman filtering algorithm, and ICP algorithm to obtain the thumb corner spatial coordinates and thumb pose;
[0090] The thumb image includes a left-eye view thumb image and a right-eye view thumb image (i.e., the left thumb image captured by the left-eye vision module in the binocular vision module and the right thumb image captured by the right-eye vision module. The left thumb image and the right thumb image are essentially two images of the same thumb captured by the left and right vision modules of the binocular vision module);
[0091] It should be noted that in the scenario where the user plays games using the mobile terminal, generally the thumb fingertip is used for screen control, and the other fingers are used to hold the mobile phone. Therefore, the present invention is only limited to tracking the spatial coordinates of the thumb corners of both hands, and it is the finger joints of the thumb; as Figure 2 shown, we define four corner points with obvious features of the nail as the key points of the thumb finger joint;
[0092] The above thumb corner coordinates in the embodiments of the present application are obtained by performing corner detection operations using a pre-trained thumb corner detection model. The above pre-trained thumb corner detection model in the embodiments of the present application is trained based on the YOLO model;
[0093] First, it is necessary to obtain the image coordinates of these 4 corner points in the binocular images. To achieve this goal, we collected thumb images of multiple people through the binocular vision module, and by means of manual annotation, annotated the image coordinates of the 4 key points of each thumb to obtain a thumb key point dataset;
[0094] Next, we use a pre-trained thumb corner detection model (such as YOLO) to train the thumb key point dataset generated in the previous step to obtain a thumb corner detection model. Based on this model, the thumb corner coordinates (the thumb corner coordinates are the image 2D coordinates Left_Thumb_Coords and Right_Thumb_Coords of the 4 key points of the thumb, that is, the image 2D coordinates of the 4 key points of the left thumb image and the image 2D coordinates of the 4 key points of the right thumb image) of any thumb image in this scenario can be automatically detected.
[0095] Step S30: Calculate the direction size ratio of the thumb projected onto the mobile phone screen based on the thumb corner spatial coordinates and the mobile phone corner spatial coordinates;
[0096] The direction size ratio includes an X-direction size ratio r_X and a Y-direction size ratio r_Y;
[0097] It should be noted that through the above steps, the position and plane equation of the mobile phone plane in the space coordinate system, as well as the position and posture of the thumb, have been obtained. However, in the AR device, the visualization is presented according to the pixel size of the mobile phone screen. Therefore, it is necessary to calculate the size ratios in the X and Y directions after the thumb is projected onto the mobile phone screen.
[0098] Step S40: After transmitting the binocular view mobile phone image to the AR device, obtain the AR display image (i.e., the display image of the AR device); based on the direction size ratio and the thumb posture, fuse the thumb image with the AR display image to obtain a fused image;
[0099] Then, take the fused image, that is, the fused image, as the final display screen, and display on the AR device a display screen in which the binocular view mobile phone image and the thumb image are synchronously mapped.
[0100] It should be noted that in the AR device of the embodiment of the present application, first, the screen of the mobile phone terminal is transmitted to the screen of the AR device without loss through signal transmission technologies such as projection, to obtain the AR display image on the screen of the AR device; further, multiply the ratios r_X and r_Y of the thumb key points (thumb image) within the range of the mobile phone screen obtained above by the pixel resolution of the mobile phone screen to obtain the landing coordinates on the mobile phone screen, and visualize them through means such as rendering (visualization means obtaining the above-mentioned fused image on the screen of the AR device); alternatively, through tools such as Unity, prefabricate a visual digital model of the thumb, and drive the model to be visually presented in the interface through the thumb posture R and T of the t-th frame obtained above (visual presentation means obtaining the above-mentioned fused image on the screen of the AR device);
[0101] In the above embodiments of the present application, the accurate position of the mobile phone screen in the three-dimensional space is first determined through the spatial coordinates of the mobile phone corner points, providing a position basis for the subsequent steps; by obtaining an image containing the thumb, the two-dimensional corner coordinates of the thumb are obtained through corner detection, and then these two-dimensional corner coordinates are processed through triangulation, Kalman filtering, and ICP algorithms to obtain the three-dimensional spatial coordinates and posture of the thumb, accurately tracking the position and posture of the thumb, providing data support for fusing and displaying the actual position of the thumb; based on the spatial coordinates of the mobile phone corner points and the spatial coordinates of the thumb corner points, the projection size ratio of the thumb on the mobile phone screen (the size ratio in the X direction and the Y direction) is calculated, converting the spatial coordinates of the thumb into relative sizes on the mobile phone screen to ensure the accurate display of the relative position of the thumb on the AR device; the binocular view mobile phone image (the binocular view mobile phone image or the mobile phone screen image) is transmitted to the AR device, and based on the direction size ratio and the thumb posture, the thumb image is fused with the AR display image to finally generate a fused image, and then the synthesized image is displayed on the AR device, realizing the synchronous mapping of the mobile phone screen and the thumb image, providing an enhanced reality experience for user interaction;
[0102] Through deep learning technology and visual geometry algorithms, the spatial positions and postures of the mobile phone terminal and the gesture are calculated in real time, and then the spatial pose of the gesture is projected onto the mobile phone terminal plane through a projection algorithm to obtain the position of the hand relative to the mobile phone terminal. These position information and the mobile phone screen image are uploaded to the AR device, enabling the AR device to simultaneously display the position of the user's finger while displaying the mobile phone screen, providing visual reference and assistance for the user's operation.
[0103] Specifically, in step S10, the mobile phone corner coordinates obtained after corner detection of the collected binocular view mobile phone image are used, and then the mobile phone corner spatial coordinates are obtained through triangulation and plane fitting using the mobile phone corner coordinates, including the following operation steps:
[0104] Step S11: Perform corner detection on the binocular view mobile phone image through a pre-trained key point detection network model to obtain the mobile phone corner coordinates corresponding to the binocular view mobile phone image;
[0105] The mobile phone corner coordinates include the mobile phone corner coordinates of the left-eye view mobile phone image and the mobile phone corner coordinates of the right-eye view mobile phone image;
[0106] It should be noted that the above mobile phone corner coordinates (i.e., the upper left corner point coordinates, lower left corner point coordinates, upper right corner point coordinates, and lower right corner point coordinates corresponding to the left binocular view mobile phone image and the upper left corner point coordinates, lower left corner point coordinates, upper right corner point coordinates, and lower right corner point coordinates corresponding to the right binocular view mobile phone image) are the 2D spatial positions (i.e., the 2D coordinates on the binocular view mobile phone image) of the 4 corner points of the binocular view mobile phone image;
[0107] Step S12: Calculate the three-dimensional coordinates corresponding to the mobile phone corner coordinates through the triangulation algorithm using the internal and external parameters of the camera;
[0108] The three-dimensional coordinates include Phone_Pt1(x1, y1, z1), Phone_Pt2(x2, y2, z2), Phone_Pt3(x3, y3, z3), Phone_Pt4(x4, y4, z4);
[0109] It should be noted that in the above embodiments of the present application, based on the internal and external parameters of the binocular camera and the mobile phone corner coordinates in the left and right images obtained from the previous model detection, the 3D positions of the 4 corners of the mobile phone are calculated through the triangulation algorithm; the triangulation algorithm is a commonly used method in spatial geometric operations. Based on the observations of the same point in space by multiple cameras at different positions, through the pixel coordinates of the point on the imaging plane of each camera and the motion relationship between the cameras, the spatial position of the point can be calculated; the above triangulation algorithm is a prior art and will not be elaborated in this application.
[0110] Step S13: Perform plane fitting on the three-dimensional coordinates to obtain the spatial coordinates of the mobile phone corners;
[0111] It should be noted that due to the influence of noise, there will be errors in the spatial positions of the 4 corners obtained by triangulation. Since the 4 corners of the mobile phone are coplanar in space (the space refers to the same mobile phone screen plane, that is, coplanar on the 3D plane), we perform plane fitting on the results obtained from the previous step to obtain the spatial coordinates of the mobile phone corners.
[0112] Specifically, in step S11, the pre-trained key point detection network model is trained based on the YOLO model; the acquisition of the pre-trained key point detection network model includes the following operating steps:
[0113] Step S111: Crawl multiple sample mobile phone contour images through the web crawler method; perform preprocessing operations on the sample mobile phone contour images to obtain preprocessed sample mobile phone contour images;
[0114] The preprocessing operations include screen size adjustment operations (i.e., size adjustment operations for sample mobile phone contour images);
[0115] It should be noted that in the above embodiments of the present application, a large number of front images of different mobile phone models are obtained through the method of web crawlers. Each image is first transformed into the image resolution size we need by methods such as resize (adjusting the size) or adding Padding (adding borders), providing an initial data basis for the training of the subsequent key point detection network model;
[0116] Step S112: Annotate the preprocessed sample mobile phone contour image to obtain a sample data set;
[0117] It should be noted that the above annotation refers to manually annotating the mobile phone corner coordinates of each sample mobile phone contour image; in the above embodiments of the present application, the image coordinates of 4 corners of the mobile phone are annotated for each image through the method of manual annotation, thereby generating a mobile phone data set, providing a direct data basis for the training of the subsequent key point detection network model;
[0118] Step S113: Establish an initial key point detection network model; train the initial key point detection network model based on the sample data set, and output the trained key point detection network model after the initial key point detection network model converges;
[0119] It should be noted that in the above embodiments of the present application, the key point detection network of deep learning (such as the YOLO model) is used to train the model with the mobile phone data set generated in the previous step to obtain a deep learning network model that can be used for detecting 4 corners of the mobile phone. Using this model, the 2D image coordinates of 4 corners of the mobile phone in the left and right images can be automatically detected for the mobile phone image taken frontally by the binocular module;
[0120] The above operations of training the initial key point detection network model based on the sample data set and the model convergence judgment conditions are all well-known common knowledge to those skilled in the art, and the present application will not elaborate.
[0121] Specifically, in step S13, the plane fitting is performed on the three-dimensional coordinates to obtain the spatial coordinates of the mobile phone corner points, including the following operation steps:
[0122] Step S131: Calculate the average point coordinate P_m(x_m, y_m, z_m) based on each of the three-dimensional coordinates;
[0123] It should be noted that the three-dimensional coordinates of the four corner points obtained in the above embodiments of the present application are: Phone_Pt1(x1, y1, z1), Phone_Pt2(x2, y2, z2), Phone_Pt3(x3, y3, z3), Phone_Pt4(x4, y4, z4); in fact, the four corner points belong to the same plane, but the equation of this plane is unknown. Therefore, in the embodiments of the present application, the average point coordinates of the three-dimensional coordinates of the four corner points are calculated, and then a prior plane (the prior plane actually exists, but its plane equation is unknown) is determined by using the three-dimensional coordinates of the four corner points and the average point coordinates, and then the spatial coordinates of the corner points are obtained by analyzing using this prior plane, as shown in the following steps.
[0124] Step S132: Construct a prior plane based on the average point coordinates P_m(x_m, y_m, z_m) (first determine the general expression of the prior plane through the average point coordinates, then construct a prior plane without knowing the parameters through the average point coordinates, then solve the plane equation parameters of the general expression of the prior plane, and finally obtain the general expression of the target prior plane, and then project the three-dimensional coordinates onto the target prior plane to obtain the spatial coordinates of the phone corner points);
[0125] It should be noted that the prior plane (the expression of the prior plane) in the above embodiments of the present application is constructed by using the general expression of the plane in space and the average point coordinates P_m(x_m, y_m, z_m). Assume that the general expression of the plane in space is: A1x + B1y + C1z + D1 = 0, and the positions of the 4 corner points obtained in the above steps in space are: Phone_Pt1(x1, y1, z1), Phone_Pt2(x2, y2, z2), Phone_Pt3(x3, y3, z3), Phone_Pt4(x4, y4, z4); at the same time, because the prior plane passes through the average point P_m(x_m, y_m, z_m) of all points, that is, it satisfies the following equation (that is, the general expression of the prior plane is): A2x_m + B2y_m + C2z_m + D2 = 0;
[0126] However, in the above embodiments of the present application, the plane equation parameters A1, B1, C1, D1 of the general expression of the plane in the assumed space and the plane equation parameters A2, B2, C2, D2 of the general expression of the constructed prior plane are unknown. Therefore, the embodiments of the present application also need to obtain the plane equation parameters of the general expression of the prior plane to obtain the general expression of the target prior plane, and the specific operations are shown in the following steps.
[0127] Step S133: Construct an error matrix H based on the error terms of each of the three-dimensional coordinates to the prior plane (or the general expression of the prior plane).
[0128] The error matrix H is:
[0129]
[0130] In the formula, (x1, y1, z1), (x2, y2, z2), (x3, y3, z3), and (x4, y4, z4) respectively represent the three-dimensional coordinates corresponding to the corner coordinates of the mobile phone; (x m , y m , z m ) represents the average coordinate of the four three-dimensional coordinates.
[0131] Step S134: Perform a singular value decomposition operation on the error matrix H to obtain a decomposed matrix H'.
[0132] The decomposed matrix H' is:
[0133] H' = UDV T ;
[0134] In the formula, U is the left singular vector matrix, D is the singular value diagonal matrix, and V T is the right singular vector matrix.
[0135] It should be noted that the error terms of each three-dimensional coordinate to the prior plane obtained in the above embodiments of the present application can be used to construct an equation, that is:
[0136]
[0137] Then, in the embodiments of the present application, in the equation is used as the matrix H to obtain the plane parameters of the prior plane, thereby providing a data basis for obtaining the corner space coordinates of each three-dimensional coordinate subsequently.
[0138] Step S135: Obtain the last column eigenvector of the right singular vector matrix V T as the plane parameter A, the plane parameter B, and the plane parameter C.
[0139] Step S136: Calculate the plane parameter D according to the three-dimensional coordinates and the prior plane.
[0140] In the above embodiments of the present application, the three-dimensional coordinates are substituted into the general expression of the obtained prior plane (with the known plane parameters A, B, and C) to calculate the plane parameter D.
[0141] Step S137: Obtain a target prior plane according to the plane parameter A, plane parameter B, plane parameter C, and the plane parameter D;
[0142] The general expression of the target prior plane is:
[0143] Ax_m + By_m + Cz_m + D = 0;
[0144] In the formula, A, B, C, and D are plane equation parameters respectively;
[0145] Step S138: Vertically project each of the three-dimensional coordinates onto the target prior plane to obtain the spatial coordinates of the mobile phone corner points corresponding to each of the three-dimensional coordinates;
[0146] The spatial coordinates of the mobile phone corner points include Phone_Pt1_proj, Phone_Pt2_proj, Phone_Pt3_proj, and Phone_Pt4_proj.
[0147] It should be noted that in the above embodiments of the present application, the coordinates of the points obtained by vertically projecting the three-dimensional coordinates onto the above target prior plane (or the intersection coordinates of the three-dimensional coordinates and the prior plane) are used as the spatial coordinates of the corner points;
[0148] In the above embodiments of the present application, the general expression of an initial plane (i.e., the general expression of the above prior plane) is constructed through the average point coordinates of each three-dimensional coordinate. Then, an error matrix H is constructed according to the error terms of each three-dimensional coordinate to the initial plane, and a singular value decomposition operation is performed on the error matrix to obtain a decomposed matrix H'. Then, according to the right singular vector matrix V in the decomposed matrix T the plane parameters A, B, and C in the general expression of the initial plane are obtained. Then, the known three-dimensional coordinates are substituted into the general expression of the initial plane with the known plane parameters A, B, and C to calculate the plane parameter D, so as to obtain the general expression of the final target prior plane. Finally, the perpendicular projection points (i.e., the projection points) of each three-dimensional coordinate and the target prior plane are used as the spatial coordinates of the mobile phone corner points;
[0149] Due to the influence of noise, there will be errors in the spatial positions of the three-dimensional coordinates corresponding to the four mobile phone corner coordinates obtained by triangulation above. Since the four mobile phone corner coordinates of the mobile phone are analyzed and recognized from the binocular view mobile phone images, their corresponding three-dimensional coordinates are coplanar in space. Therefore, the embodiment of the present application performs plane fitting on the result obtained in the previous step to obtain the plane corresponding to the four three-dimensional coordinates (i.e., the above-mentioned target prior plane); furthermore, by projecting the three-dimensional coordinates onto the target prior plane (this operation is essentially to correct the four three-dimensional coordinates), the mobile phone corner space coordinates corresponding to the four coplanar mobile phone corner coordinates in space are obtained.
[0150] Specifically, as Figure 3 shown, in step S20, the thumb corner coordinates are respectively processed through triangulation, Kalman filter algorithm, and ICP algorithm in sequence to obtain the thumb corner space coordinates and thumb posture, including the following operation steps:
[0151] Step S21: Obtain the internal and external parameters of the camera; calculate the initial thumb corner three-dimensional coordinates through triangulation according to the internal and external parameters of the camera and the thumb corner coordinates;
[0152] It should be noted that the above internal and external parameters of the camera are the internal and external parameters of the binocular vision module's camera;
[0153] In the above embodiment of the present application, according to the internal and external parameters of the binocular camera and the 2D coordinates of the 4 thumb key points in the left and right images detected by the previous model, the 3D position in the world coordinate system is calculated through the triangulation algorithm.
[0154] Step S22: Obtain the frame thumb corner coordinates of consecutive frames; calculate the frame thumb corner three-dimensional coordinates corresponding to each frame through triangulation for the frame thumb corner coordinates; perform prediction and correction operations on the frame thumb corner three-dimensional coordinates through the Kalman filter algorithm to obtain the corrected frame thumb corner three-dimensional coordinates (i.e., the above-mentioned thumb corner space coordinates);
[0155] Note that the initial thumb corner three-dimensional coordinates, frame thumb corner coordinates, and frame thumb corner three-dimensional coordinates are all different concepts; among them, the initial thumb corner three-dimensional coordinates and frame thumb corner three-dimensional coordinates are important parameters for constituting the thumb posture;
[0156] It should be noted that in the above embodiment of the present application, by recording the thumb (or thumb corner coordinates) of each frame in the continuous frames of the normal operation movement, and then continuing to calculate using the triangulation method, the three-dimensional space coordinates corresponding to the thumb corner coordinates of each frame are obtained; assuming that the three-dimensional space coordinates of the 4 key points (the 4 key points are the 4 thumb corner coordinates) of the left thumb are calculated for the t-th frame as:
[0157] Left_Thumb_Pt0, Left_Thumb_Pt1, Left_Thumb_Pt2, Left_Thumb_Pt3;
[0158] Then, further predict and correct the three-dimensional coordinates of the above 4 thumb corner points through the Kalman filtering algorithm to obtain the corrected three-dimensional coordinates of the frame thumb corner points. The specific operations are shown in the following steps.
[0159] Step S23: Calculate the thumb pose corresponding to each corrected three-dimensional coordinate of the frame thumb corner point through the ICP algorithm based on the initial three-dimensional coordinates of the thumb corner point and the corrected three-dimensional coordinates of the frame thumb corner point;
[0160] The thumb pose includes a rotation matrix R and a translation vector T.
[0161] It should be noted that in the above embodiments of the present application, by obtaining the internal and external parameters of the cameras of the binocular vision module and the 2D coordinates of the thumb corner points, the three-dimensional coordinates of the initial thumb corner points are calculated using the triangulation method, providing the initial position basis of the thumb in the world coordinate system for subsequent pose calculation; further, by obtaining the thumb corner point coordinates in consecutive frames and calculating the three-dimensional coordinates of each frame through triangulation, and using the Kalman filtering algorithm to predict and correct these three-dimensional coordinates, more accurate and stable three-dimensional coordinates of the thumb are obtained, making the determination of the thumb position more accurate; and then by using the ICP algorithm to calculate the pose of the thumb based on the initial three-dimensional coordinates and the corrected three-dimensional coordinates of the frame, accurate pose information of the thumb is provided, providing a basis for correctly displaying the position and orientation of the thumb in the AR device.
[0162] Specifically, in step S21, calculating the initial three-dimensional coordinates of the thumb corner point through the triangulation method according to the internal and external parameters of the camera and the thumb corner point coordinates includes the following operation steps:
[0163] Step S211: Obtain N frames of first thumb corner point coordinates of the thumb corner point coordinates in a stationary state on the mobile phone screen surface;
[0164] Step S212: Calculate the first thumb corner point three-dimensional coordinates by performing triangulation on the first thumb corner point coordinates according to the internal and external parameters of the camera; calculate the average coordinates of the first thumb corner point three-dimensional coordinates respectively corresponding to each of the thumb corner point coordinates; determine the average coordinates as the initial thumb corner point three-dimensional coordinates;
[0165] It should be noted that in the above embodiments of the present application, through interface guidance, the user's thumb is first placed flat on the surface of the mobile phone screen in the positive direction and kept stationary for N frames (that is, the coordinates of each thumb corner point remain stationary on the surface of the mobile phone screen). Within N frames, the first thumb corner point coordinates of the 4 key points of the thumb in each frame are recorded. Then, the triangulation method is used to calculate the three-dimensional coordinates corresponding to the first thumb corner point coordinates (that is, the above-mentioned first thumb corner point three-dimensional coordinates), and the average value of the results is obtained (that is, the above-mentioned average coordinates, and the average coordinates are three-dimensional coordinates), which is used as the initial position of this thumb (that is, the above-mentioned initial thumb corner point three-dimensional coordinates);
[0166] For example, the initial position of the left thumb is obtained: Left_Thumb_Init_Pt0, Left_Thumb_Init_Pt1, Left_Thumb_Init_Pt2, Left_Thumb_Init_Pt3.
[0167] Specifically, in step S22, the frame thumb corner point three-dimensional coordinates are predicted and corrected through the Kalman filter algorithm to obtain the corrected frame thumb corner point three-dimensional coordinates, including the following operation steps:
[0168] Step S221: Establish a motion model for each of the frame thumb corner point three-dimensional coordinates;
[0169] Step S222: Traverse each of the frame thumb corner point three-dimensional coordinates, and based on the frame thumb corner point three-dimensional coordinates corresponding to the thumb corner point coordinates in consecutive frames and the motion model, perform prediction to obtain the predicted three-dimensional coordinates corresponding to the frame thumb corner point three-dimensional coordinates of the t-th frame;
[0170] Step S223: Correct the predicted three-dimensional coordinates based on the frame thumb corner point three-dimensional coordinates of the (t - 1)-th frame to obtain the corrected thumb corner point three-dimensional coordinates;
[0171] Repeat the above operations until all the frame thumb corner point three-dimensional coordinates are traversed.
[0172] It should be noted that in the above embodiments of the present application, by designing a Kalman filter algorithm, the correction is performed on these 4 spatial position points respectively;
[0173] First, a motion model is established for each key point, and the position information of the historical frames is saved (the position information of the historical frames is the frame thumb corner point three-dimensional coordinates corresponding to the thumb corner point coordinates in the above-mentioned consecutive frames);
[0174] Taking one of the points, Left_Thumb_Pt0, as an example, using the motion model of this key point (i.e., the three-dimensional coordinates of the thumb corner point in the above frame) and the historical frame information, the position state of the t-th frame (the position state of the t-th frame is the predicted three-dimensional coordinates of the t-th frame above) Left_Thumb_Pt0_pred_hat is predicted. At the same time, using the spatial position of this key point in the t-th frame obtained by the previous triangulation (i.e., the three-dimensional coordinates of the thumb corner point in the t - 1 frame, where t - 1 frame is the previous frame. If there are no three-dimensional coordinates of the thumb corner point in the previous frame in a certain special case, then directly use the three-dimensional coordinates of the thumb corner point in the t-th frame as the modified three-dimensional coordinates of the thumb corner point) Left_Thumb_Pt0 as the observation value, the predicted state is updated to obtain the corrected position Left_Thumb_Pt0_hat;
[0175] Similarly, the corrected key point positions (the corrected key point positions are the corrected three-dimensional coordinates of the thumb corner point in the t-th frame) are obtained: Left_Thumb_Pt1_hat, Left_Thumb_Pt2_hat, Left_Thumb_Pt3_hat.
[0176] Specifically, in step S23, based on the initial three-dimensional coordinates of the thumb corner point and the corrected three-dimensional coordinates of the thumb corner point in the frame, the thumb postures corresponding to the corrected three-dimensional coordinates of the thumb corner point in each frame are calculated through the ICP algorithm, including the following operation steps:
[0177] Step S231: Calculate the initial centroid coordinates of each of the initial three-dimensional coordinates of the thumb corner point and the centroid coordinates of the frame of each of the corrected three-dimensional coordinates of the thumb corner point in the frame;
[0178] Step S232: Calculate the initial relative centroid coordinates based on the initial centroid coordinates and each of the initial three-dimensional coordinates of the thumb corner point; and calculate the frame relative centroid coordinates based on the corrected three-dimensional coordinates of the thumb corner point and the centroid coordinates of the frame;
[0179] The calculation method of the initial relative centroid coordinates is: initial relative centroid coordinates = initial three-dimensional coordinates of the thumb corner point - initial centroid coordinates;
[0180] The calculation method of the frame relative centroid coordinates is: frame relative centroid coordinates = corrected three-dimensional coordinates of the thumb corner point - centroid coordinates of the frame;
[0181] Step S233: Perform a normalization operation on each of the initial relative centroid coordinates to obtain an initial normalized point set; perform a normalization operation on each of the frame relative centroid coordinates to obtain a frame normalized point set;
[0182] Step S234: Calculate the thumb pose (R, T) corresponding to the three-dimensional coordinates of each corrected frame thumb corner point by the least squares method based on the initial normalized point set and the frame normalized point set;
[0183] The calculation method of the thumb pose (R, T) is as follows:
[0184] In the formula, R is the rotation matrix; T is the translation vector; argmin is the input value corresponding to the minimum value of the function; sp_i is the i-th initial relative centroid coordinate in the initial normalized point set; dp_i is the i-th frame relative centroid coordinate in the frame normalized point set;
[0185] It should be noted that in the above embodiments of the present application, the pose of the thumb in the t-th frame is calculated by using the ICP (Iterative Closest Point) algorithm;
[0186] In the above embodiments of the present application, the initial three-dimensional spatial positions of the thumb Left_Thumb_Init_Pt0, Left_Thumb_Init_Pt1, Left_Thumb_Init_Pt2, Left_Thumb_Init_Pt3, and the three-dimensional spatial positions of the thumb in the t-th frame Left_Thumb_Pt0_hat, Left_Thumb_Pt1_hat, Left_Thumb_Pt2_hat, Left_Thumb_Pt3_hat have been obtained;
[0187] First, subtract the centroid coordinates of each group of three-dimensional position points from their respective centroid coordinates to obtain the normalized results, the initial normalized point set Src_Pts = {sp_0, sp_1, sp_2, sp_3} and the frame normalized point set Dst_Pts = {dp_0, dp_1, dp_2, dp_3}. Then, use the least squares method to calculate the rotation matrix R and the translation vector T for the three-dimensional coordinates of each corrected frame thumb corner point, so that the positions of the two groups of three-dimensional points after transformation are the most fitting, that is, the thumb pose corresponding to the three-dimensional coordinates of each corrected frame thumb corner point is obtained.
[0188] Specifically, in step S30, based on the spatial coordinates of the thumb corner points and the spatial coordinates of the mobile phone corner points, the direction size ratio of the thumb projected onto the mobile phone screen is calculated, including the following operation steps:
[0189] Step S31: The spatial intersection points obtained by vertically projecting each of the spatial coordinates of the thumb corner points onto the target prior plane;
[0190] The spatial intersection points represent the intersection points of the spatial coordinates of each thumb corner point projected onto the target prior plane and the target prior plane;
[0191] Step S32: Calculate the left Euclidean distance dist_x of each of the spatial intersection points perpendicular to the left edge of the target prior plane, and the upper Euclidean distance dist_y of each of the spatial intersection points perpendicular to the upper edge of the target prior plane;
[0192] Step S33: Calculate the length L and the width W of the target prior plane based on the spatial coordinates of each of the mobile intersection points;
[0193] Step S34: Calculate the ratio r_X in the X direction based on the left Euclidean distance dist_x and the length L of the target prior plane; calculate the ratio r_Y in the Y direction based on the upper Euclidean distance dist_y and the width W of the target prior plane;
[0194] The calculation method of the ratio r_X in the X direction is: r_X = dist_x ÷ L;
[0195] The calculation method of the ratio r_Y in the Y direction is: r_Y = dist_y ÷ W;
[0196] It should be noted that first, in the above embodiments of the present application, the 4 key points of the thumb (i.e., the above-mentioned spatial coordinates of the thumb corner points, or the three-dimensional coordinates of the corrected frame thumb corner points):
[0197] Left_Thumb_Pt0_hat, Left_Thumb_Pt1_hat, Left_Thumb_Pt2_hat, Left_Thumb_Pt3_hat are respectively projected perpendicularly onto the spatial mobile plane (i.e., the above-mentioned target prior plane), and the spatial intersection points with the plane are calculated to obtain Left_Thumb_Pt0_proj, Left_Thumb_Pt1_proj, Left_Thumb_Pt2_proj, Left_Thumb_Pt3_proj.
[0198] Next, calculate the Euclidean distance dist_x of each spatial intersection point perpendicular to the left edge of the mobile plane and the Euclidean distance dist_y perpendicular to the upper edge of the mobile plane respectively, as Figure 4 shown; at the same time, the length L (length) and width W (width) of the mobile phone can be calculated from the 4 projection points of the mobile phone (the projection points are the above-mentioned spatial intersection points) Phone_Pt1_proj, Phone_Pt2_proj, Phone_Pt3_proj, Phone_Pt4_proj. Therefore, the ratio r_X in the X direction and the ratio r_Y in the Y direction within the range of the mobile phone can be calculated for these 4 key points of the thumb.
[0199] That is, the direction dimension ratio of the thumb projected onto the mobile phone screen can be obtained therefrom, including the ratio r_X in the X direction and the ratio r_Y in the Y direction, and output; then the operations in the above operation steps S31 - S34 are ended.
[0200] Embodiment 2
[0201] As Figure 5 shown, correspondingly, the present invention further provides a response mapping device for an AR glasses collaborative gesture operation terminal, including an acquisition and calculation unit 10, a calculation unit 20, and a fusion module 30;
[0202] Among them, the acquisition and calculation unit 10 includes a binocular vision module 11 and a calculation module 12;
[0203] Among them, the binocular vision module 11 is used to obtain a binocular view mobile phone image containing a complete mobile phone contour;
[0204] The calculation module 12 is used to obtain the mobile phone corner coordinates obtained after corner detection of the collected binocular view mobile phone image, and then use the mobile phone corner coordinates to obtain the mobile phone corner spatial coordinates through triangulation and plane fitting;
[0205] The binocular view mobile phone image includes a left-eye view mobile phone image and a right-eye view mobile phone image;
[0206] The binocular vision module 11 is further used to obtain a thumb image containing a thumb picture;
[0207] The calculation module 12 is further used to obtain the thumb corner coordinates obtained after corner detection of the collected thumb image, and then use the thumb corner coordinates to process through triangulation, Kalman filter algorithm, and ICP algorithm in sequence to obtain the thumb corner spatial coordinates and thumb posture;
[0208] The thumb image includes a left-eye view thumb image and a right-eye view thumb image;
[0209] The calculation unit 20 is used to calculate the direction dimension ratio of the thumb projected onto the mobile phone screen based on the thumb corner spatial coordinates and the mobile phone corner spatial coordinates;
[0210] The direction dimension ratio includes the X-direction dimension ratio r_X and the Y-direction dimension ratio r_Y;
[0211] The fusion module 30 is used to transmit the binocular view mobile phone image to the AR device to obtain an AR display image; fuse the thumb image and the AR display image based on the direction dimension ratio and the thumb posture to obtain a fused image.
[0212] In summary, a response mapping method and device for an AR glasses collaborative gesture operation terminal proposed in the embodiments of the present invention first determine the accurate position of the mobile phone screen in the three-dimensional space through the corner space coordinates of the mobile phone, providing a position basis for the subsequent steps; by obtaining an image containing the thumb, the two-dimensional corner coordinates of the thumb are obtained through corner detection, and then these two-dimensional corner coordinates are processed through triangulation, Kalman filtering, and ICP algorithms to obtain the three-dimensional space coordinates and posture of the thumb, accurately tracking the position and posture of the thumb, and providing data support for fusing and displaying the actual position of the thumb; based on the corner space coordinates of the mobile phone and the corner space coordinates of the thumb, the projection size ratio of the thumb on the mobile phone screen is calculated, and the space coordinates of the thumb are converted into relative sizes on the mobile phone screen to ensure the accurate display of the relative position of the thumb on the AR device; the binocular view mobile phone image is transmitted to the AR device, and based on the direction size ratio and the thumb posture, the thumb image is fused with the AR display image, and finally a fused image is generated and then the synthesized image is displayed on the AR device to realize the synchronous mapping of the mobile phone screen and the thumb image, providing an enhanced reality experience for user interaction;
[0213] Through deep learning technology and visual geometry algorithms, the spatial positions and postures of the mobile phone terminal and the gesture are calculated in real time, and then the spatial pose of the gesture is projected onto the mobile phone terminal plane through a projection algorithm to obtain the position of the hand relative to the mobile phone terminal. These position information and the mobile phone screen image are uploaded to the AR device, so that the AR device can simultaneously display the position of the user's finger while displaying the mobile phone screen, providing a visual reference and assistance for the user's operation.
[0214] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; those of ordinary skill in the art can modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A response mapping method for AR glasses in collaboration with a gesture operation terminal, characterized in that: The steps are as follows: Acquire a binocular view mobile phone image containing a complete mobile phone outline, obtain the coordinates of the mobile phone corner points after corner point detection on the acquired binocular view mobile phone image, and then use the mobile phone corner point coordinates to obtain the spatial coordinates of the mobile phone corner points by triangulation and plane fitting; A thumb image containing a thumb picture is obtained, and the coordinates of the thumb corner points are obtained after corner point detection of the collected thumb image. Then, the coordinates of the thumb corner points are processed by triangulation, Kalman filtering algorithm and ICP algorithm in sequence to obtain the spatial coordinates of the thumb corner points and the thumb posture; The directional dimension ratio of the thumb projected onto the mobile phone screen is calculated based on the spatial coordinates of the thumb corner points and the spatial coordinates of the mobile phone corner points; The directional dimension ratio includes an X-direction dimension ratio r_X and a Y-direction dimension ratio r_Y; The binocular view mobile phone image is transmitted to the AR device to obtain an AR display image; the thumb image is fused with the AR display image based on the direction size ratio and the thumb posture to obtain a fused image.
2. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 1, characterized in that: The method of obtaining the mobile phone corner point coordinates by corner point detection of the collected binocular view mobile phone image, and then obtaining the mobile phone corner point spatial coordinates by triangulation and plane fitting using the mobile phone corner point coordinates, includes the following steps: Performing corner point detection on the binocular view mobile phone image through a pre-trained key point detection network model to obtain the mobile phone corner point coordinates corresponding to the binocular view mobile phone image; The coordinates of the corner points of the mobile phone are calculated by a triangulation algorithm using the internal and external parameters of the camera to obtain the three-dimensional coordinates corresponding to the coordinates of the corner points of the mobile phone; Perform plane fitting on the three-dimensional coordinates to obtain the spatial coordinates of the corner points of the mobile phone.
3. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 2, characterized in that: The pre-trained key point detection network model is trained based on the YOLO model; the acquisition of the pre-trained key point detection network model includes the following steps: Crawling through a web crawler method to obtain a plurality of sample mobile phone outline images; performing a preprocessing operation on the sample mobile phone outline images to obtain preprocessed sample mobile phone outline images; The pre-processing operation includes a picture size adjustment operation; Annotating the preprocessed sample mobile phone contour image to obtain a sample data set; Establishing an initial key point detection network model; training the initial key point detection network model based on the sample data set, and outputting a trained key point detection network model after the initial key point detection network model converges.
4. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 3, characterized in that: The three-dimensional coordinates are plane-fitted to obtain the spatial coordinates of the mobile phone corner points; specifically, the following steps are included: Calculate the average point coordinates P_m (x_m, y_m, z_m) based on each of the three-dimensional coordinates; Constructing a priori plane based on the average point coordinates P_m(x_m, y_m, z_m); An error matrix H is constructed based on the error terms from each of the three-dimensional coordinates to the priori plane; Performing a singular value decomposition operation on the error matrix H to obtain a decomposition matrix H'; Get the last column of eigenvectors of the right singular vector matrix as plane parameter A, plane parameter B, and plane parameter C; Calculate the plane parameter D according to the three-dimensional coordinates and the priori plane; Acquire a target priori plane according to the plane parameter A, the plane parameter B, the plane parameter C and the plane parameter D; Each of the three-dimensional coordinates is vertically projected onto the target a priori plane to obtain the spatial coordinates of the mobile phone corner points corresponding to each of the three-dimensional coordinates.
5. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 4, characterized in that: The method of using the thumb corner point coordinates to process the thumb corner point spatial coordinates and thumb posture through triangulation, Kalman filter algorithm and ICP algorithm in sequence includes the following steps: Obtaining camera internal and external parameters; calculating initial thumb corner point three-dimensional coordinates through triangulation method according to the camera internal and external parameters and the thumb corner point coordinates; Acquire the frame thumb corner coordinates of continuous frames; calculate the frame thumb corner coordinates by triangulation method to obtain the frame thumb corner three-dimensional coordinates corresponding to each frame; predict and correct the frame thumb corner three-dimensional coordinates by Kalman filtering algorithm to obtain the corrected frame thumb corner three-dimensional coordinates; Based on the initial three-dimensional coordinates of the thumb corner points and the corrected three-dimensional coordinates of the frame thumb corner points, the thumb postures corresponding to the corrected three-dimensional coordinates of the frame thumb corner points are calculated by the ICP algorithm; The thumb gesture includes a rotation matrix R and a translation vector T.
6. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 5, characterized in that: The method of calculating the initial three-dimensional coordinates of the thumb corner points by triangulation according to the camera internal and external parameters and the thumb corner point coordinates includes the following steps: Obtain the first thumb corner point coordinates of the thumb corner point in N frames when the mobile phone screen surface is in a stationary state; The first thumb corner point coordinates are calculated by triangulation according to the internal and external parameters of the camera to obtain the first thumb corner point three-dimensional coordinates; the average coordinates of the first thumb corner point three-dimensional coordinates corresponding to each of the thumb corner point coordinates are calculated; and the average coordinates are determined as the initial thumb corner point three-dimensional coordinates.
7. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 6, characterized in that: The predicting and correcting operation of the frame thumb corner point three-dimensional coordinates by using a Kalman filter algorithm to obtain the corrected frame thumb corner point three-dimensional coordinates includes the following operation steps: Establishing a motion model for the three-dimensional coordinates of the thumb corner points in each frame; Traversing the three-dimensional coordinates of the thumb corner points of each frame, predicting the three-dimensional coordinates of the thumb corner points of the frame corresponding to the three-dimensional coordinates of the thumb corner points of the continuous frames and the motion model, and obtaining the predicted three-dimensional coordinates corresponding to the three-dimensional coordinates of the thumb corner points of the frame of the t-th frame; Correcting the predicted three-dimensional coordinates based on the three-dimensional coordinates of the thumb corner point in frame t-1 to obtain corrected three-dimensional coordinates of the thumb corner point; Repeat the above operation until all the three-dimensional coordinates of the thumb corner points of the frames are traversed.
8. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 7, characterized in that: The method of calculating the thumb postures corresponding to the corrected frame thumb corner point three-dimensional coordinates based on the initial thumb corner point three-dimensional coordinates and the corrected frame thumb corner point three-dimensional coordinates by using the ICP algorithm comprises the following steps: Calculate and obtain the initial centroid coordinates of each of the initial thumb corner point three-dimensional coordinates and the frame centroid coordinates of each of the corrected frame thumb corner point three-dimensional coordinates; The initial relative centroid coordinates are calculated based on the initial centroid coordinates and the three-dimensional coordinates of each initial thumb corner point; and the frame relative centroid coordinates are calculated based on the corrected three-dimensional coordinates of the thumb corner point and the frame centroid coordinates; The initial relative centroid coordinates are calculated as follows: Initial relative centroid coordinates = initial thumb corner three-dimensional coordinates - initial centroid coordinates; The frame relative centroid coordinates are calculated as follows: Frame relative centroid coordinates = corrected thumb corner three-dimensional coordinates - frame centroid coordinates; Normalizing each of the initial relative centroid coordinates to obtain an initial normalized point set; normalizing each of the frame relative centroid coordinates to obtain a frame normalized point set; Based on the initial normalized point set and the frame normalized point set, the thumb posture (R, T) corresponding to each corrected frame thumb corner point three-dimensional coordinate is calculated by the least square method; The thumb posture (R, T) is calculated as follows: Where R is the rotation matrix; T is the translation vector; argmin is the input value corresponding to the minimum value of the function; sp_i is the i-th initial relative centroid coordinate in the initial normalized point set; dp_i is the i-th frame relative centroid coordinate in the frame normalized point set.
9. The response mapping method of an AR glasses cooperative gesture operation terminal according to claim 8, characterized in that: The step of calculating the directional dimension ratio of the thumb projected onto the mobile phone screen based on the spatial coordinates of the thumb corner points and the spatial coordinates of the mobile phone corner points comprises the following steps: The spatial coordinates of each thumb corner point are respectively vertically projected onto the spatial intersection points in the target a priori plane; The spatial intersection points represent the intersection points of the spatial coordinates of each thumb corner point projected onto the target a priori plane and the target a priori plane; Calculate and obtain the left Euclidean distance dist_x of each of the spatial intersection points perpendicular to the left edge of the target a priori plane and the upper Euclidean distance dist_y of each of the spatial intersection points perpendicular to the upper edge of the target a priori plane; Calculating the length L of the target a priori plane and the width W of the target a priori plane based on the spatial coordinates of each of the mobile phone intersections; The ratio r_X in the X direction is calculated based on the left Euclidean distance dist_x and the length L of the target prior plane; the ratio r_Y in the Y direction is calculated based on the upper Euclidean distance dist_y and the width W of the target prior plane.
10. A response mapping device for AR glasses in collaboration with a gesture operation terminal, characterized in that: It includes an acquisition and calculation unit, a calculation unit and a fusion module; Wherein, the acquisition and calculation unit includes a binocular vision module and a calculation module; The binocular vision module is used to obtain a binocular view mobile phone image containing a complete mobile phone outline; The calculation module is used to obtain the coordinates of the corner points of the mobile phone after corner point detection of the collected binocular view mobile phone image, and then use the coordinates of the mobile phone corner points to obtain the spatial coordinates of the mobile phone corner points by triangulation and plane fitting; The binocular view mobile phone image includes a left-view mobile phone image and a right-view mobile phone image; The binocular vision module is also used to obtain a thumb image containing a thumb picture; The calculation module is further used to obtain the coordinates of the thumb corner points after corner point detection of the collected thumb image, and then use the coordinates of the thumb corner points to process through triangulation, Kalman filter algorithm and ICP algorithm in sequence to obtain the spatial coordinates of the thumb corner points and the thumb posture; The thumb image includes a left-eye view thumb image and a right-eye view thumb image; The calculation unit is used to calculate the directional size ratio of the thumb projected onto the mobile phone screen based on the spatial coordinates of the thumb corner points and the spatial coordinates of the mobile phone corner points; The directional dimension ratio includes an X-direction dimension ratio r_X and a Y-direction dimension ratio r_Y; the fusion module is used to transmit the binocular view mobile phone image to the AR device to obtain an AR display image; based on the directional dimension ratio and the thumb gesture, the thumb image is fused with the AR display image to obtain a fused image.
Citation Information
Patent Citations
Hand function training system based on mixed reality, and data processing method
CN110211661A
Mobile terminal and method of controlling operation thereof
US20120302289A1