Finger Pointing Coordinate Calibration Method and Device
A depth learning model-based method corrects finger pointing coordinates using a gaze-based reference system to address misalignment issues, ensuring accurate target identification in industrial and smart home applications.
Patent Information
- Application Number
- CN202510248025.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Due to the offset of the installation position of the device with camera function or human characteristics, the finger pointing may not be correctly identified by the system, resulting in an error in the operation.
The gaze reference coordinate system is constructed with the center of the human eye as the origin, and the θ and φ deviation values of finger pointing are corrected through the deep learning model, and the head posture coordinates and the spherical coordinates pointed by the finger pointing are used to achieve positioning correction of finger pointing.
Correctly identifying the intended position of the finger pointing, solving the problem of finger pointing coordinate recognition errors caused by equipment installation position offset or human characteristics, and improving the accuracy of the operation.
Smart Images

Figure CN119722537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spatial positioning, and more particularly to a method and device for correcting finger pointing coordinates. Background Art
[0002] Currently, in application scenarios such as industry and smart home, a device with a camera function is usually used to locate the current position of the operator, and then the finger pointing of the operator can be located to know the target object to be operated by the operator, so as to execute corresponding operations according to operation instructions. For example, in the field of smart home, in a modeled room, the target member is identified and located. By the finger pointing of the target member and supplemented with an instruction of "open the curtain", the curtain in the correct room can be opened. Or, the member can point the finger from a distance and supplement it with an instruction of "turn on this light", and the home system can turn on the light pointed by the member.
[0003] However, due to reasons such as the installation position deviation of the device with a camera function and human body characteristics (left and right eyes, fingers, etc.), the finger pointing of the user may not be recognized by the system to the correct intended position, which is not conducive to the execution of subsequent operations. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device for correcting finger pointing coordinates, so as to achieve the purpose of positioning and correcting finger pointing, and thus correctly identify the correct intended position of finger pointing.
[0005] In a first aspect, an embodiment of the present invention provides a method for correcting finger pointing coordinates, which includes: obtaining the attitude coordinates of the head under the current view and the spherical coordinates Q(r, θ, φ) of the current finger pointing in the line-of-sight reference coordinate system corresponding to the current view; wherein, the line-of-sight reference coordinate system is a coordinate system constructed with the center of the two eyes of the human body as the origin, the line-of-sight center direction as the x-axis, the direction perpendicular to the line connecting the two eyes as the y-axis, and the direction parallel to the line connecting the two eyes as the z-axis, and r is the distance from the finger to the origin; obtaining the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing according to the obtained attitude coordinates by using a pre-trained deep learning model corresponding to the attitude coordinates, so as to obtain the actual coordinates of the current finger pointing; wherein, the pre-trained deep learning model is trained based on the actual coordinates (θ, φ) of the positioning points selected in the view visible at different attitude coordinates in the line-of-sight reference coordinate system and the finger spherical coordinate points q(r, θ, φ) obtained by the positioning points corresponding to the finger pointing under the corresponding attitude coordinates, and is used to output the deviation values δ and ζ of θ and φ in the obtained finger spherical coordinate points from the actual coordinates (θ, φ) of the corresponding positioning points.
[0006] In a second aspect, an embodiment of the present invention further provides a device for correcting finger pointing coordinates, including units for executing the above method.
[0007] Compared with the prior art, the present invention constructs a line-of-sight reference coordinate system based on the current perspective of the user. That is, taking the center of the two eyes of the human body as the origin, the direction of the line-of-sight center as the x-axis, the direction perpendicular to the line connecting the two eyes as the y-axis, and the direction parallel to the line connecting the two eyes as the z-axis to construct the line-of-sight reference coordinate system. The attitude coordinates of the head under the current perspective and the spherical coordinates Q(r, θ, φ) of the current finger pointing under the line-of-sight reference coordinate system corresponding to the current perspective are obtained. According to the obtained attitude coordinates, the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing are obtained by using the actual coordinates (θ, φ) of the positioning points in the picture visible from the perspective corresponding to the attitude coordinates in the line-of-sight reference coordinate system and the finger spherical coordinate point q(r, θ, φ) obtained by pointing the finger at the positioning point, and are used as correction parameters, so as to obtain the actual coordinates of the current finger pointing, realize the positioning correction of the finger pointing, correctly identify the intended position of the finger pointing, and solve the problem that the finger pointing coordinates of the user cannot be correctly identified due to reasons such as the deviation of the installation position of the device with a camera function or human body characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a schematic flow chart of the finger pointing coordinate correction method provided by an embodiment of the present invention;
[0009] Figure 2 It is a schematic sub-flow chart of the finger pointing coordinate correction method provided by an embodiment of the present invention;
[0010] Figure 3 It is a schematic block diagram of the sorting finger pointing coordinate correction device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0013] It should also be understood that the term " / and" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0014] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of the finger pointing coordinate correction method provided by an embodiment of the present invention. As shown in the figure, the method includes the following steps S110-120.
[0015] S110. Obtain the attitude coordinates of the head under the current perspective and the spherical coordinates Q(r, θ, φ) of the current finger pointing in the line-of-sight reference coordinate system corresponding to the current perspective.
[0016] In the present invention, the line-of-sight reference coordinate system is a coordinate system constructed with the center of the two eyes of the human body as the origin, the line-of-sight center direction as the x-axis, the direction perpendicular to the line connecting the two eyes as the y-axis, and the direction parallel to the line connecting the two eyes as the z-axis. r is the distance from the finger to the origin. It can be understood that the line-of-sight reference coordinate system is a dynamic coordinate system established based on the user's current perspective, and its origin and axis directions will change with the movement of the head, that is, they will change with different attitude coordinates of the head. r represents the distance from the finger to the origin, θ is the angle between the projection on the xz plane and the z-axis, and φ is the angle between the projection on the xy plane and the x-axis. θ and φ together determine the direction of the finger relative to the line of sight. Preferably, in this embodiment, r can be defined as the distance from the tip of the middle finger to the origin.
[0017] In this step, the attitude coordinates of the head under the current perspective and the spherical coordinates corresponding to the current finger pointing in the current perspective can be obtained through a camera device (such as a head-mounted camera device or a camera device fixed in the scene, etc.). The attitude coordinates include 3 Euler angles (pitch angle, yaw angle, and roll angle) measured according to the current line-of-sight direction.
[0018] S120. Use the pre-trained deep learning model corresponding to the obtained attitude coordinates to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing, so as to obtain the actual coordinates of the current finger pointing; wherein, the pre-trained deep learning model is trained based on the actual coordinates (θ, φ) of the positioning points selected in the visible view of different attitude coordinates in the line-of-sight reference coordinate system and the finger spherical coordinate points q(r, θ, φ) obtained by the finger pointing corresponding to the positioning points under the corresponding attitude coordinates, and is used to output the deviation values δ and ζ of θ and φ in the obtained finger spherical coordinate points from the actual coordinates (θ, φ) of the corresponding positioning points.
[0019] As Figure 2 shown, this step specifically includes:
[0020] S121. According to the obtained attitude coordinates, search whether there is a pre-trained deep learning model in the database that uses the same attitude coordinates as the obtained attitude coordinates. If not, execute steps S122-S124; if so, execute step S125;
[0021] S122. Select multiple pre-trained deep learning models according to a preset threshold of pose coordinates, and finally select multiple deep learning models with a distance less than a preset distance threshold from the multiple pre-trained deep learning models according to the spherical coordinates Q(r, θ, φ) pointed by the current finger.
[0022] In this step, multiple deep learning models trained from the perspective of pose coordinates similar to the obtained pose coordinates are selected according to the preset threshold of pose coordinates. Then, according to θ and φ of the spherical coordinates pointed by the finger, multiple pre-trained deep learning models with the closest distance between the training coordinates and the spherical coordinates are found, that is, multiple deep learning models with the distance (the sum of the squared distances of θ and φ) between the training coordinates and the spherical coordinates within the preset distance threshold are found. It can be understood that the preset threshold of pose coordinates and the preset distance threshold can be set according to actual needs;
[0023] S123. Input the spherical coordinates Q(r, θ, φ) pointed by the current finger into multiple finally selected deep learning models respectively to calculate the corresponding deviation values δ and ζ of θ and φ;
[0024] S124. Perform comprehensive calculation on the deviation values δ and ζ calculated by multiple finally selected deep learning models according to the distance, and finally obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ).
[0025] Specifically, in this step, weights are attached to each finally selected deep learning model according to the distance, and the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) are calculated by using a weighted algorithm based on the weights of each deep learning model and the deviation values δ and ζ. That is, the weighted coefficient of each finally selected deep learning model is calculated according to the distance. The closer the distance, the larger the weighted coefficient. Then, the weighted coefficients of each deep learning model are multiplied by the corresponding δ and ζ calculated by the deep learning model and added together to obtain the angular deviation values of θ and φ in the spherical coordinates Q(r, θ, φ).
[0026] S125. Obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) pointed by the current finger by using the searched pre-trained deep learning model.
[0027] In some embodiments, the deep learning model can be trained by the following steps:
[0028] (1) Obtain the viewable image corresponding to the current pose coordinates captured by the imaging device, select multiple positioning points from the image, and obtain the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system, forming an actual coordinate training set and inputting it into the deep learning model to be trained. Specifically, when the imaging device is capturing, the user keeps the position below the neck of the body stationary and turns the head to look at the target to be measured, so as to capture the viewable image of the current view. And the process of selecting multiple positioning points from the image and obtaining the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system includes: selecting multiple positioning points from the image, and within the preset range of each positioning point, selecting multiple corresponding environmental markers (such as the corners, intersections of the lines in the modeled environmental space, etc.), and obtaining the coordinates (θ, φ) of the centers of the environmental markers corresponding to all the positioning points in the current line-of-sight reference coordinate system, so as to obtain the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system, forming an actual coordinate training set. Preferably, the number of the selected positioning points can be five or more, and the selected multiple positioning points can be evenly distributed around the perimeter and the center of the image. For example, the number of the positioning points can be 9, and the 9 positioning points form a 3*3 dot matrix. In the line-of-sight reference coordinate system, the [θ, φ] values of these 9 points from left to right and from top to bottom in front of the eyes can be [3π / 8, π / 8], [π / 2, π / 8], [5π / 8, π / 8], [3π / 8, 0], [π / 2, 0], [5π / 8, 0], [3π / 8, -π / 8], [π / 2, -π / 8], [5π / 8, -π / 8], that is, the 5th point is at the center of the field of view, and the other points are offset by an angle of π / 8 in the up, down, left, and right directions. Understandably, the coordinates (i.e., spherical coordinates) of the center of the environmental marker in the current line-of-sight reference coordinate system can be converted from the world coordinates of the center of the environmental marker in the modeled environmental space;
[0029] (2)Obtain the finger spherical coordinate points q(r, θ, φ) corresponding to different distances r between the fingertip and the origin of the current line-of-sight reference coordinate system when the finger points to each fixed point under the current perspective, and form a pointing coordinate training set and input it into the deep learning model to be trained. In this step, the user can use the middle finger or other fingers of the dominant hand to point to the environmental markers corresponding to the positioning points. When pointing, the arm and finger should be straightened, and the eyes can be opened during pointing. Or the user can use the dominant eye for positioning (left eye or right eye), as long as it is consistent during subsequent calibration. When obtaining the finger spherical coordinate points q(r, θ, φ) corresponding to different distances r when the finger points to each fixed point under the current perspective, the user slowly bends the arm while keeping the finger pointing to the environmental marker corresponding to the positioning point, gradually brings the finger closer to the head, and stops when it is less than 15 cm away from the head. During the process, keep the finger straight and always point to the environmental marker for the camera device to obtain the coordinates θ and φ of the fingertip at different distances from the origin. It can be understood that during the above process, the user should not have other body (including the head) movements except for the arm and eyes, and when looking at the environmental marker with the eyes, the head should not move;
[0030] (3)The deep learning model calculates the deviation values δ and ζ between the θ and φ of each finger spherical coordinate point at different distances r obtained and the coordinates (θ, φ) of the corresponding positioning point in the current line-of-sight reference coordinate system respectively, and fits the functional relationships between r and δ, ζ at this positioning point according to r and the calculated δ, ζ at each positioning point. In the present invention, the functional relationships between r and δ, ζ at this positioning point can be respectively fitted using the polynomial regression algorithm according to r and the calculated δ, ζ at each positioning point. Specifically, in this embodiment, the functional relationship between r and δ at this positioning point can be fitted using the formula wherein, , a is the regression coefficient; the functional relationship between r and ζ at this positioning point is fitted using the formula wherein, , b is the regression coefficient. That is, the training sample set can be composed of r corresponding to pointing to the environmental markers of multiple positioning points and the obtained δ, ζ, and the functional relationships between r and δ, ζ at the positioning point are respectively fitted using the polynomial regression algorithm. Preferably, the least squares method can also be used to solve for the regression coefficients a and b. It can be understood that any deep learning model capable of learning the mapping relationship between input data and angular deviation can be used as the initial model of the present invention, that is, the deep learning model to be trained;
[0031] (4) Re-enter all the finger sphere coordinate points q(r, θ, φ) of the multiple positioning points and the obtained δ and ζ into the above-mentioned trained deep learning model for training. The above steps can obtain the calibration parameters δ and ζ when the finger points to the environmental markers corresponding to all the positioning points. In this training step, re-entering all the finger sphere coordinate points q(r, θ, φ) and the obtained δ and ζ into the above-mentioned trained deep learning model for training can obtain a deep learning model for calculating the θ and φ deviations δ and ζ of the actual pointing coordinates of the finger in any case of r, θ, φ.
[0032] Further, in some embodiments, all the pose coordinates and the corresponding obtained spherical coordinates Q(r, θ, φ) of the current finger pointing and the corresponding deviation values are input into a pre-trained deep learning model for training. Understandably, the spherical coordinates Q(r, θ, φ) of the current finger pointing obtained by the deep learning model corresponding to all the pose coordinates and the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) can be jointly trained to obtain a unified deep learning model. The input of this deep learning model simultaneously considers the influences of finger pointing and head pose and directly outputs the angular calibration parameters δ and ζ of θ and φ. Thereafter, the newly added calibration data will be directly used for the incremental training of this unified model to continuously optimize the model performance.
[0033] In summary, the present invention constructs a line-of-sight reference coordinate system based on the current perspective of the user, obtains the pose coordinates of the head under the current perspective and the spherical coordinates Q(r, θ, φ) of the current finger pointing in the line-of-sight reference coordinate system corresponding to the current perspective. According to the obtained pose coordinates, using the actual coordinates (θ, φ) of the positioning points in the line-of-sight reference coordinate system that are visible in the picture based on the perspective corresponding to the pose coordinates and the finger sphere coordinate points q(r, θ, φ) obtained by the finger pointing to the positioning points, the pre-trained deep learning model obtains the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing as calibration parameters, so as to obtain the actual coordinates of the current finger pointing, realize the positioning correction of the finger pointing, correctly identify the intended position of the finger pointing, and solve the problem that the finger pointing coordinates of the user cannot be correctly identified currently due to reasons such as the installation position deviation of the device with a camera function or human body characteristics, that is, the deviation influences caused by the device and user characteristics (left and right eyes, fingers, etc.) can be eliminated. And in the subsequent use process, the user can continue to complete more finger pointing positioning corrections, making the model more accurate and more in line with the physical characteristics of the user.
[0034] Figure 3 is a schematic block diagram of a finger pointing coordinate correction device 300 provided by an embodiment of the present invention. As Figure 3As shown, corresponding to the finger-pointing coordinate correction method of the above embodiments, the present invention further provides a finger-pointing coordinate correction device 300. The finger-pointing coordinate correction device 300 can be configured in a server. Specifically, please refer to Figure 3 , the finger-pointing coordinate correction device 300 includes a training information acquisition unit 301, a model training and generation unit 302, a correction information acquisition unit 303, and a correction information calculation unit 304.
[0035] The training information acquisition unit 301 is used to acquire the pose coordinates of the head at different viewing angles, and acquire the corresponding visible images at different pose coordinates captured by the imaging device. It is also used to acquire the finger sphere coordinate points q(r, θ, φ) corresponding to different distances r between the fingertip and the origin of the current line-of-sight reference coordinate system when the finger points to each selected fixed point in the image at each viewing angle, so as to form a pointing coordinate training set; wherein, the line-of-sight reference coordinate system is a coordinate system constructed with the center of the two eyes of the human body as the origin, the line-of-sight center direction as the x-axis, the direction perpendicular to the line connecting the two eyes as the y-axis, and the direction parallel to the line connecting the two eyes as the z-axis, and r is the distance from the finger to the origin.
[0036] The model training and generation unit 302 is used to select multiple fixed points from each acquired image, and acquire the coordinates (θ, φ) of the multiple fixed points in the current line-of-sight reference coordinate system to form an actual coordinate training set. Based on the actual coordinate training sets corresponding to multiple viewing angles and the corresponding pointing coordinate training sets, a deep learning model for outputting the deviation values δ and ζ between θ and φ of each finger sphere coordinate point at different distances r and the coordinates (θ, φ) of the corresponding fixed point in the current line-of-sight reference coordinate system is trained. The deep learning model can respectively fit the functional relationship between r and δ, ζ for each fixed point according to r and the calculated δ, ζ in each fixed point; and each deep learning model is also trained based on all the finger sphere coordinate points q(r, θ, φ) of multiple fixed points in its corresponding viewing angle and the obtained δ, ζ. Preferably, in this embodiment, the step of selecting multiple fixed points from each acquired image and acquiring the coordinates (θ, φ) of the multiple fixed points in the current line-of-sight reference coordinate system to form an actual coordinate training set includes: selecting multiple fixed points from the image, and selecting multiple corresponding environmental markers within the preset range of each fixed point, and acquiring the coordinates (θ, φ) of the centers of the environmental markers corresponding to all fixed points in the current line-of-sight reference coordinate system to obtain the coordinates (θ, φ) of the multiple fixed points in the current line-of-sight reference coordinate system, so as to form an actual coordinate training set. And in this embodiment, the polynomial regression algorithm can be used to respectively fit the functional relationship between r and δ, ζ for each fixed point according to r and the calculated δ, ζ in each fixed point.
[0037] The calibration information acquisition unit 303 is configured to acquire the attitude coordinates of the head in the current perspective and the spherical coordinates Q(r, θ, φ) of the current finger pointing in the line-of-sight reference coordinate system corresponding to the current perspective.
[0038] The calibration information calculation unit 304 is configured to use the deep learning model corresponding to the acquired attitude coordinates generated by the model training and generation unit 302 to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing as calibration parameters, so as to obtain the actual coordinates of the current finger pointing. Specifically, in this embodiment, it is possible to search the database according to the acquired attitude coordinates to check whether there is a pre-trained deep learning model using the same attitude coordinates as the acquired attitude coordinates; if so, use the searched pre-trained deep learning model to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger pointing; if not, select multiple pre-trained deep learning models according to the attitude coordinate preset threshold, and finally select multiple deep learning models whose distance from the spherical coordinates Q(r, θ, φ) of the current finger pointing is less than the preset distance threshold from the multiple pre-trained deep learning models; input the spherical coordinates Q(r, θ, φ) of the current finger pointing into the multiple finally selected deep learning models respectively to calculate the corresponding deviation values δ and ζ of θ and φ; perform comprehensive calculation on the deviation values δ and ζ calculated by the multiple finally selected deep learning models according to the distance, and finally obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ), that is, attach weights to each finally selected deep learning model according to the distance, and calculate the angular deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) using the weighted algorithm according to the weights of each deep learning model and the deviation values δ and ζ.
[0039] Furthermore, the model training and generation unit 302 can also be configured to uniformly train the pre-trained deep learning model based on all the attitude coordinates acquired by the calibration information acquisition unit 303, the corresponding spherical coordinates Q(r, θ, φ) of all the current finger pointings, and the deviation values calculated by the calibration information calculation unit 304, so as to obtain a unified model that is compatible with the head attitude and finger pointing.
[0040] As can be seen from the above, the finger pointing coordinate calibration device 300 of the present invention can train a deep learning model based on the actual coordinates (θ, φ) of the positioning point in the line-of-sight reference coordinate system in the viewable screen captured and the finger spherical coordinate point q(r, θ, φ) obtained by the finger pointing to the positioning point, so as to obtain a model for outputting the angle deviation values δ and ζ according to any (r, θ, φ), thereby being applicable to the calibration of the position coordinates of any finger pointing, realizing the positioning calibration of the finger pointing, and correctly identifying the intended position of the finger pointing.
[0041] It should be noted that those skilled in the art can clearly understand the specific implementation processes of the above-mentioned finger-pointing coordinate correction device 300 and each unit. For corresponding descriptions, reference can be made to the foregoing method embodiments. For the convenience and brevity of description, they will not be repeated here.
[0042] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0043] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0044] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A finger-pointing coordinate correction method, characterized in that The finger-pointing coordinate correction method includes: Obtaining the attitude coordinates of the head under the current perspective and the spherical coordinates Q(r, θ, φ) of the current finger-pointing in the line-of-sight reference coordinate system corresponding to the current perspective; wherein, the line-of-sight reference coordinate system is a coordinate system constructed with the center of the two eyes of the human body as the origin, the line-of-sight center direction as the x-axis, the direction perpendicular to the line connecting the two eyes as the y-axis, and the direction parallel to the line connecting the two eyes as the z-axis, and r is the distance from the finger to the origin. Using the obtained attitude coordinates and the pre-trained deep learning model corresponding to the attitude coordinates to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger-pointing, so as to obtain the actual coordinates of the current finger-pointing; wherein, the pre-trained deep learning model is trained based on the actual coordinates (θ, φ) of the positioning points selected in the perspective-visible images taken with different attitude coordinates in the line-of-sight reference coordinate system and the finger spherical coordinate points q(r, θ, φ) obtained by the finger-pointing corresponding to the positioning points under the corresponding attitude coordinates, and is used to output the deviation values δ and ζ of θ and φ in the obtained finger spherical coordinate points from the actual coordinates (θ, φ) of the corresponding positioning points; wherein, The step of using the obtained attitude coordinates and the pre-trained deep learning model corresponding to the attitude coordinates to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger-pointing specifically includes: Searching the database according to the obtained attitude coordinates to check if there is a pre-trained deep learning model pre-trained with the same attitude coordinates as the obtained attitude coordinates. If not, selecting multiple pre-trained deep learning models according to the attitude coordinate preset threshold, and finally selecting multiple deep learning models with a distance less than the preset distance threshold from the multiple pre-trained deep learning models according to the spherical coordinates Q(r, θ, φ) of the current finger-pointing. Inputting the spherical coordinates Q(r, θ, φ) of the current finger-pointing into the multiple finally selected deep learning models respectively to calculate the corresponding deviation values δ and ζ of θ and φ. Performing comprehensive calculation on the deviation values δ and ζ calculated by the multiple finally selected deep learning models according to the distance, and finally obtaining the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ).
2. The finger pointing coordinate calibration method according to claim 1, wherein, The step of performing comprehensive calculation on the deviation values δ and ζ calculated by the multiple finally selected deep learning models according to the distance and finally obtaining the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) specifically includes: Attaching weights to each of the finally selected deep learning models according to the distance, and calculating the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) using a weighted algorithm based on the weights and deviation values δ and ζ of each deep learning model.
3. The finger-pointing coordinate correction method according to claim 1, characterized in that After searching the database according to the obtained attitude coordinates to check if there is a pre-trained deep learning model pre-trained with the same attitude coordinates as the obtained attitude coordinates, the following steps are further included: If so, using the searched pre-trained deep learning model to obtain the deviation values of θ and φ in the spherical coordinates Q(r, θ, φ) of the current finger-pointing.
4. The finger-pointing coordinate calibration method according to claim 1, characterized in that The deep learning model is trained using the following steps: Obtain the views visible in the images corresponding to different pose coordinates captured by the imaging device, select multiple positioning points from the images, and obtain the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system, forming an actual coordinate training set and inputting it into the deep learning model to be trained; Obtain the finger spherical coordinate points q(r, θ, φ) corresponding to different distances r between the fingertip and the origin of the current line-of-sight reference coordinate system when the finger points to each positioning point in the current view, forming a pointing coordinate training set and inputting it into the deep learning model to be trained; The deep learning model respectively calculates the deviation values δ and ζ between the θ and φ of each finger spherical coordinate point at different distances r obtained and the coordinates (θ, φ) of the corresponding positioning point in the current line-of-sight reference coordinate system, and respectively fits the functional relationship between r and δ, ζ for each positioning point according to r and the calculated δ, ζ in each positioning point; Re-input all the finger spherical coordinate points q(r, θ, φ) of the multiple positioning points and the obtained δ, ζ into the above-trained deep learning model for training.
5. The finger-pointing coordinate correction method according to claim 4, wherein The step of selecting multiple positioning points from the images and obtaining the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system specifically includes: Select multiple positioning points from the images, and select multiple corresponding environmental markers within the preset range of each positioning point, obtain the coordinates (θ, φ) of the centers of the environmental markers corresponding to all the positioning points in the current line-of-sight reference coordinate system, so as to obtain the coordinates (θ, φ) of the multiple positioning points in the current line-of-sight reference coordinate system, forming an actual coordinate training set.
6. The finger-pointing coordinate correction method according to claim 4, characterized in that, After the step of re-inputting all the finger spherical coordinate points q(r, θ, φ) of the multiple positioning points and the obtained δ, ζ into the above-trained deep learning model for training, it further includes: Input all the pose coordinates and the corresponding obtained spherical coordinates Q(r, θ, φ) and deviation values of all current finger points into the pre-trained deep learning model for training.
7. The finger-pointing coordinate calibration method according to claim 4, wherein The step of respectively fitting the functional relationship between r and δ, ζ for each positioning point according to r and the calculated δ, ζ in each positioning point specifically includes: Use the polynomial regression algorithm to respectively fit the functional relationship between r and δ, ζ for each positioning point according to r and the calculated δ, ζ in each positioning point.
8. The finger pointing coordinate calibration method according to claim 7, characterized in that The step of using the polynomial regression algorithm to respectively fit the functional relationship between r and δ, ζ for each positioning point according to r and the calculated δ, ζ in each positioning point specifically includes: Based on the δ calculated from r at each locus, use the formula to fit the functional relationship between r and δ at this locus, where , and a is the regression coefficient; According to ζ obtained by calculating r and in each locus, use the formula to fit the functional relationship between r and ζ at this locus, where , and b is the regression coefficient.
9. A finger-pointing coordinate correction device, characterized in that, The finger pointing coordinate correction device includes a unit for executing the method according to any one of claims 1-8.
Citation Information
Patent Citations
Method and device for training sight line deviation estimation model and method and device for correcting sight line deviation value
CN117132869A