Input method and device, near-to-eye display equipment and storage medium
By identifying palm key points on the near-eye display device and projecting the input interface, the problem of low input efficiency in the prior art is solved, and more efficient input operations are achieved.
Patent Information
- Application Number
- CN202411987702.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the input efficiency of the near-eye display device is low, especially when identifying key points on both hands, the processing workload is large and the efficiency is not high.
By determining the first area corresponding to the palm of the palm or back of the palm of the wearer of the close-eye display device, the palm key point detection is carried out, multiple palm key points are determined, and the second area projected on the palm or back of the palm is determined based on these key points. Based on the close-eye display device projecting the input interface in the second area, the wearer's input operation on the input interface is detected to determine the input information.
This method reduces the amount of calculation and improves the input efficiency of the near-eye display device, and only the key points of the palm part are needed to be identified rather than all key points of both hands.
Smart Images

Figure CN120010657A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of display technology, and in particular to an input method, apparatus, near-eye display device and storage medium. Background Art
[0002] Nowadays, near-eye display devices such as AR (Augmented Reality) glasses and VR (Virtual Reality) glasses have been applied to various scenarios. For example, AR glasses can be used to identify key points of the user's hands, so that a virtual keyboard can be projected on the user's left hand, and the user can perform corresponding input operations based on the virtual keyboard. In traditional key point recognition technology, a user's hand usually includes 21 key points. Therefore, key point recognition of both hands requires approximately 42 key point recognitions, which is a large processing workload and low efficiency. Summary of the invention
[0003] The present application provides an input method, an apparatus, a near-eye display device, and a storage medium, aiming to improve the input efficiency based on the near-eye display device.
[0004] To achieve the above objectives, the present application provides an input method based on a near-eye display device, the input method based on a near-eye display device comprising:
[0005] Determine a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device;
[0006] Perform palm key point detection in the first area to determine a plurality of palm key points;
[0007] Determine, according to the plurality of palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm;
[0008] Projecting the input interface in the second area based on the near-eye display device;
[0009] An input operation of the wearer on the input interface is detected to determine input information of the wearer according to the input operation.
[0010] In addition, to achieve the above-mentioned purpose, the present application also provides an input device, the input device comprising a memory and a processor;
[0011] The memory is used to store computer programs;
[0012] The processor is used to execute the computer program and implement the steps of the above-mentioned input method based on the near-eye display device when executing the computer program.
[0013] In addition, to achieve the above-mentioned purpose, the present application also provides a near-eye display device, which includes the input device as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned input method based on a near-eye display device.
[0015] The present application discloses an input method, apparatus, near-eye display device and storage medium. In the above embodiment, by determining a first area corresponding to the palm or back of a palm of a wearer of the near-eye display device, palm key point detection is performed in the first area to determine multiple palm key points, and based on the multiple palm key points, a second area corresponding to the projection of an input interface on the palm or back is determined, based on the near-eye display device projecting the input interface in the second area, the wearer's input operation on the input interface is detected to determine the wearer's input information based on the input operation. The entire process only needs to identify the palm key points corresponding to the palm part, and there is no need to perform key point recognition on both hands, which reduces the amount of calculation, thereby improving the input efficiency based on the near-eye display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 It is a schematic diagram of the key points of both hands;
[0018] Figure 2 is a schematic flow chart of the steps of an input method based on a near-eye display device provided in an embodiment of the present application;
[0019] Figure 3 is a schematic flow chart of the steps of another input method based on a near-eye display device provided in an embodiment of the present application;
[0020] Figure 4 is a schematic flow chart of steps for determining a first area corresponding to the palm or back of a palm of a wearer of a near-eye display device provided by an embodiment of the present application;
[0021] Figure 5 is a schematic diagram of the first region in an embodiment of the present application;
[0022] Figure 6 is a schematic diagram of the second region in an embodiment of the present application;
[0023] Figure 7 is a schematic diagram of fingertip points in an embodiment of the present application;
[0024] Figure 8 It is a schematic flow chart of the steps of determining a second area corresponding to the projection of the input interface on the palm or the back of the palm according to a plurality of palm key points provided in an embodiment of the present application;
[0025] Fig. 9 is a schematic diagram of a virtual keyboard interface display provided by an embodiment of the present application;
[0026] Fig.10 is a schematic flow chart of the steps of detecting an input operation of the wearer on the input interface to determine the input information of the wearer according to the input operation, provided by an embodiment of the present application;
[0027] Fig.11 is a schematic diagram of a flow chart of input operation through AR glasses provided in an embodiment of the present application;
[0028] Fig.12 It is a schematic block diagram of an input device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0031] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0032] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0033] Nowadays, when a user wears AR glasses, the AR glasses can recognize the key points of the user's hands, and then project a virtual keyboard on the user's left hand, and the user can perform corresponding input operations based on the virtual keyboard. In traditional key point recognition technology, a user's hand usually includes 21 key points. Therefore, key point recognition for both hands requires approximately 42 key point recognitions, for example, Figure 1 As shown, Figure 1 The red dots in the figure are key points, which requires a lot of processing work and is not efficient.
[0034] In order to solve the above problems, the embodiments of the present application provide an input method, an apparatus, a near-eye display device and a storage medium for improving the input efficiency based on the near-eye display device.
[0035] See also Figure 2 , Figure 2 This is a flowchart of an input method based on a near-eye display device provided in an embodiment of the present application. The method can be applied to an input device or a near-eye display device, and the application scenario of the method is not limited in the present application. Among them, the near-eye display device includes but is not limited to AR glasses, VR (Virtual Reality) glasses and other devices.
[0036] like Figure 2 As shown, the input method based on the near-eye display device specifically includes steps S101 to S105.
[0037] S101. Determine a first area corresponding to the palm or the back of a palm of a wearer of a near-eye display device.
[0038] For example, assuming that the right hand of the wearer of the near-eye display device is the dominant hand, when the wearer of the near-eye display device needs to perform an input operation, the right hand can be used to perform the input operation, and the left hand is used to display the virtual input interface (such as a virtual keyboard interface) projected by the near-eye display device, and the wearer of the near-eye display device performs the corresponding input operation on the input interface. In this case, in order to project the input interface on the left hand of the wearer of the near-eye display device, first, determine the first area corresponding to the palm of the wearer's left hand, or determine the first area corresponding to the back of the wearer's left hand.
[0039] Exemplarily, the near-eye display device is provided with multiple input modes, such as a first input mode, a second input mode, etc., wherein the input interface is projected in different ways in different input modes. The first input mode refers to the palm / back of palm input mode. In the first input mode, the input interface is projected on the palm or back of the wearer. If the wearer's hand moves, the input interface will also move accordingly. The second input mode refers to the normal input mode. In the second input mode, the input interface is projected in a pre-set manner, such as projected directly in front of the wearer. The current input mode of the near-eye display device can be flexibly set according to actual conditions, thereby improving the user experience.
[0040] In some embodiments, Figure 3 As shown, step S106 may be included before step S101, and step S101 may include sub-step S1011.
[0041] S106, controlling the near-eye display device to enter a first input mode based on a preset instruction;
[0042] S1011. After the near-eye display device enters the first input mode, determine the first area.
[0043] Among them, the preset instruction is a control instruction for starting the first input mode of the near-eye display device. The preset instruction can be triggered by the user performing a control operation, or the preset instruction can also be actively triggered by the near-eye display device. Only when the near-eye display device enters the first input mode, the operation of determining the first area corresponding to the palm of the wearer's left hand, or determining the first area corresponding to the back of the wearer's left hand, is performed to realize the projection of the input interface on the palm or back of the wearer's palm, thereby avoiding useless operations, saving energy consumption of the near-eye display device, and making the near-eye display device more intelligent.
[0044] In some embodiments, the input method based on the near-eye display device further includes: identifying gestures of the wearer through a camera device of the near-eye display device, and triggering the preset instruction if a preset gesture is recognized.
[0045] Among them, the camera device can be a camera that comes with the near-eye display device. Of course, it can be understood that the camera device can also be other types of devices such as a camera mounted on the near-eye display device, and no specific limitation is made in this application.
[0046] When the wearer of the near-eye display device needs to perform an input operation, the wearer performs a corresponding control operation to trigger a control instruction, and when the near-eye display device receives the control instruction, the camera device is turned on.
[0047] For example, taking the near-eye display device as AR glasses, the wearer of the AR glasses inputs a voice command such as "turn on the camera device". When the wearer receives the voice command, the camera device is turned on.
[0048] Exemplarily, the AR glasses are provided with touch keys, for example, a touch bar is provided at the temple of the AR glasses, and when the wearer of the AR glasses presses the touch bar, the camera device is controlled to be turned on.
[0049] It should be noted that the method of controlling the start of the camera device is not limited to the examples listed above, and no specific limitation is made in this application.
[0050] After the camera device is turned on, the near-eye display device recognizes the wearer's gestures through the camera device. If the wearer's gestures are recognized to match the preset gestures, the near-eye display device triggers the preset instructions to control the near-eye display device to enter the first input mode.
[0051] For example, assuming that the preset gesture action is two consecutive pinching actions, if the wearer's two consecutive pinching gesture actions are recognized, that is, the wearer has performed the preset gesture action, then the preset instruction is triggered to control the near-eye display device to enter the first input mode.
[0052] It is understandable that the preset gesture action may also be other gesture actions, which is not specifically limited in this application.
[0053] In some embodiments, Figure 4 As shown, step S101 may include sub-step S1012 and sub-step S1013.
[0054] S1012, determining at least two target points on a palm of the wearer;
[0055] S1013: Determine the first area according to at least two of the target points.
[0056] Among them, at least two target points include but are not limited to the wearer's wrist point, the base point of the middle finger, etc.
[0057] Exemplarily, after the near-eye display device enters the first input mode, the wearer's hand is detected through the camera device. For example, the near-eye display device captures an image through the camera device, and performs image recognition detection on the captured image to detect whether there is a hand in the image. If no hand is detected, the detection operation continues. If a hand is detected, gesture detection is further performed to determine at least two target points of the palm, for example, the wrist point and the base point of the middle finger of the wearer are determined.
[0058] Then, a first region is determined according to each target point, wherein each target point is included in the first region. Exemplarily, the maximum circumscribed rectangle of each target point is taken, and the maximum circumscribed rectangle of each target point is stretched to obtain the first region. For example, taking at least two target points as the wrist point and the base point of the middle finger of the wearer's left hand as an example, for example, Figure 5 As shown, Figure 5 The two red dots in the figure are the wrist point and the middle finger root point. A rectangle is defined with the wrist point and the middle finger root point as vertices, such as Figure 5 Then, take the center point of the rectangle as the center point, increase the long side of the rectangle by a part, for example, increase the long side of the rectangle by 30%, as the side length, and determine a larger rectangle, such as Figure 5 The large blue rectangular box in the image takes the larger rectangular area as the first area.
[0059] By stretching the maximum circumscribed rectangle of each target point, a first region with a larger range is obtained, thereby ensuring the reliability of subsequent palm key point detection based on the first region.
[0060] S102: Perform palm key point detection in the first area to determine a plurality of palm key points.
[0061] Different from the traditional palm key point detection, this embodiment does not need to detect the key points of the entire palm, but only needs to detect the palm key points in the first area, that is, only the palm or the back of the palm needs to be detected to determine multiple palm key points. Figure 6 As shown, Figure 6 The seven red dots in the figure are multiple palm key points.
[0062] In some embodiments, the palm key point detection is performed in the first area to determine multiple palm key points, including: cropping the first area of the current frame image taken by the camera device of the near-eye display device to obtain a target image; inputting the target image into a neural network model for classification and palm key point detection, and outputting a detection result, wherein the detection result includes coordinate information of the multiple palm key points.
[0063] The current frame image may be an image captured by a camera device of the near-eye display apparatus. It is understandable that the current frame image may also be an image extracted from a video captured by the camera device.
[0064] For example, the current frame image captured by the camera is as follows Figure 5 Taking the image shown as an example, the first area is Figure 5 The blue rectangular box in the image is cropped based on the first region to obtain the corresponding target image. For example, Figure 6The target image is shown.
[0065] The obtained target image is input into a neural network model, wherein the neural network model may be a deep learning model. The target image is subjected to image classification and palm key point detection by the neural network model, and a corresponding detection result is output. The detection result includes whether the target image is a human hand image, whether the target image is an image of an open palm posture, whether there are fingertips of another hand in the target image, coordinate information of multiple palm key points, coordinate information of fingertips of another hand, etc.
[0066] For example, the neural network model outputs the detection result as [a, b, c, d, e], where a indicates whether it is a human hand image; b indicates whether it is an image of an open palm; c indicates whether there are fingertips of another hand; d indicates the coordinate information of multiple palm key points; and e indicates the coordinate information of the fingertips of another hand.
[0067] For example, if a is the first detection result (such as the result is "0"), it means that it is not a human hand image, and the wearer's hand detection continues, such as by performing image recognition detection on the next frame of image captured by the camera device to detect whether there is a hand in the image. If a is the second detection result (such as the result is "1"), it means that it is a human hand image, and further detection is performed to determine whether the target image is an image of an open palm posture.
[0068] In some embodiments, after inputting the target image into a neural network model for classification and palm key point detection and outputting the detection result, the step further includes: if the target image is not an image of an open palm posture, obtaining the next frame image taken by the camera device; performing the first area cropping on the next frame image to obtain a new target image, and returning to the step of inputting the target image into a neural network model for classification and palm key point detection and outputting the detection result.
[0069] If b is the third detection result (for example, the result is "no"), it means that it is not an image of an open palm posture. At this time, the tracking is started, and the next frame image is cropped based on the first area to obtain a new target image. Because the position of the hand will not change too much in two adjacent frames, the hand of the next frame image will fall in the first area. Therefore, the next frame image can be directly cropped without the need to detect the wearer's hand before. After that, the new target image is continuously classified through the neural network model. For details, please refer to the above description, which will not be repeated here.
[0070] If b is the fourth detection result (for example, the result is "yes"), it means that the image is an image of an open palm posture, and the palm key point detection is performed on the target image to determine multiple palm key points and obtain the coordinate information of the multiple palm key points. Figure 6As shown, Figure 6 The seven red dots in the figure are multiple palm key points.
[0071] If c is the fifth detection result (for example, the result is "false"), it means that there is no fingertip of the other hand, and the wearer has not performed any input operation. At this time, the tracking is started. For details, please refer to the above description, which will not be repeated here. If c is the sixth detection result (for example, the result is "true"), it means that there is a fingertip of the other hand, and the wearer performs input operation through the fingertip of the other hand. At this time, the fingertip point is determined and the coordinate information of the fingertip point is obtained. For example, Figure 7 The blue dots shown in the figure are the fingertips of the other hand.
[0072] In this application, only the palm key points corresponding to the palm part and the fingertips of the other hand need to be identified, without the need to identify the key points of both hands. This is not only faster, but also has the following advantages:
[0073] (1) Human fingers are very flexible and diverse, making them difficult to identify. By removing the key point recognition of the fingers and only needing to identify the key points of the palm, it is much simpler and reduces the data processing capability requirements of the neural network model.
[0074] (2) Since only the key points of the palm need to be identified, only the part containing the palm needs to be cropped. In this case, the image size input to the neural network model can be reduced a lot. Compared with the traditional method that requires identifying the key points of the entire hand, the computational complexity of the neural network model is further reduced.
[0075] S103: Determine, based on the plurality of palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm.
[0076] The second area includes various palm key points. Figure 6 For example, based on the seven palm key points, the largest circumscribed rectangle of the seven palm key points is taken as the second area. For example, the second area is as follows: Figure 6 As shown in the green rectangular box in the figure. The second area is used as the projection area of the input interface, and the input interface is projected on it, such as projecting the virtual keyboard interface on the second area, so as to facilitate the user to watch and operate.
[0077] In some embodiments, Figure 8 As shown, step S103 may include sub-step S1031 and sub-step S1032.
[0078] S1031, determining the maximum circumscribed rectangle of the plurality of palm key points;
[0079] S1032: Enlarge the size of the maximum circumscribed rectangle to obtain the second area.
[0080] For example, still Figure 6 Take the seven palm key points in as an example, and take the maximum circumscribed rectangle of these seven palm key points, such as Figure 6 As shown in the green rectangle in the figure, the maximum circumscribed rectangle is enlarged, for example, the center point of the maximum circumscribed rectangle is used as the center point, the side length of the rectangle is increased by a part, for example, the side length of the rectangle is increased by 20%, and the maximum circumscribed rectangle is stretched to obtain the corresponding second area. By enlarging the size of the maximum circumscribed rectangle, a larger second area is obtained, and the size of the input interface projection is increased, which is more convenient for users to watch and perform input operations.
[0081] S104: Projecting the input interface in the second area based on the near-eye display device.
[0082] For example, Fig. 9 As shown, a virtual keyboard interface is projected on the palm of the wearer's left hand. In this application, the input interface is specifically displayed based on the palm or back of the wearer's palm, which enhances the intelligence of the near-eye display device and improves the user experience.
[0083] S105: Detecting an input operation of the wearer on the input interface to determine input information of the wearer according to the input operation.
[0084] Exemplarily, the target image is input into the neural network model. If the detection result output by the neural network model includes the coordinate information of the fingertips of the other hand, the input information of the wearer is determined based on the coordinate information of the fingertips of the other hand.
[0085] In some embodiments, Fig.10 As shown, step S105 may include sub-step S1051 and sub-step S1052.
[0086] S1051, obtaining coordinate information of the fingertip point corresponding to the input operation;
[0087] S1052. Determine, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determine the character corresponding to the character area as the input information of the wearer.
[0088] For example, get Figure 7 The coordinate information of the blue dot shown in FIG. Exemplarily, the coordinate information of the fingertip point includes but is not limited to 3D coordinates. According to the coordinate information of the fingertip point, the character corresponding to the character area where the fingertip point is located in the input interface is used as the input information of the wearer.
[0089] Exemplarily, the coordinate information of the fingertip point includes depth information. For the sake of distinction, the depth information of the fingertip point is referred to as first depth information below. For example, the first depth information of the fingertip point can be obtained by triangulation based on a depth camera or a binocular RGB camera.
[0090] In some embodiments, the determining, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determining the character corresponding to the character area as the input information of the wearer, includes: acquiring second depth information of the palm or the back of the palm; comparing the first depth information with the second depth information to determine whether the fingertip point is in contact with the palm or the back of the palm;
[0091] The method of determining, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determining the character corresponding to the character area as the input information of the wearer includes: if the fingertip point is in contact with the palm or the back of the palm, determining, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determining the character corresponding to the character area as the input information of the wearer.
[0092] For example, the depth information of the palm or back of the wearer is obtained by triangulation based on a depth camera or a binocular RGB camera. For the convenience of distinguishing descriptions, the depth information of the palm or back of the wearer is referred to as the second depth information below.
[0093] According to the first depth information of the fingertip point and the second depth information of the palm or the back of the palm, the first depth information is compared with the second depth information to determine whether the fingertip point is in contact with the palm or the back of the palm. If the first depth information is consistent with the second depth information, it should be noted that the first depth information is consistent with the second depth information, which means that the absolute value of the depth difference corresponding to the first depth information and the second depth information is less than or equal to the preset threshold, wherein the preset threshold can be flexibly set according to the actual situation, and is not specifically limited in this application. For example, when the fingertip point is about to touch the palm or the back of the palm, or when the fingertip point presses the palm or the back of the palm hard, the absolute value of the depth difference corresponding to the first depth information and the second depth information is less than the preset threshold. At this time, it is determined that the fingertip point is in contact with the palm or the back of the palm, and the wearer has performed an input operation. According to the coordinate information of the fingertip point, the character corresponding to the character area where the fingertip point is located in the input interface is determined as the wearer's input information. For example, the keyboard character corresponding to the fingertip point is used as input information. If the first depth information is inconsistent with the second depth information, it means that the fingertip point is not in contact with the palm or the back of the palm, and it may be the stage of moving the fingertip before the input operation. At this time, continue to obtain the coordinate information of the fingertip point. For details, please refer to the above description, which will not be repeated here. Until it is determined that the first depth information is consistent with the second depth information, the character corresponding to the position corresponding to the coordinate information of the latest fingertip point is located in the character area in the input interface, and is determined as the input information of the wearer.
[0094] Depth information is used to determine whether the fingertip is in contact with the palm or the back of the palm. Only when the fingertip is in contact with the palm or the back of the palm, the character corresponding to the character area where the fingertip is located in the input interface is determined as the wearer's input information based on the coordinate information of the fingertip. This can avoid misdetection of the user's input operation and improve the accuracy of determining the input information.
[0095] In some embodiments, after detecting the input operation of the wearer on the input interface to determine the input information of the wearer according to the input operation, it also includes: acquiring the next frame image taken by the camera device; cropping the next frame image by the first area to obtain a new target image, and returning to execute the step of inputting the target image into the neural network model, classifying the target image, and determining whether the target image is an image of a human hand and an image of an open palm posture.
[0096] When it is determined that the fingertip is in contact with the palm or the back of the palm, the character corresponding to the coordinate information of the fingertip is located in the character area in the input interface, which is determined as the wearer's input information, and then tracking is started. For details, please refer to the previous description and will not be repeated here.
[0097] Take the near-eye display device as AR glasses as an example. Fig.11As shown, the process of input operation by the AR glasses wearer through the AR glasses is as follows:
[0098] Step A, turn on the AR glasses and enter the first input mode (palm / back of palm input mode);
[0099] Step B, turn on the camera of the AR glasses and take a photo;
[0100] Step C, perform gesture detection based on the captured image to determine whether there is a hand; if so, execute Step D; if not, return to execute Step C;
[0101] Step D, make a rectangle with the wrist point and the base of the middle finger as vertices, and then use the center point of the rectangle as the center point, increase the long side of the rectangle by a part (for example, 30%) as the side length, and cut;
[0102] Step E, palm key point detection and classification based on cropping;
[0103] StepF, determine whether there is a hand in the clipping; if so, execute StepG; if not, return to execute StepC;
[0104] Step G, determine whether the palm is fully spread out during cutting; if so, execute Step H; if not, execute Step I;
[0105] Step H, obtain the coordinates of multiple palm key points, take the maximum circumscribed rectangle of multiple palm key points as the virtual keyboard interface, and project the keyboard interface on it;
[0106] StepI, enter tracking, and return to execute StepD;
[0107] Step J, determine whether there is another fingertip in the clipping; if so, execute Step K; if not,
[0108] Then return to execute Step I;
[0109] StepK, obtain the coordinates of the other fingertip, including depth information;
[0110] Step L, judging whether the fingertip is in contact with the palm according to the depth information of the fingertip and the palm; if so, executing Step M; if not, returning to executing Step I;
[0111] Step M, determine the finger tip position and obtain the corresponding keyboard input.
[0112] For example, the largest circumscribed rectangle can be obtained based on multiple palm key points, and then the center point remains unchanged and the side length is appropriately extended to obtain a large square.
[0113] Palm input on AR glasses can realize information input without the help of any external devices such as mobile phones. It is quick and convenient and improves the user experience.
[0114] In the above embodiment, by determining a first area corresponding to the palm or the back of a palm of a wearer of the near-eye display device, palm key point detection is performed in the first area to determine multiple palm key points, and based on the multiple palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm is determined, based on the near-eye display device projecting the input interface in the second area, the wearer's input operation on the input interface is detected to determine the wearer's input information based on the input operation. The entire process only needs to identify the palm key points corresponding to the palm part, and there is no need to perform key point recognition on both hands, which reduces the amount of calculation, thereby improving the input efficiency based on the near-eye display device.
[0115] See also Fig.12 , Fig.12 It is a schematic block diagram of an input device provided in an embodiment of the present application. The input device can be configured in a near-eye display device to execute the aforementioned input method based on the near-eye display device.
[0116] like Fig.12 As shown, the input device 200 may include a processor 210 and a memory 220, wherein the processor 210 and the memory 220 are connected via a bus, such as an I2C (Inter-integrated Circuit) bus.
[0117] Specifically, the processor 210 may be a micro-controller unit (MCU), a central processing unit (CPU), or a digital signal processor (DSP).
[0118] Specifically, the memory 220 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB disk, or a mobile hard disk, etc. The memory 220 stores various computer programs for the processor 210 to execute.
[0119] The processor 210 is used to run a computer program stored in the memory, and implements the following steps when executing the computer program:
[0120] Determine a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device;
[0121] Perform palm key point detection in the first area to determine a plurality of palm key points;
[0122] Determine, according to the plurality of palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm;
[0123] Projecting the input interface in the second area based on the near-eye display device;
[0124] An input operation of the wearer on the input interface is detected to determine input information of the wearer according to the input operation.
[0125] In some embodiments, when implementing the step of determining the first area corresponding to the palm or the back of a palm of a wearer of the near-eye display device, the processor 210 is configured to implement:
[0126] Determining at least two target points on a palm of the wearer;
[0127] The first area is determined according to at least two of the target points.
[0128] In some embodiments, before implementing the step of determining the first area corresponding to the palm or the back of a palm of a wearer of the near-eye display device, the processor 210 is configured to implement:
[0129] Controlling the near-eye display device to enter a first input mode based on a preset instruction;
[0130] The determining of a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device includes:
[0131] After the near-eye display device enters the first input mode, the first area is determined.
[0132] In some embodiments, the processor is further configured to implement:
[0133] The wearer's gesture action is recognized through the camera device of the near-eye display device, and if a preset gesture action is recognized, the preset instruction is triggered.
[0134] In some embodiments, when the processor 210 detects palm key points in the first area and determines a plurality of palm key points, it is configured to implement:
[0135] Performing the first region cropping on a current frame image captured by a camera device of the near-eye display device to obtain a target image;
[0136] The target image is input into a neural network model for classification and palm key point detection, and a detection result is output, wherein the detection result includes coordinate information of a plurality of palm key points.
[0137] In some embodiments, the detection result further includes whether the target image is an image of an open palm posture. After implementing the input of the target image into the neural network model for classification and palm key point detection and outputting the detection result, the processor 210 is used to implement:
[0138] If the target image is not an image of an open palm posture, acquiring the next frame of image taken by the camera device;
[0139] The first region is cropped on the next frame image to obtain a new target image, and the step of inputting the target image into a neural network model for classification and palm key point detection and outputting the detection result is returned.
[0140] In some embodiments, when the processor 210 determines the second area corresponding to the projection of the input interface on the palm or the back of the palm according to the plurality of palm key points, it is used to implement:
[0141] Determine the maximum circumscribed rectangle of the plurality of palm key points;
[0142] The maximum circumscribed rectangle is enlarged to obtain the second area.
[0143] In some embodiments, when the processor 210 detects the input operation of the wearer on the input interface to determine the input information of the wearer according to the input operation, it is configured to implement:
[0144] Obtaining coordinate information of the fingertip point corresponding to the input operation;
[0145] According to the coordinate information of the fingertip point, it is determined that the fingertip point is located in a character area in the input interface, and the character corresponding to the character area is determined as the input information of the wearer.
[0146] In some embodiments, the coordinate information of the fingertip point includes first depth information, and the processor 210 is used to implement, before implementing the step of determining, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface and determining the character corresponding to the character area as the input information of the wearer:
[0147] Acquire second depth information of the palm or the back of the palm;
[0148] Comparing the first depth information with the second depth information to determine whether the fingertip point is in contact with the palm or the back of the palm;
[0149] When the processor 210 determines, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determines the character corresponding to the character area as the input information of the wearer, it is used to implement:
[0150] If the fingertip point contacts the palm or the back of the palm, the character area where the fingertip point is located in the input interface is determined based on the coordinate information of the fingertip point, and the character corresponding to the character area is determined as the input information of the wearer.
[0151] The input device 200 can execute the input method based on the near-eye display device provided in the embodiment of the present application. Therefore, it can achieve the beneficial effects that can be achieved by the input method based on the near-eye display device provided in the embodiment of the present application. Please refer to the previous embodiment for details and will not be repeated here.
[0152] In an embodiment of the present application, a near-eye display device is also provided. The near-eye display device includes an input device. The input device can be Fig.12 Therefore, the near-eye display device can achieve the beneficial effects that can be achieved by the input method based on the near-eye display device provided in the embodiment of the present application, as detailed in the previous embodiment, which will not be repeated here.
[0153] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the input method based on the near-eye display device as described above are implemented.
[0154] The computer-readable storage medium may be an internal storage unit of the input device or near-eye display device described in the aforementioned embodiment, such as a hard disk or memory of the input device or near-eye display device. The computer-readable storage medium may also be an external storage device of the input device or near-eye display device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital card (Secure Digital Card, SD Card), a flash card (Flash Card), etc., equipped on the input device or near-eye display device.
[0155] Since the computer program stored in the storage medium can execute any one of the input methods based on the near-eye display device provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the input methods based on the near-eye display device provided in the embodiments of the present application can be achieved. Please see the previous embodiments for details and will not be repeated here.
[0156] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0157] The above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and these modifications or substitutions should be included in the protection scope of the present application.
Claims
1. An input method based on a near-eye display device, characterized in that: The input method based on the near-eye display device includes: Determine a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device; Perform palm key point detection in the first area to determine a plurality of palm key points; Determine, according to the plurality of palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm; Projecting the input interface in the second area based on the near-eye display device; An input operation of the wearer on the input interface is detected to determine input information of the wearer according to the input operation.
2. The input method based on the near-eye display device according to claim 1, characterized in that: The determining of a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device includes: Determining at least two target points on a palm of the wearer; The first area is determined according to at least two of the target points.
3. The input method based on the near-eye display device according to claim 1, characterized in that: Before determining the first area corresponding to the palm or the back of a palm of a wearer of the near-eye display device, the method includes: Controlling the near-eye display device to enter a first input mode based on a preset instruction; The determining of a first area corresponding to a palm or a back of a palm of a wearer of the near-eye display device includes: After the near-eye display device enters the first input mode, the first area is determined.
4. The input method based on the near-eye display device according to claim 3, characterized in that: The method further comprises: The wearer's gesture action is recognized through the camera device of the near-eye display device, and if a preset gesture action is recognized, the preset instruction is triggered.
5. The input method based on the near-eye display device according to claim 1, characterized in that: The detecting of palm key points in the first area to determine a plurality of palm key points includes: Performing the first region cropping on a current frame image captured by a camera device of the near-eye display device to obtain a target image; The target image is input into a neural network model for classification and palm key point detection, and a detection result is output, wherein the detection result includes coordinate information of a plurality of palm key points.
6. The input method based on the near-eye display device according to claim 5, characterized in that: The detection result also includes whether the target image is an image of an open palm posture. After inputting the target image into a neural network model for classification and palm key point detection and outputting the detection result, the method further includes: If the target image is not an image of an open palm posture, acquiring the next frame of image taken by the camera device; The first region is cropped on the next frame image to obtain a new target image, and the step of inputting the target image into a neural network model for classification and palm key point detection and outputting the detection result is returned.
7. The input method based on the near-eye display device according to claim 1, characterized in that: The step of determining, based on the plurality of palm key points, a second area corresponding to the projection of the input interface on the palm or the back of the palm comprises: Determine the maximum circumscribed rectangle of the plurality of palm key points; The maximum circumscribed rectangle is enlarged to obtain the second area.
8. The input method based on the near-eye display device according to claim 1, characterized in that: The detecting the input operation of the wearer on the input interface to determine the input information of the wearer according to the input operation includes: Obtaining coordinate information of the fingertip point corresponding to the input operation; According to the coordinate information of the fingertip point, it is determined that the fingertip point is located in a character area in the input interface, and the character corresponding to the character area is determined as the input information of the wearer.
9. The input method based on the near-eye display device according to claim 8, characterized in that: The coordinate information of the fingertip point includes the first depth information, and determining that the fingertip point is located in a character area in the input interface according to the coordinate information of the fingertip point, and determining the character corresponding to the character area as the input information of the wearer, includes: Acquire second depth information of the palm or the back of the palm; Comparing the first depth information with the second depth information to determine whether the fingertip point is in contact with the palm or the back of the palm; The step of determining, based on the coordinate information of the fingertip point, that the fingertip point is located in a character area in the input interface, and determining the character corresponding to the character area as the input information of the wearer includes: If the fingertip point contacts the palm or the back of the palm, the character area where the fingertip point is located in the input interface is determined based on the coordinate information of the fingertip point, and the character corresponding to the character area is determined as the input information of the wearer.
10. An input device, characterized in that: The input device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the steps of the input method based on a near-eye display device as described in any one of claims 1 to 9 when executing the computer program.
11. A near-eye display device, characterized in that: The near-eye display device comprises the input apparatus as claimed in claim 10.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the input method based on a near-eye display device according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Input method and apparatus, near-eye display device, and storage medium
WO2026144041A1