Human-computer interaction method, device, electronic device and storage medium
By setting up front camera or infrared sensors at the terminal, obtaining user's eyes and hands information, identifying and matching hand movements, the problem of low accuracy in interaction of 3D display screens is solved, and efficient human-computer interaction without touch is achieved.
Patent Information
- Application Number
- CN202210587106.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, the user's interaction with a 3D display screen with a 3D visual effect has low touch accuracy, poor user experience, and it is impossible to accurately locate the virtual screen in the air.
By setting a front camera or infrared sensor at the terminal, the distance between the user's eyes and the hand movements and hand movements are obtained, the hand shape is identified, the hand movements and preset movements are matched, and the coordinate position is determined to achieve 3D human-computer interaction without touch.
Improves the user's interactive experience during 3D display, reduces false touches, and realizes accurate information input based on hand movements.
Smart Images

Figure CN114895789B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a human-computer interaction method, device, electronic equipment and storage medium. Background Art
[0002] Three-dimensional (3D) technology allows users to see 3D images. 3D images are created by creating a three-dimensional image using parallax between the left and right eyes. Smart devices like mobile phones and tablets are becoming increasingly commonplace, and 3D display technology is being incorporated into these devices. Users can now view 3D images on these devices. How to further enhance the user experience through 3D display is a key research area for future human-computer interaction.
[0003] In related technologies, users generally need to touch the screen with their fingers to achieve human-computer interaction with smart terminals such as mobile phones and tablets. However, for 3D display screens with 3D visual effects, this human-computer interaction method has low touch accuracy and poor user experience. Summary of the Invention
[0004] The embodiments of the present invention provide a human-computer interaction method, device, electronic device and storage medium, which can perform human-computer interaction according to the user's hand movements when the terminal performs 3D display, without touching the terminal, thereby improving the user's interactive experience.
[0005] In a first aspect, an embodiment of the present invention provides a human-computer interaction method, which is applied in a terminal, wherein the terminal is provided with a 3D display screen, and the method includes: controlling the 3D display screen to display a 3D object to be operated; obtaining a first distance from the user's eyes to the 3D display screen, and determining a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determining a first visual plane where the user views the object to be operated based on the first screen point position; obtaining a first hand movement of the user on the first visual plane, and matching the coordinate position on the object to be operated based on the first hand movement; and obtaining input information corresponding to the object to be operated based on the coordinate position.
[0006] In some embodiments, the terminal is provided with a front camera, and obtaining the first hand movement of the user on the first visual plane includes: obtaining a second distance from the user's hand to the 3D display screen; if the second distance indicates that the user's hand is located on the first visual plane, obtaining the user's hand image, and the hand image is obtained by taking the hand image by the front camera; identifying the hand shape in the picture in the hand image; obtaining a preset target hand shape, and performing a matching analysis between the hand shape in the picture and the target hand shape; if the hand shape in the picture matches the target hand shape, determining that the hand shape in the picture is the first hand movement.
[0007] In some embodiments, matching the coordinate position on the object to be operated based on the first hand movement includes: obtaining the hand position corresponding to when the user makes the first hand movement; matching the coordinates of the hand position with the coordinates on the first visual plane to obtain the corresponding coordinate position on the object to be operated.
[0008] In some embodiments, the object to be operated includes a virtual keyboard, and obtaining the input information corresponding to the object to be operated based on the coordinate position includes: obtaining the coordinate mapping relationship between each key value on the virtual keyboard and the corresponding input position; triggering the corresponding target key value from the virtual keyboard based on the coordinate position and the coordinate mapping relationship.
[0009] In some embodiments, the terminal is provided with a front camera, and the hand position is obtained by analyzing the hand image taken by the front camera; or, the terminal is provided with an infrared sensor or an ultrasonic sensor, and the infrared sensor and the ultrasonic sensor are used to obtain the hand position.
[0010] In some embodiments, the method also includes at least one of the following: obtaining a second hand movement of the user on the first visual plane, and opening or closing the object to be operated according to the second hand movement; obtaining a third hand movement of the user on the first visual plane, and controlling the object to be operated to perform a response action according to the third hand movement, and the response action includes zooming in, zooming out, scrolling down, or turning pages.
[0011] In some embodiments, the first hand action includes one of a press down action, a click operation, a grabbing action, or a sliding action, the second hand action includes one of a press down action, a click operation, a grabbing action, or a sliding action, the third hand action includes one of a press down action, a click operation, a grabbing action, or a sliding action, and the first hand action, the second hand action, and the third hand action are different from each other.
[0012] In some embodiments, the object to be operated includes at least one of a virtual keyboard, a separate control, or a gesture control.
[0013] In some embodiments, the terminal is provided with a front camera, and obtaining the first distance from the user to the 3D display screen includes: obtaining a facial image of the user, where the facial image is captured by the front camera; performing pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes; calculating the screen pupil distance of the user in the facial image based on the first pupil position information and the second pupil position information; and calculating the first distance from the user's eyes to the 3D display screen based on the screen pupil distance.
[0014] In some embodiments, the front camera is an under-screen camera, and the under-screen camera is arranged at the center of the 3D display screen.
[0015] In some embodiments, performing pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes includes: converting the facial image into a grayscale image and binarizing the grayscale image to obtain a first preprocessed image; corroding and dilating the first preprocessed image and removing noise from the image to obtain a second preprocessed image; extracting the position of a circular area representing the user's pupil in the second preprocessed image using a circular structural element; and calculating the center point of the circular area to obtain the first pupil position information and second pupil position information of the user's eyes.
[0016] In some embodiments, calculating the first distance from the user's eyes to the 3D display screen based on the interpupillary distance of the picture includes: obtaining a preset standard interpupillary distance; obtaining the focal length of the facial image captured by the front camera, and obtaining an initial distance from the facial image to the imaging point based on the focal length; obtaining a first ratio based on the interpupillary distance of the picture and the standard interpupillary distance, and obtaining the first distance from the user's eyes to the 3D display screen based on the first ratio and the initial distance.
[0017] In some embodiments, calculating the first distance from the user's eyes to the 3D display screen based on the interpupillary distance of the picture includes: obtaining a preset distance lookup table; and obtaining the first distance from the user's eyes to the 3D display screen by looking up the distance lookup table based on the interpupillary distance of the picture.
[0018] In some embodiments, calculating the first distance from the user's eyes to the 3D display screen based on the picture pupil distance includes: obtaining a reference distance, a reference object size, and a picture size corresponding to the reference object captured by the front camera; obtaining a preset standard pupil distance; and obtaining the first distance from the user's eyes to the 3D display screen based on the reference distance, the reference object size, the picture size, the picture pupil distance, and the standard pupil distance.
[0019] In some embodiments, determining a first screen point position between the user's eyes and the 3D display screen based on the first distance includes: obtaining a negative parallax value for 3D image display by the terminal; obtaining a third distance based on the first distance and the negative parallax value; and determining a position between the user and the 3D display screen and the third distance away from the 3D display screen as the first screen point position.
[0020] In some embodiments, after determining a first screen point position between the user's eyes and the 3D display screen based on the first distance, the method includes: when the user's eyes move, obtaining a fourth distance from the user's eyes to the 3D display screen after the movement; obtaining a fifth distance based on the fourth distance and the negative parallax value; updating the first screen point position, and updating a position between the user and the 3D display screen and at a fifth distance from the 3D display screen as the first screen point position.
[0021] In a second aspect, an embodiment of the present invention further provides a human-computer interaction device, comprising: a first module for controlling a 3D display screen to display a 3D object to be operated; a second module for obtaining a first distance from a user's eyes to the 3D display screen, and determining a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determining a first visual plane where the user views the object to be operated based on the first screen point position; a third module for obtaining a first hand movement of the user on the first visual plane, and matching a coordinate position on the object to be operated based on the first hand movement; and a fourth module for obtaining input information corresponding to the object to be operated based on the coordinate position.
[0022] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the human-computer interaction method as described in the embodiment of the first aspect of the present invention is implemented.
[0023] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the human-computer interaction method as described in the embodiment of the first aspect of the present invention.
[0024] Embodiments of the present invention include at least the following advantageous effects: Embodiments of the present invention provide a human-computer interaction method, apparatus, electronic device, and storage medium. The human-computer interaction method can be applied to a terminal equipped with a 3D display screen capable of displaying a 3D image. By executing the human-computer interaction method, the 3D display screen is controlled to display a 3D object to be operated. A first distance between a user's eyes and the 3D display screen is then obtained, and a first screen point position between the user's eyes and the 3D display screen is determined based on the first distance. Because the object to be operated is displayed in 3D, the 3D object to be operated is displayed on a plane where the first screen point position is located based on a three-dimensional image virtualized by binocular parallax. Therefore, a first visual plane where the user views the object to be operated is determined based on the first screen point position. A first hand motion of the user on the first visual plane is then obtained, and a coordinate position on the object to be operated is matched based on the first hand motion. Finally, input information corresponding to the object to be operated is obtained based on the coordinate position. Embodiments of the present invention enable human-computer interaction based on the user's hand motion when the terminal is performing a 3D display. Different input information on the object to be operated is obtained based on different hand motions. Interaction can be achieved without touching the terminal, thereby improving the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic diagram of a 3D imaging principle provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a terminal provided by an embodiment of the present invention;
[0027] Figure 3 is a flowchart of a human-computer interaction method provided by one embodiment of the present invention;
[0028] Figure 4 is a schematic diagram of a first screen point position provided by an embodiment of the present invention;
[0029] Figure 5 is a schematic diagram of a human-computer interaction scenario provided by an embodiment of the present invention;
[0030] Figure 6 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0031] Figure 7 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0032] Figure 8 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0033] Figure 9is a schematic diagram of realizing human-computer interaction through a 3D displayed virtual keyboard provided by one embodiment of the present invention;
[0034] Figure 10 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0035] Figure 11 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0036] Figure 12a is a schematic diagram of a facial image provided by one embodiment of the present invention;
[0037] Figure 12b is a schematic diagram of a facial image provided by another embodiment of the present invention;
[0038] Figure 13 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0039] Figure 14 1 is a schematic diagram of obtaining pupil positions by processing a facial image according to an embodiment of the present invention;
[0040] Figure 15 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0041] Figure 16 is a schematic diagram of relative lens (imaging point) imaging provided by one embodiment of the present invention;
[0042] Figure 17 is a schematic diagram of calculating a first distance using a triangle principle provided by an embodiment of the present invention;
[0043] Figure 18 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0044] Figure 19 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0045] Figure 20 is a schematic diagram of obtaining a first distance according to a reference system provided by one embodiment of the present invention;
[0046] Figure 21 is a schematic diagram of obtaining a first distance according to a reference system provided by another embodiment of the present invention;
[0047] Figure 22 is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0048] Figure 23is a flowchart of a human-computer interaction method provided by another embodiment of the present invention;
[0049] Figure 24 is a schematic structural diagram of a human-computer interaction device provided by one embodiment of the present invention;
[0050] Figure 25 It is a structural diagram of an electronic device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0052] It should be understood that in the description of the embodiments of the present invention, "multiple" (or multiple) means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, and "above," "below," and "within" are understood to include the number itself. The terms "first," "second," and so on are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features.
[0053] At present, smart terminals such as mobile phones and tablets are gradually entering people's lives, and 3D display technology has also been gradually applied to terminals. Users can watch 3D images on the terminals. 3D images are three-dimensional spatial images virtualized through the parallax of the left and right eyes. Specifically, the terminal can display 3D images. For humans, the 3D stereoscopic sense comes from two visual images received in the brain. The brain combines the similarities of the two images, and the subtle differences will guide the user to feel the sense of space. The two very different images will be combined into a single three-dimensional image, thereby realizing the 3D display of the terminal.
[0054] In 3D display, eye convergence refers to the angle between the eyes and the observed object. The larger the angle, the closer the object seems. Conversely, the smaller the angle, the farther the object seems. Parallax images refer to the images seen by the left and right eyes respectively. All 3D images or videos contain pairs of parallax images, which enter the user's left and right eyes separately but simultaneously. For example, when the target object is tilted to the right in the left eye image and to the left in the right eye image, the focus of the user's eyes will be guided to fall behind the 3D display screen. This phenomenon is called positive parallax. When each pair of parallax images is overlaid on the 3D display screen, the focus of the user's eyes will be guided to fall on the 3D display screen. This phenomenon is called positive parallax. Figure 1As shown, when the target object is tilted to the left in the left eye image and to the right in the right eye image, the user's eye focus will be guided to fall in front of the 3D display screen. This phenomenon is called negative parallax. Both positive and negative parallax have certain values. The terminal can adjust the ratio of the image entering or exiting the screen by setting the positive or negative parallax value.
[0055] How to further enhance the user's interactive experience based on 3D display will become the research direction of human-computer interaction in the future.
[0056] In the related art, users generally need to touch the screen with their fingers to achieve human-computer interaction with smart terminals such as mobile phones and tablets. However, the applicant has found that for 3D display screens with 3D visual effects, this human-computer interaction method has low touch accuracy and poor user experience. This is because in terminals capable of 3D display, the 3D picture displayed by the 3D display screen does not have its screen points on the screen. As mentioned in the above embodiment, when a user watches a 3D picture with negative parallax, the picture viewed by the user's eyes should be in the area between the 3D display screen and the user's eyes, which is a virtual picture. Therefore, when the user reaches out to touch the screen, it is easy to cause accidental touches.
[0057] The applicant further discovered that the relevant 3D technology is unable to determine the position of the virtual image in the air because the terminal cannot know how far the human eye is from the screen, and therefore cannot know the position of the 3D displayed virtual image outside the screen. Continuing to use the touch operation in the relevant technology will cause false touches, seriously reducing the user experience.
[0058] Based on this, the embodiments of the present invention provide a human-computer interaction method, device, electronic device and storage medium, which can perform human-computer interaction according to the user's hand movements when the terminal performs 3D display, without touching the terminal, thereby improving the user's interactive experience.
[0059] The terminal in the embodiments of the present invention may be a mobile terminal device or a non-mobile terminal device. Mobile terminal devices may include mobile phones, tablet computers, laptop computers, PDAs, vehicle-mounted terminal devices, wearable devices, ultra-mobile personal computers, netbooks, personal digital assistants, etc. Non-mobile terminal devices may include personal computers, televisions, ATMs, or self-service kiosks, etc. The embodiments of the present invention are not particularly limited.
[0060] The terminal may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a mobile communication module, a wireless communication module, an audio module, a speaker, a receiver, a microphone, an earphone interface, a sensor module, buttons, a motor, an indicator, a front camera, a rear camera, a display, and a subscriber identification module (SIM) card interface. The terminal may implement a shooting function through the front camera, rear camera, video codec, GPU, display, and application processor.
[0061] The front camera or rear camera is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP (image signal processor) to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal may include 1 or N front cameras, where N is a positive integer greater than 1.
[0062] The terminal implements display functions through a GPU, display screen, and application processor. The GPU is a microprocessor for image processing that connects the display screen and application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or modify display information.
[0063] A display screen is used to display images, videos, and the like. It includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED).
[0064] In one embodiment, the display screen in the embodiment of the present invention is a 3D display screen that can display 3D images. It may be referred to as the display screen hereinafter. The display screen may be a naked-eye 3D display screen. The naked-eye 3D display screen may process multimedia data and split it into two parts, left and right. For example, a 2D video is cropped into two parts, and the light refraction direction of the two parts is changed. After viewing, the user's eyes can see a 3D image. The resulting 3D image has a negative parallax and can be displayed between the user and the display screen to achieve a naked-eye 3D viewing effect. Alternatively, the display screen is a 2D display screen, and the terminal may be equipped with an external 3D grating film to refract the outgoing light of the 2D display screen, so that the user can see a 3D display effect after viewing the display screen through the 3D grating film.
[0065] For example, Figure 2 As shown, in the embodiment of the present invention, the terminal is taken as an example of a mobile phone. For example, a front camera 11 is provided on the front panel 10 of the mobile phone, which can obtain image information. A display screen 12 is also provided on the front panel 10, which can display the picture. It can be understood that when the terminal is a 3D visual training terminal, or a mobile phone that can display 3D pictures, the 3D picture can be displayed through the display screen 12.
[0066] It should be noted that the front camera 11 in the embodiment of the present invention can be arranged on the same surface as the display screen 12, and the position of the front camera 11 is fixed. In one embodiment, the front camera 11 can be perpendicular to the display screen 12 or not perpendicular to the display screen 12. The front camera 11 can be located inside the display screen 12 or around the display screen 12, such as Figure 2As shown, in addition, the front camera 11 can also have a vertical distance from the display screen 12, so that the front camera 11 and the display screen 12 are not in the same plane. The terminal can perform parameter verification according to the front camera 11 with different settings to realize the human-computer interaction method and control method in the embodiments of the present invention.
[0067] The following introduces the human-computer interaction method, device, electronic device and storage medium in the embodiments of the present invention. First, the human-computer interaction method in the embodiments of the present invention is introduced.
[0068] Reference Figure 3 As shown, an embodiment of the present invention provides a human-computer interaction method, which is applied in a terminal. The human-computer interaction method may include but is not limited to steps S101 to S104.
[0069] Step S101: Control the 3D display screen to display a 3D object to be operated.
[0070] Step S102: Obtain a first distance between the user's eyes and the 3D display screen, determine a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determine a first visual plane where the user views the object to be manipulated based on the first screen point position.
[0071] Step S103: Acquire a first hand motion of the user on a first visual plane, and match a coordinate position on the object to be operated according to the first hand motion.
[0072] Step S104: obtaining input information corresponding to the object to be operated according to the coordinate position.
[0073] It should be noted that in the human-computer interaction method in the embodiment of the present invention, a 3D object to be operated is first displayed on a 3D display screen. The 3D display screen is the display screen described in the above embodiment, hereinafter referred to as the display screen. The object to be operated is a screen for 3D display. In one embodiment, the object to be operated can be at least one of a virtual keyboard, a separate control, or a gesture control. For example, when the object to be operated is a virtual keyboard, the terminal can display the 3D virtual keyboard through the display screen. The terminal in the embodiment of the present invention sets the 3D screen to be displayed off-screen, so the virtual keyboard has negative parallax, and the virtual keyboard will be displayed between the user's eyes and the display screen. When the object to be operated is a separate control or a gesture control, it can also be displayed off-screen. No further details are given here. It can be understood that the separate space or gesture control in the embodiment of the present invention can be any content page, an independent software interface, or a software interface based on gesture operation, for example, an article reading interface, an image reading interface, etc. It is any interface that can perform human-computer interaction and is not specifically limited here.
[0074] like Figure 4 As shown, after the 3D object to be operated is displayed, the terminal obtains a first distance from the user's eyes to the 3D display screen, and determines a first screen point position between the user's eyes and the 3D display screen based on the first distance. The first screen point position is the intersection point of the sight lines of the object to be operated seen by the user's left and right eyes, and is also the off-screen position of the 3D displayed object to be operated. The user will view a 3D virtual interface on the plane at this position. Therefore, the first visual plane where the user views the object to be operated can be determined based on the first screen point position. When human-computer interaction is subsequently performed, as shown in FIG. Figure 5 As shown, the first hand movement of the user on the first visual plane is obtained. The first hand movement is the hand movement made by the user on the corresponding first visual plane based on the 3D object to be operated seen. Coordinate matching is performed based on the first hand movement to match the coordinate position on the object to be operated, so as to finally obtain input information corresponding to the object to be operated based on the coordinate position. The input information is information generated and input into the terminal in response to the user's first hand movement, thereby realizing information input based on the user's hand movement.
[0075] It can be understood that since the terminal in the embodiment of the present invention is a terminal that performs 3D display, the user makes relevant hand movements at the position of the corresponding object to be operated after viewing the object to be operated on the first visual plane, and the terminal obtains the user's first hand movement on the first visual plane to realize the input of information. Since the user makes the first hand movement on the viewing plane, false touches can be reduced. The embodiment of the present invention can perform human-computer interaction according to the user's hand movements when the terminal performs 3D display by executing the human-computer interaction method, without touching the terminal, thereby improving the user's interactive experience.
[0076] Reference Figure 6 As shown, in one embodiment, the terminal is provided with a front camera, and the above step S102 may also include but is not limited to steps S201 to S204.
[0077] Step S201: Acquire a second distance from the user's hand to the 3D display screen.
[0078] Step S202: If the second distance indicates that the user's hand is located on the first visual plane, an image of the user's hand is obtained, where the hand image is captured by a front camera.
[0079] Step S203: Recognize the hand shape in the hand image.
[0080] Step S204: obtaining a preset target hand shape, and performing a matching analysis between the hand shape in the picture and the target hand shape.
[0081] Step S205 : If the hand shape in the picture matches the target hand shape, the hand shape in the picture is determined to be a first hand motion.
[0082] It should be noted that, in the embodiment of the present invention, the image is acquired through the front camera provided on the terminal, and the hand shape is recognized. Specifically, the present embodiment first acquires the distance from the user's hand to the display screen, which is the second distance. It can be understood that when the user's hand appears on the first visual plane, it is determined that human-computer interaction is required. When the second distance indicates that the user's hand is located on the first visual plane, the user's hand image is acquired through the front camera to analyze the gestures made by the user based on the hand image, and the user's hand shape in the hand image is recognized to obtain the screen hand shape. The screen hand shape is the hand shape made by the user recognized by the terminal from the user's hand image taken by the front camera. Based on this, the embodiment of the present invention matches and analyzes the screen hand shape with the preset target hand shape. The target hand shape is used to determine whether the user has made the correct hand movement. When the screen hand shape matches the target hand shape, the screen hand shape is determined to be the first hand movement.
[0083] For example, in one embodiment, the target hand shape is a shape corresponding to a click action. The embodiment of the present invention obtains the hand shape on the screen through identification. When the hand shape on the screen is a sliding action or a clapping action, it cannot be the same as the click action shape of the target hand shape, so it cannot be matched. When the hand shape on the screen is also a shape corresponding to a click action, the two match, so the hand shape on the screen is determined to be the first hand action. Here, the first hand action is the click action, thereby realizing human-computer interaction according to the specific hand actions made by the user.
[0084] It should be noted that, in an embodiment of the present invention, image processing can be performed on the hand image to identify the shape of the hand in the picture. The shape of the hand in the picture can be identified by edge contour extraction method, centroid finger and other multi-feature combination method and finger joint tracking method. For example, the acquired hand image is processed into a grayscale image, and then noise reduction processing is performed to extract the edge contour of the hand. The hand shape is distinguished from other objects due to its unique shape. The gesture recognition algorithm combined with geometric moment and edge detection calculates the distance between images by setting the weight of the feature to realize the recognition of gestures; the multi-feature combination rule analyzes the posture or trajectory of the gesture according to the physical characteristics of the hand, and combines the gesture shape with the fingertip features to realize the recognition of gestures; the finger joint tracking method mainly constructs a two-dimensional or three-dimensional model of the hand, and then tracks it according to the position changes of the human hand joints. It is mainly used for dynamic trajectory tracking and finally recognizes the shape of the hand in the picture.
[0085] It is understandable that in the embodiment of the present invention, the hand image can also be input into a preset neural network model, and the neural network model processes the hand image and identifies the corresponding hand shape in the picture. It is understandable that the neural network model can be trained by a large number of hand shapes and sample images in the sample, and the loss value is continuously optimized. Finally, a neural network model with higher accuracy can be obtained to realize the recognition of the hand shape in the picture in the embodiment of the present invention. No specific limitation is made compared to the embodiment of the present invention.
[0086] Reference Figure 7 As shown, in one embodiment, the above step S103 may also include but is not limited to steps S301 to S302.
[0087] Step S301: Obtain the hand position corresponding to the user performing the first hand movement.
[0088] Step S302 : Match the coordinates of the hand position with the coordinates on the first visual plane to obtain the coordinate position corresponding to the object to be operated.
[0089] It should be noted that, in an embodiment of the present invention, when matching the corresponding coordinate position according to the user's first hand movement, the coordinate matching is performed according to the hand position of the user when making the first hand movement. Specifically, in an embodiment of the present invention, when the user makes the first hand movement, the corresponding hand position is obtained. The hand position can be expressed in the form of a coordinate. According to the coordinates of the hand position, it is matched with the coordinates on the first visual plane, so that the coordinate position on the object to be operated corresponding to when the user makes the first hand movement can be obtained.
[0090] It can be understood that since the user makes the first hand movement on the first visual plane, the coordinate point representing the hand position will be located on the first visual plane. Therefore, based on the coordinates of the hand position, its position on the first visual plane can be obtained, thereby matching the coordinate position on the object to be operated in the first visual plane.
[0091] Reference Figure 8 As shown, in one embodiment, the object to be operated includes a virtual keyboard, and the above step S104 may also include but is not limited to steps S401 to S402.
[0092] Step S401: Obtain the coordinate mapping relationship between the key value of each key on the virtual keyboard and the corresponding input position.
[0093] Step S402 : triggering the corresponding target key value from the virtual keyboard according to the coordinate position and the coordinate mapping relationship.
[0094] It should be noted that if Figure 9As shown, when the substitute operation object in the embodiment of the present invention is a virtual keyboard, in the process of obtaining input information according to the coordinate position, the coordinate mapping relationship between the key values of each key on the virtual keyboard and the corresponding input position is first obtained. It can be understood that the embodiment of the present invention establishes a coordinate mapping relationship in advance according to the position of the presented virtual keyboard, so that a certain coordinate above can correspond to the corresponding input key, and then triggers the corresponding target key value from the virtual keyboard according to the coordinate position and the coordinate mapping relationship. For example, in the coordinate mapping relationship, the position corresponding to the coordinate (x, y) is the "D" key on the virtual keyboard. When the coordinate position obtained in the above embodiment is (x, y), it can be matched that the target key is the "D" key, so that the input information is the input key value of the "D" key, and finally the user completes the input of the corresponding target key value by clicking the "D" key of the virtual keyboard on the first visual plane, thereby realizing human-computer interaction operation.
[0095] It should be noted that the terminal is provided with a front camera. In the above embodiment, the hand position is obtained by analyzing the hand image taken by the front camera. For example, based on the hand image of the user taken by the front camera, the coordinate position of the hand in the image is analyzed to obtain the coordinate representation of the hand position. Alternatively, the terminal is provided with an infrared sensor or an ultrasonic sensor. The infrared sensor and the ultrasonic sensor are used to obtain the hand position. The infrared sensor or the ultrasonic sensor sends the obtained hand position information to the background processing, and after analysis, it is converted into a coordinate representation on the first visual plane, thereby obtaining the corresponding hand position when the user makes the first hand movement.
[0096] Reference Figure 10 As shown, in one embodiment, the human-computer interaction method in the embodiment of the present invention may also include but is not limited to steps S501 to S502.
[0097] Step S501: Acquire a second hand motion of the user on the first visual plane, and open or close the object to be operated according to the second hand motion.
[0098] Step S502: Acquire a third hand motion of the user on the first visual plane, and control the object to be operated to perform a response action according to the third hand motion. The response action includes zooming in, zooming out, scrolling down, or turning pages.
[0099] It should be noted that the human-computer interaction method in the embodiment of the present invention, in addition to obtaining the input information corresponding to the object to be operated based on the first hand movement made by the user, can also perform corresponding operations according to different hand movements of the user. For example, the embodiment of the present invention can obtain the second hand movement of the user on the first visual plane, and open or close the object to be operated according to the second hand movement. Similarly, the determination of the second hand movement can be the same as the determination method of the first hand movement in the above embodiment. It can also be obtained through image recognition and matched with other preset target hand shapes to determine that the user has made the second hand movement. Human-computer interaction through the second hand movement can include multiple methods, such as the user can make a waving gesture to close the object to be operated, and open the object to be operated through a waving gesture.
[0100] Not only that, the human-computer interaction method in the embodiment of the present invention can also obtain the user's third hand movement on the first visual plane, and control the object to be operated to perform a response action according to the third hand movement. The response action includes zooming in, zooming out, sliding and scrolling or turning pages. Similarly, the determination of the third hand movement can be the same as the determination method of the first hand movement in the above embodiment. It can also be obtained through image recognition and matched with other preset target hand shapes to determine that the user has made the third hand movement. Human-computer interaction through the third hand movement can include multiple methods, such as the fingers making a "pinching"-like shrinking movement to control the interface to be operated to shrink the interface, or when the object to be operated is a slidable page, such as an article reading interface, the user makes a hand movement of sliding and scrolling or turning the page, which can control the object to be operated to slide and scroll the page or turn the page. Compared with the embodiment of the present invention, there is no specific limitation.
[0101] It should be noted that the first hand action includes one of a pressing action, a clicking operation, a grabbing action or a sliding action, the second hand action includes one of a pressing action, a clicking operation, a grabbing action or a sliding action, and the third hand action includes one of a pressing action, a clicking operation, a grabbing action or a sliding action, and the first hand action, the second hand action and the third hand action are different from each other.
[0102] Reference Figure 11 As shown, in one embodiment, the terminal is provided with a front camera, and the above step S102 may also include but is not limited to steps S601 to S604.
[0103] Step S601: Acquire a facial image of the user, where the facial image is captured by a front-facing camera.
[0104] Step S602: Perform pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes.
[0105] Step S603 : Calculate the inter-pupillary distance of the user in the facial image based on the first pupil position information and the second pupil position information.
[0106] Step S604: Calculate a first distance from the user's eyes to the 3D display screen based on the interpupillary distance of the image.
[0107] It should be noted that the human-computer interaction method in the embodiment of the present invention can be based on camera distance measurement. By obtaining a facial image of the user, which is captured by a front camera, pupil recognition is performed on the facial image to determine first pupil position information and second pupil position information of the user's eyes. The screen pupil distance of the user in the facial image is calculated based on the first pupil position information and the second pupil position information of the user's eyes. The screen pupil distance can be determined based on the number of pixels of the display screen. Finally, based on the screen pupil distance, the viewing distance from the user's eyes to the display screen is calculated, which is the first distance. The embodiment of the present invention can achieve ranging through the front camera set on the terminal, and the user's pupil position is obtained by analyzing the image obtained by the front camera, and the required screen pupil distance is calculated. The screen pupil distance is the pupil distance that represents the user in the acquired image, and the viewing distance from the user's eyes to the display screen can be calculated based on the screen pupil distance. The ranging cost is low and ranging can be achieved without the need to set up other additional sensors.
[0108] It is understandable that the terminal obtains the facial image through the front camera, and the pupil distance of the user in the facial image is the pupil distance of the screen. During the shooting process of the front camera, if the camera is used as a reference, the position of the user relative to the camera can change at any time. Through the camera imaging, the size of the image of the user at different distances is different, such as Figure 12a and Figure 12b As shown, in Figure 12a In the image, the user is closer to the camera, so the size of the user's face in the image is larger, so the pupil distance is larger. Figure 12b In the image, the user is farther away from the camera, the size of the user's face in the formed image is smaller, and therefore the pupil distance is smaller. In the embodiment of the present invention, the size of the pupil distance in the user image taken by the front camera can determine the first distance from the user's eyes to the display screen.
[0109] For example, the viewing distance can be calculated using a preset formula using a reference object of known size at a known distance. For example, a 10-centimeter reference object (e.g., a ruler) is placed 50 centimeters in front of the display screen. Based on the parameters of the front camera, the 10-centimeter object will appear as a specific size (determined by the number of pixels) in the image captured by the front camera. Given that the image contains a 6.3-centimeter target object (i.e., the pupils), the size of the object in the photo is known to be fixed. Therefore, the distance from the target object to the display screen can be calculated.
[0110] It should be noted that the pupillary distance is the distance between the pupils of the user's eyes, also known as the interpupillary distance. It refers to the length between the centers of the pupils of the two eyes. The normal interpupillary distance for an adult is between 58-64mm. The interpupillary distance itself is determined by the individual's genetics and development, so the interpupillary distance varies at different ages. For a certain user, the interpupillary distance is fixed. Therefore, the distance between the user and the terminal can be determined based on the size of the interpupillary distance in the facial image, and the first distance from the user's eyes to the display screen can be calculated based on this.
[0111] It should be noted that in the embodiment of the present invention, the user's image is recognized by the front camera, and ranging can be achieved without setting up additional sensor equipment. The design cost is low, and no additional hardware settings are required. It can be applied to a terminal with a front camera. It can be understood that the processor of the terminal can execute the method in the embodiment of the present invention, obtain the image through the front camera, and finally calculate by the processor to achieve accurate ranging.
[0112] It should be noted that the facial image in the embodiment of the present invention can be obtained by directly recognizing the user's face by the front camera. In one embodiment, the facial image is obtained by cropping the image obtained by the front camera. For example, the terminal obtains an image through the front camera. The image contains the user's face and may also contain some other debris, which may interfere with pupil recognition. Therefore, the embodiment of the present invention crops the image to obtain a facial image by cropping the user's facial area to improve the accuracy of pupil position recognition.
[0113] In one embodiment, the front camera is an under-screen camera and the display screen is an OLED screen, so the front camera can be set below the display screen. Specifically, the under-screen camera is set at the center of the display screen. By setting it at the center of the display screen and obtaining the user's facial image here, the user's pupil distance can be measured more accurately, thereby achieving higher-precision ranging.
[0114] In addition, when the front camera is an under-screen camera, it can also better capture the user's hand image and better identify the interaction between the user's hand and the 3D object to be operated.
[0115] Reference Figure 13 As shown, in one embodiment, the above step S602 may also include but is not limited to steps S701 to S704.
[0116] Step S701 : converting the facial image into a grayscale image and performing binarization processing on the grayscale image to obtain a first pre-processed image.
[0117] Step S702 : performing erosion and dilation processing on the first pre-processed image and removing noise in the image to obtain a second pre-processed image.
[0118] Step S703 : extracting the position of the circular area representing the user's pupil in the second pre-processed image using the circular structuring element.
[0119] Step S704: Calculate the center point of the circular area to obtain first pupil position information and second pupil position information of the user's eyes.
[0120] It should be noted that, in the embodiment of the present invention, the facial image is processed to obtain the pupil position information of the user's face. Specifically, Figure 14 As shown, the facial image is first converted into a grayscale image, and the grayscale image is binarized to obtain a first preprocessed image. To obtain the eyeballs in the binarized first preprocessed image, an opening operation can be performed on the image through a circular structuring element. The image is first corroded and expanded. After corrosion and expansion, there is still noise in the central circular area, and this noise needs to be removed to obtain a second preprocessed image. Finally, the position of the circular area representing the user's pupil in the second preprocessed image is extracted based on the circular structuring element. The circular area will be marked at the corresponding position in the entire facial image. Therefore, the center point of the circular area of the user's left and right eyes can be calculated to obtain the first pupil position information and the second pupil position information of the user's two eyes.
[0121] It can be understood that after the first pupil position information and the second pupil position information are obtained in the embodiment of the present invention, the screen pupil distance can be calculated. For example, in one embodiment, the first pupil position information and the second pupil position information obtained are both coordinate information. The screen pupil distance can be obtained by calculation based on the two coordinate information. The screen pupil distance is the pupil distance in the user's facial image captured by the front camera, and is not the user's pupil distance in the display. The first distance from the user's eyes to the display screen can be calculated based on the size of the screen pupil distance.
[0122] Reference Figure 15 As shown, in one embodiment, the above step S604 may also include but not be limited to steps S801 to S803.
[0123] Step S801: Obtain a preset standard pupil distance.
[0124] Step S802 , obtaining the focal length of the facial image captured by the front camera, and obtaining the initial distance from the facial image to the imaging point according to the focal length.
[0125] Step S803 : obtaining a first ratio according to the interpupillary distance of the picture and the standard interpupillary distance, and obtaining a first distance from the user's eyes to the 3D display screen according to the first ratio and the initial distance.
[0126] It should be noted that in calculating the first distance from the user's eyes to the display screen based on the pupillary distance in the image, specifically, embodiments of the present invention first obtain a preset standard pupillary distance. The standard pupillary distance is the user's actual pupillary distance. The standard pupillary distance can be a default setting, for example, 63 mm. Alternatively, the standard pupillary distance can be user-entered, so the user can accurately enter the pupillary distance. Alternatively, through big data and artificial intelligence analysis, it can be found that the pupillary distances of people of different ages and genders are different. By replacing this data analysis conclusion with the pupillary distance of 63 mm for adults, a more accurate pupillary distance and a more accurate first distance can be obtained. Subsequently, the focal length of the facial image captured by the front camera is obtained, and the initial distance from the facial image to the imaging point is obtained based on the focal length. Finally, a first ratio is obtained based on the pupillary distance in the image and the standard pupillary distance. The first distance from the user's eyes to the display screen can be obtained based on the first ratio and the initial distance.
[0127] It is understandable that each camera should have a certain field of view (FOV) and focal length when shooting. The focal length and field of view of each camera are one-to-one corresponding and can be obtained through public means or measured. The field of view is the angle between the two ends of the camera's cone of view, and the focal length is the distance from the camera lens to the internal "sensor". However, in actual cameras, the sensor is behind the lens. For simplicity, it can be assumed that the lens is in front of the sensor. Relative to the lens mirror, we can get Figure 16 In the picture shown, for example, the plane where the sensor is located is the plane where the facial image is located. The resulting facial image is equivalent to being above the plane where the lens is located. The position of the lens can be described as the imaging point in the embodiment of the present invention. The plane where the imaging point is located is below the plane where the facial image is located and is arranged parallel to it. Therefore, the position of the plane where the facial image is located relative to the plane where the imaging point is located can be obtained based on the focal length. In one embodiment, the distance between the plane where the facial image is located and the plane where the imaging point is located can be obtained based on the focal length, which is defined as the initial distance.
[0128] It can be understood that the plane where the facial image is located corresponds to the plane where the display screen is located, which is determined by the wide angle and focal length of the front camera. In one embodiment, the plane where the facial image is located is the plane where the display screen is located. Alternatively, the plane where the display screen is located can be obtained by adding or subtracting a small distance from the plane where the facial image is located. This can be calculated in advance based on the physical parameters of the front camera used and applied in subsequent processing. The embodiment of the present invention takes the plane where the facial image is located as the plane where the display screen is located as an example.
[0129] It should be added that, in the embodiment of the present invention, the initial distance is obtained based on the focal length, and the initial distance can also be obtained by obtaining the field of view angle of the shooting. However, since the field of view angle and the focal length are in a one-to-one correspondence, the focal length is obtained for processing as an example. It should be noted that the initial distance can be calculated based on the characteristics of the camera imaging, or it can be measured in advance. However, it can be understood that each different focal length will correspond to an initial distance, and no specific limitation is made here.
[0130] It is understandable that if Figure 17 As shown, based on the characteristics of camera imaging, a line segment of the user's actual pupillary distance forms a triangle with the imaging point, and a line segment of the screen pupillary distance is located within the triangle and is parallel to the line segment of the actual pupillary distance. In one embodiment, the triangle formed by the line segment of the screen pupillary distance and the imaging point and the triangle formed by the line segment of the actual pupillary distance and the imaging point are similar triangles. Since the initial distance is known and a first ratio can be obtained between the screen pupillary distance and the actual pupillary distance, a first distance between the user's eyes and the display screen can be obtained based on the initial distance and the first ratio.
[0131] It should be noted that, in the process of obtaining the first distance from the user's eyes to the display screen, calculation is performed based on the characteristics of the triangle. In one embodiment, the definition is Figure 17 The triangle between the line segment where the pupil distance in the middle picture is located and the imaging point is the first triangle, and the definition is Figure 17 The triangle between the line segment where the actual pupil distance is located and the imaging point is the second triangle, and the first triangle and the second triangle are similar triangles.
[0132] like Figure 17 In the example, the distance from the user's eye to the imaging point can be obtained based on the first ratio and the initial distance. The first ratio is the screen pupil distance Q divided by the actual pupil distance K. The initial distance H0 is then divided by the first ratio to obtain the distance H1. Finally, the initial distance H0 is subtracted from the distance H1 to obtain the first distance H from the user's eye to the display screen. The calculation formula for H is as follows:
[0133] H=H1-H0 (1)
[0134] H1=H0 / (Q / K) (2)
[0135] In one embodiment, the present invention can obtain a more accurate first distance based on the rotation angle of the user's face for correction. Specifically, when a user views a display screen, they may view it at a certain angle. For the front camera, the image obtained is a two-dimensional plane image. The angle of rotation of the user's face cannot be distinguished based solely on the image. If the pupil distance is directly calculated at this time, it will cause errors, resulting in inaccurate distance measurement. Therefore, geometric calculations can be performed based on the specific position of the front camera to correct the parameters. In addition, when the pupil is not directly in front of the front camera, correction is also performed based on geometric principles. When the display screen and the front camera are not on the same plane, correction can also be performed based on the distance difference. Compared with the embodiment of the present invention, no specific limitations are imposed.
[0136] Reference Figure 18 As shown, in one embodiment, the above step S604 may also include but not limited to steps S901 to S902.
[0137] Step S901: Obtain a preset distance lookup table.
[0138] Step S902 : obtaining a first distance from the user's eyes to the 3D display screen by looking up the distance lookup table according to the pupil distance of the picture.
[0139] It should be noted that the first distance in the embodiment of the present invention can also be obtained by querying a preset distance lookup table. Specifically, in the embodiment of the present invention, a mapping relationship table from the pupil distance of the screen to the first distance can be pre-established. In the process of calculating the first distance, the preset distance lookup table is first obtained, and the first distance from the user's eye to the display screen is obtained by querying the distance lookup table based on the measured pupil distance of the screen.
[0140] It should be noted that the distance lookup table in the above embodiment can be calculated based on the data in the sample. It can be understood that when it is necessary to reduce errors by measuring the user's facial proportions, the distance lookup table can also be established based on the rotation angle. No specific restrictions are made here.
[0141] Reference Figure 19 As shown, in one embodiment, the above step S604 may also include but not be limited to steps S1001 to S1003.
[0142] Step S1001 , obtaining a reference distance, a reference object size, and a picture size corresponding to the reference object photographed by a front camera.
[0143] Step S1002: Obtain a preset standard pupil distance.
[0144] Step S1003 : obtaining a first distance between the user's eyes and the 3D display screen according to the reference distance, the reference object size, the screen size, the screen pupil distance, and the standard pupil distance.
[0145] It should be noted that, in an embodiment of the present invention, the first distance can also be obtained by establishing a reference system. Specifically, in an embodiment of the present invention, a reference object is first placed in front of the terminal, and the reference distance from the reference object to the display screen and the object size of the reference object are measured. The reference object is imaged by the front camera, and the size of the reference object in the image is calculated in the image to obtain the screen size. Subsequently, a reference system can be established based on this. Therefore, by obtaining a preset standard pupillary distance, the first distance from the user's eyes to the display screen can be obtained based on the reference distance, the reference object size, the screen size, the screen pupillary distance and the standard pupillary distance.
[0146] Specifically, the present invention first obtains a first coefficient by dividing the standard pupillary distance by the screen pupillary distance, then obtains a second coefficient by assuming that the reference object size is within the screen size, and finally obtains a third coefficient by dividing the reference distance by the second coefficient. Finally, the first distance from the user's eyes to the display screen is obtained by multiplying the third coefficient by the first coefficient. In addition, the screen size and the screen pupillary distance can both be calculated based on the pixels of the display screen. For example, Figure 20 and Figure 21 In the example, when the object size of the standard reference object is 10 cm, the reference distance is 50 cm, the screen size is AB, and the standard pupillary distance is 6.3 cm, the screen pupillary distance is ab. At this time, the first coefficient is 6.3 ÷ ab, and the second coefficient is 10 ÷ AB. Finally, the formula for the first distance h can be established as follows:
[0147] 50÷(10÷AB)=h÷(6.3÷ab) (3)
[0148] Since the screen size AB and the screen pupil distance ab are known, the first distance h can be obtained according to formula (3).
[0149] Reference Figure 22 As shown, in one embodiment, the above step S102 may also include but not be limited to steps S1101 to S1103.
[0150] Step S1101: Obtain a negative disparity value for a terminal to display a 3D image.
[0151] Step S1102 : Obtain a third distance according to the first distance and the negative parallax value.
[0152] Step S1103 : A position between the user and the 3D display screen and at a third distance from the 3D display screen is determined as a first screen point position.
[0153] It should be noted that, in the embodiment of the present invention, the first screen point position can be determined according to the set negative disparity value. The human-computer interaction method in the embodiment of the present invention first obtains the negative disparity value of the terminal for 3D image display, as described above. Figure 1 As shown, when the target object is deviated to the left in the left-eye image and to the right in the right-eye image, the focal length (convergence point) of the user's eyes will be guided to fall in front of the display screen. This phenomenon is called negative parallax. The stereoscopic effect can be seen due to the existence of parallax. The larger the negative parallax, the closer it will be to the audience. The present invention can set a required negative parallax value, which represents the percentage of the 3D virtual image out-of-screen distance to the first distance between the user's eyes and the display screen to determine the position where the 3D picture appears. Subsequently, the third distance is obtained based on the first distance and the negative parallax value, wherein the first distance is multiplied by the percentage represented by the negative parallax value. Finally, the position between the user and the 3D display screen and the third distance away from the 3D display screen is determined as the first screen point position.
[0154] Specifically, when the percentage represented by the negative parallax value in the embodiment of the present invention is set to -20%, it means that the 3D picture will appear at a position 20% between the user's eyes and the display screen. When the terminal obtains the distance from the user's eyes to the display screen as 50 cm through the front camera, it can be obtained that the distance from the 3D picture viewed by the user to the display screen is 10 cm. Therefore, after obtaining the position where the 3D picture is displayed according to the negative parallax value, the terminal can obtain the user's hand movements at this position, thereby realizing virtual interaction.
[0155] For another example, when the negative parallax value in an embodiment of the present invention is set to -20%, and the distance between the user's eyes and the terminal is 50 cm, the terminal obtains the hand movements of the user at a distance of 10 cm from the terminal or the 3D display screen. This is because the user can view the 3D image at a distance of 10 cm from the 3D display screen according to the set negative parallax value, and then the input interface of the terminal (such as a virtual keyboard) can appear as a 3D interactive image at this position. After the user views the input interface at this position, he or she makes gestures at this position, and the terminal then obtains the hand movements of the user at this position, thereby realizing a set of virtual interaction methods based on 3D display.
[0156] Reference Figure 23 As shown, in one embodiment, after the above step S102, steps S1201 to S1203 may also be included but not limited to.
[0157] Step S1201: After the user's eyes move, a fourth distance between the user's eyes and the 3D display screen after the movement is acquired.
[0158] Step S1202: Obtain a fifth distance according to the fourth distance and the negative parallax value.
[0159] Step S1203 : updating the first screen point position, and updating the position between the user and the 3D display screen and at a fifth distance from the 3D display screen as the first screen point position.
[0160] It should be noted that when the terminal recognizes that the distance between the user's eyes and the terminal has changed, the position of the 3D image will also change. The terminal will update the position of the 3D image in real time based on the negative parallax value and the distance between the user's eyes and the display screen, thereby obtaining the user's hand movements at the correct position and completing the virtual interaction based on the 3D display. Specifically, when the user's eyes move, the fourth distance between the user's eyes and the 3D display screen after the movement is obtained. The fourth distance is the distance from the current user to the display screen after the movement. The fifth distance is obtained based on the fourth distance and the negative parallax value to update the first screen point position, and the position between the user and the 3D display screen and the fifth distance from the 3D display screen is updated to the first screen point position.
[0161] Reference Figure 24 As shown, an embodiment of the present invention further provides a human-computer interaction device, the device comprising:
[0162] The first module 2401 is configured to control the 3D display screen to display a 3D object to be operated.
[0163] The second module 2402 is used to obtain a first distance between the user's eyes and the 3D display screen, determine a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determine a first visual plane where the user views the object to be manipulated based on the first screen point position.
[0164] The third module 2403 is used to obtain a first hand motion of the user on the first visual plane, and match the coordinate position of the object to be operated according to the first hand motion.
[0165] The fourth module 2404 is configured to obtain input information corresponding to the object to be operated according to the coordinate position.
[0166] It should be noted that the human-computer interaction device in the embodiments of the present invention can implement the human-computer interaction method in any of the above-mentioned embodiments. The human-computer interaction device can be a terminal device such as a mobile phone, a tablet computer, or a 3D vision training terminal. The human-computer interaction device controls a 3D display screen to display a 3D object to be operated, then obtains a first distance between the user's eyes and the 3D display screen, and determines a first screen point position between the user's eyes and the 3D display screen based on the first distance. Since the object to be operated is displayed in 3D, the 3D object to be operated is displayed on the plane where the first screen point position is located based on a three-dimensional spatial image virtualized by binocular parallax. Therefore, the first visual plane where the user views the object to be operated is determined based on the first screen point position. Then, a first hand gesture of the user on the first visual plane is obtained, and a coordinate position of the object to be operated is matched based on the first hand gesture. Finally, input information corresponding to the object to be operated can be obtained based on the coordinate position. In the embodiments of the present invention, human-computer interaction can be performed based on the user's hand gesture when the terminal is performing 3D display. Different input information on the object to be operated is obtained based on different hand gestures. Interaction can be achieved without touching the terminal, thereby improving the user's interactive experience.
[0167] It should be noted that the above-mentioned first module 2401, second module 2402, third module 2403 and fourth module 2404 can be various functional modules on the terminal. In one embodiment, the above-mentioned modules can all be various functional modules in the processor and can all be executed by the processor set on the terminal. The hardware structure of the human-computer interaction device in the embodiment of the present invention is only embodied in the form of functional modules, and does not represent a limitation on the embodiment of the present invention.
[0168] Figure 25 The electronic device 2500 provided by an embodiment of the present invention is shown. The electronic device 2500 includes: a processor 2501, a memory 2502, and a computer program stored in the memory 2502 and executable on the processor 2501. When the computer program is executed, it is used to execute the above-mentioned human-computer interaction method.
[0169] The processor 2501 and the memory 2502 may be connected via a bus or other means.
[0170] Memory 2502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs, such as the human-computer interaction method described in the embodiments of the present invention. Processor 2501 implements the human-computer interaction method by executing the non-transitory software programs and instructions stored in memory 2502.
[0171] The memory 2502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; and the data storage area may store and execute the above-mentioned human-computer interaction method. In addition, the memory 2502 may include a high-speed random access memory 2502, and may also include a non-transitory memory 2502, such as at least one storage device memory device, a flash memory device or other non-transitory solid-state memory device. In some embodiments, the memory 2502 may optionally include a memory 2502 remotely arranged relative to the processor 2501, and these remote memories 2502 may be connected to the electronic device 2500 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0172] The non-transient software program and instructions required to implement the above human-computer interaction method are stored in the memory 2502. When executed by one or more processors 2501, the above human-computer interaction method is executed, for example, Figure 3 Steps S101 to S104 of the method, Figure 6 Steps S201 to S205 of the method, Figure 7 Steps S301 to S302 of the method, Figure 8 Steps S401 to S402 of the method, Figure 10 Steps S501 to S502 of the method, Figure 11 Steps S601 to S604 of the method, Figure 13 Steps S701 to S704 of the method, Figure 15 Steps S801 to S803 of the method, Figure 18 Steps S901 to S902 of the method, Figure 19 Steps S1001 to S1003 of the method, Figure 22 Steps S1101 to S1103 of the method, Figure 23 Method steps S1201 to S1203.
[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0174] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, storage device storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0175] It should also be understood that the various implementations provided in the embodiments of the present invention can be arbitrarily combined to achieve different technical effects. The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above implementation. Those skilled in the art can make various equivalent modifications or substitutions under the same conditions without violating the spirit of the present invention.
Claims
1. A human-computer interaction method, applied in a terminal, wherein the terminal is provided with a 3D display screen and a front camera, characterized in that: The method comprises: Controlling the 3D display screen to display a 3D object to be operated; Acquire a facial image of the user, where the facial image is captured by the front camera; Performing pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes; Calculating an inter-pupillary distance of the user in the facial image based on the first pupil position information and the second pupil position information; calculating a first distance between the user's eyes and the 3D display screen based on the interpupillary distance of the screen, determining a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determining a first visual plane where the user views the object to be operated based on the first screen point position; Acquire a second distance from the user's hand to the 3D display screen; If the second distance indicates that the user's hand is located on the first visual plane, obtaining an image of the user's hand, where the image of the hand is captured by the front camera; Recognizing a hand shape in the hand image; Obtaining a preset target hand shape, and performing a matching analysis between the hand shape in the image and the target hand shape; If the hand shape on the screen matches the target hand shape, determining that the hand shape on the screen is a first hand gesture, obtaining a hand position corresponding to when the user performs the first hand gesture, and matching the coordinates of the hand position with the coordinates on the first visual plane to obtain a coordinate position corresponding to the object to be operated; The input information corresponding to the object to be operated is obtained according to the coordinate position.
2. The human-computer interaction method according to claim 1, characterized in that: The object to be operated includes a virtual keyboard, and obtaining input information corresponding to the object to be operated according to the coordinate position includes: Obtaining a coordinate mapping relationship between the key value of each key on the virtual keyboard and the corresponding input position; A corresponding target key value is triggered from the virtual keyboard according to the coordinate position and the coordinate mapping relationship.
3. The human-computer interaction method according to claim 1, characterized in that: The hand position is obtained by analyzing the hand image taken by the front camera; Alternatively, the terminal is provided with an infrared sensor or an ultrasonic sensor, and the infrared sensor and the ultrasonic sensor are used to obtain the hand position.
4. The human-computer interaction method according to claim 1, wherein: The method further comprises at least one of the following: Acquiring a second hand motion of the user on the first visual plane, and opening or closing the object to be operated according to the second hand motion; Acquire a third hand motion of the user on the first visual plane, and control the object to be operated to perform a response action according to the third hand motion, where the response action includes zooming in, zooming out, scrolling down, or turning pages.
5. The human-computer interaction method according to claim 4, characterized in that: The first hand action includes one of a press down action, a click operation, a grabbing action or a sliding action, the second hand action includes one of a press down action, a click operation, a grabbing action or a sliding action, and the third hand action includes one of a press down action, a click operation, a grabbing action or a sliding action, and the first hand action, the second hand action and the third hand action are different from each other.
6. The human-computer interaction method according to claim 1, characterized in that: The object to be operated includes at least one of a virtual keyboard, a separate control, or a gesture control.
7. The human-computer interaction method according to any one of claims 1 and 3, characterized in that: The front camera is an under-screen camera, and the under-screen camera is arranged at the center of the 3D display screen.
8. The human-computer interaction method according to claim 1, wherein: The performing pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes includes: Converting the facial image into a grayscale image and performing binarization processing on the grayscale image to obtain a first preprocessed image; Performing erosion and dilation processing on the first preprocessed image and removing noise in the image to obtain a second preprocessed image; extracting a position of a circular area representing the user's pupil in the second preprocessed image using a circular structuring element; The center point of the circular area is calculated to obtain first pupil position information and second pupil position information of the user's eyes.
9. The human-computer interaction method according to claim 7, characterized in that: The calculating a first distance from the user's eyes to the 3D display screen according to the interpupillary distance of the picture includes: Get the preset standard pupil distance; Obtaining a focal length of the facial image captured by the front camera, and obtaining an initial distance from the facial image to an imaging point based on the focal length; A first ratio is obtained according to the screen pupil distance and the standard pupil distance, and a first distance from the user's eyes to the 3D display screen is obtained according to the first ratio and the initial distance.
10. The human-computer interaction method according to claim 7, characterized in that: The calculating a first distance from the user's eyes to the 3D display screen according to the interpupillary distance of the picture includes: Obtain a preset distance lookup table; The first distance from the user's eyes to the 3D display screen is obtained by looking up the distance lookup table according to the pupil distance of the picture.
11. The human-computer interaction method according to claim 7, characterized in that: The calculating a first distance from the user's eyes to the 3D display screen according to the interpupillary distance of the picture includes: Obtaining a reference distance, a reference object size, and a picture size corresponding to the reference object photographed by the front camera; Get the preset standard pupil distance; A first distance from the user's eyes to the 3D display screen is obtained according to the reference distance, the reference object size, the screen size, the screen pupil distance, and the standard pupil distance.
12. The human-computer interaction method according to claim 1, characterized in that: The determining a first screen point position between the user's eyes and the 3D display screen according to the first distance includes: Obtaining a negative disparity value for 3D image display by the terminal; Obtaining a third distance according to the first distance and the negative parallax value; A position between the user and the 3D display screen and at a third distance from the 3D display screen is determined as a first screen point position.
13. The human-computer interaction method according to claim 12, characterized in that: After determining a first screen point position between the user's eye and the 3D display screen according to the first distance, the method includes: When the user's eyes move, obtaining a fourth distance from the user's eyes to the 3D display screen after the movement; Obtaining a fifth distance according to the fourth distance and the negative parallax value; The first screen point position is updated, and a position between the user and the 3D display screen and at a fifth distance from the 3D display screen is updated as the first screen point position.
14. A human-computer interaction device, characterized in that: include: The first module is used to control the 3D display screen to display the 3D object to be operated; The second module is used to obtain a facial image of the user, where the facial image is captured by a front-facing camera; Performing pupil recognition on the facial image to determine first pupil position information and second pupil position information of the user's eyes; Calculating an inter-pupillary distance of the user in the facial image based on the first pupil position information and the second pupil position information; calculating a first distance between the user's eyes and the 3D display screen based on the interpupillary distance of the screen, determining a first screen point position between the user's eyes and the 3D display screen based on the first distance, and determining a first visual plane where the user views the object to be operated based on the first screen point position; The third module is used to obtain a second distance between the user's hand and the 3D display screen; if the second distance indicates that the user's hand is located on the first visual plane, obtain the user's hand image, and the hand image is captured by the front camera; identify the hand shape in the hand image; obtain a preset target hand shape, and perform a matching degree analysis between the hand shape in the screen and the target hand shape; if the hand shape in the screen matches the target hand shape, determine that the hand shape in the screen is a first hand action, obtain the hand position corresponding to the user when making the first hand action, and match the coordinates of the hand position with the coordinates on the first visual plane to obtain the coordinate position corresponding to the object to be operated; The fourth module is used to obtain input information corresponding to the object to be operated according to the coordinate position.
15. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the human-computer interaction method according to any one of claims 1 to 13 when executing the computer program.
16. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the human-computer interaction method according to any one of claims 1 to 13.
Citation Information
Patent Citations
3D virtual touch control man-machine interaction method based on stereoscopic vision
CN104714646A