Human-computer interaction method and system, processing equipment and computer readable storage medium

By fusing the hand image and the virtual view input multi-view image into the network, the problem of hand space information deviation in mobile XR devices is solved, and a natural human-computer interaction experience is achieved and the effect of improving user experience is achieved.

CN120045053APending Publication Date: 2025-05-27BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311594923.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When existing mobile XR devices realize natural human-computer interactive experience, there is a deviation in hand space information, resulting in poor user experience.

Method used

By fusing the hand image and the virtual view input multi-view image into the network, a fused image of the virtual view and the hand image is generated, overcoming the hand space information deviation and realizing a natural human-computer interactive experience.

Benefits of technology

It improves the user experience and realizes the effect of hand-eye integration, so that gesture control can more truly reflect people's interactive behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045053A_ABST
    Figure CN120045053A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction method and system, processing equipment and a computer readable storage medium. The method comprises the following steps: controlling a head-mounted display to present a virtual view; receiving a hand image collected by a gesture recognition sensor; detecting whether fusion display is needed; when fusion display is needed, inputting the hand image and the virtual view into a multi-view image fusion network to obtain a fusion image of the virtual view and the hand image; and controlling the head-mounted display to present the fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to intelligent display technology, and particularly to a human-computer interaction method, system, processing device, and computer-readable storage medium. Background Art

[0002] Extended Reality (XR) refers to an environment that combines the real world and the virtual world and enables human-computer interaction through computer technology and wearable devices. XR technologies include Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), etc., and can be widely applied in many fields such as entertainment, gaming, healthcare, advertising, industry, online education, and engineering. Among them, VR applications are computer-generated 3D environments that allow users to be completely immersed in the virtual world presented by the head-mounted device without seeing the real environment. AR applications superimpose computer-generated virtual information and images on the real world and can be experienced through electronic devices such as smartphones and AR glasses. MR applications are comprehensive applications of AR+VR, which mix the real world and the virtual world to generate a new visual and interactive environment that contains both physical entities and virtual information. The characters and objects in the real world and the virtual world can cross the real boundary to create a more complex and exciting experience. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0004] The present disclosure provides a human-computer interaction method, including: controlling a head-mounted display to present a virtual view; receiving a hand image collected by a gesture recognition sensor; detecting whether fusion display is required; when fusion display is required, inputting the hand image and the virtual view into a multi-view image fusion network to obtain a fusion image of the virtual view and the hand image; and controlling the head-mounted display to present the fusion image.

[0005] An embodiment of the present disclosure further provides a processing device, including: a processor and a memory storing a computer program that can run on the processor, wherein the processor implements the steps of the human-computer interaction method as described above when executing the program.

[0006] An embodiment of the present disclosure further provides a human-computer interaction system, including: a gesture recognition sensor, a head-mounted display, and the processing device as described above, wherein: the gesture recognition sensor is used to capture a hand image of a user; the head-mounted display is used to present a virtual view or a fusion image.

[0007] Embodiments of the present disclosure also provide a computer-readable storage medium storing executable instructions that, when executed by a processor, can implement the human-computer interaction method as described in any one of the above.

[0008] The human-computer interaction method, system, processing device, and computer-readable storage medium according to the embodiments of the present disclosure detect whether fusion display is required; when fusion display is required, input the hand image and the virtual view into a multi-view image fusion network to obtain a fused image of the virtual view and the hand image, overcome the deviation of hand spatial information, achieve a natural human-computer interaction experience, and improve the user experience.

[0009] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. Other advantages of the present disclosure can be realized and obtained by the solutions described in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to provide an understanding of the technical solutions of the present disclosure and constitute a part of the specification. They are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.

[0011] Figure 1 It is a schematic flowchart of a human-computer interaction method according to an exemplary embodiment of the present disclosure.

[0012] Figure 2 It is a schematic diagram of the usage method of a head-mounted display and a neck-worn device according to an exemplary embodiment of the present disclosure.

[0013] Figure 3 It is a schematic diagram of a head interaction method according to an exemplary embodiment of the present disclosure.

[0014] Figure 4 It is a schematic diagram of the positions of key hand nodes according to an exemplary embodiment of the present disclosure.

[0015] Figure 5A and Figure 5B It is a schematic diagram of two gesture interaction methods according to an exemplary embodiment of the present disclosure.

[0016] Figure 6 It is a schematic diagram of a human-computer interaction method that integrates the human eye fixation area and control gestures according to an exemplary embodiment of the present disclosure.

[0017] Figure 7A It is a schematic diagram of the structure of a multi-view image fusion network according to an exemplary embodiment of the present disclosure.

[0018] Figure 7B is Figure 7A a schematic diagram of the structure of the flow processing module in

[0019] Figure 8 This is a schematic diagram of the result of multi-view image fusion according to an exemplary embodiment of the present disclosure.

[0020] Figure 9 This is a schematic diagram of the structure of a processing device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0021] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it will be apparent to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope of the embodiments described in the present disclosure. Although many possible combinations of features are shown in the drawings and discussed in the detailed implementation manners, many other combination ways of the disclosed features are also possible. Unless specifically restricted, any feature or element of any embodiment can be combined with any other feature or element in any other embodiment, or can replace any other feature or element in any other embodiment.

[0022] The present disclosure includes and contemplates combinations with features and elements known to those of ordinary skill in the art. The embodiments, features, and elements already disclosed in the present disclosure can also be combined with any conventional features or elements to form unique inventive solutions defined by the claims. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure can be implemented alone or in any suitable combination. Therefore, the embodiments are not subject to other restrictions except those made according to the appended claims and their equivalent replacements. In addition, various modifications and changes can be made within the scope of the protection of the appended claims.

[0023] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not depend on the specific order of the steps described herein, the method or process should not be limited to the specific order of the steps described. As will be understood by those of ordinary skill in the art, other step orders are possible. Therefore, the specific order of the steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can easily understand that these orders can be changed and still remain within the spirit and scope of the embodiments of the present disclosure.

[0024] As Figure 1 shown, the embodiments of the present disclosure provide a human-computer interaction method, including:

[0025] Step 101, control the head-mounted display to present a virtual view;

[0026] Step 102, receive the hand image collected by the gesture recognition sensor;

[0027] Step 103, detect whether fused display is required;

[0028] Step 104, when fused display is required, input the hand image and the virtual view into a multi-view image fusion network to obtain a fused image of the virtual view and the hand image, and control the head-mounted display to present the fused image.

[0029] The human-computer interaction method of the embodiments of the present disclosure overcomes the deviation of hand spatial information by detecting whether fused display is required; when fused display is required, inputting the hand image and the virtual view into a multi-view image fusion network to obtain a fused image of the virtual view and the hand image, realizing a natural human-computer interaction experience and improving the user experience.

[0030] In some exemplary embodiments, detecting whether fused display is required includes:

[0031] Display a hand-shaped button on the virtual view, and the hand-shaped button is used for the user to set whether fused display is required;

[0032] Detect whether an operation of clicking the hand-shaped button is captured;

[0033] When an operation of clicking the hand-shaped button is captured, switch the current state of whether fused display is required.

[0034] In this embodiment, the hand-shaped button can be a virtual button on the current virtual view. When an operation of clicking the hand-shaped button is captured, switch the current state of whether fused display is required, specifically including: switching the state of requiring fused display to the state of not requiring fused display, or switching the state of not requiring fused display to the state of requiring fused display. In some other examples, the hand-shaped button can include two virtual buttons on the current virtual view: one is a display hand button and the other is a close hand button. When the display hand button is clicked, the system state is set to require fused display, and when the close hand button is clicked, the system state is set to not require fused display.

[0035] The human-computer interaction method of the embodiments of the present disclosure can also use other methods to detect whether fusion display is required. For example, a fusion control button (physical button) can be set on the head-mounted display. When the user presses the fusion control button, the system switches the current state of whether fusion display is required. That is, if the system state is that fusion display is not required before the user presses, after the user presses, the system state switches to that fusion display is required; if the system state is that fusion display is required before the user presses, after the user presses, the system state switches to that fusion display is not required. Of course, two fusion control buttons (physical buttons) can also be set on the head-mounted display, one is the display hand button and the other is the close hand button.

[0036] In some other examples, the current state of whether fusion display is required can also be switched by some special control gestures (such as clenching the fist twice, etc.). The embodiments of the present disclosure do not limit this.

[0037] In the embodiments of the present disclosure, as Figure 2 shown, a head-mounted display (HMD), that is, a head-mounted display, is worn on the user's head and displays images in front of the user's eyes. The head-mounted display generally forms the shape of goggles or the frame shape of large glasses. The head-mounted display usually includes a display device and a headband. The display device has a high-resolution liquid crystal display panel and is usually called glasses because it is made in the shape of eyes, and is used to display the same (or different) video images to the left and right eyes; the headband is used to connect the display device and wear the display device on the user's head and face.

[0038] In the embodiments of the present disclosure, the playback mode of the head-mounted display can be a 3D playback mode. At this time, the virtual view presented by the head-mounted display is a 3D virtual view; the playback mode of the head-mounted display can also be a 2D playback mode. At this time, the virtual view presented by the head-mounted display is a 2D virtual view. In the embodiments of the present disclosure, the head-mounted display can present a 3D virtual view or a 2D virtual view according to the user's settings.

[0039] In the embodiments of the present disclosure, the head-mounted display includes a projector mechanism ( Figure 2 not shown in the figure) for projecting or displaying frames including left and right images to the user's eyes, so as to provide a 3D virtual view to the user.

[0040] In some exemplary embodiments, the parallax images can be provided to the left and right eyes in sequence according to the time series. Since the images obtained by our left and right eyes are different, a sense of space is generated in the brain. By using the head-mounted display, the left and right eyes can observe images with slight differences, so as to observe a three-dimensional virtual view. At this time, assuming that the refresh rate of the head-mounted display is 120 Hz, then when playing the frame sequence, the playback frame sequence of the left and right eyes is 60 Hz each.

[0041] In some exemplary embodiments, the head-mounted display may include a head motion tracking sensor ( Figure 2 not shown in the figure), and the X, Y, Z axes and the front and back sides can be tracked through the head motion tracking sensor.

[0042] In some exemplary embodiments, the head motion tracking sensor may include a gyroscope and an accelerometer. The gyroscope is used to measure the rotation angle of the head, and the accelerometer is used to measure the acceleration of the head. The gyroscope and the accelerometer can be integrated into a six-axis inertial sensor. In some other exemplary embodiments, the head motion tracking sensor may further include a magnetometer, which can be used to measure the direction of the head relative to the earth's magnetic field. The gyroscope, the accelerometer and the magnetometer can be integrated into a nine-axis sensor.

[0043] In some exemplary embodiments, the method further includes:

[0044] Obtaining data from the head motion tracking sensor, determining head motion parameters according to the data of the head motion tracking sensor, and controlling the playback of the virtual view according to the head motion parameters.

[0045] For example, in the 3D playback mode, by turning the head, the user can freely observe all corners of the virtual environment. This feature of free observation is closer to the observation experience of the real world. At this time, the user can choose to exit the 3D playback mode and enter the 2D playback mode. As Figure 3 shown, in the 2D playback mode, the system can appropriately reduce the data reporting rate of the head motion tracking sensor because the system does not need to simulate the observation experience of the real world at this time. In the 2D playback mode, obtaining data from the head motion tracking sensor, judging the head motion direction through head tracking, and video switching, volume adjustment, etc. can be performed according to the head motion direction. For example, when the user looks down, a playback progress bar will appear on the virtual view interface. At this time, when the user turns the head from left to right, the video will fast forward by the corresponding progress; correspondingly, when the user turns the head from right to left, the video will rewind by the corresponding progress; when the user turns the head from bottom to top, the volume can be increased, and when the user turns the head from top to bottom, the volume can be decreased. After the adjustment is completed, the user can choose to exit the 2D playback mode and re-enter the 3D playback mode to view the three-dimensional virtual view. In the 3D playback mode, the system needs to increase the data reporting rate of the head motion tracking sensor because the system needs to simulate the observation experience of the real world at this time.

[0046] In some exemplary embodiments, the method further includes: obtaining the user's eye image collected by the eye tracking sensor, determining the area of the human eye's fixation according to the eye image, and controlling the playback of the virtual view according to the area of the human eye's fixation.

[0047] In the embodiments of the present disclosure, the head-mounted display may include an eye-tracking sensor ( Figure 2 not shown in the figure), which tracks the position and movement of the user's eyes through the eye-tracking sensor, and then determines the human eye fixation area. In some exemplary embodiments, the eye-tracking sensor may include a red-green-blue (RGB) camera and an infrared (IR) camera with an infrared light source. The position of the pupil center is determined through the user image captured by the RGB camera, and the position of the infrared light spot is detected through the human eye image with the infrared light spot captured by the IR camera. Through the positions of the pupil center and the two infrared light spots, the three-dimensional line of sight of both eyes is calculated using the pupil corneal reflection algorithm, and the area where the two lines of sight intersect in space is the human eye fixation area.

[0048] In some exemplary embodiments, the virtual view is controlled for playback according to the human eye fixation area, including: performing high-definition rendering on the virtual view in the human eye fixation area and performing low-definition rendering on the virtual view in the area not fixated by the human eye. The embodiments of the present disclosure reduce the bandwidth of the display by performing high-definition display locally. In some embodiments, the high-definition rendering described in the embodiments of the present disclosure refers to that the format of the output image frame is above 1080P (p means Progressive, progressive scanning), and correspondingly, the low-definition rendering described in the embodiments of the present disclosure refers to that the format of the output image frame is below 1080P.

[0049] Exemplarily, in some embodiments, there may be two eye-tracking sensors, and the two eye-tracking sensors respectively track the user's left eye and right eye. The information collected by the eye-tracking sensors can be used to adjust the rendering of the image to be projected and / or to adjust the projection of the image through the projection system of the HMD based on the direction and angle of the user's eye fixation. For example, in some embodiments, the content of the image in the human eye fixation area can be rendered with more details and can be at a higher resolution than the content in the area not fixated by the human eye, which allows the available image data processing time to be spent on the content viewed by the central area of the retina of the eye rather than on the content viewed by the peripheral area of the eye. Similarly, the content of the image in the area not fixated by the human eye can be compressed more than the content in the human eye fixation area.

[0050] In some exemplary embodiments, the eye-tracking sensor can also be used to track the dilation of the user's pupils. In some embodiments, the brightness of the projected image can be adjusted based on the dilation of the user's pupils determined by the eye-tracking sensor.

[0051] In the embodiments of the present disclosure, the gesture recognition sensor is used to track the position, movement, and posture of the user's hand, finger, and / or arm. By tracking the position, movement, and posture of the user's hand, finger, and / or arm, the operations performed by the user in the three-dimensional environment are determined, including but not limited to playing video games, navigating menus, controlling media playback, etc. For example, in some embodiments, the detected position, movement, and posture of the user's hand, finger, and / or arm can be used to simulate the movement of the hand, finger, and / or arm of the user avatar in the virtual space in the virtual view. As another example, the detected hand and finger postures of the user can be used to determine the interaction between the user and the virtual content in the virtual space, including but not limited to the posture of manipulating virtual objects, gestures for interacting with virtual user interface elements displayed in the virtual space, etc.

[0052] In the embodiments of the present disclosure, as Figure 2 shown, the gesture recognition sensor can be located on the neck-worn device.

[0053] Currently, in the operation of many games and applications on XR devices, phenomena such as screen delay and motion feedback occur, such as motion experience and simulation games, etc.; the screen is not as delicate and vivid as that of PC-level XR games. If people are to socialize, play games, and work in the metaverse, it is very difficult for existing mobile XR devices to meet the requirements of being both lightweight and comfortable and having fast computing and rendering speeds. Limited by processor capabilities, device space size, heat dissipation and other indicators, XR devices have now encountered an industry ceiling in the balance between lightweight and high performance. The present disclosure sets up a neck-worn device. As Figure 2 shown, the neck-worn device can be connected to the head-mounted display by wire or wirelessly. The neck-worn device can send the demand exceeding the local computing power to the cloud for processing, without increasing the weight of the head-mounted display and can also power the head-mounted display, improving the device performance and achieving lightweight. In addition, compared with the gesture recognition sensor set on the traditional head-mounted display, the gesture recognition sensor set on the neck-worn device in the present disclosure has a lower shooting range, and the position of the user's hand can be lower when in use, which is not easy to cause fatigue due to long-term lifting.

[0054] However, since the neck-worn device is located on the neck during use, and the relative position between the head and the neck cannot be fixed, the user's vision and gesture space do not belong to the same coordinate system and need to be fused. The present disclosure fuses the hand image and the virtual view into a multi-view image fusion network. The multi-view image fusion network is an end-to-end learnable architecture. The multi-view image fusion network first extracts the features on all input images, then warps the calculated features to the reference view, and finally fuses the information from all images to generate a higher-quality fusion result, obtaining a fusion image of the virtual view and the hand image, overcoming the deviation of hand space information, realizing a natural human-computer interaction experience, and improving the user's experience.

[0055] In some exemplary embodiments, the method further includes: mapping the hand image to the user's visual coordinate system, detecting the user's control gesture through a gesture recognition algorithm, and performing playback control on the virtual view according to the user's control gesture.

[0056] In the embodiments of the present disclosure, the user's control gesture is detected through a gesture recognition algorithm. The detection method can be real-time detection or detection according to preset rules. For example, when the user's hand reaches into a preset space, detection starts, that is, detecting the user's hand includes two situations: detected and undetected; among them, real-time detection of the user's hand is a better solution.

[0057] In some exemplary embodiments, an image of the hand captured by a gesture recognition sensor is obtained, 21 key nodes of the hand are distinguished according to light and darkness, and then the three-dimensional coordinates of each key node are obtained. In this embodiment, the specific process of detecting the user's control gesture through a gesture recognition algorithm can be implemented by using existing gesture recognition methods, and the present disclosure does not limit this.

[0058] In one exemplary embodiment, multiple joint points of the hand can be selected as key nodes. For example, the joint points that mainly work when the hand grasps an object can be selected as key nodes. Exemplarily, as Figure 4 shown, 21 parts of the hand can be selected as the key nodes of the hand, and the position data of the key nodes are correspondingly established. Among them, the 21 parts of the hand include 4 parts for each of the 5 fingers. Among them, the fingertip is 1 part, and 3 joints correspond to 3 parts, and 1 part at the wrist. The position data of the 21 key nodes of the hand are established in the three-dimensional space (i.e., the user's visual coordinate system) where the virtual view is located. For example, they can be P1(x1, y1, z1), P2(x2, y2, z2),..., P21(x21, y21, z21); where the corresponding P1 is the fingertip of the thumb, P2 is the joint connection of the first phalanx of the thumb, P3 is the joint connection of the second phalanx of the thumb, P4 is the end of the thumb, P5 is the fingertip of the index finger, P6 is the joint connection of the first phalanx of the index finger, P7 is the joint connection of the second phalanx of the index finger, P8 is the end of the index finger, P9 is the fingertip of the middle finger, P10 is the joint connection of the first phalanx of the middle finger, P11 is the joint connection of the second phalanx of the middle finger, P12 is the end of the middle finger, P13 is the fingertip of the ring finger, P14 is the joint connection of the first phalanx of the ring finger, P15 is the joint connection of the second phalanx of the ring finger, P16 is the end of the ring finger, P17 is the fingertip of the little finger, P18 is the joint connection of the first phalanx of the little finger, P19 is the joint connection of the second phalanx of the little finger, P20 is the end of the little finger, and P21 is the end of the palm.

[0059] In some exemplary embodiments, the user's control gestures may include, but are not limited to, any one or more of the following gestures: click gesture, swipe gesture, pan gesture, pinch gesture, rotate gesture, long - press gesture, fist gesture, etc.

[0060] Among them, the click gesture can be a single click or multiple clicks, and can be made with one finger or multiple fingers. The click gesture is the most commonly used gesture.

[0061] The sliding direction of the swipe gesture can be in any direction: up, down, left, or right. Exemplarily, when viewing photos or icons, we can use the swipe gesture to slide from one photo or icon to another. The movement direction of the hand is recognized through multiple consecutive frames of images to control the movement of the interface photos or icons. The arrangement of photos or icons can be switched in two ways: rotation in the horizontal direction and rotation in the vertical direction. As Figure 5A and Figure 5B shown, when the user's control gesture is a left - to - right swipe, it can control multiple photos or icons on the current screen to switch from left to right. When the user's control gesture is a bottom - to - top swipe, it can control multiple photos or icons on the current screen to switch from bottom to top. The foremost icon is the current selection, and the selection is confirmed through a click gesture or a fist gesture.

[0062] The pan gesture, which can also be called the drag gesture, requires that one or more fingers remain pressed on the view when the user pans the view.

[0063] The pinch gesture requires two fingers to touch the view simultaneously. When the two fingers approach each other, the view shrinks; when the two fingers move away from each other, the view enlarges.

[0064] The rotate gesture requires two fingers to touch the view simultaneously. When the user's fingers move in a circular motion relative to each other, the corresponding view rotates in the same direction and at the same speed.

[0065] To successfully trigger the long - press gesture, one or more fingers need to press on the view for no less than a pre - set minimum press duration (for example, it can be 0.5 seconds), and the distance that the finger moves during the long - press should be less than the pre - set allowable movement distance (for example, it can be 10 pixel points).

[0066] The fist gesture requires four fingers to curl towards the palm, with the thumb wrapped outside or inside the four fingers. The fist gesture can be used to confirm a selection, switch the state of whether fusion display is required, etc.

[0067] In some exemplary embodiments, as Figure 6 shown, the method further includes:

[0068] Obtain the human eye fixation area and the user's control gestures, and perform playback control on the virtual view according to the human eye fixation area and the user's control gestures.

[0069] Exemplarily, the virtual view includes one or more virtual target objects. Performing playback control on the virtual view according to the human eye fixation area and the user's control gestures includes:

[0070] When the human eye fixates on a virtual target object, perform high-definition rendering on the human eye fixation area and perform low-definition rendering on the area outside the human eye fixation area;

[0071] Detect whether the user's hand grasps the virtual target object according to the hand image;

[0072] When the user's hand grasps the virtual target object, calculate the motion vector of the user's hand, and update the position of the virtual target object according to the calculated motion vector;

[0073] When the human eye does not fixate on the virtual target object, perform low-definition rendering on the entire virtual view.

[0074] The multi-view image fusion network of the present disclosure takes two or more images that are misaligned in time and / or space as inputs and produces a fusion result that matches the user's visual coordinate system. This network consists of three modules: a feature extraction module, a warping module, and a fusion module. This network provides an end-to-end CNN architecture that can merge images in multiple different coordinate systems and uses a cascaded pyramid to simultaneously learn the features of multiple resolution levels of streams.

[0075] In some exemplary embodiments, the multi-view image fusion network includes: a feature extraction module, a warping module, and a fusion module, where:

[0076] The feature extraction module is configured to receive the hand image and the virtual view, construct an image pyramid according to the hand image and the virtual view, extract features from each layer of the image pyramid, and form a feature pyramid;

[0077] The warping module is configured to perform the following operations on multiple layers of the feature pyramid from bottom to top: obtain an initial optical flow prediction from the next layer of the feature pyramid through bilinear upsampling, warp the features of the current layer of the feature pyramid with the initial optical flow prediction, and input the warped features and the target image features into a residual prediction network to obtain a corrected optical flow prediction value;

[0078] The fusion module is configured to perform the following operations on multiple layers of the feature pyramid from bottom to top: splice the warped features of the current layer of the feature pyramid with the upsampled features of the next layer of the feature pyramid, and perform an upsampling operation on the spliced features.

[0079] An image pyramid is a set of images composed of sub-images of an image at different resolutions, which is generated by continuously reducing the sampling rate of an image. The smallest image may have only one pixel. As Figure 7A shown, the leftmost column constitutes the image pyramids of two input images (hand image and virtual view). In the embodiments of the present disclosure, the image pyramid is a set of images arranged in an inverted pyramid shape and with gradually decreasing resolution from top to bottom. The top of the image pyramid is the high-resolution image (original image) to be processed, while the bottom is its low-resolution approximate image. When moving towards the bottom of the pyramid, both the size and resolution of the image continuously decrease. Each time moving down one level, the width and height of the image are reduced to half of the original.

[0080] In some exemplary embodiments, the feature pyramid is arranged sequentially from layer 0 to layer N - 1, where the weights of the residual prediction networks from the second layer to layer N - 1 are the same, and N is the number of layers of the feature pyramid.

[0081] Figure 7A It is a schematic structural diagram of the multi-view image fusion network of the embodiments of the present disclosure. Among them, on the left is the feature extraction module (using the image pyramid to build the feature pyramid), in the middle is the warping module (warping the features of one of the feature pyramids), and on the right is the fusion module (concatenating the warped features with the target image features). As Figure 7A shown, A0, A1, A2... and F0, F1, F2... represent kernels, and t k and s k represent features. A0, A1, A2... are sequentially applied to the feature extraction kernels at each level of the image pyramid, and each level corresponds to a feature extraction kernel. For each level k, we concatenate the features obtained by applying A0 at the current level, A1 at the previous level, and A2... at the level before the previous level, so as to generate the feature s k of the source image and the feature t k of the target image. Therefore, for all levels except the two levels of A0 and A1, we have the same number of feature channels (the number of feature channels at the A0 level is generally 1, the number of feature channels at the A1 level is generally 2, and the number of feature channels at the remaining levels is 3). This enables the warping module to share these levels, so as to predict features. The source feature s k is warped to the target feature t k , thus generating w(s k ). Then these aligned features are concatenated and fused with the information from the coarser pyramid levels to generate the fused output.

[0082] In a Convolutional Neural Network (CNN), the convolution operation is performed on two matrices. CNN mainly completes feature extraction through convolution operations. Image convolution operations are mainly achieved by setting various feature extraction filter matrices (convolution kernels, usually matrices with a size of 3x3 or 5x5), and then using this convolution kernel to'slide' on the original image matrix (an image is actually a matrix composed of pixel values) to implement the convolution operation. This disclosure uses a new cascaded feature extraction architecture that ensures the meaning of the filters at each shared level is the same. First, an image pyramid is constructed, and features are extracted from the Figure 7A cascaded arrangement shown. Each block An for n = 0, 1, 2 represents two 3×3 convolutions, each convolution having 2 n+4 filters, filtering layer by layer. These blocks are repeated for all pyramid levels in the figure (the finest pyramid level is represented by zero, with the highest resolution). And for each level k≥2, the extracted features have the same size. This is in sharp contrast to traditional encoder structures and other traffic prediction methods, where the number of filters grows with each downsampling (shrinking the image), and the benefit of this is to reduce the data volume and reduce redundancy.

[0083] The warping module is repeatedly applied from the coarsest level k = N to the finest level k = 0 for estimating the optical flow. In the embodiments of this disclosure, the optical flow is an image variable. As Figure 7B shown, w represents the warping operation, f′ k is the estimated value of the flow at level k (the coarsest level f′ N = 0), the superscript ′ represents the derivative operation, P k is the learnable residual flow prediction module, Δf k is the correction flow, equivalent to a correction coefficient, and f k is the fine flow at level k, equivalent to the warping coefficient.

[0084] The warping module follows the idea of residual flow prediction, but the weights are shared at most levels. For each pyramid level k, the optical flow prediction f′ k at level k is obtained from level k + 1 through bilinear upsampling (enlarging the image). Then this initial estimate f′ k is used to warp the image features s k at level k. Then the warped features and the target image features t k are input into the residual flow prediction network P k , and P k can only predict a small correction Δf k to improve f k . Then the fine flow f k is upsampled to obtain f′k-1 and repeat the process until we reach the optimal level k = 0.

[0085] The residual flow prediction network P in the embodiments of the present disclosure k is a serial application of five 2D convolutions: 3×3×32, 3×3×64, 1×1×64, 1×1×16, and 1×1×2. All layers except the last one use ReLU (a neural activation function) for activation. This module is small because it only performs small residual corrections. There are two key differences between the warping module of the present disclosure and the existing structures: 1) Weight sharing: For k≥2, the residual flow prediction module P k uses shared weights and learns simultaneously (synchronously) at all resolution levels, which can prevent excessive computational complexity and overfitting; 2) End-to-end training: Instead of being trained to minimize the loss with respect to the ground truth optical flow, the present disclosure is trained to generate image warping by penalizing the difference between the warped image and the target.

[0086] The input to the fusion module of the present disclosure is a feature pyramid. Each level in the feature pyramid is constructed by concatenating the warped image features and the target image features and applying two 3×3 convolutions with filters of size 2 k+4 and ReLU activators, where k is a certain level in the pyramid. We denote these convolutions as Figure 7A F in k . The upsampling from level k + 1 to level k is performed by nearest neighbor sampling and then a 2×2×2 k+4 convolution. In the fusion stage, each level F 0 , F 1 …F N uses independently learned weights, enabling autonomous learning and correction. During the fusion process, the size needs to be recalculated for each level to convert all sizes to the sizes of the source image and the target image in the first level.

[0087] The multi-view image fusion network of the present disclosure is trained using two losses: one is the reconstruction loss between the final image and the real image, and the other is the warping loss between the intermediate warped image and the real image. By training using the two losses, the training efficiency can be improved and the model convergence speed can be accelerated.

[0088] An exemplary multi-view image fusion result is as Figure 8 shown. By inputting the hand image and the virtual view into the multi-view image fusion network, the present disclosure obtains a fused image of the virtual view and the hand image, overcomes the deviation of the hand spatial information, realizes a natural human-computer interaction experience, and improves the user experience.

[0089] An embodiment of the present disclosure also provides a processing device, which may include a processor and a memory storing a computer program that can run on the processor. When the processor executes the computer program, the steps of the human-computer interaction method described in any one of the previous items in the present disclosure are implemented.

[0090] As Figure 9 shown, in one example, the processing device may include: a processor 910, a memory 920, a bus system 930, and a transceiver 940. Among them, the processor 910, the memory 920, and the transceiver 940 are connected through the bus system 930. The memory 920 is used to store instructions, and the processor 910 is used to execute the instructions stored in the memory 920 to control the transceiver 940 to send signals. Specifically, the transceiver 940 can receive the hand image collected by the gesture recognition sensor under the control of the processor 910. The processor 910 controls the head-mounted display to present a virtual view; detects whether fusion display is required; when fusion display is required, inputs the hand image and the virtual view into a multi-view image fusion network to obtain a fusion image of the virtual view and the hand image; controls the head-mounted display to present the fusion image.

[0091] It should be understood that the processor 910 may be a central processing unit (CPU), and the processor 910 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] The memory 920 may include a read-only memory and a random access memory, and provide instructions and data to the processor 910. A part of the memory 920 may also include a non-volatile random access memory. For example, the memory 920 may also store information about the device type.

[0093] In addition to including a data bus, the bus system 930 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, in Figure 9 all kinds of buses are labeled as the bus system 930.

[0094] In the implementation process, the processing performed by the processing device can be completed by the integrated logic circuit of the hardware in the processor 910 or the instructions in the form of software. That is, the method steps of the embodiments of the present disclosure can be embodied as being executed and completed by the hardware processor, or by a combination of the hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 920, and the processor 910 reads the information in the memory 920 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0095] The embodiments of the present disclosure also provide a human-computer interaction system, including a gesture recognition sensor, a head-mounted display, and a processing device. The processing device can be the processing device described in any embodiment of the present disclosure. The gesture recognition sensor is used to collect the hand image of the user; the head-mounted display is used to present a virtual view or a fused image.

[0096] In an exemplary embodiment, the head-mounted display further includes a head motion tracking sensor, and the X, Y, Z axes and the front and back sides are tracked through the head motion tracking sensor.

[0097] In an exemplary embodiment, the head-mounted display may include an eye tracking sensor, and the position and movement of the user's eyes are tracked through the eye tracking sensor, and then the human eye fixation area is determined.

[0098] In an exemplary embodiment, the gesture recognition sensor can be located on a neck-worn device.

[0099] The embodiments of the present disclosure also provide a computer-readable storage medium, which stores executable instructions. When the executable instructions are executed by a processor, the human-computer interaction method provided in any of the above embodiments of the present disclosure can be implemented. The human-computer interaction method can be used to control the head-mounted display provided in the above embodiments of the present disclosure to perform output virtual view playback control, which can achieve the effect of hand-eye coordination during human-computer interaction, and at the same time make the gesture control more truly reflect people's interaction behaviors, improve the user's immersion, overcome the deviation of hand space information, realize a natural human-computer interaction experience, and improve the user's usage experience. The method of driving the human-computer interaction system to perform human-computer interaction by executing the executable instructions is basically the same as the human-computer interaction method provided in the above embodiments of the present disclosure, and will not be elaborated here.

[0100] In the description of the embodiments of the present disclosure, it should be understood that the orientation or positional relationships indicated by the terms "middle", "upper", "lower", "front", "rear", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present disclosure.

[0101] In the description of the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "mounted", "connected", and "coupled" shall be construed broadly. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the meanings of the above terms in the present disclosure can be understood accordingly.

[0102] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components in cooperation. Some or all of the components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.

[0103] Although the embodiments disclosed in the present disclosure are as described above, the content described is only an embodiment adopted for the convenience of understanding the present disclosure and is not intended to limit the present disclosure. Any person skilled in the art within the scope of the present disclosure may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present disclosure. However, the protection scope of the present disclosure shall still be subject to the scope defined by the appended claims.

Claims

1. A human-computer interaction method, It is characterized in that include: Control the head mounted display to present a virtual view; receiving a hand image acquired by a gesture recognition sensor; Detect whether fusion display is needed; When fusion display is required, the hand image and the virtual view are input into a multi-view image fusion network to obtain a fused image of the virtual view and the hand image, and the head mounted display is controlled to present the fused image.

2. The human-computer interaction method according to claim 1, It is characterized in that The head mounted display includes an eye tracking sensor, and the method further includes: The user's eye image captured by the eye tracking sensor is obtained, a human eye gaze area is determined according to the user's eye image, and playback of the virtual view is controlled according to the human eye gaze area.

3. The human-computer interaction method according to claim 2, It is characterized in that The method further includes: mapping the hand image to a user visual coordinate system, detecting a user's control gesture through a gesture recognition algorithm, and controlling the playback of the virtual view according to the user's control gesture.

4. The human-computer interaction method according to claim 3, It is characterized in that The virtual view includes one or more virtual target objects, and the virtual view is played and controlled according to the human eye gaze area and the user's control gesture, including: When the human eye is gazing at the virtual target object, high-definition rendering is performed on the area where the human eye is gazing, and low-definition rendering is performed on the area outside the area where the human eye is gazing; Detecting whether the user's hand grabs a virtual target object according to the hand image; When the user's hand grabs the virtual target object, calculating the motion vector of the user's hand, and updating the position of the virtual target object according to the calculated motion vector; When the human eye is not looking at the virtual target object, the entire virtual view is rendered in low definition.

5. The human-computer interaction method according to claim 1, It is characterized in that Whether the detection needs to be fused and displayed includes: Displaying a hand-shaped button on the virtual view, wherein the hand-shaped button is used for the user to set whether a fusion display is required; Detecting whether the hand-shaped button is clicked; When the operation of clicking the hand-shaped button is captured, the state of whether the fusion display is currently required is switched.

6. The human-computer interaction method according to claim 1, It is characterized in that The gesture recognition sensor is located on a neck-mounted device, and the neck-mounted device is connected to the head-mounted display via wired or wireless communication.

7. The human-computer interaction method according to claim 1, It is characterized in that The multi-view image fusion network includes: a feature extraction module, a distortion module and a fusion module, wherein: The feature extraction module is configured to receive the hand image and the virtual view, construct an image pyramid according to the hand image and the virtual view, and extract features from each layer of the image in the image pyramid to form a feature pyramid; The warping module is configured to perform the following operations on the multi-layer feature pyramid from bottom to top: obtain an initial optical flow prediction from the next layer of the feature pyramid through bilinear upsampling, use the initial optical flow prediction to warp the features of the current layer of the feature pyramid, input the warped features and the target image features into the residual prediction network, and obtain a modified optical flow prediction value; The fusion module is configured to perform the following operations on the multi-layer feature pyramid from bottom to top: splicing the distorted features of the current layer feature pyramid with the upsampled features of the next layer feature pyramid, and performing an upsampling operation on the spliced ​​features.

8. The human-computer interaction method according to claim 7, It is characterized in that The feature pyramid is arranged downward from the 0th layer to the N-1th layer, wherein the weights of the residual prediction networks from the second layer to the N-1th layer are the same, and N is the number of layers of the feature pyramid.

9. A processing device, It is characterized in that include: A processor and a memory storing a computer program executable on the processor, wherein the processor implements the steps of the human-computer interaction method according to any one of claims 1 to 8 when executing the computer program.

10. A human-computer interaction system, It is characterized in that include: A gesture recognition sensor, a head mounted display, and a processing device as claimed in claim 9, wherein: The gesture recognition sensor is used to collect the user's hand image; The head mounted display is used to present a virtual view or a fused image.

11. A computer-readable storage medium, It is characterized in that Computer executable instructions are stored, and the computer executable instructions are used to execute the human-computer interaction method according to any one of claims 1 to 8.