A multi-device interaction system and method based on an augmented reality device

By using augmented reality devices to model and render virtual models of multiple screen devices, the problems of inflexible interaction and inaccurate resource positioning between multiple devices are solved, and unified interaction and efficient resource operation between multiple devices are achieved.

CN115756170BActive Publication Date: 2026-01-23INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211490160.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2026-01-23
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

Existing technologies have problems such as difficulty in single-handed operation, both hands being occupied, inability to interact at a distance, inflexible switching between devices, and inaccurate resource positioning in multi-device interaction, especially with low interaction efficiency between various types and sizes of screen devices.

Method used

Augmented reality devices are used to model multiple screen devices to obtain a unified spatial coordinate system. Virtual models are rendered using augmented reality devices, and eye tracking and gestures are used to realize interaction between multiple devices, supporting precise positioning and operation of fine-grained resources.

Benefits of technology

It enables flexible and unified interaction between multiple devices, supports precise resource operation between devices of different types and sizes, and improves user interaction efficiency and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115756170B_ABST
    Figure CN115756170B_ABST
Patent Text Reader

Abstract

Provided are a multi-device interaction system and method based on an augmented reality device, the system including an augmented reality device and a plurality of screen-type devices, the augmented reality device being configured to obtain spatial coordinates of the plurality of screen-type devices and model the plurality of screen-type devices; obtain spatial coordinates of in-screen resources of the plurality of screen-type devices based on relative positions of the in-screen resources and the spatial coordinates of the corresponding screen-type devices; and render virtual models of the plurality of screen-type devices and the in-screen resources thereof when interacting with the plurality of screen-type devices, and operate the virtual models of the in-screen resources of the plurality of screen-type devices through the augmented reality device to achieve interaction between the plurality of screen-type devices.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of human-computer interaction, and in particular to a multi-device interaction system and method based on an augmented reality device. BACKGROUND

[0002] In contemporary life, digital devices based on screen interaction such as computers and mobile phones are the most frequently used devices by users. When interacting with such devices, the main input medium is an interactive screen (such as gestures such as direct points and strokes on the screen) or peripherals attached to it (such as mice, keyboards, and styluses), and the output mainly relies on feedback through the screen and the mapping of the input (such as the movement of gestures, the movement of the mouse, and the matching of the coordinate system on the screen). Although such methods have been widely used, there are still some problems:

[0003] (1) For gesture-based screen interaction, there are difficulties in single-handed operation, both hands are occupied, and other situations that prevent users from successfully completing interaction tasks, reducing the efficiency of user interaction; and for touch-based gesture interaction, long-distance interaction cannot be completed;

[0004] (2) For screen interaction based on peripherals, users need to coordinate one or even several peripherals to complete interaction tasks, and interaction tasks cannot be separated from fixed external devices, reducing the flexibility of interaction;

[0005] (3) For the current multi-device working environment, the interaction mode centered on a specific device cannot timely, flexibly, and uniformly switch work between multiple devices.

[0006] Based on the above problems, researchers have begun to improve the corresponding interaction methods. Some research has enhanced the space of gesture-based screen interaction by detecting fingers close to the screen through self-capacitive touch screens. There are also external devices on the market that integrate multiple functions, such as using Bluetooth to bind multiple devices to one external device and quickly switch between multiple devices through shortcut keys. However, the existing touch gesture-based and peripheral-based screen interaction methods still have the following limitations:

[0007] (1) Traditional methods centered on a specific device are difficult to meet the needs of users in a multi-device working environment;

[0008] (2) Screen devices vary in size (such as large display screens, 13-inch tablets, and 6-inch mobile phones), and the types of electronic resources within the screen are diverse and relatively small in size (such as pictures, icons, text, and videos within the screen), making it difficult to accurately obtain screen devices and their internal electronic resources;

[0009] (3) The existing device interaction mode is relatively single, which limits the user's interaction possibility, and the introduction of a new interaction mode often needs to add relatively complex or diverse sensing devices, which cannot be directly reused among multiple devices. SUMMARY

[0010] In view of the above problems of the prior art, the present application provides a multi-device interaction system based on an augmented reality device, which comprises an augmented reality device and a plurality of screen devices, the augmented reality device being configured to obtain the spatial coordinates of the plurality of screen devices and model the plurality of screen devices; obtain the spatial coordinates of the in-screen resources of the plurality of screen devices based on the relative positions of the in-screen resources of the plurality of screen devices and the spatial coordinates of the corresponding screen devices; and when interacting with the plurality of screen devices, render a virtual model of the plurality of screen devices and their in-screen resources, and operate the virtual model of the in-screen resources of the plurality of screen devices through the augmented reality device to realize the interaction among the plurality of screen devices.

[0011] In one embodiment, the augmented reality device is further configured to obtain the spatial coordinates of the key points of the plurality of screen devices; and model the plurality of screen devices based on the key points.

[0012] In one embodiment, the augmented reality device further comprises a depth camera and an RGB camera, which are configured to obtain the depth and position information of the key points of the plurality of screen devices to calculate the spatial coordinates of the corresponding screen device key points.

[0013] In one embodiment, a patch for emitting a wireless signal is arranged on the key points of the plurality of screen devices, and an antenna array for detecting the wireless signal is arranged on the augmented reality device, and the augmented reality device detects and calculates the spatial coordinates of the corresponding screen device key points through the antenna array.

[0014] In one embodiment, the in-screen resources of the plurality of screen devices are divided into one or more minimum resources, and the relative positions of the one or more minimum resources in the corresponding screen are obtained, and the types, source addresses and relative positions of the one or more minimum resources in the screen are sent to the augmented reality device.

[0015] In one embodiment, when the in-screen resources of the screen device are updated, the relative positions of the minimum resources in the screen are updated, and then the types, source addresses and relative positions of the updated minimum resources in the screen are sent to the augmented reality device.

[0016] In one embodiment, the resource transmission is performed between the augmented reality device and the plurality of screen-based devices through a Socket connection, the plurality of screen-based devices are server ends, and the augmented reality device is a client end.

[0017] In one embodiment, the augmented reality device controls the plurality of screen-based devices by detecting hand movements and / or eye movements of a user.

[0018] In one embodiment, the augmented reality device is a head-mounted display.

[0019] The application also provides a multi-device interaction method for the above-mentioned augmented reality device-based multi-device interaction system, the method comprising:

[0020] The augmented reality device obtains spatial coordinates of the plurality of screen-based devices in the environment and models the plurality of screen-based devices;

[0021] The spatial coordinates of the in-screen resources of the plurality of screen-based devices are obtained based on the relative positions of the in-screen resources of the plurality of screen-based devices and the spatial coordinates of the corresponding screen-based devices; and

[0022] When interacting with the plurality of screen-based devices, a virtual model of the plurality of screen-based devices and their in-screen resources is rendered, the virtual model of the in-screen resources of the plurality of screen-based devices is operated through the augmented reality device, and the interaction between the plurality of screen-based devices is realized.

[0023] The augmented reality device-based multi-device interaction system and method of the application utilize the augmented reality device to model the environment to obtain a unified spatial coordinate system, calculate the spatial coordinates of the plurality of screen-based devices in the spatial coordinate system, establish a unified spatial coordinate system for the plurality of screen-based devices, and facilitate the subsequent calculation of background data when the plurality of screen-based devices are interacted with. The method can also obtain resource information of different types and sizes in different screen-based devices, including resource types, data and coordinates, and map the resource information to coordinate points in the unified spatial coordinate system, thereby supporting fine-grained resource interaction. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 A schematic diagram of an augmented reality device-based multi-device interaction system according to one embodiment of the application is shown.

[0025] Figure 2 A flowchart of an augmented reality device-based multi-device interaction method according to one embodiment of the application is shown.

[0026] Figure 3 A modeling result of a screen-based device according to one embodiment of the application is shown.

[0027] Figure 4 A diagram showing the picture 1 being grabbed from computer A to computer B is shown.

[0028] Figure 5 A diagram showing the pictures 1-3 being grabbed from computer A to computer B is shown. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application with specific embodiments and in connection with the drawings. It should be noted that the embodiments given by the present application are only for illustration, and do not limit the protection scope of the present application.

[0030] In the past, there were few devices that a user could deal with simultaneously in the use process, so that the "device-centered" interaction mode with a device having its own set of interaction systems could work well. However, now there can be 3-5 intelligent screen devices in the environment, and the user has an urgent need to switch among the devices. However, it is a challenging problem to establish a unified interaction enhancement system among devices of different types, different sizes and even different underlying operation logics. Based on this, the present application proposes a multi-device interaction system and method based on an augmented reality device, that is, a user-centered multi-device interaction mode through mixed reality connection of multiple devices.

[0031] Firstly, the terms and concepts applied in the present application are explained. The screen device is a digital device for interaction based on a screen, which can be, for example, a personal computer (PC), a smart television, a tablet computer and a mobile phone, etc. The augmented reality device is a device with augmented reality function, which can be, for example, a head-mounted display (HMD), a mobile phone, smart glasses, etc. The spatial coordinates of the screen device refer to the spatial coordinates of the screen of the screen device relative to the head-mounted display (i.e. the coordinate origin). The in-screen resource refers to the content displayed on the screen of the screen device, which can be, for example, text, picture, file, etc. The relative position of the in-screen resource refers to the position of the in-screen resource relative to the screen of the screen device, that is, the position on the screen of the screen device. The spatial coordinates of the in-screen resource refer to the spatial coordinates of the in-screen resource relative to the head-mounted display (i.e. the coordinate origin).

[0032] Figure 1 A diagram showing a multi-device interaction system based on an augmented reality device according to an embodiment of the present application is shown. Figure 1The multi-device interaction system based on the augmented reality device in the application comprises an augmented reality device 101 worn by a user and a plurality of screen devices 102-105, i.e. a personal computer 102, a smart television 103, a tablet computer 104 and a mobile phone 105. The augmented reality device 101 is used to detect and calculate the spatial coordinates of the plurality of screen devices in the environment, and model the plurality of screen devices; obtain the spatial coordinates of the in-screen resources based on the relative positions of the in-screen resources of the screen devices and the spatial coordinates of the corresponding screen devices; and when interacting with the plurality of screen devices, render a virtual model of the plurality of screen devices and the in-screen resources thereof by using mixed reality, and realize the interaction between the plurality of screen devices by operating the virtual model of the in-screen resources of the plurality of screen devices through the augmented reality device.

[0033] Figure 2 A flow chart of a multi-device interaction method based on an augmented reality device according to an embodiment of the application is shown. The method comprises:

[0034] Step S1: using the augmented reality device to detect and calculate the spatial coordinates of the plurality of screen devices, and model the plurality of screen devices.

[0035] According to an embodiment of the application, step S1 comprises the following sub-steps:

[0036] Step S11: using the augmented reality device to detect the plurality of screen devices in the environment, and obtain the spatial coordinates of the key points of the plurality of screen devices.

[0037] The key points of the screen devices refer to points that can locate the position of the screen of the screen devices. For example, in the embodiment of a rectangular screen, the key points of the screen devices can be the four corners of the rectangular screen or three corners of the rectangular screen. In the following, the case that the key points are the four corners of the rectangular screen is taken as an example for illustration.

[0038] Preferably, the initial detection position of the augmented reality device is taken as a fixed anchor point (i.e. the coordinate origin) for calculating the spatial coordinates of the corresponding key points of the screen devices relative to the fixed anchor point. When the position of the augmented reality device changes with the position of the user, the spatial coordinates of the fixed anchor point and the key points of the screen devices will not be affected, and therefore it is not necessary to repeatedly detect the spatial coordinates of the screen devices in the environment. The spatial coordinates of the screen devices in the environment can be re-detected and calculated when the position of the screen devices changes. Preferably, the spatial coordinates of the screen devices in the environment can also be re-detected and calculated when the augmented reality device is re-worn.

[0039] In one embodiment, a visual algorithm can be used to detect and recognize the screen-based device in the environment. In another embodiment, a two-dimensional code can be posted on the screen-based device in advance or displayed in a corner of the screen-based device, and the screen-based device in the environment can be detected and recognized by scanning the two-dimensional code on the screen-based device through the augmented reality device.

[0040] In one embodiment, the augmented reality device further comprises a depth camera and an RGB camera for obtaining depth and position information of the plurality of screen-based device key points and calculating spatial coordinates of the corresponding screen-based device key points. Taking the Hololens 2 generation helmet of Microsoft Corporation as the augmented reality device, after the user wears the Hololens 2 generation helmet, the RGB camera and the depth camera equipped with the Hololens 2 generation helmet are used to detect the screen-based device in the environment, and the depth and position information of the screen-based device key points are obtained to calculate the spatial coordinates of the corresponding screen-based device key points.

[0041] According to one embodiment of the present application, positioning the spatial coordinates of the key points can be divided into the following three steps:

[0042] (1) Real-time acquisition of RGB photo stream of the plurality of screen-based devices: the RGB camera and the depth camera equipped with the Hololens 2 generation helmet are used to detect the screen-based device in the environment, and the real-time photo stream of the plurality of screen-based devices captured by the RGB camera is obtained through the PhotoCapture provided in the open source software package "Mixed Reality Toolkit" provided by the Hololens developer Microsoft Corporation.

[0043] (2) Positioning the corresponding two-dimensional pixel points (X, Y) of the key points (such as the four corners of the screen rectangular frame of the screen-based device) of the plurality of screen-based devices in the RGB photo: the RGB photo is processed by using the OpenCV algorithm package. The detection of the screen rectangular frame is mainly to extract the edge, and the brightness of the display part of the screen-based device is usually higher than that of the surrounding environment, so the picture can be thresholded. The cvtColor algorithm is used to convert the RGB photo into a grayscale image, the medianBlur algorithm is used for median filtering, the threshold algorithm is used to convert the image into a binary image, the Canny algorithm is used for edge detection, the findContour algorithm is used to extract the rectangular contour, the contourArea algorithm and the approxPolyDP algorithm are used to extract the largest contour and surround the contour with a polygon, and the convexHull algorithm is used to find the convex hull. Based on the above processing, the two-dimensional pixel coordinates of the key points of the screen-based device in the RGB photo can be accurately obtained.

[0044] (3) Obtain the three-dimensional space coordinates of the corresponding two-dimensional pixel coordinates of the key points in the Hololens 2 helmet: the Hololens 2 helmet encapsulates the depth information in the Mixed Reality Toolkit. The X and Y parts of the three-dimensional space coordinates are obtained by calling ConvertPixelCoordsToScaledCoords of PhotoCaptureFrame, and the Z value is obtained by the collision point information of SpatialAwareness. The three are combined to obtain the three-dimensional space coordinates (X, Y, Z) of the corresponding key points.

[0045] In another embodiment, the screen-type device in the environment can be detected by the augmented reality device, and the spatial coordinates of the corresponding screen-type device key points are calculated using the antenna array. For example, a wireless signal is emitted by a sticker attached to the screen-type device, the signal is received by the augmented reality device, and the spatial coordinates of the corresponding screen-type device key points are calculated. In this embodiment, the sticker is attached to the screen-type device key point and emits a wireless signal, and the augmented reality device has an antenna array for detecting the wireless signal. Preferably, positioning the three-dimensional space coordinates of the key points can be divided into the following two steps:

[0046] (1) The antenna array on the augmented reality device obtains the azimuth angle R and distance L of the sticker on the screen-type device key point: the sticker sends a wireless signal, which is received and decoded by the antenna array, reads the IQ (in-phase component and quadrature component) data values received by different antennas, calculates the phase of the wireless signal from the sending end to the receiving end through the data values, and then obtains the phase difference received between the antennas. The phase difference data is processed and input into a super-resolution algorithm such as the multiple signal classification (MUSIC) algorithm, and the maximum value of the spectral function is obtained in the spatial spectral domain. The angle corresponding to the spectral peak is the estimated value of the direction angle of the sticker (the azimuth angle and the pitch angle can be estimated at the same time). In addition, the received signal strength decoded from the data packet can be used to represent the estimated distance of the screen-type device, i.e., using RSSI data to represent.

[0047] (2) Calculate the three-dimensional space coordinates of the key points according to the azimuth angle and the distance: taking the Hololens helmet as an example, the current coordinates and rotation angle of the Hololens helmet are obtained through Main.Camera.Transform.Position and Main.Camera.Transform.Rotation. Taking the Hololens helmet as the center, the rotation angle is the angle R calculated in the first step, and the rotation radius is the distance L calculated in the first step. The three-dimensional space coordinates of the sticker can be calculated.

[0048] Step S12: model the screen-type device based on the key points, and store the screen-type device related information in the spatial model.

[0049] Figure 3 The modeling result of the screen-based device according to one embodiment of the present application is shown. As shown, the screen 310 of the screen-based device is represented by a solid line box, and the screen 310 is modeled based on the key points (i.e. the four corners) of the screen 310, and the modeling result (i.e. the virtual model) 320 is represented by a dashed line box. For clarity, the dashed line box of the virtual model 320 does not coincide with the screen 310, but in actual operation, the dashed line box of the virtual model 320 preferably coincides with the screen 310. Figure 3

[0050] The screen-based device related information includes the positions of the key points, the modeling information of the screen-based device, the identity information of the screen-based device (e.g. the network communication address of the device), etc., and is used to inform the augmented reality device how to connect with the screen-based device. For example, the identity information of the screen-based device can be identified through the two-dimensional code pasted on the screen-based device, the identity information of the screen-based device can be identified through the two-dimensional code in the corner of the screen, or the identity information of the screen-based device can be identified through the signal emitted by the smart sticker.

[0051] When the augmented reality device interacts with the screen-based device that has been modeled, the virtual model of the screen-based device (e.g. the dashed line box in FIG. 3) can be rendered by using the mixed reality. When the screen-based device is observed through the augmented reality device (e.g. the HMD), the virtual model that coincides with the screen can be seen to provide visual feedback for the user interaction and improve the interaction accuracy. Figure 3

[0052] Therefore, the user wears the augmented reality device (e.g. the HMD), and establishes a unified coordinate system in the real space based on this, so that each screen-based device has spatial coordinate information in the augmented reality device. The user does not need to add additional sensing channels in the environment or on the specific device, and only needs to wear the augmented reality device to detect screen-based devices of various sizes in different environments, which helps the user to better complete the multi-device collaborative work.

[0053] Preferably, all the screen-based devices are connected in the same local area network.

[0054] Step S2: obtaining the accurate spatial coordinates of the in-screen resource based on the relative positions of the in-screen resource and the spatial coordinates of the screen-based device.

[0055] Obtaining the accurate spatial coordinates of the in-screen resource can better help the user to complete more diverse interactive tasks. The traditional multi-device interactive system can only support file-level positioning, and more fine-grained resources (such as pictures, texts, videos, etc. in a file) are challenging to be accurately positioned due to their large number, small area, and compact spacing.

[0056] ​​The application proposes a method for obtaining accurate spatial coordinates of in-screen resources based on the relative positions of the in-screen resources and the spatial coordinates of the screen-based devices, and realizes fine-grained resource acquisition and management. Specifically, in step S1, the spatial coordinates of the screen-based devices are obtained, the relative positions of the in-screen resources of each screen-based device are calculated and acquired by the screen-based device itself, and the corresponding resource types, source file addresses and other information are acquired and sent to the augmented reality device in combination with the corresponding screen-based device related information to ensure the accuracy of the acquired relative positions of the resources. The augmented reality device combines the spatial coordinates of the screen-based devices and the relative positions of the in-screen resources, and after calculation, the spatial coordinates of the in-screen resources in a unified spatial coordinate system can be obtained, laying a foundation for subsequent interaction.

[0057] According to one embodiment of the application, the screen-based device divides the resources in its screen into one by one minimum resources, and then forms a resource list of the types and source addresses of the minimum resources, and sends the resource list and the relative positions of the resources in the screen to the augmented reality device in combination with the corresponding screen-based device related information. In the present application, the in-screen minimum resource refers to the smallest operable content in the screen. The in-screen minimum resource is different according to the used resource browser. In general, the in-screen minimum resource is the smallest resource unit that can be browsed by the current resource browser and visible to the human eye. The type of the minimum resource can be, for example, a resource type identifier such as text, picture and file.

[0058] For example, when using a desktop, the minimum resource is each software and file icon on the desktop. When using a word software, the minimum resource is each word and each picture in the current page.

[0059] In one embodiment, taking a browser as a resource browser, the browser can be considered as full screen, and the position of the resource in the browser represents the relative position of the resource in the screen. The web source code is obtained by the content or post algorithm in the requests library, and is decoded by the decode algorithm. The web source code is parsed by the BeautifulSoup, XPath and requests-html algorithms to obtain the types and source addresses of each minimum resource in the browser, and form a resource list. The resource list is traversed, and the horizontal displacement and vertical displacement of each minimum resource relative to the screen are obtained by the Element.offsetParent and Element.offsetParent algorithms, so that the relative position of each minimum resource in the screen can be obtained.

[0060] In another embodiment, the types, source addresses and relative positions of each minimum resource in the screen can be analyzed by a full-screen screenshot and a computer vision algorithm.

[0061] The screen device sends its resource list, the relative position of the resources in the screen, and the corresponding screen device identity information to the augmented reality device. In one embodiment, the resource transmission is conducted through the establishment of a Socket connection between the augmented reality device and the screen device. The screen device is the server end, and the augmented reality device is the client end. One client end can actively connect to multiple server ends, so the augmented reality device can communicate with multiple screen devices at the same time. When the connection is established, the augmented reality device actively sends a message to the screen device to maintain the connection. When the resource list of the screen device is updated, the resource list and the relative position of the resources in the screen are actively sent to the augmented reality device.

[0062] Based on the spatial coordinates of the multiple screen devices obtained in step S1 and the relative position of each minimum resource in the screen obtained in step S2, the spatial coordinates of the resources in the screen can be obtained through simple coordinate calculation.

[0063] Step S3: When interacting with the multiple screen devices, a virtual model of the multiple screen devices and the resources in the screens is rendered by using mixed reality, and the virtual model of the resources in the screens is operated between the multiple screen devices by using the augmented reality device, so as to realize the interaction between the multiple screen devices.

[0064] For the sake of clarity, in the following, the head-mounted display is taken as an example of the augmented reality device for detailed description. The head-mounted display is a good sensing platform, and many commercial head-mounted displays have sensing functions such as gesture recognition and eye movement tracking, and the shape characteristics and wearing position of the head-mounted display have good expansion potential.

[0065] The head-mounted display has the existing eye movement and hand movement tracking capabilities, and the corresponding raw data can be obtained. In a unified spatial coordinate system, the eye movement and the hand movement can be calculated by using corresponding vectors. Taking the eye movement as an example, the eye movement is controlled by muscles and has a corresponding movement range, so the limit value of the eye movement can be obtained. The limit value is proportionally calculated with the size of the screen device, so the range of the eye movement can be matched with the screen device. When the user moves the eyeball from top to bottom, a spatial vector with length and direction can be calculated from the starting point to the ending point. Based on the vector, the movement of the control pointer (similar to the mouse cursor) on the device can be represented by the eye movement.

[0066] Hand and eye movements can be used as another dimension of input to expand the interaction space of a user with a screen-based device in space. A head-mounted display-centric unified coordinate space is used to expand the interaction of a user with screen-based digital devices such as personal computers (PCs) and mobile phones, in combination with hand movements (e.g., hover, grip, select) and eye movements. The present invention calculates the eye movements and hand movements of a user in the same coordinate space, and maps the eye movements and hand movements to an operation pointer in the corresponding screen-based device, and completes the corresponding interaction task in combination with specific eye movements (e.g., a forceful blink) and hand movements (e.g., a grip), and can achieve seamless switching operations between multiple devices.

[0067] The present invention is based on a head-mounted display unified space coordinate system. When a user interacts with a screen-based device, the system calculates the interaction plane size of the screen-based device and the eye movement and hand movement range, and adaptively matches the eye movement with the movement in the interaction plane of the screen-based device. This method can achieve the following: when a user looks at or points to a screen-based device, the user can operate the screen-based device; when a user rotates the eyes or moves the hands on different sizes of interaction planes, the corresponding pointers on the interaction planes can be matched with the eye or hand movements.

[0068] Figure 4 A schematic diagram of grabbing picture 1 from computer A to computer B is shown. In combination with Figure 4 Take grabbing a picture from computer A to computer B as an example to explain the interaction between multiple screen-based devices.

[0069] Computer A screen has pictures 1-6. Based on steps S1 and S2, the head-mounted display has obtained the spatial coordinates of pictures 1-6. According to the spatial coordinates of pictures 1-6, a virtual model X that a user can touch and interact with is rendered in the head-mounted display. Figure 4 The virtual model X in the head-mounted display is a dashed box that frames the picture. In actual applications, it can also be a three-dimensional box that can frame the picture, and the user can freely design as needed.

[0070] The user can directly grab the virtual model of picture 1 through hand movements, or can select the virtual model of picture 1 through eye movements at a distance, and then directly grab it through hand movements. At this time, the head-mounted display background will record "select picture 1 on computer A", and transfer picture 1 from computer A to the head-mounted display.

[0071] The user grabs picture 1 to position Y of computer B with eyes, and then releases the hand. The spatial coordinates of position Y can be obtained directly by the software package of the head-mounted display (for example, the position of the hand or the position of the eye gaze). The relative position of position Y in the screen of computer B is calculated, and then picture 1 is transmitted from the head-mounted display to computer B and displayed at position Y. The head-mounted display can subsequently delete picture 1 thereon. In an embodiment, the head-mounted display can also notify computer A to delete picture 1 thereon.

[0072] In another embodiment, the head-mounted display can select multiple minimum resources according to eye movements and form a virtual model. Figure 5 The schematic diagram of grabbing pictures 1-3 from computer A to computer B is shown. The user can select pictures 1-3 by eye movements, and the head-mounted display renders a virtual model X that the user can touch and interact according to the user's selection. Pictures 1-3 are transmitted from computer A to computer B, and other steps are the same as the embodiment shown, which will not be described here. Figure 4 The schematic diagram of grabbing pictures 1-3 from computer A to computer B is shown. The user can select pictures 1-3 by eye movements, and the head-mounted display renders a virtual model X that the user can touch and interact according to the user's selection. Pictures 1-3 are transmitted from computer A to computer B, and other steps are the same as the embodiment shown, which will not be described here.

[0073] In an embodiment, the present application is based on eye movements and gestures and two types of digital devices, personal computers and mobile phones, and provides the following four types of application scenarios and interactive applications in the corresponding interactive scenarios to demonstrate the application potential of the present application. However, the application scenarios of the present application are only examples, and any other application scenarios can be implemented by those skilled in the art as needed.

[0074] Scenario 1: Application scenario based on eye movements and personal computers:

[0075] a) Use eye movements to move the pointer, and use deliberate blinking to indicate confirmation of selection;

[0076] b) Automatically lock the screen when the line of sight leaves the personal computer, and automatically unlock when the line of sight moves into the personal computer;

[0077] c) The line of sight enters the personal computer to automatically match the external device (for example, keyboard, mouse, etc.), wherein the external device is used to assist the operation of the personal computer, which can be connected to multiple personal computers through wireless or Bluetooth for input to the personal computer; when a personal computer is selected by eye movements, the personal computer is automatically connected to the external device and receives the input of the external device; for example, when the line of sight moves from personal computer A to personal computer B, personal computer B can be automatically connected to the external device; thus, a set of external devices can be used for multiple screen devices;

[0078] d) Text editing: in a text editing task, the line of sight is used to control the text editing point, and the page up and down is controlled, so that the user does not need to frequently use the mouse and frequently switch between the keyboard and the mouse.

[0079] Scenario 2: Application scenario based on eye movement and mobile phone

[0080] a) Use eye movement to move the pointer, and use deliberate blinking to indicate confirmation of selection;

[0081] b) Automatically lock the screen when the line of sight leaves the mobile phone, and automatically unlock the mobile phone when the line of sight moves into the mobile phone;

[0082] c) Assist single-handed operation: when the user operates the mobile phone with one hand, some areas located in the corners of the screen are difficult to click, the corresponding resources can be relocated using the line of sight movement, and a resource is selected using the line of sight, and then the line of sight is moved to another location on the screen.

[0083] Scenario 3: Application scenario based on gestures and personal computers

[0084] a) Move and copy resources between multiple personal computer devices through a grab-and-release gesture; for example, a picture in personal computer A can be grabbed while facing personal computer A, and then released into personal computer B while facing personal computer B, to realize copying of resources between personal computer A and personal computer B; in another embodiment, the grab-and-release operation can also be performed without facing the personal computer, because the spatial positions of all resources are known;

[0085] b) Flip pages by sliding up and down through gestures, and temporarily render resources outside the screen through mixed reality functions.

[0086] Scenario 4: Application scenario based on gestures and mobile phones

[0087] a) Move and copy resources between multiple mobile phones through a grab-and-release gesture.

[0088] In the above examples, based on the unified spatial coordinate system of the head-mounted display, combined with two types of input modes of eye movement and gestures, four specific interactive application scenarios are proposed, which can improve the interaction efficiency and experience of users.

[0089] In one embodiment, multiple screen devices are centered on the head-mounted display, resources in one screen device are copied to the head-mounted display, and the resources are copied from the head-mounted display to another screen device, realizing interaction between multiple screen devices. In another embodiment, multiple screen devices and the head-mounted display are connected in the same local area network, so that under the control of the head-mounted display, resources can be directly interacted between multiple screen devices.

[0090] The multi-device interaction system and method based on the augmented reality device of the present application utilizes the augmented reality device to model the environment to obtain a unified spatial coordinate system, and calculates the spatial coordinates of the plurality of screen devices under the spatial coordinate system, to establish a unified spatial coordinate system for the plurality of screen devices, facilitating the subsequent calculation of the background data when the plurality of screen devices are interacted. The method can also obtain different types and sizes of resource information in different screen devices, including resource type, data and spatial coordinates, and map the resource information to coordinate points under the unified spatial coordinate, to support the interaction of fine-grained resources.

[0091] Although the present application has been described by way of preferred embodiments, it is not intended to limit the present application to the embodiments described herein, and various modifications and changes can be made without departing from the scope of the present application.

Claims

1. A multi-device interaction system based on augmented reality devices, the system comprising augmented reality devices and multiple screen-type devices, wherein: The augmented reality device is used for: Obtain the spatial coordinates of the plurality of screen devices and model the plurality of screen devices, wherein the spatial coordinates of the screen devices refer to the spatial coordinates of the screen devices relative to the augmented reality device; The spatial coordinates of the resources within the screen of the multiple screen devices are obtained based on the relative positions of the resources within the screen and the corresponding spatial coordinates of the screen devices. The relative positions of the resources within the screen refer to their positions on the screen, and the spatial coordinates of the resources within the screen refer to the spatial coordinates of the resources within the screen relative to the augmented reality device. as well as When interacting with the multiple screen devices, a virtual model of the multiple screen devices and their on-screen resources is rendered. The augmented reality device operates the virtual model of the on-screen resources between the multiple screen devices to realize the interaction between the multiple screen devices. The interaction includes moving or copying on-screen resources from one screen device to another screen device. The screen resources of the screen-type device are divided into one or more minimum resources, the relative positions of the one or more minimum resources within the corresponding screen are obtained, and the type, source address, and relative position of the one or more minimum resources within the screen are sent to the augmented reality device.

2. The multi-device interaction system based on augmented reality devices according to claim 1, wherein, The augmented reality device is also used for: Obtain the spatial coordinates of key points of the multiple screen devices; Model the multiple screen-type devices based on the aforementioned key points.

3. The multi-device interaction system based on augmented reality devices according to claim 2, wherein, The augmented reality device also includes a depth camera and an RGB camera, used to obtain the depth and position information of key points of the multiple screen devices, so as to calculate the spatial coordinates of the corresponding key points of the screen devices.

4. The multi-device interaction system based on augmented reality devices according to claim 2, wherein, The multiple screen devices are equipped with patches for transmitting wireless signals at key points, and the augmented reality device is equipped with an antenna array for detecting the wireless signals. The augmented reality device detects and calculates the spatial coordinates of the corresponding key points of the screen devices through the antenna array.

5. The multi-device interaction system based on augmented reality devices according to claim 1, wherein, When the resources within the screen of a screen-type device are updated, the relative position of the minimum resource within its screen is updated, and then the type, source address, and relative position of the updated minimum resource within the screen are sent to the augmented reality device.

6. The multi-device interaction system based on augmented reality devices according to claim 1, wherein, Resource transfer is performed between the augmented reality device and the plurality of screen-type devices via a Socket connection, wherein the plurality of screen-type devices are the server side and the augmented reality device is the client side.

7. The multi-device interaction system based on augmented reality devices according to claim 1, wherein, The augmented reality device controls the multiple screen devices by detecting the user's hand movements and / or eye movements.

8. The multi-device interaction system based on augmented reality devices according to any one of claims 1-7, wherein, The augmented reality device is a head-mounted display.

9. A multi-device interaction method for a multi-device interaction system based on augmented reality devices according to any one of claims 1-8, the method comprising: The augmented reality device obtains the spatial coordinates of multiple screen-type devices in the environment and models the multiple screen-type devices, wherein the spatial coordinates of the screen-type devices refer to the spatial coordinates of the screen of the screen-type devices relative to the augmented reality device; The augmented reality device obtains the spatial coordinates of the resources within the screen of the multiple screen devices based on the relative positions of the resources within the screen and the corresponding spatial coordinates of the screen devices. The relative positions of the resources within the screen refer to their positions on the screen, and the spatial coordinates of the resources within the screen refer to the spatial coordinates of the resources within the screen relative to the augmented reality device. When the augmented reality device interacts with the multiple screen devices, it renders virtual models of the multiple screen devices and their on-screen resources. The augmented reality device then manipulates the virtual models of the on-screen resources between the multiple screen devices to achieve interaction between them. The interaction includes moving or copying on-screen resources from one screen device to another. as well as The method further includes: The resources within the screen of a screen-type device are divided into one or more minimum resources. The relative positions of the one or more minimum resources within the corresponding screen are obtained, and the type, source address, and relative position of the one or more minimum resources within the screen are sent to the augmented reality device.

Citation Information

Patent Citations

  • Screen-oriented augmented reality interaction method and device and storage medium

    CN113961107A