Method for triggering actions in a metaverse or virtual world
The method uses eye-tracking to map gaze vectors and set interaction conditions, addressing the lack of secure avatar interactions in the metaverse by detecting genuine social intentions and preventing unwanted interactions, ensuring realistic and safe user experiences.
Patent Information
- Application Number
- JP2025529205
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-12-03
AI Technical Summary
Current metaverse technologies lack a reliable method for enabling secure, intentional, and bilaterally authorized interactions between avatars without manual tools, while preventing unwanted interactions and preserving realistic social interactions without physical boundaries, especially for users with disabilities affecting their arms or hands.
A method that utilizes eye-tracking devices to map gaze vectors of avatars, detecting eye contact and triggering actions based on predetermined conditions, such as setting a social interaction time or gaze aversion time to prevent unwanted interactions.
Enables secure, intentional, and realistic interactions between avatars by detecting genuine social intentions, preventing unwanted interactions, and accommodating users with disabilities, thus enhancing user safety and interaction authenticity in virtual environments.
Smart Images

Figure 2025539143000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for triggering actions in a metaverse or virtual world. [Background technology]
[0002] "Virtual world" as used herein means a virtual, mixed, or augmented reality world that is accessible through a virtual, mixed, or augmented reality headset that provides a user with a computer-generated virtual reality experience with which the user can interact. The user enters the virtual world through their avatar and can control objects and perform a series of actions.
[0003] To enable more immersive and realistic participation in virtual worlds, users can use head-mounted displays (HMDs), as mentioned above, which can display images through a display device and reproduce sound through speakers built into the device.
[0004] The HMD may also include an eye-tracking module as an auxiliary input means, which tracks eye movements when the user moves their eyes without moving their head, making it possible to detect which object the user is paying attention to.
[0005] The metaverse is a network of integrated 3D virtual worlds, or computing environments, that provide users with an immersive experience. Typically, the metaverse is accessed by users using a virtual reality headset. Users navigate the metaverse using eye movements, feedback controllers, or voice commands, although these are not required.
[0006] The academic paper "A Metaverse: Taxonomy, Components, Applications, and Open Challenges" (Sang-Min Park et al.) (Reference 1), published in January 2022, describes the concept, architecture, and content of the metaverse and provides a comprehensive analysis of the current state of the technology, as well as directions and open challenges for implementing an immersive metaverse.
[0007] First, it is important to emphasize the following: "The metaverse differs from augmented reality (AR) and virtual reality (VR) in three ways. First, while research on VR focuses on physical approaches and rendering, the metaverse is more about services with sustainable content and social meaning. Second, the metaverse does not necessarily have to use AR or VR technology. Even if a platform does not support VR or AR, it can still be a metaverse application. Finally, having a scalable environment that can accommodate a large number of people is essential to strengthen social meaning." (Reference 1)
[0008] Therefore, metaverse applications can also be accessed by users through a regular personal computer (PC) without using a specific head-mounted device such as a VR headset.
[0009] There are also devices known as eye-tracking devices, which may be in the form of glasses and can be used as a means of accessing the Metaverse world displayed on a regular PC screen. These eye-tracking glasses typically include sensors aimed at the wearer's eyes to acquire data about the eyes. This data is processed and output as pupil coordinates and gaze direction. This gaze direction can be displayed on a computer device with a corresponding display, and another user can view the gaze direction in the wearer's field of view through live streaming over the Internet. Thus, the user's actual gaze point can be determined using the glasses, along with so-called field of view video acquired by an additional field of view camera on the glasses in the user's field of view, and can be streamed over the Internet to a second user remotely connected to the gaze tracking device.
[0010] Typically, users interact with the Metaverse through their avatars, which are their alter egos and active agents in the Metaverse. Avatars are anthropomorphic computer representations of users, typically in the form of three-dimensional (3D) models. The avatars can be defined by users to represent aspects of their behavior, persona, beliefs, interests, and social status.
[0011] A computing environment implementing a Metaverse world allows for the creation of an avatar and allows for customization of the character's appearance. For example, a user can customize an avatar with hairstyle, skin color, physique, etc. Additionally, an avatar may have clothing, accessories, emotes, animations, etc.
[0012] As far as we know, virtual reality has limitations and only allows you to "travel in a virtual world," and it seems to be primarily focused on simulation and entertainment. Verma (Reference 4) adds: "The Metaverse has no clear boundaries; it is the product of the convergence of multiple technologies, including augmented reality (AR) and virtual reality (VR). In the Metaverse, users can purchase or even develop digital objects, places, and non-transferable fiat currencies (NFTs). While virtual reality is typically limited to a limited number of people, such as the number of players in a game, the Metaverse is considered an open virtual environment, where users can freely move, enjoy, and interact with anyone across the Internet, at no cost. The Metaverse is envisioned as a shared digital space that users can experience via the World Wide Web. The Metaverse is constantly changing, blending real and virtual experiences using technologies such as augmented reality (AR), providing users with a 'virtual modality of reality,' which is constantly accessible and impacts the real world in a variety of ways." (Reference 4) In contrast, virtual reality (VR) operates intermittently, functioning only when the user desires that particular experience. When the headset is turned off, the virtual world ceases to evolve and remains static.
[0013] (prior art) M. Kaur et al. (Reference 2) have the following to say about the metaverse: "Metaverse technology has been called the next great revolution of the Internet. The metaverse is a virtual environment in which users can create avatars and recreate experiences from the real world, or physical world, on a virtual platform... (omitted)...By 2028, the metaverse market is projected to reach US$814.2 billion, growing at a compound annual growth rate (CAGR) of 43.8% during the forecast period. The global metaverse market is expanding due to growing interest in areas such as socialization, entertainment, and creativity."
[0014] In addition, J. Goldston et al. (Reference 3) state the following about NVIDIA's Omniverse: "One metaverse will be built for community gatherings and games, while another will be built for scientists, creators, and businesses. One driver of innovation in the metaverse will be the creator economy. Developers and creators will have access to a wide variety of tools that will enable them to bring innovative products to market like never before. In fact, creators will likely create more tools in the virtual world than in the real world. One AI program designed for these virtual world builders is NVIDIA's Omniverse. Omniverse provides a user-friendly server backend and gives users access to an inventory of 3D assets in Universal Scene Description (USD) format, enabling artists and developers to collaborate, test, design, and visualize remotely in real time. Assets from this inventory can be used in a variety of ways, and Omniverse includes tools to assist artists, such as plugins for 3D digital content creation (DCC), PhysX 5.0, an RTX-based real-time rendering engine, and a built-in Python interpreter and extension system." (Jon Peddie) Research, 2021). Finally, all Omniverse tools are built as plugins, making it easy for artists and developers to customize the product for their own use cases.
[0015] One example of an Omniverse use case is Bayerische Motoren Werke AG (BMW), a German multinational manufacturer of luxury cars. Today, BMW produces one new car every minute, and to meet demands for continuous improvement and innovation, it needs to simulate complex production scenarios to increase production speed, improve agility, and optimize efficiency.
[0016] Furthermore, metaverse games such as ROBLOX are now well-known, and Sang-Min Park et al. state that ROBLOX "is used by two-thirds of 9- to 12-year-olds in the US, and has 150 million monthly active users (MAU) as a representative game of the metaverse" (Reference 1). "ROBLOX is also used to develop simulations of urban environments in the classroom, incorporating virtual pathways into the experience and illustrating the city's sculptural heritage. ... Although education and entertainment norms have often been considered separate worlds, ROBLOX is also being used as an educational tool in the classroom from the perspectives of motivation, problem-solving, and STEM."
[0017] DECENTRALAND is a metaverse designed around the cryptocurrency MANA, which can be used to trade items and virtual real estate. The virtual gaming platform runs on the Ethereum blockchain.
[0018] Other metaverse worlds currently exist, and more are expected to be developed in the future. In all cases, interactions between avatars (typically humanoid avatars) serve as "alter egos" of real users in the virtual world. However, there have already been reports of inappropriate behavior between users acting through avatars in the metaverse. The December 16, 2021, MIT Technology Review article titled "The Metaverse Already Has a Groping Problem" noted, "According to Meta, on November 26, a beta tester reported being groped by a stranger on Horizon Worlds." On December 1, Meta revealed that the tester had posted about their experience in the Horizon Worlds beta test group on Facebook.
[0019] Furthermore, on February 3, 2022, USA TODAY TECH published an article titled "Sexual Harassment in the Metaverse? Woman Claims Rape in Virtual World," which reported that in a December post, the victim wrote, "Within 60 seconds of joining, I was verbally and sexually harassed by three to four male avatars with male voices, essentially subjecting my avatar to a virtual gang rape."
[0020] In response, Meta announced on March 15, 2022, that it has added a personal boundary feature to its Metaverse platform. This feature creates an invisible boundary that prevents other avatars from coming within a radius of approximately four feet. Users can choose from three boundary settings, giving the community customizable control to determine how they want to interact in their VR experiences, but it is not possible to completely disable this invisible physical boundary that prevents unwanted interactions.
[0021] Therefore, this measure always limits close interaction between avatars and does not faithfully reproduce the situation in the real world, where there are no physical boundaries around the human body.
[0022] Furthermore, Patent Document 1 describes a method that allows a user to start a chat with a specific avatar when the user gazes at the avatar using a head-mounted device, but does not propose any solution to the problem that the user of the avatar may be able to prevent this action.
[0023] On the other hand, Patent Document 2 addresses how to scale an avatar in the user's physical world, i.e., a one-to-one mapping method between a user and an avatar in augmented reality technology. Patent Document 2 particularly focuses on automatically scaling the dimensions of the avatar to increase and maximize direct eye contact based on the user's eye height and minimize strain on the user's neck (see Figure 11A and paragraphs 153 and 170). However, Patent Document 2 does not address social interactions between avatars in the Metaverse or other virtual worlds.
[0024] Thus, the Metaverse still faces safety challenges due to the lack of a means to prevent unwanted interactions. The only measure implemented by Meta is the physical "personal boundary," which is perceived as artificial and unrealistic, limiting all interactions between users acting as avatars in the Metaverse world. [Prior art documents] [Patent documents]
[0025] [Patent Document 1] European Patent No. 3491781 [Patent Document 2] International Publication No. 2021 / 202783 Summary of the Invention [Problem to be solved by the invention]
[0026] (Object of the present invention) One of the objectives of a first aspect of the present invention is to provide a method for providing consent to trigger actions and / or status changes to other avatars when interacting with them in a virtual world without using (effectively eliminating) manual tools / devices such as a mouse, hand tools, controllers, etc.
[0027] A second object of the present invention is to provide a reliable method for establishing secure, intentional, bilaterally authorized interactions between avatars with the simultaneous consent of both avatars representing corresponding users.
[0028] A third object of the present invention is to provide a method for preventing unwanted / unwanted interactions between users while preserving realistic and spontaneous interactions without introducing physical boundaries.
[0029] A fourth object of the present invention is to further provide a method for determining the degree of interaction between two avatars, such as simple gaze, willingness to interact, or even willingness to avoid interaction.
[0030] A fifth object of the present invention is to provide a method that is accessible to people with diseases that affect the arms and / or hands.
[0031] A further object of the present invention is to provide a method that allows realistic interactions between avatar users in a safer manner than in the prior art.
[0032] Another object of the present invention is to provide a method that can solve all of the problems associated with the prior art disclosed herein. [Means for solving the problem]
[0033] The following summarizes the technical aspects of the present invention to achieve the main objectives.
[0034] A first aspect of the present invention relates to a method for triggering status changes and / or specific actions between two avatars operating in a virtual world, which may be a metaverse or a virtual / mixed / augmented reality world. After mapping two gaze vectors of the two avatars in the metaverse or virtual world, the method detects whether eye contact has been established between the two avatars and, if so, triggers a subsequent action, such as allowing social interaction between them, depending on that condition.
[0035] This approach avoids problems arising from unwanted interactions, concerns about the safety of avatars in the metaverse, and the need to introduce physical boundaries that can make virtual environments unrealistic.
[0036] According to a second aspect, the invention relates to a method for setting a predetermined social interaction time and the possibility of triggering further social interactions as conditions, which prevents simple gaze from being mistaken for eye contact.
[0037] According to a third aspect, the invention relates to a method for setting a predetermined gaze aversion time and setting the possibility of triggering further social interactions as a condition. This feature can prevent unwanted social interactions by malicious avatars.
[0038] Further, other aspects of the present invention relate to additional method features as set out in the dependent claims herein.
[0039] The structural and functional features of the present invention and its advantages over the prior art will become more apparent from a consideration of the subclaims as well as the following description taken in conjunction with the accompanying drawings, in which preferred, but not limiting, schematic embodiments of the method, system and apparatus of the present invention are shown. [Brief explanation of the drawings]
[0040] [Figure 1A]1 illustrates a first preferred embodiment of a system architecture according to the present invention. [Figure 1B] 1 illustrates a first preferred embodiment of a system architecture according to the present invention. [Figure 2a] 1 shows a flowchart of a method according to a first preferred embodiment of the present invention and its variants; [Figure 2b] 1 shows a flowchart of a method according to a first preferred embodiment of the present invention and its variants; [Figure 3A] 2 illustrates a second preferred embodiment of a system architecture according to the present invention. [Figure 3B] 2 illustrates a second preferred embodiment of a system architecture according to the present invention. [Figure 4a] 3 shows a flowchart of a method according to a second preferred embodiment of the present invention and its variants; [Figure 4b] 3 shows a flowchart of a method according to a second preferred embodiment of the present invention and its variants; [Figure 5] 1 illustrates the functioning of the method according to the invention within a virtual world. [Figure 6a] 1 shows an example of a region of interest according to the method of the present invention. [Figure 6b] 1 shows an example of a region of interest according to the method of the present invention. [Figure 6c] 1 shows an example of a region of interest according to the method of the present invention. [Figure 6d] 1 shows an example of a region of interest according to the method of the present invention. [Figure 7] A schematic representation of gaze behavior consisting of an initial fixation, a saccade, and a second fixation. [Figure 8] A schematic representation of gaze behavior consisting of an initial fixation, a saccade, and a second fixation. [Figure 9a] 3 shows a flow chart of a further preferred embodiment of the method according to the invention; [Figure 9b] 3 shows a flow chart of a further preferred embodiment of the method according to the invention; DETAILED DESCRIPTION OF THE INVENTION
[0041] This disclosure describes a method for triggering status changes and / or specific actions between two avatars operating in a metaverse or virtual world (which may be a virtual / mixed / augmented reality world). The metaverse or these virtual worlds (i.e., virtual / mixed / augmented reality worlds) is a system of multiple computers connected by wire or wirelessly to a network. In some examples, the network may be in the form of a local area network (LAN), a wide area network (WAN), a wired network, a wireless network, a personal area network, or a combination thereof, and may encompass the Internet, similar to current metaverse architectures. As previously mentioned, in the so-called metaverse, each user controls a humanoid avatar.
[0042] A typical scenario in the Metaverse is during a coffee break at a virtual seminar. Participants may want to network over a drink and, via their avatars, attempt to strike up a conversation with someone who holds an intriguing job title or works at a company of particular interest. The first and most basic interaction is to establish eye contact with the person of interest (especially if they don't know the person yet). If the person returns the gaze, that is, eye contact is established, the next step could be a deeper interaction, such as a conversation or the exchange of detailed information about the job.
[0043] A similar situation can occur when an avatar is walking down the street in a virtual world: a particular person catches the eye of those around them, and the avatar instinctively turns its gaze toward that person, hoping that the gaze will be returned in kind, in order to initiate a deeper social interaction.
[0044] Conversely, if a user is walking in an unsafe area and senses that malicious avatars are staring at them, they may wish to avoid any eye contact with the malicious targets and thus avoid any interaction at all. At the very least, the user may only make a very brief, preliminary eye contact to verify that the malicious avatar is indeed staring at them, and then cut off further eye contact with the malicious avatar.
[0045] Due to technological limitations in the consumer market, current consumer VR headsets lack an eye-tracking module that can detect the user's gaze direction and replicate it on a corresponding avatar in the virtual world, so this type of social interaction is technically substituted by having the user click or point at another avatar with their device (mouse, pad, etc.) to request a specific connection and, in some cases, obtain more information about that person.
[0046] The present invention aims to implement an eye-contact based method to trigger automatic actions / status changes on avatars and improve social interactions between users acting in the metaverse or virtual world.
[0047] In the metaverse and virtual world, all objects, including avatars, are located according to metaverse world coordinates, and the metaverse simulation engine controls the state of the virtual environment and maintains global information about object locations within the metaverse.
[0048] It is also known that avatars are represented as 3D meshes, i.e., mathematically modeled humanoid entities, and the positions of the avatar's face, eyes, nose, mouth, and even body parts are known.
[0049] Each avatar has a virtual camera with known parameters (e.g., focal length) that renders the Metaverse 3D scene from its viewpoint. The avatar's virtual camera is linked to the avatar's gaze vector, changing its position and orientation within the Metaverse world coordinate system.
[0050] The present invention covers two system architecture scenarios corresponding to two different device systems. In the first scenario, the system includes first and second wearable devices 1, 2 worn by first and second users, and eye-tracking devices (eye-tracking glasses / smart glasses) equipped with an eye-tracking module and a front-facing camera, capable of detecting the gaze direction of each of the first and second users (FIGS. 1a, 1b). The system further includes first and second display devices 10, 20 that are part of the first and second computing devices 11, 21 and visible to each user wearing the eye-tracking glasses / smart glasses, and one or more servers 3 that provide virtual scenes 12, 22 of the virtual world shown on the display devices 10, 20 based on the respective virtual scenes 12, 22 of the first and second users. The bidirectional arrows in FIGS. 1a and 1b represent bidirectional communication between the first and second computing devices 11, 21, the server 3, and the first and second wearable devices 1, 2.
[0051] In the second scenario (FIGS. 3a and 3b), the system includes first and second wearable devices 1 and 2, i.e., first and second VR headsets 1 and 2. These VR headsets 1 and 2 are equipped with eye-tracking modules that can detect where each user is looking on the display devices within the VR headsets worn by the first and second users, thereby identifying the gaze direction of each of the first and second users. The first and second VR headsets 1 and 2 further include first and second display devices 10 and 20 integrated into the VR headsets. The first and second display devices 10 and 20 are visible to each user wearing the VR headsets. The first and second VR headsets 1 and 2 can be connected to one or more servers 3 via the Internet or a local LAN. The servers 3 provide virtual worlds displayed on the display devices 10 and 20 based on virtual scenes 12 and 22 corresponding to the first and second users. The bidirectional arrows in FIGS. 3a and 3b indicate bidirectional communication between the server 3 and the first and second wearable devices 1 and 2.
[0052] The eye tracking device 1 may include a frame having at least one receiving opening / lens receiving opening for the disc-shaped structure, and a U-shaped portion in which a right eye acquisition sensor and a left eye acquisition sensor are preferably disposed, which detect the position of the user's eyes and continuously determine the gaze direction during use.
[0053] Additionally, a U-shaped portion of the frame may be provided to position the eye-tracking device 1 on the person's nose.
[0054] A third mixed scenario is a system in which a first user wears an eye-tracking device and a second user wears a VR headset (or vice versa), where the eye-tracking device uses the method according to the first scenario above and the VR headset uses the method according to the second scenario described herein.
[0055] The designation of "right", "left", "upper", or "lower" indicates a predetermined wearing direction when the eye-gaze tracking device 1 is worn by a person.
[0056] As described above, in a preferred configuration, the right eye acquisition sensor is disposed on the right nose frame of the eye tracking device, and the left eye acquisition sensor is disposed on the left nose frame of the eye tracking device. The binocular acquisition sensors are configured as digital cameras and may include objective lenses. In a preferred configuration, each eye acquisition camera is configured to observe one eye of the person wearing the eye tracking device 1 and generate an eye video or individual images including individual eye images.
[0057] According to a preferred embodiment of the eye tracking device 1, at least one field of view camera is arranged on the frame of the eye tracking device, preferably on the U-shaped portion of the frame. The field of view camera is arranged to record a field of view image, which includes individual field of view images and successive field of view images. This allows the recordings of the two eye acquisition cameras and the recordings of the at least one field of view camera to be correlated and input in the field of view image of the corresponding gaze point. It is also possible to arrange a larger number of field of view cameras on the eye tracking device 1.
[0058] In carrying out the method of the present invention, a non-glasses-shaped eye-tracking module may be used, which comprises at least two eye sensors (one for each eye) and the field of view camera described above, and thus can be used in any kind of eye-tracking device.
[0059] Preferably, the eye tracking device 1 includes electronic components such as a data processing unit and a data interface. The data processing unit may be connected to a right eye acquisition sensor and a left eye acquisition sensor. The eye tracking device 1 may further include an energy storage unit for supplying power to the right eye acquisition sensor, the left eye acquisition sensor, the data processing unit, and the data interface.
[0060] In particular, according to a preferred embodiment of the eye-tracking device 1, the electronic components, including the processor and connected storage medium, may be located on the side of the frame of the eye-tracking device 1. With this configuration, all recording, initial analysis, and storage of the recorded footage may be performed by the eye-tracking device 1 itself, or by a computer device 2 connected to the eye-tracking device 1.
[0061] The data processing unit also includes a data memory. Preferably, the data processing unit is a combination of a microcontroller or processor and RAM. The data processing unit is connected to the data interface in a signal-transmitting manner. The data interface and the data processing unit can also be integrated into a single piece of hardware, for example, by an ASIC or FPGA. The interface can be designed as a wireless interface (such as Bluetooth or IEEE802.x) or a wired interface (such as USB). The eye tracking device 1 is provided with a socket compatible with, for example, a microUSB. Additional sensors can also be integrated into the eye tracking device 1 and connected to the data processing unit. The data processing unit and the data interface are at least indirectly connected to the energy storage unit in the circuit, and are further connected in a signal-transmitting manner to the field of view camera, the right eye acquisition sensor, and the left eye acquisition sensor.
[0062] The gaze vector in the real world can also be obtained using a stationary eye-tracking device, which is a fixed device mounted at a known, fixed position relative to the display. In this case, the so-called first and second computer displays provide the user's gaze vector relative to the head coordinate system. As mentioned above, this method is particularly suitable for the eye-tracking glasses introduced in the first scenario.
[0063] In the second scenario, the VR headset functions as a head-mounted device, such as goggles, and includes at least a head-mounted display for stereoscopic vision, capable of presenting separate images to each eye, stereo sound, and tracking sensors for detecting head movement.
[0064] VR headsets are strapped to the user's head and cover both eyes, visually immersing the user in the content being viewed. Users can select and browse 3D content using eye gaze gestures, or they can use hand controllers such as gloves. These controllers and eye gaze controls track the user's physical movements and appropriately position simulated images and videos on the display device to alter perception.
[0065] The VR headset may further include optional devices such as wired / wireless connections, sensors that detect and transmit user movements to a computer or smartphone, audio headphones, and a camera, which are used to enhance the user experience.
[0066] The first scenario is more complex than the second because multiple reference coordinate systems must be transformed to locate the user's gaze vector in the virtual world. These coordinate systems include the real-world coordinate system, the eye-tracking coordinate system (head frame), the display coordinate system (XY plane—the display visible to the user), the metaverse virtual camera coordinate system, the avatar head coordinate system, and the metaverse world coordinate system. The display coordinate system is particularly noteworthy. The display is assumed to be a rectangular display with a known width and height. The X and Y axes of the display coordinate system are aligned along the edges of the display, and the Z axis is positioned so that the XYZ axes form a left-handed coordinate system. The display's position in real-world coordinates is uniquely determined by its position in the image plane (XY plane) and its orientation within the real-world coordinate system.
[0067] In the first and second scenarios, a method for triggering an action in the metaverse or virtual world, which may be provided by the server 3 according to the present invention, includes the following steps, and is based on the assumption of a humanoid avatar (see FIG. 2a):
[0068] (Step 100) A first avatar 13 and a second avatar 23 are placed in the same virtual environment within the metaverse or virtual world, and the first and second avatars 13, 23 are able to view each other in the virtual environment through their respective virtual cameras. A virtual scene 12 including the second avatar 23 is rendered on a first display device 10 viewable by a first user based on the virtual camera of the first avatar, and a virtual scene 22 including the first avatar 13 is rendered on a second display device 20 viewable by a second user based on the virtual camera of the second avatar. (This step is performed by a simulation engine of the metaverse. The simulation engine renders the 3D scene from the avatar's virtual camera on the display device of the corresponding user.)
[0069] (Step 200) While the first gaze tracking device 1 acquires data of a first gaze vector 14 of a first user, the second gaze tracking device 2 acquires data of a second gaze vector 24 of a second user.
[0070] (Step 500) Within the metaverse or virtual world, the coordinates of the first gaze vector 14 are mapped onto the first avatar 13, while the coordinates of the second gaze vector 24 are mapped onto the second avatar 23.
[0071] (Step 550) Move the eyes of the first avatar 13 and the second avatar 23 in the virtual environment according to the data of the first gaze vector 14 and the second gaze vector 24, respectively (see FIGS. 2B and 4B). This step is optional and applicable to all embodiments of the present invention and both system configuration scenarios.
[0072] (Step 600) Within the metaverse or virtual world, a predetermined first region of interest 612 is defined on the face of a first avatar, while a second region of interest 622 is defined on the face of a second avatar.
[0073] (Step 700) Trigger an action in the first avatar 13 and the second avatar 23 when the first gaze vector 14 is directed toward the second region of interest 622 on the face of the second avatar and at the same time the second gaze vector 24 is directed toward the first region of interest 612 on the face of the first avatar.
[0074] It is worth noting that step 200 can be performed by a portable eye-tracking device as described earlier in this specification, or by a stationary eye-tracking system with stereo cameras mounted at known fixed positions relative to the display device. The stereo cameras can use image recognition technology to identify the user's face, eyes, and pupils, and then calculate the gaze vector using the stereo data acquired from the cameras.
[0075] The regions of interest 612, 622 mentioned in step 600 are preferably regions that include both eyes of the avatars 13, 23. In the method of the present invention, any region of interest that satisfies this condition is a valid candidate. Preferably, the regions of interest 612, 622 may be defined as the smallest convex hull that encompasses both eyes on the avatar mesh / model (see FIG. 6a). A match between the gaze vector and the corresponding region of interest 612, 622 on the other avatar indicates a willingness to initiate a social interaction between two users acting through the avatars. The regions of interest 612, 622 may be defined as a social triangle, an imaginary inverted isosceles triangle on the avatar's face that includes both eyes and has the vertex of the triangle's equal sides centered at the mouth or chin (FIG. 6b). In this case, a match between the gaze vector and the region of interest may suggest an emotional engagement with the other avatar in addition to a mere willingness to initiate a social interaction. Furthermore, the region of interest 612, 622 may be defined as an imaginary inverted isosceles triangle with its base at the center of the forehead and its vertex at the lowest point of the nose (FIG. 6d) or the midpoint between the eyebrows (FIG. 6c) of the avatar the user is looking at, where a gaze vector matching the region of interest 612, 622 indicates a desire (even if accompanied by anxiety or caution) for business or a formal interaction.
[0076] As described above, the regions of interest 612, 622 in the metaverse or virtual world can be defined very precisely because the coordinates of the entire humanoid shape of the avatar are known. Therefore, whether a convex hull region, a social triangle, or a formal triangle is selected, the regions of interest 612, 622 can be uniquely defined using specific points on each model avatar.
[0077] Furthermore, a boundary region may be set around the regions of interest 612, 622. This is to increase the probability of eye contact and to compensate for positional shifts and mismatches due to depth differences caused by the regions of interest 612, 622 being 3D curved surfaces.
[0078] A further preferred embodiment of the method according to the present invention solves the problem of how to make a first avatar 13 (and therefore its user) aware that another avatar, i.e., a second avatar 23, sees it and wants to interact with it, since an important opportunity for interaction may be missed without being noticed. This problem can be solved by the following steps:
[0079] If the second gaze vector 24 is directed toward the first region of interest 612 on the face of the first avatar, trigger a stimulus action in the first avatar 13. The above steps may be implemented in any embodiment disclosed herein, with the roles of the first avatar 13 and the second avatar 23 reversed to match the wording of the corresponding method steps.
[0080] This function aims to enable the first avatar 13 or the second avatar 23, and thus the corresponding user, to intuitively recognize that they are being watched by others. The stimulus action may be displaying a special sign / symbol on the first display device 10, changing a specific status, or any other means that allows the first user to clearly sense that the second avatar 23 operated by the second user is watching their first avatar 13.
[0081] Furthermore, when implementing steps 650 and 850 as a condition for triggering an advanced interaction action, it may be necessary to establish eye contact twice between both avatars 13, 23. In this case, the following steps may be added to all embodiments disclosed in this invention:
[0082] After first eye contact is established with the first gaze vector 14 pointing at the second region of interest 622 on the face of the second avatar and simultaneously the second gaze vector 24 pointing at the first region of interest 612 on the face of the first avatar, if second eye contact is re-established between both avatars 13, 23, an advanced interaction action is triggered in the first avatar 13 and the second avatar 23. This technical feature aims to safely and robustly detect a truly clear intention of social interaction (advanced interaction action) between two users via their respective avatars.
[0083] In the first scenario, in order to take into account the relative position of the eye-tracking devices 1, 2 (so-called eye-tracking system pose) with respect to the display devices 10, 20 of the computing devices 11, 21, the method further comprises the following steps:
[0084] (Step 300) Determining the pose of the first wearable device 1 and the second wearable device 2 (both of which are eye tracking devices 1, 2) relative to the real world coordinate system, where the eye tracking system pose refers to the position and orientation of the eye tracking system in the real world coordinate system.
[0085] In order to determine the orientation of the eye-tracking devices 1 and 2, it is necessary to know the orientation of the devices in the real-world coordinate system. There are several options for obtaining the orientation of the eye-tracking device.
[0086] The first option is to use the front camera of the eye-tracking glasses to obtain the pose of the glasses relative to the display device. This can be achieved by displaying a specific marker (e.g., an Aruco marker) on (or near) the display device and using image recognition techniques to obtain the pose of the marker. This information can then be used to map the gaze to the coordinate system of the display device. To achieve the same goal, image recognition techniques can also be used to detect the pose of the display device relative to the camera frame, based on specific algorithms.
[0087] As a second example, there is a method of acquiring the gaze vector using a stationary eye tracking device. The stationary eye tracking device is attached at a known fixed position relative to the display device. Since the orientation of the user's eyes and the orientation of the eye tracking device are known relative to the display device, the detected gaze vector can be mapped to the display device coordinate system using a transformation matrix.
[0088] Another problem to be solved is how to account for ocular parallax correction, which can be corrected using offset data between the vertical position of the front camera of the eye tracking device 1, 2 and the user's eyes, and by knowing the distance between the eye tracking glasses 1, 2 and the display device 10, 20. Note that this distance is known from the pose of the eye tracking glasses obtained by one of the methods described above.
[0089] Thus, the relevant steps are to calculate the distance between the eye tracking glasses 1, 2 and the corresponding display devices 10, 20 and then perform the parallax correction. Furthermore, the following steps relate to the first scenario:
[0090] (Step 400) A step of converting (by applying an extrinsic matrix to perform coordinate transformation) the first and second gaze vectors 14, 24 of the first and second users determined by the first and second wearable devices 1, 2 (both eye tracking devices 1, 2) into the coordinate systems of the first and second display devices 10, 20 of the first and second computer devices 11, 21.
[0091] In the first and second scenarios of the present invention, and generally in all embodiments of the present invention, the following additional steps can be performed to define the eye contact duration, so as to prevent the eye contact from being mistaken for a mere gaze phenomenon and to recognize a real interest in a positive social interaction. The corresponding steps are as follows (see Fig. 9a):
[0092] (Step 650) Following step 600, determining the eye contact time, i.e., the time during which the first gaze vector 14 and the second gaze vector 24 are simultaneously directed towards the corresponding first area of interest 612 and second area of interest 622, respectively, and then setting a predetermined social interaction time corresponding to the actual willingness to engage in social interaction.
[0093] (Step 850) A step of triggering an interaction action to the first avatar 13 and the second avatar 23 when the first gaze vector 14 is directed toward the second area of interest 622 on the face of the second avatar and simultaneously the second gaze vector 24 is directed toward the first area of interest 612 on the face of the first avatar and the eye contact time matches a predetermined social interaction time.
[0094] The occurrence of a gaze phenomenon is determined when the eye contact duration exceeds a predetermined social interaction duration, so by applying the above steps, the gaze phenomenon is detected and avoided.
[0095] In other situations, one may wish to prevent unwanted interactions altogether. For this purpose, a gaze aversion time can be defined, i.e., a short gaze to determine whether the avatar the user is looking at is a potential target for interaction. For this purpose, the corresponding steps that can be performed in all embodiments of the present invention are as follows (see FIG. 9b):
[0096] (Step 670) Following step 600, a step of determining the time during which the first gaze vector 14 and the second gaze vector 24 are simultaneously directed toward the corresponding first area of interest 612 and second area of interest 622, respectively, as the eye contact time, and then setting a predetermined gaze aversion time to prevent social interaction.
[0097] (Step 800) Triggering an avoidance action in the first avatar 13 and the second avatar 23 when the first gaze vector 14 is directed toward the second region of interest 622 on the face of the second avatar and simultaneously the second gaze vector 24 is directed toward the first region of interest 612 on the face of the first avatar and the eye contact time matches a predetermined gaze avoidance time.
[0098] In the present invention, a preferred solution is to define a gaze aversion time event criterion when the gaze vectors 14, 24 stabilize over the respective corresponding regions of interest 612, 622 and coincide with a predetermined eye contact time, preferably a time period of 0.5 seconds or more but less than 2 seconds (more precisely, 0.5 seconds≦t<2 seconds), according to step 600. Furthermore, a preferred solution is to define a social interaction time event criterion when the gaze vectors 14, 24 stabilize over the respective corresponding regions of interest 612, 622 and coincide with a predetermined eye contact time, preferably a time period of 2 seconds or more but less than 4 seconds (more precisely, 2 seconds≦t≦4 seconds), according to step 600.
[0099] The method according to the invention can also be carried out on the display device of a smartphone or any computing device equipped with a screen, in particular a touch screen.
[0100] The method of the present invention includes the well-known concept of gaze (fixation) that can be used to set the gaze time. One definition of this concept can be easily understood as shown in Figures 7 and 8 and the following paragraphs.
[0101] 7 and 8, exemplary viewpoints 37, 38 are examined and compared by a comparison device for conformance with at least a first fixation criterion 25. The comparison device can be any suitable device, but is particularly preferably a device that includes an integrated electronic logic module, in particular a processor, a microprocessor, and / or a programmable logic controller. It is also particularly preferred that the comparison device be implemented in a computer.
[0102] The comparator processes so-called visual coordinates (hereinafter abbreviated as VCOs), which are calculated based on the above correlation function between the visual field image 79 and the eye image 78, although other methods and techniques can also be used to calculate the VCOs.
[0103] The first fixation criterion 25 may be any criterion capable of distinguishing between fixations and saccades. In a preferred embodiment of the method of the present invention, the first fixation criterion 25 is a predefinable first distance 39 centered on a first viewpoint 37. A first relative distance 44 between the first viewpoint 37 and a second viewpoint 38 is determined. If the first relative distance 44 is less than the first distance 39, the first and second viewpoints 37, 38 are assigned to a first fixation 48. Therefore, as long as the second viewpoint 38 following the first viewpoint 37 remains within the foveal region 34 of the first viewpoint 37, i.e., within the ordered perception region of the first viewpoint 37, ordered perception continues to be uninterrupted and the first fixation criterion 25 is satisfied. This is therefore the first fixation 48. In a particularly preferred embodiment of the method of the present invention, the first distance 39 is a first visual angle 41. The first visual angle 41 describes the region 34 assigned to central vision, specifically within a radius of 0.5° to 1.5°, preferably approximately 1°. The distance between the first viewpoint 37 and the second viewpoint 38 is defined as a first relative angle 42. Based on visual coordinates calculated using an eye-tracking system, saccades and fixations 48, 49 can be easily and accurately identified. Figure 7 shows, for example, a first fixation 48 formed from a series of four viewpoints 37, 38, 69, and 70, along with a first distance 39, a first visual angle 41, a first relative distance 44, and a first relative angle 42. A first circle 43 having a radius of the first distance 39 is defined around each viewpoint 37, 38, 69, and 70, indicating that the subsequent viewpoints 38, 69, and 70 are within the first circle 43 having a radius of the first distance 39 from the previous viewpoints 37, 38, and 69, respectively. This satisfies the desired first fixation criterion 25. In order to adapt to different viewing objects, different subjects and / or conditions, in further update embodiments of the present invention, the first fixation criterion 25, in particular the first distance 39 and / or the first viewing angle 41, may be predefined.
[0104] FIG. 8 illustrates a sequence of viewpoints in which viewpoints 37, 38, 69, 70, 71, 72, 73, 74, and 75 do not all meet first fixation criterion 25. The first four viewpoints 37, 38, 69, and 70 meet first fixation criterion 25 and together form first fixation 48. None of the three viewpoints 71, 72, and 73 following first fixation 48 meet first fixation criterion 25. Only fourth viewpoint 74 following first fixation 48 meets first fixation criterion 25 compared to third viewpoint 73 following first fixation 48. Therefore, third viewpoint 73 following first fixation 48 becomes the first viewpoint 73 of second fixation 49. Second fixation 49 consists of three viewpoints: viewpoints 73, 74, and 75. FIGS. 7 and 8 are merely illustrative; fixations 48 and 49 may occur in natural environments in a variety of viewpoint arrangements. Furthermore, the area between the final viewpoint 70 of the first fixation 48 and the first viewpoint 73 of the second fixation 49 forms a saccade (a region of perceptual discontinuity), and the angle between the final viewpoint 70 of the first fixation 48 and the first viewpoint 73 of the second fixation 49 is referred to as the first saccade angle 52.
[0105] Once the viewpoints 37 and 38 have been assigned as saccades or fixations 48, 49, they can be output for further evaluation, processing, and display. In particular, the first viewpoint 37 and the second viewpoint 38 can be output and marked as a first fixation 48 or a first saccade. Below are additional definitions of fixations and saccades that can be used and implemented in the present method to mark fixation events of the present invention:
[0106] Saccades are rapid eye movements reaching up to 500° per second, while fixations involve the eyes remaining relatively still for approximately 200–300 milliseconds (Reference 5).
[0107] Fixations are eye movements that stabilize the retina on a stationary object, while saccades are rapid eye movements that reposition the fovea to a new location in the visual environment (Reference 5).
[0108] A distinction is made between a period in which a region of the visual scene is held on the fovea as a “fixation” and a period in which a rapid change in eye position that moves a region of the visual scene toward the fovea as a “saccade” (Reference 5).
[0109] A saccade is defined as a movement of the gaze direction of the wearer of the eye-tracking device by more than a certain angle within a certain time (i.e., exceeding a minimum angular velocity), where the cutoff criterion can be specified in units of angular velocity.
[0110] The method of the present invention may trigger multiple types of actions, including actions, avoidance actions, interaction actions, and advanced interaction actions, which may include changing the avatar's status, causing the avatar to make a facial gesture, highlighting the avatar's nickname, setting cookie acceptance, granting consent to certain privacy settings, etc. An interaction action may consist of initiating the first stage of interaction, such as opening a chat box between the two avatars to allow users to chat and exchange preliminary information via the avatars, or displaying the users' real names and countries of location.
[0111] Advanced interaction actions may allow access to another communication channel between users via voice messages or video content. If the eye-tracking devices 1, 2 and the VR headset are equipped with speakers and microphones, these may be automatically activated to enable the sending and receiving of voice data within the system, allowing users to speak and hear each other.
[0112] Further actions that can be triggered include allowing "physical contact" between avatars in the metaverse or virtual world (e.g., handshakes, hugs, etc.), or automatically changing the privacy settings of a particular avatar, meaning automatically displaying details about the user controlling the other avatar after eye contact is established, and changing settings regarding whether or not to receive commercial offers, advertising, and technical cookies.
[0113] In contrast, avoidance actions may include blocking further eye contact with other avatars or blocking any possibility of physically approaching the avatar in the metaverse / virtual world.
[0114] The invention further relates to a server 3 comprising a VR headset 1, 2, a computing device 11, 21, a processor and a computer readable storage medium connected to said processor, the computer readable storage medium storing computer executable instructions which, when executed, configure the processor to perform corresponding steps of the methods previously described herein.
[0115] The invention further relates to an eye-tracking device 1, 2 comprising a processor and a computer-readable storage medium connected to said processor, the computer-readable storage medium storing computer-executable instructions which, when executed, configure the processor to perform corresponding steps of some of the methods previously described herein, in particular the following steps:
[0116] (Step 100) A step of making a first avatar 13 and a second avatar 23 visible to each other in the same virtual environment within the metaverse or virtual world, the first and second avatars 13, 23 visible to each other through their respective virtual cameras, rendering a virtual scene 12 including the second avatar 23 on a first display device 10 viewable by a first user based on the virtual camera of the first avatar, and rendering a virtual scene 22 including the first avatar 13 on a second display device 20 viewable by a second user based on the virtual camera of the second avatar.
[0117] (Step 200) Providing data of a first gaze vector 14 of a first user by a first gaze tracking device 1 and providing data of a second gaze vector 24 of a second user by a second gaze tracking device 2.
[0118] (Step 500) In the metaverse or virtual world, a step of mapping the coordinates of the first gaze vector 14 onto a first avatar 13 and mapping the coordinates of the second gaze vector 24 onto a second avatar 23.
[0119] (Step 600) Identifying a predetermined first region of interest 612 on the face of a first avatar and a second region of interest 622 on the face of a second avatar within a virtual environment.
[0120] (Step 700) Triggering an action in the first avatar 13 and the second avatar 23 when the first gaze vector 14 is directed toward a second region of interest 622 on the face of the second avatar and at the same time the second gaze vector 24 is directed toward a first region of interest 612 on the face of the first avatar.
[0121] The eye tracking devices 1, 2 defined above may be further configured to perform all other technical features and methods according to embodiments described herein.
[0122] The subject of the present invention is a computer-readable storage medium connected to a processor, the computer-readable storage medium having stored thereon computer-executable instructions that, when executed, configure the processor to perform the corresponding steps of the methods previously described herein, according to all embodiments described and disclosed herein.
[0123] The present invention provides a system for triggering status changes and / or specific actions between two avatars operating in a virtual world, which may be a metaverse or a virtual / mixed / augmented reality world. The system includes at least first and second wearable devices (1, 2), a processing unit capable of processing eye-tracking data from the wearable devices (1, 2), and a computing system connectable to the processing unit and configured to host virtual worlds displayed on first and second display devices according to virtual scenes (12, 22) of a first and second user, respectively. The computing system may be a server device (3) equipped with the processing unit, or may include multiple servers or computer devices. The computing system may also be implemented and operate as a server hosting a virtual world, or may include such a server. A system consisting of one or more computers may be configured to perform specific operations or processes by software, firmware, hardware, or a combination thereof installed on the one or more computers. The software, firmware, hardware, or a combination thereof, when executed, causes the system to perform the processes. Furthermore, one or more computer programs may be configured to perform specific operations or processes by instructions. The instructions, when executed by a processor of the system, cause the system to perform the process.
[0124] Other embodiments include corresponding computer systems, computer-readable storage media, or devices configured to perform the steps of the methods described herein, as well as computer programs stored on one or more computer-readable storage media or computer storage devices.
[0125] In a preferred embodiment, the computing system is connected to a processing unit connectable to or forming part of the wearable device 1,2.
[0126] A processing unit may operate in a client role when connected to a computer system acting as a server. The client and server are typically remote from each other and typically interact through a communications network such as a TCP / IP data network. The client-server relationship is effected by software running on the respective devices.
[0127] Furthermore, the system is typically configured to perform the various processes described herein. In a preferred embodiment, the system comprises a first wearable device 1 and a second wearable device 2 as described herein, and at least one processing unit configured to perform the steps of the methods described herein in all preferred embodiments described.
[0128] Furthermore, in the system, the processing unit may be provided by a desktop computer or a server, or the processing unit may be integrated into the wearable devices 1, 2 described in all embodiments herein.
[0129] <References> (Reference 1) Sang-Min Park, et al. (2022). A Metaverse: Taxonomy, Components, Applications, and Open Challenges. (Reference 2) M. Kaur, et al. (2021). Metaverse Technology and the current Market. (Reference 3) Justin Goldston et al. (2022). The Metaverse as Digital Leviathan: a case study of Bit. Country: (Reference 4) Smita. Verma (2022). Metaverse Vs. Virtual Reality: a detailed comparison. https: / / www.blockchain-council.org / Metaverse / Metaverse-vs-virtual-reality / (Reference 5) Roy S. Hessels (2017). Noise-robust fixation detection in eye movement data: Identification by two-means clustering (I2MC)
Claims
1. A method for triggering actions in a metaverse or virtual world in which a first user wearing a first wearable device (1) and a second user wearing a second wearable device (2) each act through a humanoid avatar, comprising: a step (100) of placing a first avatar (13) and a second avatar (23) in the same virtual environment within the metaverse or virtual world, the first and second avatars (13, 23) being visible to each other through their respective virtual cameras, causing the virtual camera of the first avatar to render a virtual scene (12) including the second avatar (23) on a first display device (10) viewable by the first user, and causing the virtual camera of the second avatar to render a virtual scene (22) including the first avatar (13) on a second display device (20) viewable by the second user; receiving (200) first gaze vector data of the first user by a first gaze tracking device (1) and second gaze vector data of the second user by a second gaze tracking device (2); a step (500) of mapping coordinates of the first gaze vector onto the first avatar (13) and mapping coordinates of the second gaze vector onto the second avatar (23) in the metaverse or virtual world, the first avatar (13) corresponding to the first user and the second avatar (23) corresponding to the second user; defining (600) a predetermined first region of interest (612) on the face of the first avatar and a second region of interest (622) on the face of the second avatar within the virtual environment; Triggering (700) an action in the first avatar (13) and the second avatar (23) when the first gaze vector (14) is directed toward the second region of interest (622) on the face of the second avatar and the second gaze vector (24) is directed toward the first region of interest (612) on the face of the first avatar; A method comprising:
2. The first wearable device (1) and the second wearable device (2) are VR headsets equipped with an eye-tracking module; The method of claim 1.
3. A step (300) of determining the orientation of the first wearable device (1) and the second wearable device (2), which are eye-tracking devices (1, 2), relative to a real-world coordinate system; transforming the first and second line-of-sight vectors (14, 24) into the coordinate systems of the respective display devices (10, 20) of the first and second computer devices (11, 21); further comprising: The method of claim 1.
4. and calculating a distance between the eye-tracking glasses (1, 2) and the corresponding display devices (10, 20) and then calculating a parallax correction. The method of claim 3.
5. The method further includes a step of moving the eyeballs of the first avatar (13) and the second avatar (23) based on the data of the first and second gaze vectors (14, 24), respectively.
5. The method according to any one of claims 1 to 4.
6. The region of interest (612, 622) is located within the virtual environment. a convex hull that contains both eyes of the avatar (13, 23); a social triangle defined on the face of the avatar (13, 23) as an inverted isosceles triangle that includes both eyes and has a vertex common to the equal sides of the triangle located at the center of the mouth of the avatar (13, 23); or a formal triangle defined as an imaginary inverted isosceles triangle with its base at the center of the forehead and its vertex common to the equal sides located at the lowest point of the avatar's (13, 23) nose or the midpoint between the eyebrows; It is defined as 6. The method according to any one of claims 1 to 5.
7. determining an eye contact time during which the first and second gaze vectors (14, 24) are simultaneously directed toward the corresponding first and second regions of interest (612, 622), and setting a predetermined gaze aversion time (670) to prevent social interaction; triggering (800) an avoidance action in the first avatar (13) and the second avatar (23) when the first gaze vector (14) is directed toward the second region of interest (622) on the face of the second avatar and simultaneously the second gaze vector (24) is directed toward the first region of interest (612) on the face of the first avatar, and the eye contact time matches the predetermined gaze aversion time; further comprising:
7. The method according to any one of claims 1 to 6.
8. The gaze avoidance time is 0.5 seconds≦t<2 seconds. The method of claim 7.
9. determining an eye contact time as a time during which the first and second gaze vectors (14, 24) are simultaneously directed toward the corresponding first and second regions of interest (612, 622), and then setting a predetermined social interaction time (650) corresponding to an actual willingness to engage in social interaction; triggering (850) an interaction action in the first avatar (13) and the second avatar (23) when the first gaze vector (14) is directed toward a second region of interest (622) on the face of the second avatar and simultaneously the second gaze vector (24) is directed toward a first region of interest (612) on the face of the first avatar, and the eye contact time matches the predetermined social interaction time; further comprising:
9. The method according to any one of claims 1 to 8.
10. the social interaction time is 2 seconds≦t≦4 seconds; 10. The method of claim 9.
11. the method further comprises the step of triggering an advanced interaction action in the first avatar (13) and the second avatar (23) when second eye contact is established again between the two avatars (13, 23) after first eye contact has been established with the first gaze vector (14) directed toward a second region of interest (622) on the face of the second avatar and simultaneously the second gaze vector (24) directed toward a first region of interest (612) on the face of the first avatar; 11. The method according to claim 9 or 10.
12. A computer-readable storage medium containing computer-executable instructions, The computer-executable instructions, when executed, configure a processor to perform the method of any one of claims 1 to 11. A computer-readable storage medium.
13. a processor; A computer-readable storage medium according to claim 12, coupled to the processor; A server device (3) comprising: The computer-readable storage medium has computer-executable instructions stored thereon; The computer-executable instructions, when executed, configure the processor to perform corresponding steps of the method of any one of claims 1 to 11. Server device (3).
14. a processor; The computer-readable storage medium of claim 12 coupled to the processor; A wearable device (1, 2) comprising: The computer-readable storage medium has computer-executable instructions stored thereon; The computer-executable instructions, when executed, configure the processor to perform corresponding steps of the method of any one of claims 1 to 11. Wearable devices (1, 2).
15. a processor; a computer-readable storage medium coupled to the processor; An eye tracking device (1, 2) comprising: The computer-readable storage medium has computer-executable instructions stored thereon; The computer-executable instructions, when executed, configure the processor to perform a method for triggering actions in a metaverse or virtual world in which a first user wearing a first wearable device (1) and a second user wearing a second wearable device (2) each act through a humanoid avatar; The method comprises: a step (100) of making a first avatar (13) and a second avatar (23) visible in the same virtual environment within the metaverse or virtual world, the first and second avatars (13, 23) visible to each other through their respective virtual cameras, rendering a virtual scene (12) including the second avatar (23) on a first display device (10) viewable by the first user through the virtual camera of the first avatar, and rendering a virtual scene (22) including the first avatar (13) on a second display device (20) viewable by the second user through the virtual camera of the second avatar; providing (200) first gaze vector data of the first user by a first gaze tracking device (1) and second gaze vector data of the second user by a second gaze tracking device (2); a step (500) of mapping coordinates of the first gaze vector onto the first avatar (13) and mapping coordinates of the second gaze vector onto the second avatar (23) in the metaverse or virtual world, the first avatar (13) corresponding to the first user and the second avatar (23) corresponding to the second user; Identifying (600) a predetermined first region of interest on the face of the first avatar and a second region of interest on the face of the second avatar within the virtual environment; Triggering (700) an action in the first avatar (13) and the second avatar (23) when the first gaze vector (14) is directed toward the second region of interest (622) on the face of the second avatar and the second gaze vector (24) is directed toward the first region of interest (612) on the face of the first avatar; Including, and / or The method comprises the additional steps and features of any one of claims 3 to 11. Eye tracking device (1, 2).
Citation Information
Patent Citations
Private communication by gazing at avatar
EP3491781A1
Avatar customization for optimal gaze discrimination
WO2021202783A1