Display method and apparatus, device, medium, and product

By acquiring users' eye-tracking data and target gestures, and dynamically displaying local perspective areas, the problem of perspective areas disrupting immersion in virtual environments is solved, enabling flexible interaction between users and the real world and personalized perspective area updates.

WO2026086718A1PCT designated stage Publication Date: 2026-04-30VIVO MOBILE COMM CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2025-10-20
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

The current method of displaying perspective areas in virtual environments can easily disrupt the user's immersion, especially in immersive experiences such as watching movies and playing games, leading to user fatigue.

Method used

By acquiring the user's eye-tracking data, the gaze position is determined, and when the target gesture is captured, a local perspective area associated with the gaze position is displayed on the screen to display the real scene. The perspective area is dynamically updated to meet the user's interaction needs.

Benefits of technology

Without affecting user immersion or reducing fatigue, it enables flexible interaction between users and the real world, and the dynamic changes in the perspective area meet personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025128634_30042026_PF_FP_ABST
    Figure CN2025128634_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of virtual reality, and discloses a display method and apparatus, a device, a medium, and a product. The display method is applied to an extended reality device. The method comprises: acquiring first eye movement data of a user; on the basis of the first eye movement data, determining a first gaze position corresponding to a gaze point of the user on a display screen of the extended reality device; and when a target gesture of the user is captured, displaying, in a region associated with the first gaze position in the display screen, a first perspective region corresponding to the target gesture, the first perspective region being used for displaying a real scene.
Need to check novelty before this filing date? Find Prior Art

Description

Display methods, devices, equipment, media and products

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411507046.2, filed on October 25, 2024, entitled “Display Method, Apparatus, Device, Medium and Product”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application belongs to the field of virtual reality technology, specifically relating to a display method, device, equipment, medium, and product. Background Technology

[0004] With the development of technology, Extended Reality (XR) devices support the function of enabling Video See-Through (VST) mode in a virtual environment. In VST mode, users can see the scene of the real world and interact with it.

[0005] Currently, the main way to display transparent areas in a virtual environment is by tapping the XR device or touching a virtual button to activate the full transparent mode.

[0006] This approach involves a complete switch from the virtual world to the real world, which can easily disrupt the user's sense of immersion in immersive experiences such as watching movies and playing games. Summary of the Invention

[0007] The purpose of this application is to provide a display method, apparatus, device, medium, and product that can flexibly display partial perspective areas without affecting the user's immersion, so as to meet the user's interaction needs with the real world.

[0008] In a first aspect, embodiments of this application provide a display method applied to an extended reality device, the display method comprising:

[0009] Obtain the user's first eye movement data;

[0010] Based on the first eye movement data, determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device;

[0011] When a user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen. The first perspective area is used to display the real scene.

[0012] Secondly, embodiments of this application provide a display device for use in an extended reality device, the display device comprising:

[0013] The acquisition module is used to acquire the user's first eye movement data;

[0014] The determination module is used to determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device based on the first eye movement data;

[0015] The display module is used to display a first perspective area corresponding to the target gesture in an area associated with the first gaze position on the display screen when the user's target gesture is captured. The first perspective area is used to display the real scene.

[0016] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.

[0017] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first aspect.

[0018] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface, the communication interface and the processor being coupled together, the processor being used to run programs or instructions to implement the steps of the method as described in the first aspect.

[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method as described in the first aspect.

[0020] In this embodiment, the user's first eye-tracking data is acquired; based on the first eye-tracking data, a first gaze position corresponding to the user's gaze point on the extended reality device's display screen is determined; when the user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen, and the first perspective area is used to display the real scene. That is, this embodiment can display a partial perspective area based on the user's gaze position and target gesture, displaying the real scene through the partial perspective area and the virtual scene through other areas. This does not affect the user's immersion and eliminates the need for the user to operate the physical buttons of the extended reality device, thus preventing user fatigue. Furthermore, since the user's eye-tracking data changes, the gaze position determined based on the eye-tracking data also changes accordingly; that is, the partial perspective area can change dynamically, allowing for flexible updates to the partial perspective area. Attached Figure Description

[0021] Figure 1 is a flowchart of a display method provided in an embodiment of this application;

[0022] Figure 2 is a schematic diagram of a first gaze position provided in an embodiment of this application;

[0023] Figure 3 is a flowchart of another display method provided in an embodiment of this application;

[0024] Figure 4 is a schematic diagram of a movement trajectory provided in an embodiment of this application;

[0025] Figure 5 is a schematic diagram of a pinching gesture provided in an embodiment of this application;

[0026] Figure 6 is a schematic diagram of a first partial perspective region provided in an embodiment of this application;

[0027] Figure 7 is a schematic diagram of a confirmation gesture provided in an embodiment of this application;

[0028] Figure 8 is a schematic diagram of a second partial perspective region provided in an embodiment of this application;

[0029] Figure 9 is a logical schematic diagram of a display method provided in an embodiment of this application;

[0030] Figure 10 is a schematic diagram of a display device provided in an embodiment of this application;

[0031] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0032] Figure 12 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0035] As mentioned above, the current method for displaying transparent areas in virtual environments is mainly achieved by tapping the XR device or touching a virtual button to activate the full transparent mode. This method involves a complete switch from the virtual world to the real world, which can easily disrupt the user's immersion in immersive experiences such as watching movies and playing games.

[0036] Therefore, embodiments of this application provide a display method, apparatus, device, medium, and product that can flexibly display partial perspective areas without affecting the user's immersion, thereby meeting the user's interaction needs with the real world.

[0037] The display methods, apparatus, devices, media, and products provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0038] The display method provided in this application can be applied to XR devices, which may include, for example, augmented reality (AR) devices, virtual reality (VR) devices, mixed reality (MR) devices, etc. XR devices may include, for example, head-mounted devices. In some embodiments, an XR device may include VR glasses.

[0039] Figure 1 is a flowchart of a display method provided in an embodiment of this application. As shown in Figure 1, the display method may include the following steps:

[0040] S110, Obtain the user's first eye movement data.

[0041] S120. Based on the first eye movement data, determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device.

[0042] S130. When the user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen.

[0043] The first perspective area is used to display the real scene.

[0044] In this embodiment, the user's first eye-tracking data is acquired; based on the first eye-tracking data, a first gaze position corresponding to the user's gaze point on the extended reality device's display screen is determined; when the user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen, and the first perspective area is used to display the real scene. That is, this embodiment can display a partial perspective area based on the user's gaze position and target gesture, displaying the real scene through the partial perspective area and the virtual scene through other areas. This does not affect the user's immersion and eliminates the need for the user to operate the physical buttons of the extended reality device, thus preventing user fatigue. Furthermore, since the user's eye-tracking data changes, the gaze position determined based on the eye-tracking data also changes accordingly; that is, the partial perspective area can change dynamically, allowing for flexible updates to the partial perspective area.

[0045] The above steps are explained in detail below:

[0046] In S110, the first eye movement data is the user's eye movement data, which can be collected by the eye tracking sensor built into the XR device. The eye movement data may include, for example, the user's pupil size, eye movement time, eye sac direction, blinking, and other data.

[0047] In some embodiments, eye-tracking sensors can be used for eye-tracking calibration before collecting user eye movement data to improve the accuracy and reliability of the initial eye-tracking data. The specific calibration process is not limited in the embodiments of this application.

[0048] In practical applications, the eye-tracking function of XR devices can be enabled by default, meaning the eye-tracking sensor is on by default. Alternatively, to save power, the eye-tracking sensor can be activated only when needed. In this case, users can activate the XR's eye-tracking sensor via touch, voice, or gestures. Once activated, the eye-tracking sensor can collect the user's eye movement data.

[0049] In S120, the first gaze position is the location of the user's gaze point on the XR device's display screen. This means that when a user needs to interact with the real world through a specific area, they can do so by gazing at that area. When the XR device detects the user gazing at an area, it can determine the gaze point's location on the display screen based on the user's first eye movement data, providing a basis for subsequently displaying local perspective areas. In this way, the user does not need to manually operate the XR's physical knobs, thus avoiding operational fatigue.

[0050] For example, an XR device can determine the spatial location information of the gaze point based on the first eye movement data, and then perform coordinate transformation on the spatial location information of the gaze point to obtain the position of the gaze point on the display screen.

[0051] In S130, the target gesture can be a gesture associated with parameters such as the size and shape of the local perspective area. That is, the user can control the size and shape of the local perspective area through the target gesture to meet the user's needs. The target gesture can be a static gesture or a dynamic gesture.

[0052] Different target gestures can correspond to different local perspective areas. That is, the local perspective areas in the embodiments of this application can be updated as the target gesture changes, thus improving the flexibility of the local perspective areas.

[0053] For example, for a target hand gesture where the palm bends inward, with the thumb and the other four fingers facing each other, resembling a "C", the corresponding local perspective area will be circular, elliptical, or semi-circular. Similarly, for a target hand gesture where the left and right thumbs touch, and the left and right index fingers touch, resembling a "triangle", the corresponding local perspective area will be triangular.

[0054] The user's gestures can be obtained by recognizing the user's hand image; the specific recognition process is not limited in this application embodiment. The hand image can be captured by a camera, which can be a camera on an XR device or exist independently of the XR device and communicate with it. Taking the camera as an example of a camera on an XR device, in practical applications, the camera can be turned on by default or turned on when needed. In this case, the user can turn on the XR device's camera through touch, voice, or other methods. Once the camera is turned on, it can capture the user's hand image.

[0055] The area associated with the first gaze position can be an area on the display screen that includes the first gaze position, or an area that does not include the first gaze position but whose distance from the first gaze position meets a preset distance condition. The preset distance condition can be, for example, that the farthest distance from the first gaze position is not greater than a distance threshold. The size of the distance threshold can be set according to actual needs.

[0056] When the region contains the first gaze position, the first gaze position can be the center point of the region or any point in the region other than the center point.

[0057] The first perspective area is a partial perspective area on the display screen; that is, in addition to displaying the first perspective area, the display screen also displays other areas. Users can see the real world through the first perspective area, thus interacting with it, while through the other areas, they can see virtual scenes from the virtual world. In this way, interaction between the user and the real world is achieved without disrupting the user's immersion in the virtual world.

[0058] In some embodiments, the above-described S120 may include the following steps:

[0059] Determine the user's eye movement behavior based on first eye movement data;

[0060] When eye movement includes fixation, the fixation point is transformed from the spatial coordinate system to the screen coordinate system based on the spatial location information of the fixation point and the transformation relationship between the spatial coordinate system and the screen coordinate system. This yields the first fixation position of the fixation point on the display screen, where the screen coordinate system is the coordinate system in which the display screen is located.

[0061] A user's eye movement behavior can include fixation and non-fixation behaviors. As shown in Figure 2, fixation behavior refers to the user's gaze at a certain area of ​​the display screen 200. For example, the user's eye movement spatial location information can be determined based on the first eye movement data, and the user's eye movement behavior can be determined based on the eye movement spatial location information.

[0062] For example, if the duration for which the user's eye-tracking spatial position information remains unchanged is not less than a duration threshold, the user's eye-tracking behavior can be determined as fixation; if the duration for which the user's eye-tracking spatial position information remains unchanged is less than the duration threshold, the user's eye-tracking behavior can be determined as non-fixation. The duration threshold can be set according to actual needs, and the unit is milliseconds (ms).

[0063] The aforementioned spatial coordinate system is the coordinate system containing the spatial location information of the gaze point, while the screen coordinate system is the coordinate system containing the display screen. The transformation relationship between the spatial coordinate system and the screen coordinate system can be determined during eye-tracking calibration. Based on the transformation relationship between the spatial coordinate system and the screen coordinate system, and combined with the position information of the gaze point in one coordinate system, the position information of the gaze point in the other coordinate system can be obtained. For example, in this embodiment, based on the spatial location information of the gaze point and combined with the above transformation relationship, the position of the gaze point in the screen coordinate system can be obtained, that is, the first gaze position A of the gaze point on the display screen 200.

[0064] This application embodiment can determine whether a user is gazing based on the user's eye movement data. When the user is gazing, the gazing point in the spatial coordinate system is transformed into the screen coordinate system based on the spatial location information of the gazing point and the transformation relationship between the spatial coordinate system and the screen coordinate system. This provides a basis for the subsequent display of the local perspective area. Moreover, because the user's gazing area can change dynamically, the final displayed local perspective area can also change dynamically, thus improving the flexibility of the local perspective area.

[0065] Figure 3 is a flowchart of another display method provided by an embodiment of this application. The difference between Figure 3 and Figure 1 is that S130 in Figure 1 can be refined into S310-S320 in Figure 3.

[0066] S310, in response to a target gesture, displays the movement trajectory of the hand on the display screen.

[0067] For example, the target gesture can be a gesture that allows the user to draw a certain area on the display screen, such as including but not limited to tapping gestures, swiping gestures, and knuckle selection gestures. The user does not need to gaze at the previously determined first gaze position while drawing the area based on the target gesture.

[0068] After detecting the user's target gesture, the XR device can acquire and display the hand's movement trajectory. In some embodiments, as shown in FIG4, the XR device can display a movement trajectory 401 on the display screen 200 in response to the user's tap gesture. The movement trajectory 401 shown in FIG4 can form a closed area, and the closed area does not include the first gaze position A. In actual applications, the closed area can also include the first gaze position A, and the movement trajectory 401 can also form a non-closed area.

[0069] In some embodiments, when displaying the movement trajectory 401, the XR device may also display the drawing order of the movement trajectory 401. For example, in Figure 4, if the user is moving in a clockwise direction, the XR device may display an arrow representing the clockwise direction.

[0070] S320. Based on the initial region parameters of the region formed by the movement trajectory, display a first perspective region centered on the first gaze position on the display screen.

[0071] The initial region parameters are the region parameters of the area formed by the movement trajectory, used to characterize the regional features of the area. For example, the initial region parameters may include, but are not limited to, parameters such as the area and shape of the region.

[0072] Understandably, when the region formed by the movement trajectory is a closed region, the initial region parameter can be the region parameter of that closed region. When the region formed by the movement trajectory is a non-closed region, the initial region parameter can be the region parameter corresponding to the circumscribed shape of that non-closed region. Here, the circumscribed shape can be the smallest bounding rectangle, circle, etc., of the non-closed region, or it can be a shape obtained by expanding the smallest bounding shape according to a certain step size. The step size can be set according to actual needs.

[0073] Once the region parameters of the area formed by the movement trajectory are determined, the first perspective region corresponding to the region parameters can be displayed with the first gaze position determined above as the center.

[0074] For example, a first perspective area with the same regional parameters as the hand-drawn area can be displayed, that is, the shape, area and other parameters of the first perspective area are the same as the shape, area and other parameters of the user's hand-drawn area.

[0075] For example, a first perspective region that has a mapping relationship with the region parameters of the hand-drawn region can also be displayed. This mapping relationship can be used to adjust parameters such as the shape and area of ​​the hand-drawn region. For example, the first perspective region has a more regular shape and a larger area than the hand-drawn region.

[0076] This application embodiment can display the movement trajectory of the hand based on the user's target gesture, making it easier for the user to understand the information of the hand-drawn graphic. At the same time, it can also redraw and display the first perspective area with the first gaze position as the center based on the area parameters of the hand-drawn graphic. That is, this application embodiment can dynamically update the first perspective area according to the user's gaze position and target gesture, improving the flexibility of the first perspective area.

[0077] Understandably, an XR device can only display a partially transparent area when its partial perspective feature is enabled, allowing users to interact with the real world through that area. Therefore, to enhance the user's interaction with the real world, the XR device's partial perspective feature must be enabled first, followed by displaying the user's movement trajectory, and then revealing the partially transparent area.

[0078] Based on this, in some embodiments, the target gesture may include a first gesture and a second gesture. The first gesture is used to enable the partial perspective function of the XR device. For example, the first gesture may include a pinch gesture, that is, the tips of the thumb and index finger touch, and the other fingers can be spread out, similar to the "OK" gesture. Of course, the first gesture may also be other gestures.

[0079] The second gesture is used to draw a certain area on the display screen. For example, the second gesture can be a tap gesture, a swipe gesture, a knuckle selection gesture, etc.

[0080] For example, the above S310 may include the following steps:

[0081] In response to the first gesture, and if the first gaze position remains unchanged within a preset duration, the partial perspective function of the augmented reality device is activated;

[0082] In response to the second gesture, the movement trajectory corresponding to the second gesture is displayed on the display screen.

[0083] The fact that the first gaze position remains unchanged for a preset duration means that the user is constantly looking at the first gaze position.

[0084] For example, the partial perspective function of the XR device is turned off by default. When the XR device captures the user's first gesture and the user is still looking at the first gaze position, it indicates that the user has a need to interact with the real world. At this time, the XR device can turn on the partial perspective function to provide a basis for the subsequent display of the partial perspective area.

[0085] If the first gesture is not captured, or if the user's first gesture is captured but the user does not look at the first gaze position, the partial perspective function of the XR device will not be activated, thus avoiding user misoperation.

[0086] For example, as shown in Figure 5, when the XR device captures the user's pinch gesture 501 and the user is still looking at the first gaze position A, the partial perspective function can be activated.

[0087] Once the partial perspective function of the XR device is enabled, the XR device can display the movement trajectory corresponding to the second gesture on the display screen based on the captured second gesture. The display process can be referred to in the above embodiment.

[0088] In this embodiment, the XR partial perspective function is activated only when the user's first gesture is captured and the user keeps looking at the same position during the capture of the first gesture. After the partial perspective function is activated, the corresponding movement trajectory is displayed based on the captured second gesture. This can effectively avoid user misoperation and meet the user's need to interact with the real world.

[0089] Taking the initial region parameters including at least one of the following: the shape and area of ​​the region, as an example, the above S320 may include the following steps:

[0090] When the user's third gesture is captured, the target region parameters of the first perspective region are determined based on the initial region parameters and the reference mapping relationship. The reference mapping relationship is used to characterize the mapping relationship between the initial region parameters and the target region parameters.

[0091] Based on the target area parameters, the first perspective area is displayed on the display screen with the first gaze position as the center.

[0092] The third gesture is used to confirm the hand-drawn area. That is, when the user confirms that the hand-drawn area is completed, the third gesture can be executed. The embodiments of this application do not specifically limit the third gesture. For example, the third gesture can be a pinching gesture similar to "OK".

[0093] For example, when the XR device detects a third gesture, the first gaze position can be de-displayed.

[0094] The mapping relationship between the initial region parameters and the target region parameters can be set according to actual needs. For example, this mapping relationship can be a mapping from an irregular shape to a regular shape. That is, the initial region parameter is an irregular shape, and the target region parameter can be a regular shape. The shape types of the two can be the same or different. For example, the initial region parameter is an irregular quadrilateral, and the target region parameter can be a regular circle, or the initial region parameter is an irregular rectangle, and the target region parameter is a regular rectangle, and so on.

[0095] For example, the mapping relationship can also be a multiple relationship between areas, that is, the area corresponding to the target area parameter can be enlarged or reduced relative to the area corresponding to the initial area parameter, and the enlargement or reduction factor can be set according to actual needs.

[0096] Taking the hand-drawn area shown in Figure 4 as an example, assuming that the mapping relationship includes irregular shapes to regular shapes and the area is magnified by 1.5 times, then according to the mapping relationship, the first perspective area 601 with a regular shape and an area magnified by 1.5 times as shown in Figure 6 can be displayed.

[0097] For example, as shown in Figure 7, when the XR device detects a confirmation gesture 701, the first gaze position A can be de-displayed, allowing the user to interact with the real world through the first perspective area 601. For instance, through the first perspective area 601, the user can locate items such as a water cup or a mobile phone.

[0098] When the user's confirmation gesture is captured, the embodiments of this application can obtain the target area parameters of the first perspective area based on the area parameters of the hand-drawn area and the mapping relationship. Then, based on the target area parameters, a first perspective area with a regular shape and appropriate area is obtained with the first gaze position as the center, thus satisfying the user's local perspective needs.

[0099] In some embodiments, the display method may further include the following steps:

[0100] Obtain the user's second eye movement data;

[0101] Based on the second eye movement data, the second fixation position corresponding to the user's fixation point on the display screen is determined. The second fixation position is different from the first fixation position.

[0102] If the user's fourth gesture is not captured, the second perspective area is displayed on the display screen in the area associated with the second gaze position, based on the target area parameters of the first perspective area.

[0103] Second eye-tracking data can be the eye movement data of a user looking at another area. That is, while wearing an XR device, the user can adjust the area of ​​focus as needed, thereby dynamically adjusting the position of the area of ​​focus without the need for the user to operate physical buttons, reducing user fatigue.

[0104] The second gaze position is the location of the user's gaze point on the display screen, determined based on the second eye-tracking data. The process for determining the second gaze position can be found in the same manner as the first gaze position, and will not be repeated here for the sake of brevity. For example, as shown in Figure 8, the second gaze position B is a different position on the display screen than the first gaze position A.

[0105] The fourth gesture can be a gesture that matches the second gaze position B, used to determine a new local perspective area. For example, the fourth gesture may include a tap gesture, a swipe gesture, a knuckle selection gesture, etc.

[0106] For example, if a user is detected looking at a second gaze position B, but a fourth gesture is not captured, a second perspective region can be displayed in the area associated with the second gaze position B based on the region parameters of the previously displayed local perspective region. For instance, if the previously displayed local perspective region was a first perspective region, the second perspective region can be displayed centered on the second gaze position B based on the target region parameters of the first perspective region. That is, as shown in Figure 8, the shape and area of ​​the second perspective region 801 are the same as those of the first perspective region 601.

[0107] In this embodiment, when the user's gaze point changes but no gesture matching the new gaze point is captured, a new perspective area can be automatically displayed in the area associated with the new gaze point based on the area parameters of the previously displayed local perspective area. In other words, this embodiment can dynamically update the local perspective area according to the change in the user's intention (change in gaze point), thereby reducing the user's operations and improving the flexibility of the local perspective area.

[0108] In some embodiments, after "determining the second gaze position corresponding to the user's gaze point on the display screen based on the second eye-tracking data", the display method may further include the following steps:

[0109] If a user’s fourth gesture is captured and the fourth gesture is different from the target gesture, a third perspective area corresponding to the fourth gesture is displayed in the area associated with the second gaze position on the display screen.

[0110] The fourth gesture is different from the target gesture, which means that the hand-drawn area corresponding to the fourth gesture has different area parameters than the hand-drawn area corresponding to the target gesture. Taking the area parameters including shape and area as an example, the hand-drawn area corresponding to the fourth gesture is different from the hand-drawn area corresponding to the target gesture in at least one of the shapes and areas.

[0111] For example, when the user's gaze point changes and a fourth gesture of the user is captured, and the fourth gesture is different from the target gesture, the XR device can update the local perspective area according to the new gaze point and the new gesture to obtain a third perspective area.

[0112] For example, based on the area parameters of the area corresponding to the fourth gesture, a third perspective area can be displayed with the second gaze position as the center.

[0113] For example, if the fourth gesture is the same as the target gesture, the fourth perspective area can be displayed with the second gaze position as the center, based on the target area parameters of the first perspective area.

[0114] The embodiments of this application can dynamically update the local perspective area according to changes in the gaze point and gestures, thereby improving the flexibility of the local perspective area and meeting the personalized needs of users.

[0115] The display method of this application embodiment will be described below with reference to FIG9.

[0116] S1. Enable the eye tracking and gesture tracking functions of the XR device, that is, enable the eye tracking sensor and camera of the XR device.

[0117] S2. Acquire the user's eye movement data through an eye-tracking sensor.

[0118] S3. Determine if the user is gazing based on eye movement data.

[0119] S4. Transform the gaze point from the spatial coordinate system to the screen coordinate system to obtain the gaze position of the gaze point on the display screen.

[0120] S5. Capture the user's gestures and recognize the captured gestures.

[0121] S6. When the user performs a pinch gesture, enable the partial perspective function of the XR device.

[0122] S7. When a user performs a tap gesture, determine the area parameters of the hand-drawn area.

[0123] S8. Based on the region parameters, display the local perspective region centered on the gaze position in S4.

[0124] S9. The user gazes at another area and obtains a new gaze position.

[0125] S10. Whether the user's gesture corresponding to the new gaze position is captured. If yes, execute S11; otherwise, execute S12.

[0126] S11. Using the new gaze position as the center, generate and display a local perspective region with the same region parameters as the previous local perspective region.

[0127] S12. Using the new gaze position as the center, generate and display the local perspective region corresponding to the region parameters corresponding to the new gesture.

[0128] This application's embodiments display a local perspective area based on eye-tracking and gestures, enabling user interaction with the real world without affecting user immersion or fatigue. Furthermore, the local perspective area can be flexibly updated according to eye-tracking position and / or gestures, meeting users' personalized needs.

[0129] It should be noted that the display method provided in this application embodiment can be executed by a display device or a processing module within the display device for executing the display method. This application embodiment uses the execution of the display method by a display device as an example to illustrate the display device provided in this application embodiment.

[0130] Figure 10 is a schematic diagram of the structure of a display device provided in an embodiment of this application.

[0131] As shown in Figure 10, the display device 1000 may include:

[0132] Module 1001 is used to acquire the user's first eye movement data;

[0133] The determining module 1002 is used to determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device based on the first eye movement data;

[0134] The display module 1003 is used to display a first perspective area corresponding to the target gesture in an area associated with the first gaze position on the display screen when the user's target gesture is captured. The first perspective area is used to display the real scene.

[0135] In this embodiment, the user's first eye-tracking data is acquired; based on the first eye-tracking data, a first gaze position corresponding to the user's gaze point on the extended reality device's display screen is determined; when the user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen, and the first perspective area is used to display the real scene. That is, this embodiment can display a partial perspective area based on the user's gaze position and target gesture, displaying the real scene through the partial perspective area and the virtual scene through other areas. This does not affect the user's immersion and eliminates the need for the user to operate the physical buttons of the extended reality device, thus preventing user fatigue. Furthermore, since the user's eye-tracking data changes, the gaze position determined based on the eye-tracking data also changes accordingly; that is, the partial perspective area can change dynamically, allowing for flexible updates to the partial perspective area.

[0136] In some possible implementations of the embodiments of this application, the determining module 1002 is specifically used for:

[0137] Determine the user's eye movement behavior based on first eye movement data;

[0138] When eye movement includes fixation, the fixation point is transformed from the spatial coordinate system to the screen coordinate system based on the spatial location information of the fixation point and the transformation relationship between the spatial coordinate system and the screen coordinate system. This yields the first fixation position of the fixation point on the display screen, where the screen coordinate system is the coordinate system in which the display screen is located.

[0139] In some possible implementations of the embodiments of this application, the display module 1003 is specifically used for:

[0140] In response to a target gesture, the movement trajectory of the hand is displayed on the screen;

[0141] Based on the initial region parameters of the area formed by the movement trajectory, a first perspective region centered on the first gaze position is displayed on the display screen.

[0142] In some possible implementations of the embodiments of this application, the target gesture includes a first gesture and a second gesture;

[0143] Display module 1003 is specifically used for:

[0144] In response to the first gesture, and if the first gaze position remains unchanged within a preset duration, the partial perspective function of the augmented reality device is activated;

[0145] In response to the second gesture, the movement trajectory corresponding to the second gesture is displayed on the display screen.

[0146] In some possible implementations of the embodiments of this application, the initial region parameters include at least one of the following: the shape and area of ​​the region;

[0147] The determining module 1002 is further configured to, upon capturing a user's third gesture, determine the target region parameters of the first perspective region based on the initial region parameters and the reference mapping relationship, wherein the reference mapping relationship is used to characterize the mapping relationship between the initial region parameters and the target region parameters;

[0148] Display module 1003 is specifically used for:

[0149] Based on the target area parameters, the first perspective area is displayed on the display screen with the first gaze position as the center.

[0150] In some possible implementations of the embodiments of this application, the acquisition module 1001 is further configured to acquire the user's second eye-tracking data;

[0151] The determining module 1002 is also used to determine the second gaze position corresponding to the user's gaze point on the display screen based on the second eye movement data, wherein the second gaze position is different from the first gaze position;

[0152] The display module 1003 is also used to display a second perspective region in the area associated with the second gaze position on the display screen, based on the target region parameters of the first local perspective region, in the absence of capturing the user's fourth gesture.

[0153] In some possible implementations of the embodiments of this application, the display module 1003 is further configured to, after the determining module 1002 determines the second gaze position corresponding to the user's gaze point on the display screen based on the second eye-tracking data, and if the user's fourth gesture is captured and the fourth gesture is different from the target gesture, display a third perspective area corresponding to the fourth gesture in the area associated with the second gaze position on the display screen.

[0154] This application's embodiments display a local perspective area based on eye-tracking and gestures, enabling user interaction with the real world without affecting user immersion or fatigue. Furthermore, the local perspective area can be flexibly updated according to eye-tracking position and / or gestures, meeting users' personalized needs.

[0155] The display device in this application embodiment can be a device or a component in an electronic device, such as an integrated circuit or a chip. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0156] The electronic device in this application embodiment can be a terminal with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0157] The display device provided in this application embodiment can implement each process in the display method embodiments of Figures 1 to 9 and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0158] As shown in Figure 11, this application embodiment also provides an electronic device 1100, including a processor 1101 and a memory 1102. The memory 1102 stores programs or instructions that can run on the processor 1101. When the program or instructions are executed by the processor 1101, they implement the various steps of the above-described display method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0159] It should be noted that the electronic devices in the embodiments of this application include the mobile terminals and non-mobile terminals mentioned above.

[0160] Figure 12 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0161] The electronic device 1200 includes, but is not limited to, components such as: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.

[0162] Those skilled in the art will understand that the electronic device 1200 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 1210 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The structure of the electronic device 1200 shown in Figure 12 does not constitute a limitation on the electronic device 1200. The electronic device 1200 may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0163] The processor 1210 is used to acquire the user's first eye movement data; and based on the first eye movement data, to determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device.

[0164] Display unit 1206 is used to display a first perspective area corresponding to the target gesture in an area associated with a first gaze position on the display screen when the user's target gesture is captured. The first perspective area is used to display the real scene.

[0165] In this embodiment, the user's first eye-tracking data is acquired; based on the first eye-tracking data, a first gaze position corresponding to the user's gaze point on the extended reality device's display screen is determined; when the user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen, and the first perspective area is used to display the real scene. That is, this embodiment can display a partial perspective area based on the user's gaze position and target gesture, displaying the real scene through the partial perspective area and the virtual scene through other areas. This does not affect the user's immersion and eliminates the need for the user to operate the physical buttons of the extended reality device, thus preventing user fatigue. Furthermore, since the user's eye-tracking data changes, the gaze position determined based on the eye-tracking data also changes accordingly; that is, the partial perspective area can change dynamically, allowing for flexible updates to the partial perspective area.

[0166] In some possible implementations of embodiments of this application, the processor 1210 is specifically used for:

[0167] Determine the user's eye movement behavior based on first eye movement data;

[0168] When eye movement includes fixation, the fixation point is transformed from the spatial coordinate system to the screen coordinate system based on the spatial location information of the fixation point and the transformation relationship between the spatial coordinate system and the screen coordinate system. This yields the first fixation position of the fixation point on the display screen, where the screen coordinate system is the coordinate system in which the display screen is located.

[0169] In some possible implementations of the embodiments of this application, the display unit 1206 is specifically used for:

[0170] In response to a target gesture, the movement trajectory of the hand is displayed on the screen;

[0171] Based on the initial region parameters of the area formed by the movement trajectory, a first perspective region centered on the first gaze position is displayed on the display screen.

[0172] In some possible implementations of the embodiments of this application, the target gesture includes a first gesture and a second gesture;

[0173] Display unit 1206 is specifically used for:

[0174] In response to the first gesture, and if the first gaze position remains unchanged within a preset duration, the partial perspective function of the augmented reality device is activated;

[0175] In response to the second gesture, the movement trajectory corresponding to the second gesture is displayed on the display screen.

[0176] In some possible implementations of the embodiments of this application, the initial region parameters include at least one of the following: the shape and area of ​​the region;

[0177] Processor 1210, specifically used for:

[0178] When the user's third gesture is captured, the target region parameters of the first perspective region are determined based on the initial region parameters and the reference mapping relationship. The reference mapping relationship is used to characterize the mapping relationship between the initial region parameters and the target region parameters.

[0179] Display unit 1206 is specifically used for:

[0180] Based on the target area parameters, the first perspective area is displayed on the display screen with the first gaze position as the center.

[0181] In some possible implementations of embodiments of this application, the processor 1210 is specifically used for:

[0182] Obtain the user's second eye movement data;

[0183] Based on the second eye movement data, the second fixation position corresponding to the user's fixation point on the display screen is determined. The second fixation position is different from the first fixation position.

[0184] Display unit 1206 is specifically used for:

[0185] If the user's fourth gesture is not captured, the second perspective area is displayed on the display screen in the area associated with the second gaze position, based on the target area parameters of the first perspective area.

[0186] In some possible implementations of the embodiments of this application, the display unit 1206 is further configured to display a third perspective area corresponding to the fourth gesture in the area associated with the second gaze position on the display screen after the processor 1210 determines the second gaze position corresponding to the user's gaze point on the display screen based on the second eye-tracking data, and if the user's fourth gesture is captured and the fourth gesture is different from the target gesture.

[0187] This application's embodiments display a local perspective area based on eye-tracking and gestures, enabling user interaction with the real world without affecting user immersion or fatigue. Furthermore, the local perspective area can be flexibly updated according to eye-tracking position and / or gestures, meeting users' personalized needs.

[0188] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processing unit (GPU) 12041 and a microphone 12042. The GPU 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1207 includes a touch panel 12071 and at least one of other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.

[0189] The memory 1209 can be used to store software programs and various data. The memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1209 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0190] Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.

[0191] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described display method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0192] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0193] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described display method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0194] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0195] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes shown in the above-described method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be described again here.

[0196] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0198] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A display method applied to an extended reality device, the method comprising: Obtain the user's first eye movement data; Based on the first eye-tracking data, determine the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device; When a user's target gesture is captured, a first perspective area corresponding to the target gesture is displayed in the area associated with the first gaze position on the display screen. The first perspective area is used to display the real scene.

2. The method according to claim 1, wherein, Determining the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device based on the first eye-tracking data includes: The user's eye movement behavior is determined based on the first eye movement data; When the eye movement behavior includes gaze behavior, the gaze point is transformed from the spatial coordinate system to the screen coordinate system based on the spatial location information of the gaze point and the transformation relationship between the spatial coordinate system and the screen coordinate system, so as to obtain the first gaze position of the gaze point on the display screen, where the screen coordinate system is the coordinate system in which the display screen is located.

3. The method according to claim 1, wherein, The step of displaying a first perspective area corresponding to the target gesture in the area associated with the first gaze position on the display screen when the user's target gesture is captured includes: In response to the target gesture, the movement trajectory of the hand is displayed on the display screen; Based on the initial region parameters of the area formed by the movement trajectory, a first perspective region centered on the first gaze position is displayed on the display screen.

4. The method according to claim 3, wherein, The target gesture includes a first gesture and a second gesture; The step of displaying the hand movement trajectory on the display screen in response to the target gesture includes: In response to the first gesture, and if the first gaze position remains unchanged within a preset time, the partial perspective function of the extended reality device is activated; In response to the second gesture, the movement trajectory corresponding to the second gesture is displayed on the display screen.

5. The method according to claim 3, wherein, The initial region parameters include at least one of the following: the shape and area of ​​the region; The step of displaying a first perspective region centered on the gaze position on the display screen based on the initial region parameters of the region formed by the movement trajectory includes: Upon capturing the user's third gesture, the target region parameter of the first perspective region is determined based on the initial region parameter and the reference mapping relationship, wherein the reference mapping relationship is used to characterize the mapping relationship between the initial region parameter and the target region parameter; Based on the target area parameters, the first perspective area is displayed on the display screen with the first gaze position as the center.

6. The method according to any one of claims 1-5, further comprising: Acquire the user's second eye movement data; Based on the second eye-tracking data, a second gaze position corresponding to the user's gaze point on the display screen is determined, and the second gaze position is different from the first gaze position; If the user's fourth gesture is not captured, a second perspective region is displayed on the display screen in the area associated with the second gaze position, based on the target region parameters of the first perspective region.

7. The method according to claim 6, wherein after determining the second gaze position corresponding to the user's gaze point on the display screen based on the second eye-tracking data, the method further comprises: If the user's fourth gesture is captured and the fourth gesture is different from the target gesture, a third perspective area corresponding to the fourth gesture is displayed in the area associated with the second gaze position on the display screen.

8. A display device for use in an augmented reality device, the device comprising: The acquisition module is used to acquire the user's first eye movement data; The determining module is used to determine, based on the first eye-tracking data, the first gaze position corresponding to the user's gaze point on the display screen of the extended reality device; The display module is configured to, upon capturing a user's target gesture, display a first perspective area corresponding to the target gesture in an area associated with the first gaze position on the display screen, wherein the first perspective area is used to display a real scene.

9. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as claimed in any one of claims 1 to 7.

10. A readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as claimed in any one of claims 1 to 7.

11. A computer program product stored in a storage medium, the program product being executed by at least one processor to implement the steps of the method as claimed in any one of claims 1 to 7.

12. A chip comprising a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method as claimed in any one of claims 1 to 7.

13. An electronic device configured to perform the steps of the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for displaying real object in head-mounted display and head-mounted display thereof

    CN110275619A

  • Local perspective method and device of virtual reality equipment and virtual reality equipment

    CN112462937A

  • Arrangement of virtual objects

    CN116917850A

  • Target area selection method and device, equipment and storage medium

    CN118131970A

  • Space calibration method and device, equipment and storage medium

    CN118628570A