Interaction method and apparatus, and device, storage medium and program product
By using eye tracking and gesture interaction through smart glasses, the images of objects in virtual reality are automatically adjusted to the optimal position and perspective, solving the problem that users need to physically move to view images of objects in virtual reality, and achieving a low-cost, highly intelligent interactive experience.
Patent Information
- Application Number
- PCT/CN2025/103956
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2025-06-26
- Publication Date
- 2026-01-15
AI Technical Summary
In virtual reality scenarios, users need to physically move to view images of objects, which increases fatigue, increases interaction costs, and makes it difficult to guarantee the best viewing angle, affecting the user experience and the diversity and intelligence of interaction methods.
Through eye tracking and gesture interaction using smart glasses, interactive prompts are activated in response to the duration of eye contact. Target distances are generated based on the characteristics of the object image, and the display area is located and adjusted to move the object image to the optimal position and viewing angle.
It reduces interaction costs, increases the diversity and intelligence of interaction methods, enhances user experience, and reduces the physical and cognitive burden on users.
Smart Images

Figure CN2025103956_15012026_PF_FP_ABST
Abstract
Description
Interaction methods, devices, equipment, storage media, and program products
[0001] This application claims priority to Chinese Patent Application No. 202410939993.2, filed on July 12, 2024, the contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of artificial intelligence technology, and more specifically, to an interaction method, apparatus, device, storage medium, and program product. Background Technology
[0003] With the rapid development of artificial intelligence, virtual reality technology is gradually being applied to everyday life. Users can interact with images of objects in virtual reality scenes by wearing smart glasses to select items.
[0004] In realizing the present invention, the inventors discovered at least the following problems in the related technology: In virtual reality scenarios, if a user wants to view an image of an object, the user generally needs to walk in front of the image to view it. This not only increases user fatigue but also places requirements on the application scenarios of virtual reality. As a result, current interaction methods not only suffer from insufficient diversity and low intelligence, but also reduce the user experience and increase interaction costs. Summary of the Invention
[0005] In view of the above, this disclosure provides an interaction method, apparatus, device, storage medium, and program product.
[0006] One aspect of this disclosure provides an interaction method, comprising: activating an interactive cue point on an object image in a virtual space in response to a target object's gaze lingering on the object image for a duration exceeding a predetermined duration; generating a target distance between the object image and the object image based on attribute characteristics of the object image in response to an interaction operation between the target object and the interactive cue point; locating a target position in the virtual space based on the target distance; moving the object image to the target position; and adjusting the target display area of the object image.
[0007] According to an embodiment of this disclosure, the attribute features of the item image include the size of the item image in multiple directions; generating a target distance between the item image and the target object based on the attribute features of the item image includes: generating the volume of the item image based on the size of the item image in multiple directions; matching the volume of the item image with the volume in the target distance generation strategy to obtain a target volume; and obtaining the target distance based on the distance in the target distance generation strategy that has a mapping relationship with the target volume.
[0008] According to embodiments of this disclosure, the target distance generation strategy is constructed as follows: a field of view of the target object is constructed based on the object's line of sight; an orthographic projection area of the object image within the field of view is generated based on the object image's dimensions in multiple directions, wherein the orthographic projection area is mapped to the volume of the object image; an orthographic projection area ratio is generated based on the ratio between the orthographic projection area and the area of the field of view; and the target distance generation strategy is constructed based on the orthographic projection area ratio within a predetermined range.
[0009] According to embodiments of this disclosure, the method further includes: calibrating an eye-tracking component in the smart glasses in response to the target object activating the smart glasses, and calibrating an interaction operation recognition component in the smart glasses; responding to the gaze of the target object using the calibrated eye-tracking component, and responding to the interaction operation of the target object using the calibrated interaction operation recognition component.
[0010] According to embodiments of this disclosure, the above-mentioned calibration of the eye-tracking component in the smart glasses includes: invoking multiple gaze reference points of the eye-tracking component; and calibrating the eye-tracking component in response to the interaction between the target object and the gaze reference points.
[0011] According to an embodiment of this disclosure, adjusting the target display area of the item image includes: locating a display center point on the target display area; determining a line-of-sight concentration point located directly in front of the target object based on the line of sight of the target object; and adjusting the display center point based on the line-of-sight concentration point so that the target display area is located directly in front of the line of sight of the target object.
[0012] According to an embodiment of this disclosure, the above-mentioned item image includes multiple display areas, and the above-mentioned target display area is obtained by: obtaining the number of image features included in each of the multiple display areas; sorting each of the above-mentioned display areas according to the number of image features to obtain a sorting result; and obtaining the above-mentioned target display area based on the display area with the most image features in the sorting result.
[0013] According to embodiments of this disclosure, the method further includes: determining candidate display areas based on the sorting results; and displaying the candidate display areas in response to the interaction of the target object with the display areas.
[0014] Another aspect of this disclosure provides an interactive device, comprising: an activation module, configured to activate an interactive prompt point on the object image in response to a target object's gaze lingering on the object image in a virtual space for a duration exceeding a predetermined time; a first generation module, configured to generate a target distance between the object image and the target object based on attribute characteristics of the object image in response to the interaction between the target object and the interactive prompt point; a positioning module, configured to locate a target position in the virtual space based on the target distance; and an adjustment module, configured to move the object image to the target position and adjust the target display area of the object image.
[0015] Another aspect of this disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described above.
[0016] Another aspect of this disclosure provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described above.
[0017] Another aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method described above. Attached Figure Description
[0018] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1A schematically illustrates an exemplary system architecture to which interactive methods and apparatus can be applied according to embodiments of the present disclosure;
[0020] Figure 1B schematically illustrates an application scenario of the interaction method according to an embodiment of the present disclosure;
[0021] Figure 2 schematically illustrates a flowchart of an interaction method according to an embodiment of the present disclosure;
[0022] Figure 3 schematically illustrates the architecture of an interactive system according to an embodiment of the present disclosure;
[0023] Figure 4 schematically illustrates a block diagram of an interactive device according to an embodiment of the present disclosure; and
[0024] Figure 5 schematically illustrates a block diagram of an electronic device suitable for implementing an interaction method according to an embodiment of the present disclosure. Detailed Implementation
[0025] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0028] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0029] In the embodiments of this disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security. In the embodiments of this disclosure, user authorization or consent has been obtained before acquiring or collecting user personal information.
[0030] In virtual reality (VR) scenarios, if a user wants to view an image of an object, they need to walk to it, which is not only physically demanding but also has high interaction costs. If the object is infinitely far from the user, viewing it becomes virtually impossible. For example, in current VR technology, users need to physically move to approach or navigate around obstacles to carefully observe the image, which not only increases user fatigue but is also impractical in certain scenarios. For instance, if the simulated real-world scene is large enough (e.g., a large warehouse), then the user needs to use a space that matches the size of the virtual warehouse to deploy the virtual warehouse. Overall, current interaction methods suffer from the following problems.
[0031] The interaction is laborious: the technology requires users to make a lot of physical movements, which increases user fatigue, especially during long-term use, and may cause discomfort.
[0032] High barrier to entry for interaction: For users unfamiliar with virtual reality devices, learning and adapting to this interaction method requires time and effort, increasing the difficulty of use for users.
[0033] Not the best viewpoint: When users view items by physically moving around, it is difficult to guarantee that they will always get the best viewpoint, which affects the user experience and the display effect of items.
[0034] In view of this, embodiments of this disclosure provide an interaction method, apparatus, device, storage medium, and program product for use when a user wants to view an image of an item in an e-commerce VR (Virtual Reality) scene, an interactive prompt point is provided on the item image. Through eye tracking and gesture interaction using smart glasses, the user looks at the item model and simultaneously clicks, triggering an interaction. The entire space and the item are then displayed in the user's field of vision at the optimal angle and position, reducing the user's interaction threshold and cost, improving the user experience, and increasing the diversity and intelligence of interaction methods. Specifically, the method includes: activating the interactive prompt point of the item image in response to the target object's gaze lingering on the item image in the virtual space for a duration exceeding a predetermined time; generating a target distance between the item image and the target object based on the attribute characteristics of the item image in response to the interaction operation between the target object and the interactive prompt point; locating the target position in the virtual space based on the target distance; moving the item image to the target position and adjusting the target display area of the item image.
[0035] Figure 1A schematically illustrates an exemplary system architecture 100 to which interactive methods and devices can be applied according to embodiments of the present disclosure. It should be noted that Figure 1A is merely an example of a system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not imply that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0036] As shown in Figure 1A, the system architecture 100 according to this embodiment may include a smart terminal 101, a network 102, a server 103, and an object image 104. The network 102 serves as a medium for providing communication links between the smart terminal 101 and the server 103, and between the server 103 and the object image 104. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.
[0037] Users can use smart terminal 101 to interact with server 103 via network 102 to view object images 104. Smart terminal 101 may include electronic devices for viewing virtual reality environments, including but not limited to smart glasses. Various communication client applications may be installed on smart terminal 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0038] Item image 104 can be an image from a virtual scene. For example, if the virtual reality scene is a large furniture warehouse, then the item image could be a chair from within that large furniture warehouse.
[0039] Server 103 can be a server that provides various services, such as a backend management server that supports the operations interacted by the user using the smart terminal 101 (for example only). The backend management server can analyze and process the received interactive operations and other data, and move the processing results (such as object images, web pages, information, or data obtained or generated based on the interactive operations) to the target location so that the user can view the object image from the best position and at the best viewing angle without having to move.
[0040] Figure 1B schematically illustrates an application scenario of the interaction method according to an embodiment of the present disclosure.
[0041] As shown in Figure 1B, users can interact with interactive prompts on an object image (using a chair as an example) by wearing smart glasses, so that the object image moves to the optimal distance position for the user.
[0042] It should be noted that the interaction method provided in this embodiment can generally be executed by server 103. Correspondingly, the interaction device provided in this embodiment can generally be located in server 103. The interaction method provided in this embodiment can also be executed by a server or server cluster that is different from server 103 and capable of communicating with smart terminal 101, item image 104, and / or server 103. Correspondingly, the interaction device provided in this embodiment can also be located in a server or server cluster that is different from server 103 and capable of communicating with smart terminal 101, item image 104, and / or server 103. Alternatively, the interaction method provided in this embodiment can also be executed by smart terminal 101, or by other smart terminals different from smart terminal 101. Correspondingly, the interaction device provided in this embodiment can also be located in smart terminal 101, or in other smart terminals different from smart terminal 101.
[0043] For example, virtual reality scenes and images of objects within those scenes can be stored either in the smart terminal 101 or on an external storage device and imported into the smart terminal 101. The smart terminal 101 can then execute the interaction methods provided in this embodiment locally, or send responses such as interaction requests to other terminal devices, servers, or server clusters, which in turn execute the interaction methods provided in this embodiment.
[0044] It should be understood that the number of smart terminals, networks, servers, and object images in Figure 1A is merely illustrative. Depending on the implementation requirements, any number of smart terminals, networks, servers, and object images can be included.
[0045] Figure 2 schematically illustrates a flowchart of an interaction method according to an embodiment of the present disclosure.
[0046] As shown in Figure 2, the method includes operations S201~S204.
[0047] In operation S201, in response to the target object's gaze lingering on the object image in the virtual space for a duration exceeding a predetermined time, the interactive cue point of the object image is activated.
[0048] In operation S202, in response to the interaction between the target object and the interactive prompt point, the target distance between the item image and the target object is generated based on the attribute characteristics of the item image.
[0049] In operation S203, the target location is determined in virtual space based on the target distance.
[0050] In operation S204, the item image is moved to the target position, and the target display area of the item image is adjusted.
[0051] Optionally, the target object may include objects using smart glasses, such as people or smart robots.
[0052] Optionally, the smart glasses can be configured with an eye-tracking component and an interaction recognition component. The eye-tracking component can track the gaze direction of the target object in real time and capture the target object's gaze point. When the target object's gaze rests on an image of an object for more than a predetermined time (e.g., 1 second), the server can recognize the image of the object as the target object image and activate the interactive prompt point associated with the target object image.
[0053] Optionally, the interactive operation recognition component can be used to correspond to the interactive operation between the target object and the interactive prompt point. For example, the gesture recognition component can recognize click gestures or the voice recognition component can recognize voice. The interactive operation recognition component can transmit the interactive operation to the server. When the server confirms the target object's intent based on the interactive operation, it triggers the display process of the item image.
[0054] Optionally, the display process of the item image may include determining the target distance (i.e., the optimal distance) between the item image and the target object, and determining the target display area (i.e., the optimal display area) of the item image.
[0055] In one embodiment, when determining the target distance, the target distance can be generated based on the attribute features of the object image. For example, the volume of the object image can be obtained based on its dimensions in multiple directions, and the target distance can be generated based on this volume.
[0056] Optionally, the target distance generated by the above operations can be used to locate the position of the object image in virtual space. It should be noted that the position of the target object in virtual space can remain unchanged; only the object image needs to be moved.
[0057] Optionally, the target display area of the item image can be adjusted before, simultaneously with, or after moving the item image to the target position, so that the target display area is aligned with the line of sight of the target object. For example, the target display area may be located directly in front of the target object, thus facilitating the target object to view the item image at the optimal distance from the best viewing angle. In one embodiment, the target display area may be the area capable of displaying the most features of the item image, such as the front display area of the item image, or it may be a display area that includes multiple surfaces of a three-dimensional item image, such as the front, top, and sides. The specific area can be adaptively adjusted according to actual needs.
[0058] According to embodiments of this disclosure, an interactive cue point on an item image in virtual space is activated in response to the target object's gaze lingering on the item image for a duration exceeding a predetermined time. In response to the target object's interaction with the cue point, a target distance is generated between the item image and the target object based on the item image's attribute characteristics. The target position is located in virtual space based on the target distance. The item image is then moved to the target position, and its target display area is adjusted. Because the interactive cue point is automatically activated based on the gaze duration during the interaction, and the target distance between the item image and the target object is generated based on the target object's interaction with the cue point, and the item image is moved towards the target object and its target display area is adjusted, the item image is displayed to the target object from the optimal position and perspective. This at least partially overcomes the problems of low interactive experience and high interaction costs in related technologies, thereby achieving the technical effects of increasing the diversity and intelligence of interaction methods, improving the interactive experience, and reducing interaction costs.
[0059] Optionally, before responding to the gaze of the target object, the following operations may also be performed: in response to the target object activating the smart glasses, calibrating the eye-tracking component in the smart glasses and calibrating the interaction operation recognition component in the smart glasses; using the calibrated eye-tracking component to respond to the gaze of the target object, and using the calibrated interaction operation recognition component to respond to the interaction operation of the target object.
[0060] Optionally, the target object activates the smart glasses, for example, by wearing the smart glasses or by pressing a button.
[0061] Optionally, in response to the target object activating the smart glasses, the eye-tracking component and the interaction recognition component in the smart glasses can be calibrated. This allows operation S201 to be performed using the calibrated eye-tracking component, and operation S202 to be performed using the calibrated interaction recognition component.
[0062] Optionally, the process of calibrating the eye-tracking component may include the following operations: invoking multiple gaze reference points of the eye-tracking component; and calibrating the eye-tracking component in response to the interaction between the target object and the gaze reference points.
[0063] Optionally, gaze reference points can be pre-set and distributed throughout the virtual space, such as at the top, bottom, or edges, to ensure comprehensive tracking of the target object's gaze. The specific distribution can be adaptively adjusted based on the actual virtual scene. Calibration of the eye-tracking components can be achieved through interaction between the target object's gaze and these gaze reference points (e.g., viewing these gaze reference points).
[0064] Optionally, the calibration of the interactive operation recognition component can refer to that of the eye-tracking component. For example, reference gestures, such as clicks and swipes, can be pre-set. The interactive operation recognition component can be calibrated through the interaction between the target object's gestures and these reference gestures (e.g., making the same gesture as the click reference gesture, making the same gesture as the swipe reference gesture, etc.). Alternatively, reference text can be pre-set. The interactive operation recognition component can be calibrated through the interaction between the target object's speech and these reference texts (e.g., reading out these reference texts, etc.).
[0065] According to embodiments of this disclosure, by calibrating the eye-tracking component and interaction recognition component of smart glasses based on a gaze reference point and a reference gesture or reference text, and then using the calibrated smart eye-tracking component and interaction recognition component to respond to the gaze and interaction of a target object, the accuracy of the response to gaze and interaction can be improved. Furthermore, since there may be differences in the gaze or interaction of different target objects, the above method can achieve different calibrations of the eye-tracking component and interaction recognition component according to different target objects, thereby improving the response efficiency and accuracy of gaze and interaction.
[0066] Optionally, after the target object clicks on the item image, the item image display process can be triggered, and the target distance between the item image and the target object can be generated based on the attribute characteristics of the item image.
[0067] In one embodiment, the attribute features of the item image may include the dimensions of the item image in multiple directions. The process of generating a target distance between the item image and the target object based on the attribute features of the item image may include the following operations: generating the volume of the item image based on its dimensions in multiple directions; matching the volume of the item image with the volume in the target distance generation strategy to obtain a target volume; and obtaining the target distance based on the distance in the target distance generation strategy that has a mapping relationship with the target volume.
[0068] Optionally, the target distance generation strategy described above can be divided into three categories based on volume. For example, for small-volume object images (within 0.125 cubic meters), the target distance can be within 0.5 meters. For medium-volume object images (0.125 to 1 cubic meter), the target distance can be within 1.5 meters. For large-volume object images (above 1 cubic meter), the target distance can be within 2 meters. By matching the volume of the object image with the volume in the target distance generation strategy, the volume that matches the volume of the object image in the target distance generation strategy is taken as the target volume (or it can be a target volume range). The target distance is obtained based on the distance mapped to this target volume. For example, if the object image volume is 0.1 cubic meters, the target distance can be within 0.5 meters of the target object.
[0069] Optionally, the target distance generation strategy described above can be constructed as follows: construct the field of view of the target object based on the target object's line of sight; generate the orthographic projection area of the object image within the field of view based on the size of the object image in multiple directions, wherein the orthographic projection area has a mapping relationship with the volume of the object image; generate the orthographic projection area ratio based on the ratio between the orthographic projection area and the area of the field of view; and construct the target distance generation strategy based on the orthographic projection area ratio within a predetermined range.
[0070] Optionally, the target distance generation strategy can be constructed based on the visual comfort of the viewer at different distances and the detail display effect of the object image. Smaller object images require users to move closer to observe them in order to see the details clearly, while larger object images require an appropriate distance to display the entire object image within the overall field of vision.
[0071] Optionally, based on the dimensions of the object image in multiple directions, the orthographic projection area of the object image within the field of view can be generated, and this orthographic projection area can be mapped to the volume of the object image. Based on the ratio between the orthographic projection area and the area of the field of view, an orthographic projection area percentage is generated, and based on the orthographic projection area percentage within a predetermined range (e.g., 20%~30%), a target distance generation strategy is constructed.
[0072] In one embodiment, the target distance is determined based on the orthographic projection area of the object image in the field of view. The distance between the object image and the user when the orthographic projection area of the object image in the field of view occupies 20%-30% of the field of view is taken as the target distance.
[0073] Optionally, the process of displaying an object image may also include adjusting the target display area of the object image. Specifically, this process may include the following operations: locating a display center point on the target display area; determining a focal point of vision directly in front of the target object based on the target object's line of sight; and adjusting the display center point based on the focal point of vision so that the target display area is located directly in front of the target object's line of sight.
[0074] Optionally, the optimal viewing angle for the target object to view the item image can be the direction directly opposite the item image, i.e., the target display area (i.e., the main display surface) where the target object sees the item image from the front.
[0075] Optionally, the item image may include multiple display areas, and the target display area described in the above operation can be obtained by: obtaining the number of image features included in each of the multiple display areas; sorting each display area according to the number of image features to obtain a sorting result; and obtaining the target display area based on the display area with the most image features in the sorting result.
[0076] Optionally, image features may include, for example, item labels, item patterns, and item descriptions. Each display area may include a different number of image features. Based on the number of image features in each display area, the display areas can be sorted in order of display sequence, for example, by the number of image features from most to least. The display area with the most image features in the sorted result is then selected as the target display area. Showing this target display area to the target object allows them to gain a more comprehensive understanding of the item image.
[0077] Optionally, when displaying the target display area, a display center point can be positioned on the target display area. Based on the target object's line of sight, determine the focal point of their gaze directly in front of the target object. Adjust the display center point according to this focal point, aligning it with the focal point of their gaze. This ensures the item image is positioned directly in front of the target object, guaranteeing that the target object can see the item image in the target display area from the front, thus ensuring the target object receives the best viewing experience.
[0078] Optionally, when displaying an image of an item, the interaction method may further include the following operations in addition to the above operations: determining candidate display areas based on the sorting results; and displaying the candidate display areas in response to the target object's interaction with the display areas.
[0079] Optionally, candidate display areas can be determined sequentially based on the number of image features in the sorting results. For example, the region with the second-highest number of image features can be designated as the first candidate display area; the region with the second-highest number of image features can be designated as the second candidate display area, and so on, resulting in multiple candidate display areas. By responding to the target object's interactive operations on the display area, such as swiping in one direction, the first candidate display area can be displayed, and swiping in another direction can display the second candidate display area, thus allowing the target object to understand more information about the object image from multiple perspectives.
[0080] According to the embodiments of this disclosure, through the above steps, it is possible to automatically move the image of an object to the optimal viewing distance and angle of the target object in a VR virtual reality environment, which significantly reduces the interaction cost and the interactive experience, while improving the diversity of interaction methods and the level of intelligence of interaction.
[0081] Figure 3 schematically illustrates the architecture of an interactive system according to an embodiment of the present disclosure.
[0082] As shown in Figure 3, the interactive system of this embodiment can be triggered by smart glasses configured on the target object. The interactive system may include an eye-tracking and gesture interaction module 301, an object image display module 302, a spatial positioning module 303, and a display module 304.
[0083] Optionally, the target subject may wear smart glasses. The interaction system activates and calibrates the eye-tracking and gesture interaction module 301 to ensure accurate response to the target subject's gaze and gestures.
[0084] The eye-tracking component of the eye-tracking and gesture interaction module 301 can respond in real time to the gaze direction and gaze point of the target object. When the target object's gaze rests on an object image for more than a preset time (e.g., 1 second), the interaction system recognizes the object image as the target object image and activates the relevant interactive prompt point.
[0085] The gesture interaction component of the eye-tracking and gesture interaction module 301 allows the target object to interact through predefined gestures (e.g., a click gesture). The gesture interaction component responds to the target object's gesture signal and transmits it to the interaction system for processing. After the transaction system confirms the target object's interaction intent, it triggers the item image display process.
[0086] The item image display process can be executed by the item image display module 302. The display may include the following operations:
[0087] The target distance (optimal viewing distance) can be calculated by the trading system based on the volume of the item image. The specific target distance generation strategy described above will not be repeated here.
[0088] The target display area is adjusted (optimal viewing angle calculation). The optimal viewing angle is typically the direction directly facing the product, i.e., the main display surface of the product that the user sees from the front. The transaction system, through spatial positioning module 303, can determine the position and orientation of the item image and the target object in virtual space. For example, it locates the center point of the target display area to determine the position and orientation of the item image in virtual space, and determines the focal point of the line of sight directly in front of the target object based on the target object's line of sight, thereby determining the position and orientation of the target object in virtual space.
[0089] The display module 304 can adjust the position and orientation of the product in the virtual space based on the calculated optimal viewing distance and angle. For example, it can move the product image to a position determined in the virtual space based on the target distance, and adjust the orientation of the product image, that is, adjust the display center point according to the focal point of the line of sight, so that the target display area of the product image is directly in front of the target object's line of sight, thereby ensuring that the target object gets the best viewing experience.
[0090] Optionally, when moving an object image, the surrounding environmental features can also be moved along with it. For example, the surrounding environment can be cropped at a predetermined size, centered on the object image. The cropped shape can be a circle, quadrilateral, or polygon, etc. The cropped surrounding environmental features and the object image are then moved together to the optimal display position in the virtual space determined based on the target distance. This allows the target object to not only understand the object image but also gain a better understanding of the selection and the surrounding environment. Understandably, the surrounding environmental features can be images of objects of a similar type to the object image, or the environment in which the object image is applicable, etc., and can be adaptively adjusted according to actual needs.
[0091] Through these adjustments, the image of the object can be presented in front of the target object at the most suitable size and angle, thereby reducing the physical and cognitive burden on the target object.
[0092] Optionally, the interactive system can also enable the viewing of different object images in smart glasses by responding to the switching operation of the target object on the object image.
[0093] Through the above process, the interactive system can adjust the orientation and position of the object image to ensure the target object receives the best viewing experience, reduce interaction costs, improve the interactive experience, increase the diversity of interaction methods, and enhance the intelligence of the interaction. This interactive system can utilize the understanding and application of human visual comfort and the display effect of product details.
[0094] It should be noted that, unless it is explicitly stated that there is a sequential order of execution between different operations, or that there is a sequential order of execution between different operations in terms of technical implementation, the execution order between multiple operations may not be significant, and multiple operations may be executed simultaneously.
[0095] Figure 4 schematically illustrates a block diagram of an interactive device according to an embodiment of the present disclosure.
[0096] As shown in Figure 4, the interactive device 400 includes an activation module 410, a first generation module 420, a positioning module 430, and an adjustment module 440.
[0097] The activation module 410 is used to activate the interactive prompt point of the item image in response to the target object's gaze lingering on the item image in the virtual space for a period of time exceeding a predetermined time.
[0098] The first generation module 420 is used to respond to the interaction between the target object and the interactive prompt point, and generate the target distance between the item image and the target object based on the attribute characteristics of the item image.
[0099] The positioning module 430 is used to locate the target position in virtual space based on the target distance.
[0100] The adjustment module 440 is used to move the item image to the target position and adjust the target display area of the item image.
[0101] According to embodiments of this disclosure, an interactive cue point on an item image in virtual space is activated in response to the target object's gaze lingering on the item image for a duration exceeding a predetermined time. In response to the target object's interaction with the cue point, a target distance is generated between the item image and the target object based on the item image's attribute characteristics. The target position is located in virtual space based on the target distance. The item image is then moved to the target position, and its target display area is adjusted. Because the interactive cue point is automatically activated based on the gaze duration during the interaction, and the target distance between the item image and the target object is generated based on the target object's interaction with the cue point, and the item image is moved towards the target object and its target display area is adjusted, the item image is displayed to the target object from the optimal position and from the optimal perspective. This at least partially overcomes the problems of low interactive experience and high interaction costs in related technologies, thereby achieving the technical effects of increasing the diversity and intelligence of interaction methods, improving the interactive experience, and reducing interaction costs.
[0102] According to embodiments of this disclosure, the first generation module may include a generation unit, a matching unit, and a mapping unit.
[0103] The generation unit is used to generate the volume of the object image based on the dimensions of the object image in multiple directions.
[0104] The matching unit is used to match the volume of the object image with the volume in the target distance generation strategy to obtain the target volume.
[0105] The mapping unit is used to obtain the target distance based on the distance that has a mapping relationship with the target volume in the target distance generation strategy.
[0106] According to embodiments of this disclosure, the interactive device may further include a first building module, a second generation module, a third generation module, and a second building module.
[0107] The first building module is used to construct the field of view of the target object based on the target object's line of sight.
[0108] The second generation module is used to generate the orthographic projection area of the object image within the field of view based on the size of the object image in multiple directions, wherein the orthographic projection area has a mapping relationship with the volume of the object image.
[0109] The third generation module is used to generate the proportion of the orthographic projection area based on the ratio between the orthographic projection area and the area of the field of view.
[0110] The second construction module is used to construct a target distance generation strategy based on the proportion of the orthographic projection area within a predetermined range.
[0111] According to embodiments of this disclosure, the interactive device may further include a calibration module and a response module.
[0112] The calibration module is used to calibrate the eye-tracking component and the interactive operation recognition component in the smart glasses in response to the target object activating the smart glasses.
[0113] The response module is used to respond to the gaze of the target object using the calibrated eye-tracking component and to respond to the interactive operation of the target object using the calibrated interactive operation recognition component.
[0114] According to embodiments of this disclosure, the calibration module may include a calling unit and a calibration unit.
[0115] The calling unit is used to call multiple gaze reference points of the eye-tracking component.
[0116] The calibration unit calibrates the eye-tracking components in response to the interaction between the target object and the gaze reference point.
[0117] According to embodiments of this disclosure, the adjustment module may include a positioning unit, a determining unit, and an adjustment unit.
[0118] The positioning unit is used to locate the center point of the display on the target display area.
[0119] The determining unit is used to determine the focal point of the line of sight directly in front of the target object, based on the line of sight of the target object.
[0120] The adjustment unit is used to adjust the display center point according to the focal point of the view, so that the target display area is located directly in front of the target object's line of sight.
[0121] According to embodiments of this disclosure, the interactive device may further include an acquisition module, a sorting module, and a result module.
[0122] The acquisition module is used to acquire the number of image features included in each of the multiple display areas.
[0123] The sorting module is used to sort each display area according to the number of image features to obtain the sorting result.
[0124] The results module is used to obtain the target display area based on the display area with the most image features in the sorting results.
[0125] According to embodiments of this disclosure, the interactive device may further include a determining module and a display module.
[0126] The determination module is used to determine the candidate display area based on the sorting results.
[0127] The display module is used to respond to the target object's interactive operations on the display area and display candidate display areas.
[0128] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0129] For example, any plurality of the activation module 410, the first generation module 420, the positioning module 430, and the adjustment module 440 may be combined into one module / unit / subunit, or any one of these modules / units / subunits may be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits may be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present disclosure, at least one of the activation module 410, the first generation module 420, the positioning module 430, and the adjustment module 440 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the activation module 410, the first generation module 420, the positioning module 430, and the adjustment module 440 may be implemented at least partially as a computer program module that can perform corresponding functions when the computer program module is run.
[0130] It should be noted that the interactive device part in the embodiments of this disclosure corresponds to the interactive method part in the embodiments of this disclosure. The specific description of the interactive device part is referred to in the interactive method part, and will not be repeated here.
[0131] Figure 5 schematically illustrates a block diagram of an electronic device suitable for implementing an interaction method according to an embodiment of the present disclosure. The electronic device shown in Figure 5 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.
[0132] As shown in FIG. 5, an electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0133] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0134] According to embodiments of this disclosure, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.
[0135] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0136] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0137] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0138] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0139] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the interactive methods provided in the embodiments of this disclosure.
[0140] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0141] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0142] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0144] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. An interaction method, comprising: In response to the target object's gaze lingering on the image of an item in the virtual space for a duration exceeding a predetermined time, an interactive prompt point on the image of the item is activated; In response to the interaction between the target object and the interactive prompt point, a target distance between the item image and the target object is generated based on the attribute features of the item image; Based on the target distance, locate the target position in the virtual space; Move the item image to the target location and adjust the target display area of the item image.
2. The method according to claim 1, wherein, The attribute features of the item image include the size of the item image in multiple directions; Based on the attribute features of the item image, a target distance is generated between the item image and the target object, including: The volume of the item image is generated based on its dimensions in multiple directions; The volume of the object image is matched with the volume in the target distance generation strategy to obtain the target volume; The target distance is obtained based on the distance that has a mapping relationship with the target volume in the target distance generation strategy.
3. The method according to claim 2, wherein, The target distance generation strategy is constructed in the following manner: Construct the field of vision of the target object based on its line of sight; Based on the dimensions of the object image in multiple directions, the orthographic projection area of the object image within the field of view is generated, wherein the orthographic projection area has a mapping relationship with the volume of the object image; The orthographic projection area ratio is generated based on the ratio between the orthographic projection area and the area of the field of view. The target distance generation strategy is constructed based on the proportion of the orthographic projection area within the predetermined range.
4. The method according to claim 1, wherein, The method further includes: In response to the target object activating the smart glasses, the eye-tracking component in the smart glasses is calibrated, and the interactive operation recognition component in the smart glasses is calibrated; The calibrated eye-tracking component responds to the gaze of the target object, and the calibrated interaction recognition component responds to the interaction of the target object.
5. The method according to claim 4, wherein, The calibration of the eye-tracking component in the smart glasses includes: Invoke multiple gaze reference points of the eye-tracking component; The eye-tracking component is calibrated in response to the interaction between the target object and the gaze reference point.
6. The method according to claim 1, wherein, The adjustment of the target display area of the item image includes: Locate the center point of the display on the target display area; Based on the line of sight of the target object, determine the focal point of the line of sight located directly in front of the target object; Adjust the display center point according to the focal point of the gaze so that the target display area is located directly in front of the target object's line of sight.
7. The method according to claim 6, wherein, The item image includes multiple display areas, and the target display area is obtained in the following way: Obtain the number of image features included in each of the multiple display areas; Based on the number of image features, each display area is sorted to obtain a sorting result; The target display area is obtained from the display area with the most image features in the sorting results.
8. The method according to claim 7, wherein, The method further includes: Based on the sorting results, candidate display areas are determined; In response to the target object's interactive operation on the display area, the candidate display area is displayed.
9. An interactive device, comprising: The activation module is used to activate the interactive prompt point of the item image in response to the target object's gaze lingering on the item image in the virtual space for a period of time exceeding a predetermined time. The first generation module is used to generate a target distance between the item image and the target object in response to the interaction operation between the target object and the interaction prompt point, based on the attribute features of the item image. The positioning module is used to locate the target position in the virtual space based on the target distance; The adjustment module is used to move the item image to the target position and adjust the target display area of the item image.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 8.
11. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Human-computer interaction method and device, storage medium and electronic equipment
CN116974373A
Interaction method, head-mounted display device, electronic device and readable storage medium
CN117453037A
Interaction method and device, equipment, storage medium and program product
CN118672403A
Smart logistics vehicle and method of controlling the same
KR1020240126314A