Methods, Systems, and Media for Object Grouping and Manipulation in Immersive Environments
By introducing handle interface elements and optional indicators in an immersive environment, combined with stacked thumbnail representation, the tedious problems of object grouping and manipulation in an immersive environment are solved, achieving more efficient space use and intuitive interactive experience.
Patent Information
- Application Number
- CN202210120150.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-05-22
AI Technical Summary
In an immersive environment, object grouping and manipulation are cumbersome and difficult, making it difficult to use space efficiently and provide intuitive interaction mechanisms.
By introducing handle interface elements and optional indicators in an immersive environment, users are allowed to interact with group video objects through natural gestures such as gripping gestures, and represent group video objects by generating stacked thumbnail representations.
It realizes more efficient use of space in an immersive environment, grouping video objects to reduce processing needs, and provides an intuitive interaction mechanism, improving the interaction efficiency between users and video objects.
Smart Images

Figure CN114582377B_ABST
Abstract
Description
[0001] Division Explanation
[0002] This application is a divisional application of Chinese Patent Application No. 201980005556.0 with an application date of May 22, 2019. Technical Field
[0003] The disclosed subject matter relates to methods, systems, and media for object grouping and manipulation in immersive environments. Background Art
[0004] Many users enjoy watching video content in immersive environments, such as virtual reality content, augmented reality content, three-dimensional content, 180-degree content, or 360-degree content, which can provide an immersive experience for viewers. For example, a virtual reality system can generate an immersive virtual reality environment for a user, where the user can interact with one or more virtual objects. In a more specific example, a device such as a virtual reality headset device or a head-mounted display device can be used to provide the immersive virtual reality environment. In another example, an augmented reality system can generate an immersive augmented reality environment for a user, where computer-generated content (e.g., one or more images) can be superimposed on the user's current view (e.g., using the camera of a mobile device).
[0005] It should be noted that users can navigate and / or interact with the immersive environment in various ways. For example, a user can use hand movements to interact with virtual objects in the immersive environment. In another example, a user can operate a controller such as a ray-based input controller to interact with virtual objects in the immersive environment by pointing at the objects and / or to select an object by pressing a button located on the controller. However, placing, organizing, clustering, manipulating, or otherwise interacting with groups of objects in the immersive environment remains a cumbersome and difficult task.
[0006] Therefore, there is a desire to provide new methods, systems, and media for object grouping and manipulation in immersive environments. Summary of the Invention
[0007] Methods, systems, and media for object grouping and manipulation in immersive environments are provided.
[0008] According to some embodiments of the disclosed subject matter, a method of interacting with immersive video content is provided, the method comprising: displaying a plurality of video objects in an immersive environment; detecting, via a first input, that a first video object has been virtually positioned above a second video object; in response to detecting that the first video object has been virtually positioned above the second video object, generating a group video object comprising the first video object and the second video object, wherein the group video object comprises a handle interface element for interacting with the group video object and an optional indicator representing the first video object and the second video object; displaying, in the immersive environment, the group video object together with the handle interface element and the optional indicator, and one or more remaining video objects, wherein the group video object replaces the first video object and the second video object within the immersive environment; and in response to detecting a selection of the optional indicator, displaying a user interface for interacting with the group video object.
[0009] In some embodiments, the immersive environment is a virtual reality environment generated in a head-mounted display device operating in a physical environment, and the handle interface element is a three-dimensional handle element that is interacted with by detecting a grasping gesture performed by a hand in the virtual reality environment.
[0010] In some embodiments, the first video object is represented by a first thumbnail representation, the second video object is represented by a second thumbnail representation, and the group video object is represented by a stacked thumbnail representation, wherein (i) the first thumbnail representation and the second thumbnail representation are automatically aligned to generate the stacked thumbnail representation, and (ii) in response to detecting that the first video object is interacted with and virtually positioned above the second video object, the first thumbnail representation is in a top position of the stacked thumbnail representation.
[0011] In some embodiments, the optional indicator indicates the number of video objects included in the group video object.
[0012] In some embodiments, the user interface includes an option for creating a playlist that includes the first video object and the second video object in the group video object, and wherein, after the group video object is selected, the first video object and the second video object are played back in the immersive environment.
[0013] In some embodiments, the user interface includes an option for rearranging the order of at least the first video object and the second video object associated with the group video object.
[0014] In some embodiments, the user interface includes an option for removing at least one of the first video object and the second video object from the group video object.
[0015] In some embodiments, the user interface includes an option to remove a group of video objects, as well as a first video object and a second video object, from the immersive environment.
[0016] In some embodiments, the method further includes: in response to detecting a specific hand interaction with a handle interface element, rendering a grid of video objects included within the group of video objects, wherein each video object in the grid of video objects is modifiable.
[0017] In some embodiments, the method further includes: in response to detecting a specific hand interaction with a handle interface element, removing the group of video objects, as well as the first video object and the second video object, from the immersive environment.
[0018] According to some embodiments of the disclosed subject matter, a system for interacting with immersive video content is provided. The system includes a memory and a hardware processor that, when executing computer-executable instructions stored in the memory, is configured to: display a plurality of video objects in an immersive environment; detect, via a first input, that a first video object has been virtually positioned above a second video object; in response to detecting that the first video object has been virtually positioned above the second video object, generate a group of video objects that includes the first video object and the second video object, wherein the group of video objects includes a handle interface element for interacting with the group of video objects and an optional indicator representing the first video object and the second video object; display, in the immersive environment, the group of video objects along with the handle interface element and the optional indicator, as well as one or more remaining video objects, wherein the group of video objects replaces the first video object and the second video object within the immersive environment; and in response to detecting a selection of the optional indicator, display a user interface for interacting with the group of video objects.
[0019] According to some embodiments of the disclosed subject matter, there is provided a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method for interacting with immersive video content, the method comprising: displaying a plurality of video objects in an immersive environment; detecting, via a first input, that a first video object has been virtually positioned above a second video object; in response to detecting that the first video object has been virtually positioned above the second video object, generating a group video object comprising the first video object and the second video object, wherein the group video object comprises a handle interface element for interacting with the group video object and an optional indicator representing the first video object and the second video object; displaying, in the immersive environment, the group video object together with the handle interface element and the optional indicator, and one or more remaining video objects, wherein the group video object replaces the first video object and the second video object within the immersive environment; and in response to detecting a selection of the optional indicator, displaying a user interface for interacting with the group video object.
[0020] According to some embodiments of the disclosed subject matter, there is provided a system for generating immersive video content, the system comprising: means for displaying a plurality of video objects in an immersive environment; means for detecting, via a first input, that a first video object has been virtually positioned above a second video object; means for generating, in response to detecting that the first video object has been virtually positioned above the second video object, a group video object comprising the first video object and the second video object, wherein the group video object comprises a handle interface element for interacting with the group video object and an optional indicator representing the first video object and the second video object; means for displaying, in the immersive environment, the group video object together with the handle interface element and the optional indicator, and one or more remaining video objects, wherein the group video object replaces the first video object and the second video object within the immersive environment; and means for displaying, in response to detecting a selection of the optional indicator, a user interface for interacting with the group video object.
[0021] The subject matter described in this specification can be implemented in particular embodiments to achieve one or more of the following advantages. By providing a handle interface element and an optional indicator, first and second video objects can be grouped in an immersive environment to more effectively utilize the space within the immersive environment, while ensuring that the user can easily and efficiently interact with the grouped video objects. Grouping the video objects by replacing the first and second video objects with a group video object provides the additional advantage of reducing the processing power required to display the video objects. The handle interface element provides an intuitive mechanism for the user to interact with the group video object using natural gestures such as grasping gestures, shaking gestures, etc. This avoids the need to provide dedicated user interface elements within the immersive environment for each potential interaction, thereby allowing the grouped video objects to be represented and interacted with in a computationally more efficient manner that more effectively utilizes the real estate within the immersive environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The various objects, features, and advantages of the disclosed subject matter can be more fully understood when considered in conjunction with the following drawings, in which like reference numerals refer to like elements, and in reference to the following detailed description of the disclosed subject matter.
[0023] Figure 1 Illustrative examples of processes for object grouping and manipulation in an immersive environment in accordance with some embodiments of the disclosed subject matter are shown.
[0024] Figure 2A Illustrative examples of video objects in an immersive environment in accordance with some embodiments of the disclosed subject matter are shown.
[0025] Figure 2B Illustrative examples of the generated group video objects with a handle interface element and an optional indicator element in an immersive environment in accordance with some embodiments of the disclosed subject matter are shown.
[0026] Figure 2C Illustrative examples of interactions with the generated group video objects in an immersive environment using a handle interface element (e.g., detecting gestures using the handle interface element, receiving selections of the handle interface element using a ray-based controller, etc.) in accordance with some embodiments of the disclosed subject matter are shown.
[0027] Figure 2D Illustrative examples of adding additional video objects from an immersive environment to the group video object in accordance with some embodiments of the disclosed subject matter are shown.
[0028] Figure 2EIllustrative examples of user interfaces presented in response to interaction with optional indicator elements of a group video object according to some embodiments of the disclosed subject matter are shown.
[0029] Figure 2F Illustrative examples of creating a playlist from a group video object in an immersive environment according to some embodiments of the disclosed subject matter are shown.
[0030] Figure 2G Illustrative examples of a grid interface presented to modify (e.g., add, remove, rearrange, etc.) video objects included in a group video object according to some embodiments of the disclosed subject matter are shown.
[0031] Figure 2H Illustrative examples of a grid interface presented to modify (e.g., add, remove, rearrange, etc.) video objects included in a group video object according to some embodiments of the disclosed subject matter are shown.
[0032] Figure 3A Illustrative examples of an immersive environment including a group video object in an empty state for receiving one or more video objects according to some embodiments of the disclosed subject matter, where a handle interface element provides an interface for interacting with or manipulating the group video object.
[0033] Figure 3B Illustrative examples of an immersive environment including a group video object in a thumbnail play state for playing back video objects included in the group video object according to some embodiments of the disclosed subject matter, where a handle interface element provides an interface for interacting with or manipulating the group video object.
[0034] Figure 3C Illustrative examples of an immersive environment including a group video object in a thumbnail play state and additional content information according to some embodiments of the disclosed subject matter, where a handle interface element provides an interface for interacting with or manipulating the group video object.
[0035] Figure 4 A schematic diagram of an illustrative system suitable for implementing the mechanisms for object grouping and manipulation in an immersive environment described herein according to some embodiments of the disclosed subject matter is shown.
[0036] Figure 5 Illustrative examples according to some embodiments of the disclosed subject matter that can be used in Figure 4 the server and / or user equipment are shown. Detailed Description
[0037] According to various embodiments, mechanisms for object grouping and manipulation in an immersive environment (which may include methods, systems, and media) are provided.
[0038] In some embodiments, the mechanisms described herein may provide grouping interactions for creating and / or manipulating group video objects that include one or more video objects in an immersive environment. For example, in an immersive environment that includes a plurality of video objects each represented by a thumbnail representation, the mechanisms may receive a user interaction that places the thumbnail representation of a first video object on top of the thumbnail representation of a second video object. In response to receiving the user interaction that places the thumbnail representation of the first video object on top of the thumbnail representation of the second video object, the mechanisms may create a group video object that includes the first video object and the second video object, where the group video object may be represented by a stacked thumbnail representation.
[0039] It should be noted that any suitable method may be used to generate the stacked thumbnail representation. For example, in response to receiving the user interaction that places the thumbnail representation of the first video object on top of the thumbnail representation of the second video object, the mechanisms may automatically align the thumbnail representations into a stacked thumbnail representation that represents the video objects included in the group video object. In another example, in response to receiving the user interaction that places the thumbnail representation of the first video object on top of the thumbnail representation of the second video object, the mechanisms may present an animation showing the first video object and the second video object being merged into the group video object.
[0040] It should also be noted that the stacked thumbnail representation may represent the group video object in any suitable manner. For example, the thumbnail representation of the first video object that may be selected and manipulated above the thumbnail representation of the second video object may be positioned at the top of the stacked thumbnail representation, thereby presenting the thumbnail representation of the first video object as the first thumbnail representation in the stacked thumbnail representation. In another example, the first video object that may be selected and manipulated above the second video object may be sorted as the last video object in the group video object represented by the stacked thumbnail representation. In yet another example, the stacked thumbnail representation may include a mosaic thumbnail view of each thumbnail representation or a screenshot of each video object included in the group video object. In another example, the stacked thumbnail representation may rotate each thumbnail representation of each video object included in the group video object.
[0041] In some embodiments, a stacked thumbnail representation of a group of video objects can be presented concurrently with a handle interface element. The handle interface element can, for example, allow a user to manipulate the group of video objects within an immersive environment. For example, the handle interface element can be interacted with (e.g., using a grasping hand motion) within the immersive environment to move the group of video objects from one location to another. In another example, in response to receiving a particular gesture, such as a grasping hand motion followed by a shaking hand motion, the mechanisms can cause the group of video objects to expand, thereby presenting a thumbnail representation corresponding to the video objects included within the group of video objects. In yet another example, in response to receiving a particular gesture, such as a grasping hand motion (e.g., a fist gesture with the palm facing down) followed by a palm-up gesture, the mechanism can cause the group of video objects to be removed by ungrouping the group of video objects, thereby presenting each thumbnail representation of each video object included within the group of video objects for interaction, respectively.
[0042] It should be noted that the handle interface element can be used in one or more playback states within the immersive environment. For example, the handle interface element can be presented with a video object in an empty playback state that allows a user to place one or more thumbnail representations of the video object onto the video object that is currently in the empty playback state. In this example, the handle interface element can provide the user with the ability to move the video object in the empty playback state to an area, for example, where the user can place one or more thumbnail representations of the video object onto the video object in the empty playback state. In another example, the handle interface element can be presented with a video object in a playback state that allows the user to move the video object that is currently playing back and / or add additional video objects to the group of video objects.
[0043] In some embodiments, a stacked thumbnail representation of a group video object may be presented concurrently with a counter element, any other suitable optional identifier element, or any other suitable functional visible element that provides an entry point for interacting with the group video object. For example, the counter element may indicate the number of video objects included within the group video object. Continuing with this example, in response to receiving a user selection of the counter element corresponding to the group video object, these mechanisms may provide options for interacting with the group video object and / or each of the video objects included within the group video object - such as deleting the group video object, converting the group video object into a playlist object that includes the video objects included within the group video object, reordering or otherwise arranging the video objects included within the group video object, presenting a grid view of the video objects included within the group video object, presenting details associated with each of the video objects included within the group video object, removing at least one of the video objects included within the group video object, providing a rating associated with at least one of the video objects included within the group video object, and so on.
[0044] Note that although the embodiments described herein generally relate to manipulating and / or interacting with group video objects that contain one or more videos, this is merely illustrative. For example, in some embodiments, these mechanisms may be used to manipulate and / or interact with virtual objects corresponding to suitable content items (e.g., video files, audio files, television shows, movies, live-streamed media content, animations, video game content, graphics, documents, and / or any other suitable media content). In another example, in some embodiments, these mechanisms may be used to manipulate groups of applications represented by application icons in an operating system of an immersive environment. Continuing with this example, multiple application icons may be placed into a group application object, where the group application object is presented concurrently with a handle interface element for manipulating the group application object and a counter element for indicating the number of applications included within the group application object and for interacting with the group application object. In yet another example, in some embodiments, these mechanisms may be used to manipulate a collection of content in an immersive environment and / or otherwise interact with the collection of content in the immersive environment. Continuing with this example, multiple content files may be placed into a group content object, where the group content object is presented concurrently with a handle interface element for manipulating the group content object and a counter element for indicating the number of content files included within the group content object and for interacting with the group content object.
[0045] In conjunction with Figures 1 - 5 These and other features for object grouping and manipulation in an immersive environment are further described.
[0046] Turning to Figure 1, illustrative examples of processes for object grouping and manipulation in immersive environments are shown in accordance with some embodiments of the disclosed subject matter. In some embodiments, the blocks of process 100 may be performed by any suitable device, such as a virtual reality headset, a head-mounted display device, a gaming console, a mobile phone, a tablet computer, a television, and / or any other suitable type of user device.
[0047] At 102, process 100 may provide an immersive environment in which a user may interact with one or more virtual objects. For example, a user wearing a head-mounted display device immersed in an augmented reality and / or virtual reality environment may explore the immersive environment and interact with virtual objects in the immersive environment, among other things, via various different types of inputs. These inputs may include, for example, physical interactions, which include, for example, physical movement and / or manipulation of the head-mounted display device and / or an electronic device separate from the head-mounted display device, and / or gestures, arm gestures, head movements, and / or head and / or eye directed gazes, etc. The user may implement one or more of these different types of interactions to perform a particular action to move virtually within the virtual environment or move from a first virtual environment to a second virtual environment. Movement within the virtual environment or from one virtual environment to another may include moving the features of the virtual environment relative to the user while the user remains stationary to generate the perception of movement within the virtual environment.
[0048] In a more specific example, the immersive environment may include one or more virtual video objects corresponding to videos (e.g., videos available for playback), and the user may interact with one or more of these virtual video objects. As Figure 2A shown, a plurality of virtual video objects 210, 220, and 230 may be displayed within the immersive environment 200. Also as Figure 2A shown, each of the virtual video objects 210, 220, and 230 may be represented by a thumbnail representation. The thumbnail representation may include, for example, a representative image (e.g., a screenshot) of the video corresponding to the virtual video object, a title of the video corresponding to the virtual video object, etc. It should be noted that the thumbnail representation may include any suitable content, such as metadata associated with the video corresponding to the virtual video object, creator information associated with the video corresponding to the virtual video object, keywords associated with the video corresponding to the virtual video object, etc. It should be noted that the thumbnail representation may be displayed in any suitable manner. For example, in some embodiments, each virtual video object may be displayed as a volumetric thumbnail representation.
[0049] In such an immersive environment, the user may interact with one or more of these video objects. For example, as Figures 2A - 2HAs shown, a user may manipulate one or more of these video objects. In a more specific example, the user may direct a virtual beam or ray extending from a handheld electronic device connected to a head-mounted display device toward a virtual video object to select or identify the virtual video object. Continuing with this example, the user may activate a manipulation device or button of the handheld electronic device to indicate selection of the virtual video object. In some embodiments, the user may provide a specific gesture (e.g., a grasping gesture) to physically grasp a handle interface element for manipulating the virtual video object.
[0050] In some embodiments, at 104, process 100 may detect via a first input that a first video object has been virtually positioned above a second video object. For example, as Figure 2A shown, process 100 may detect via a suitable input that video object 210 has been selected and virtually positioned above video object 220. As described above, the input may include manipulation of the head-mounted display device and / or an electronic device separated from the head-mounted display device. For example, the user may direct a virtual beam or ray extending from a handheld electronic device connected to the head-mounted display device toward the first video object to identify the virtual video object, provide a grasping gesture to select the first video object, and provide a dragging gesture to place the first video object above the second video object.
[0051] In some embodiments, in response to detecting that a first video object has been virtually positioned above a second video object, process 100 may generate, at 106, a group video object that includes the first video object and the second video object, and display, at 108, the group video object in place of the first video object and the second video object.
[0052] For example, as Figure 2B shown, process 100 may generate a group video object 240 represented by a stacked thumbnail while continuing to display the remaining video objects, such as video object 230. In a more specific example, also as Figure 2B shown, the first video object 210 may be represented by a first thumbnail, the second video object 220 may be represented by a second thumbnail, and the group video object 240 may be represented by a stacked thumbnail in which the first thumbnail and the second thumbnail may be automatically aligned to generate the stacked thumbnail. The stacked thumbnail representing the group video object 240 may include any suitable number of layers - for example, two layers, to indicate that it is a group video object that includes multiple video objects, one layer corresponding to each video object included in the group video object, and so on.
[0053] In another example, as Figure 2DAs shown, in response to detecting that the third video object 230 has been virtually positioned above the group video object 240, the process 100 can generate an updated group video object 260 represented by a stacked thumbnail, while continuing to display any remaining video objects, where the group video object 260 includes the first video object 210, the second video object 220, and the third video object 230.
[0054] Alternatively, in some embodiments, in addition to the first video object and the second video object, the generated group video object can also be displayed in an immersive environment. For example, this can allow a user to manipulate video objects within the immersive environment to create different group video objects, each of which can include one or more of the same video objects. Continuing with this example, these group video objects can be converted into playlists, each playlist including one or more of the video objects displayed in the immersive environment, where some playlists may include the same content items.
[0055] It should be noted that the group video object can be represented in any suitable manner. For example, in response to detecting that the first video object 210 is virtually positioned above the second video object 220, the thumbnail representation of the first video object 210 can be arranged at the top position of the group video object. In another example, in response to detecting that the first video object 210 is virtually positioned above the second video object 220, the thumbnail representation of the first video object 210 can be sorted as the last video object in the group video object represented by a stacked thumbnail. In yet another example, in response to detecting that the first video object 210 is virtually positioned above the second video object 220, the stacked thumbnail representation can include a mosaic thumbnail view of each thumbnail representation or a screen capture of each video object included in the group video object. In another example, the stacked thumbnail representation can rotate each thumbnail representation of each video object included in the group video object. It should be noted that the process 100 can use any suitable criterion to arrange the video objects included in the group video object (e.g., the popularity of each video, the number of views of each video, the rating of each video, etc.).
[0056] In some embodiments, the group video object can be displayed together with a handle interface element. For example, as Figure 2B shown, the handle interface element 245 can be a three-dimensional handle element displayed together with the group video object. The handle interface element 245 can allow a user to interact with the group video object by receiving a specific gesture using the handle interface element 245. For example, using the handle interface element 245, the corresponding group video object can be manipulated - for example, picked up, thrown, repositioned, placed on a dashboard interface, etc. Figure 2CAn illustrative example is shown, where in response to receiving a grasping gesture or any other suitable gesture to indicate interaction with the group video object 240, the handle interface element 245 can be used to virtually move the group video object 240 from a first position to a second position.
[0057] Return reference Figure 1 , the process 100 can detect at 110 that a specific gesture has been received using the handle interface element, and in response to receiving the specific gesture, the process 100 can determine and perform a specific manipulation action at 112. For example, as Figure 2C shown, the handle interface element 245 can be used to virtually move the group video object from one virtual position to another virtual position. It should be noted that the handle interface element 245 can respond to different gestures. For example, in some embodiments, in response to receiving a gesture of palm flipping, the handle interface element 245 can cause the group video object to be deleted, where the video objects included in the group video object are displayed in the immersive environment. In another example, in some embodiments, in response to receiving a shaking or side-to-side gesture when interacting with the handle interface element 245, the handle interface element 245 can cause the group video object to expand to display the video objects included in the group video object. In yet another example, in some embodiments, in response to receiving an up-and-down gesture when interacting with the handle interface element 245, the handle interface element 245 can cause the last added video object to be removed from the group video object (e.g., using an animation where the video object pops out of the group video object).
[0058] It should be noted that the handle interface element 245 can be displayed with the group video object in any suitable state.
[0059] Figure 3A An illustrative example of a group video object 310 in an empty state for placing video objects in an immersive environment 300 according to some embodiments of the present invention is shown, where the handle interface element 245 provides an interface for interacting with or manipulating the group video object 310. Similarly, as Figure 3A shown, the user navigating the immersive environment can be prompted to interact with the group video object that is currently in an empty state, where the group video object indicates its empty state by displaying a message such as "Place video here". Continuing with this example, the user can interact with video objects or other suitable virtual objects and place, toss, or otherwise move these video objects onto the currently empty group video object 310. In response, the video objects can be added to the group video object 310 (e.g., where the thumbnail representations of each video object are automatically aligned as a stacked thumbnail representation).
[0060] Figure 3BIllustrates an illustrative example of the group video object 320 in the thumbnail state where a video object has been added to the group video object 310 in an empty state according to some embodiments of the disclosed subject matter. Similar to Figure 3A , the group video object 320 can continue to be displayed together with the handle interface element 245, which provides an interface for interacting with or manipulating the group video object 320. As Figure 3A shown, the user can interact with the group video object 320, where the video object included in the group video object 320 can be played back in the thumbnail state. For example, in response to selecting the group video object 320, the group video object 320 can switch between playing back and pausing one or more videos included in the group video object 320. Figure 3B In addition, in some embodiments, a playback option interface 322 can be displayed to allow the user to modify the playback controls of the video - for example, play, pause, fast - forward, rewind, repeat, increase volume, decrease volume, etc. It should be noted that the playback option interface 322 can include any suitable playback options, such as a timeline where the user can use any suitable gesture or input to manipulate to select a specific playback position of the video included in the group video object. It should also be noted that the playback option interface 322 can include any suitable navigation options for navigating within the videos included in the group video object (e.g., navigate to the previous video, navigate to the next video, automatically scroll through the videos included in the group video object, etc.).
[0061] In some embodiments, the thumbnail state of the group video object 320 can also include an additional media information interface 322 and / or a related media interface 324.
[0062] For example, as
[0063] shown, in response to selecting Figure 3C the additional media information interface 322, a user interface 332 can be presented that includes any suitable information related to the video being played back. In a more specific example, as Figure 3B shown, Figure 3CAs shown, the user interface 332 may include title information corresponding to the video being played back in the thumbnail state, options for rating the video being played back in the thumbnail state (e.g., thumbs up option, thumbs down option, the number of thumbs up ratings received, the number of thumbs down ratings received, etc.), an option to download the video being played back in the thumbnail state, an option to add the video being played back in the thumbnail state to a playlist, an option to queue the video being played back in the thumbnail state for later playback, an option to subscribe to the channel associated with the content creator of the video being played back in the thumbnail state, an option to subscribe to the channel including the video being played back in the thumbnail state, release information corresponding to the video being played back in the thumbnail state, a detailed description of the video being played back in the thumbnail state (e.g., provided by the content creator), etc. It should be noted that the user interface 332 may include any suitable information related to the video being played back in the thumbnail state, such as comments provided by the viewing user.
[0064] In another example, also as Figure 3C shown, in response to selecting Figure 3B the relevant media interface 324, a user interface 336 may be presented that includes content items related to the video being played back (e.g., Videos A through F). Continuing with this example, the user may interact with one of the relevant content items, e.g., to play back the relevant video in the thumbnail state, add the relevant video to a group video object (e.g., by providing a grasping gesture to one of the relevant video objects and tossing the relevant video object onto the group video object), receive additional information about the relevant content item, and so on.
[0065] Returning to reference Figure 1 , in addition to or instead of displaying the group video object and the handle interface element, process 100 may also present, at 108, the group video object and an optional indicator element for interacting with the group video object. For example, as Figure 2B shown, an optional indicator element 250 may be presented in the upper right corner of the group video object 240, where the optional indicator element 250 may provide an entry point for the user to view, modify the group video object and the video objects included within the group video object, and / or otherwise interact with the group video object and the video objects included within the group video object when selected.
[0066] In some embodiments, the optional indicator element 250 may be displayed as a video counter of the number of video objects included in the group video object. For example, as Figure 2A and Figure 2BAs shown, in response to detecting that the first video object 210 has been virtually positioned above the second video object 220, process 100 may generate a group video object 240, where an optional indicator element 250 is positioned at the upper right corner of the group video object 240, and where the optional indicator element 250 is represented as a video count of 2 to indicate that there are two videos included in the group video object. In another example, as Figure 2D shown, in response to detecting that the first video object 230 has been virtually positioned above the group video object 240, process 100 may generate or update a group video object 260, where an optional indicator element 250 is positioned at the upper right corner of the group video object 260, and where the optional indicator element 250 is represented as a video count of 3 to indicate that there are three videos present in the group video object.
[0067] It should be noted that the optional indicator element 250 may be positioned at any suitable location of the group video object 240. For example, in some embodiments, the optional indicator element 250 may be centered along the top boundary of the group video object 240. In another example, in some embodiments, the optional indicator element 250 may be positioned at the upper left corner of the group video object 240.
[0068] It should also be noted that in some embodiments, the stacked thumbnail representation of the group video object 240 may remain the same size or the same volume, while the optional indicator element 250 may grow or shrink to indicate the number of video objects included in the group video object 240. Alternatively, in some embodiments, the stacked thumbnail representation of the group video object 240 may remain relatively the same size while expanding in depth to roughly indicate the number of video objects included in the group video object 240 (e.g., in a stacked thumbnail representation with ten layers instead of two).
[0069] Return Figure 1 , in some embodiments, in response to detecting a selection of the optional indicator element at 114, process 100 may display, at 116, a user interface including one or more options for interacting with the group video object. For example, in response to receiving a suitable gesture or a suitable input received using an electronic device separate from the head-mounted display device, the process may display a corresponding user interface for interacting with the group video object within an immersive environment. In a more specific example, in response to receiving a gesture of a finger pressing on the optional indicator element on the group video object, the user interface may slide out from the group video object for interacting with the group video object.
[0070] For example, as Figure 2EAs shown, in response to detecting a user interaction that selects the optional indicator element 250, process 100 can present a user interface 270 that includes options for creating a playlist for the videos included in the group video object 240, an option for deleting the group video object 240, and an option for displaying a grid view of the video objects included in the group video object 240.
[0071] In some embodiments, in response to detecting a user interaction that selects the option for creating a playlist for the videos included in the group video object 240, the group video object 240 can be converted into a playlist object 280. For example, as Figure 2F shown, in response to detecting a user interaction that selects the option for creating a playlist for the videos included in the group video object 240, the video objects included in the group video object 240 are converted into a playlist object 280, where the stacked thumbnail representation of the group video object 240 that is currently identified as the topmost video by video A (e.g., "Title A" of video A on the stacked thumbnail representation) is replaced by a playlist title (e.g., "List A - B"). Additionally, in some embodiments, the playlist object 280 can include additional metadata associated with each video in the playlist (e.g., title information, creator information, timing information, source information, etc.).
[0072] In some embodiments, in response to detecting a user interaction that selects the option for deleting the group video object 240, the group video object 240 can be removed from the immersive environment. For example, the group video object 240 and the video objects included within the group video object 240 can be removed from the immersive environment. In another example, the group video object 240 can be removed, and the video objects included within the group video object 240 can be displayed separately in the immersive environment. In yet another example, the group video object 240 can be presented in an empty state (e.g., with a "Place video here" message), and the video objects that were previously included within the group video object 240 can be positioned in a remote area of the immersive environment (e.g., thrown aside).
[0073] In some embodiments, in response to detecting a user interaction that selects the option for displaying a grid view of the video objects included in the group video object 240, the group video object 240 can provide a detailed user interface that shows the videos included within the group video object 240.
[0074] For example, as Figure 2GAs shown, in response to detecting a user interaction selecting an option to display a grid view of video objects included in group video object 240, group video object 240 can expand horizontally to provide a detailed user interface 290 showing video A and video B included within group video object 240. Continuing with this example, user interface 290 can provide a scrollable grid view of the video objects included in group video object 240, where the user can manipulate user interface 290 to scroll through the video objects included in group video object 240 in sequence.
[0075] In another example, as Figure 2H shown, in response to detecting a user interaction selecting an option to display a grid view of video objects included in group video object 240, detailed user interface 295 shows video A and video B included within group video object 240 and shows the number of videos included within group video object 240.
[0076] In these user interfaces, the user can view each video and / or additional information related to each video included within group video object 240, provide input for rearranging the order of the videos included within group video object 240, provide input for removing at least one video from group video object 240, convert group video object 240 to a playlist object, etc. For example, in some embodiments, gestures can be received for manipulating group video object 240 by removing and / or rearranging the videos included in group video object 240 via direct manipulation of the video objects using the user's hand.
[0077] In some embodiments, in response to detecting a user interaction selecting an optional indicator element 250 in a detailed user interface (e.g., Figure 2G detailed user interface 290 of Figure 2H detailed user interface 295 of Figure 2G detailed user interface 290 of Figure 2H detailed user interface 295 of
[0078] Turning to Figure 4, an illustrative example 400 of hardware for object grouping and manipulation in an immersive environment that can be used in accordance with some embodiments of the disclosed subject matter is shown. As shown, the hardware 400 can include a content server 402, a communication network 404, and / or one or more user devices 406, such as user devices 408 and 410.
[0079] The content server 402 can be any suitable server for storing media content and / or providing media content to the user devices 406. For example, in some embodiments, the content server 402 can store media content, such as videos, television shows, movies, live streaming content, audio content, animations, video game content, graphics, and / or any other suitable media content. In some embodiments, the content server 402 can transmit the media content to the user devices 406, for example, via the communication network 404. In some embodiments, the content server 402 can store video content (e.g., live video content, computer-generated video content, and / or any other suitable type of video content) associated with any suitable information to be used by a client device (e.g., the user device 406) to render the video content as immersive content. In some embodiments, the content server 402 can send virtual objects represented by thumbnail representations of content items such as videos.
[0080] In some embodiments, the communication network 404 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 404 can include the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any one or more of any other suitable communication networks. The user devices 406 can be connected to the communication network 404 via one or more communication links (e.g., communication link 412), and the communication network 404 can be connected to the content server 402 via one or more communication links (e.g., communication link 414). The communication link can be any communication link suitable for transmitting data between the user device 406 and the content server 402, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.
[0081] The user device 406 can include any one or more user devices suitable for requesting video content, rendering the requested video content as immersive video content (e.g., rendering as virtual reality content, three-dimensional content, 360-degree video content, 180-degree video content, and / or any other suitable manner), and / or performing any other suitable functions. For example, in some embodiments, the user device 406 can include a mobile device, such as a mobile phone, a tablet computer, a wearable computer, a laptop computer, a virtual reality headset, a vehicle (e.g., an automobile, a boat, an airplane, or any other suitable vehicle) information or entertainment system, and / or any other suitable mobile device and / or any suitable non-mobile device (e.g., a desktop computer, a gaming console, and / or any other suitable non-mobile device). As another example, in some embodiments, the user device 406 can include a media playback device, such as a television, a projector device, a gaming console, a desktop computer, and / or any other suitable non-mobile device.
[0082] In a more specific example where the user device 406 is a head-mounted display device worn by the user, the user device 406 can include a head-mounted display device connected to a portable handheld electronic device. The portable handheld electronic device can be, for example, a controller, a smart phone, a joystick, or another portable handheld electronic device that can be paired with and communicate with the head-mounted display device to interact in the immersive environment generated by the head-mounted display device and, for example, be displayed to the user on the display of the head-mounted display device.
[0083] It should be noted that the portable handheld electronic device can be operably coupled or paired with the head-mounted display device via, for example, a wired connection or a wireless connection - such as, for example, a WiFi or Bluetooth connection. Such pairing or operable coupling of the portable handheld electronic device and the head-mounted display device can provide communication between the portable handheld electronic device and the head-mounted display device and data exchange between the portable handheld electronic device and the head-mounted display device. For example, this can allow the portable handheld electronic device to be used as a controller for communicating with the head-mounted display device to interact in the immersive virtual environment generated by the head-mounted display device. For example, the manipulation of the portable handheld electronic device and / or the input received on the touch surface of the portable handheld electronic device and / or the movement of the portable handheld electronic device can be converted into corresponding selections, or movements, or other types of interactions in the virtual environment generated and displayed by the head-mounted display device.
[0084] It should also be noted that, in some embodiments, the portable handheld electronic device may include a housing, and the internal components of the device are accommodated in the housing. A user-accessible user interface may be provided on the housing. The user interface may include, for example, a touch-sensitive surface configured to receive user touch inputs, touch-and-drag inputs, etc. The user interface may also include user-operated devices, such as, for example, actuating triggers, buttons, knobs, toggle switches, joysticks, etc.
[0085] It should be further noted that, in some embodiments, the head-mounted display device may include a housing coupled to a frame, wherein the audio output device includes, for example, speakers mounted in earphones also coupled to the frame. For example, the front portion of the housing may rotate away from the base of the housing such that some of the components accommodated in the housing are visible. The display may be mounted on the inward-facing side of the front portion of the housing. In some embodiments, a lens may be mounted in the housing between the user's eyes and the display when the front portion is in the closed position against the base of the housing. The head-mounted display device may include: a sensing system including various sensors; and a control system including a processor and various control system devices to facilitate the operation of the head-mounted display device.
[0086] For example, in some embodiments, the sensing system may include an inertial measurement unit that includes various different types of sensors, such as, for example, accelerometers, gyroscopes, magnetometers, and other such sensors. The position and orientation of the head-mounted display device can be detected and tracked based on the data provided by the sensors included in the inertial measurement unit. Subsequently, the detected position and orientation of the head-mounted display device can allow the system to detect and track the user's head gaze direction and head gaze movement as well as other information related to the position and orientation of the head-mounted display device.
[0087] In some embodiments, the head-mounted display device may include a gaze tracking device that includes, for example, one or more sensors to detect and track the eye gaze direction and movement. The images captured by the sensors may be processed to detect and track the direction and movement of the user's eye gaze. The detected and tracked eye gaze may be processed as a user input to translate into corresponding interactions in an immersive virtual experience. The camera may capture still and / or moving images, which may be used to assist in tracking the physical position of the user and / or other external devices communicatively / operatively coupled to the head-mounted display device. The captured images may also be displayed to the user on the display in a see-through mode.
[0088] Although the content server 402 is illustrated as one device, in some embodiments, any suitable number of devices may be used to perform the functions performed by the content server 402. For example, in some embodiments, multiple devices may be used to implement the functions performed by the content server 402. In a more specific example, in some embodiments, a first content server may store media content items and respond to requests for media content, and a second content server may generate thumbnail representations of virtual objects corresponding to the requested media content items.
[0089] Although in Figure 4 two user devices 408 and 410 are shown to avoid overcomplicating the drawings, in some embodiments, any suitable number of user devices and / or any suitable type of user device may be used.
[0090] In some embodiments, the content server 402 and the user device 406 may be implemented using any suitable hardware. For example, in some embodiments, any suitable general-purpose computer or special-purpose computer may be used to implement the content server 402 and the user device 406. For example, a special-purpose computer may be used to implement a mobile phone. Any such general-purpose computer or special-purpose computer may include any suitable hardware. For example, as illustrated in the example hardware 500 of Figure 5 such hardware may include a hardware processor 502, a memory and / or storage 504, an input device controller 506, an input device 508, a display / audio driver 510, a display / audio output device 512, a communication interface 514, an antenna 516, and a bus 518.
[0091] In some embodiments, the hardware processor 502 may include any suitable hardware processor, such as a microprocessor, a microcontroller, a digital signal processor, special logic, and / or any other suitable circuitry, to control the operation of the general-purpose computer or special-purpose computer. In some embodiments, the hardware processor 502 may be controlled by a server program stored in the memory and / or storage 504 of a server (e.g., such as the content server 402). For example, in some embodiments, the server program may cause the hardware processor 502 to transmit media content items to the user device 206, transmit instructions for rendering a video stream as immersive video content and / or performing any other suitable actions. In some embodiments, the hardware processor 502 may be controlled by a computer program stored in the memory and / or storage 504 of the user device 406. For example, the computer program may cause the hardware processor 502 to render a video stream as immersive video content and / or perform any other suitable actions.
[0092] In some embodiments, the memory and / or storage 504 can be any suitable memory and / or storage for storing programs, data, media content, and / or any other appropriate information. For example, the memory and / or storage 504 can include random access memory, read-only memory, flash memory, hard disk storage, optical media, and / or any other suitable memory.
[0093] In some embodiments, the input device controller 506 can be any suitable circuitry for controlling and receiving inputs from one or more input devices 508. For example, the input device controller 506 can be circuitry for receiving inputs from a touch screen, from a keyboard, from a mouse, from one or more buttons, from a speech recognition circuit, from a microphone, from a camera, from an optical sensor, from an accelerometer, from a temperature sensor, from a near-field sensor, and / or any other type of input device.
[0094] In some embodiments, the display / audio driver 510 can be any suitable circuitry for controlling and driving outputs to one or more display / audio output devices 512. For example, the display / audio driver 510 can be circuitry for driving a touch screen, a flat panel display, a cathode ray tube display, a projector, one or more speakers, and / or any other suitable display and / or rendering device.
[0095] The communication interface 514 can be any suitable circuitry for engaging with one or more communication networks such as Figure 4 the network 404 shown. For example, the interface 514 can include network interface card circuitry, wireless communication circuitry, and / or any other suitable type of communication network circuitry.
[0096] In some embodiments, the antenna 516 can be any suitable one or more antennas for wireless communication with a communication network (e.g., communication network 204). In some embodiments, the antenna 516 can be omitted.
[0097] In some embodiments, the bus 518 can be any suitable mechanism for communicating between two or more components 502, 504, 506, 510, and 514.
[0098] According to some embodiments, any other suitable components can be included in the hardware 500.
[0099] In some embodiments, Figure 1 at least some of the above boxes of the process can be implemented or executed in any order or sequence that is not limited to the order and sequence shown and described in connection with the figures. Similarly, Figure 1Some of the above boxes may be implemented or executed substantially simultaneously, or in parallel, when appropriate, to reduce latency and processing time. Additionally or alternatively, some of the above boxes of the process may be omitted. Figure 1 of the process.
[0100] In some embodiments, any suitable computer-readable medium may be used to store instructions for performing the functions and / or processes described herein. For example, in some embodiments, the computer-readable medium may be transient or non-transient. For example, non-transient computer-readable media may include media such as non-transient forms of magnetic media (such as hard disks, floppy disks, and / or any other suitable magnetic media), non-transient forms of optical media (such as optical discs, digital video discs, Blu-ray discs, and / or any other suitable optical media), non-transient forms of semiconductor media (such as flash memory, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and / or any other suitable semiconductor media), any suitable media that is not transient or lacking permanence during transmission, and / or any suitable tangible media. As another example, transient computer-readable media may include signals in a network, wires, conductors, optical fibers, circuits, any suitable media that is transient and lacking permanence during transmission, and / or any suitable intangible media.
[0101] In cases where the systems described herein collect or utilize personal information about a user, the user may be provided with an opportunity to control whether programs or features collect user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or user's current location). Additionally, before storing or using certain data, certain data may be processed in one or more ways to remove personal information. For example, the user's identity may be processed so that no personal identity information can be determined for that user, or the user's geographical location may be generalized (such as generalized to the city, zip code, or state level) in the case of obtaining location information so that the user's specific location cannot be determined. Thus, the user can control how information about the user is collected and how information about the user is used by the content server.
[0102] Accordingly, methods, systems, and media for object grouping and manipulation in immersive environments are provided.
[0103] Although the invention has been described and illustrated in the foregoing illustrative embodiments, it should be understood that the present disclosure is by way of example only, and that many changes may be made to the details of the embodiments of the invention without departing from the spirit and scope of the invention, which is defined only by the appended claims. The features of the disclosed embodiments may be combined and rearranged in various ways.
Claims
1. A method for interaction in an immersive environment, comprising: displaying a plurality of virtual objects in the immersive environment; generating a group of virtual objects associated with a media information interface and a related media interface, the group of virtual objects including a handle interface element for interacting with the group of virtual objects and a playback option interface associated with the group of virtual objects and the handle interface element; displaying the group of virtual objects, the handle interface element, and the related media interface in the immersive environment; and displaying a content item in response to detecting a selection of the related media interface.
2. The method according to claim 1, wherein, the handle interface element is a three-dimensional handle element.
3. The method according to claim 1, wherein, the handle interface element interacts by detecting a gesture performed in the immersive environment.
4. The method according to claim 1, wherein, the immersive environment is a virtual reality environment generated in a head-mounted display device operating in a physical environment.
5. The method according to claim 1, wherein, a user interface is displayed for interacting with the content item, and wherein the user interface includes a plurality of content items adjacent to at least a first virtual object.
6. The method according to claim 5, wherein, the first virtual object is represented by a first thumbnail representation, and the second virtual object is represented by a second thumbnail representation.
7. The method according to claim 6, wherein, the group of virtual objects is represented by a stacked thumbnail representation, in which the first thumbnail representation and the second thumbnail representation are aligned to generate the stacked thumbnail representation, and the first thumbnail representation is virtually positioned above the second thumbnail representation.
8. The method according to claim 6, wherein, the group of virtual objects is represented by a stacked thumbnail representation, wherein the first thumbnail representation and the second thumbnail representation are aligned to generate the stacked thumbnail representation.
9. The method according to claim 6, wherein, the user interface includes an option for rearranging the order of at least one of the first virtual object or the second virtual object associated with the group of virtual objects or removing at least one of the first virtual object or the second virtual object associated with the group of virtual objects.
10. The method according to claim 1, further comprising: presenting a grid of virtual objects included in the group of virtual objects in response to detecting an interaction with the group of virtual objects.
11. A system for interaction in an immersive environment, comprising: a memory; and a hardware processor configured to, when executing computer-executable instructions stored in the memory: display a plurality of virtual objects in the immersive environment; generate a group of virtual objects associated with a media information interface and a related media interface, the group of virtual objects including a handle interface element for interacting with the group of virtual objects and a playback option interface associated with the group of virtual objects and the handle interface element; Display the set of virtual objects, the handle interface element, and the associated media interface in the immersive environment; and In response to detecting a selection of the associated media interface, display a content item.
12. The system according to claim 11, wherein, The immersive environment is a virtual reality environment generated in a head-mounted display device operating in a physical environment.
13. The system according to claim 11, wherein, The handle interface element is a three-dimensional handle element.
14. The system according to claim 11, wherein, The handle interface element interacts by detecting gestures performed in the immersive environment.
15. The system according to claim 11, wherein, Display a user interface for interacting with the content item, and wherein the user interface includes a plurality of content items adjacent to at least a first virtual object.
16. The system according to claim 15, wherein, The first virtual object is represented by a first thumbnail representation, and the second virtual object is represented by a second thumbnail representation.
17. The system according to claim 16, wherein, The set of virtual objects is represented by a stacked thumbnail representation in which the first thumbnail representation and the second thumbnail representation are aligned to generate the stacked thumbnail representation, and the first thumbnail representation is virtually positioned above the second thumbnail representation.
18. The system according to claim 16, wherein, The user interface includes options for rearranging the order of at least one of the first virtual object or the second virtual object associated with the set of virtual objects or removing at least one of the first virtual object or the second virtual object associated with the set of virtual objects.
19. The system according to claim 11, wherein, The hardware processor is further configured to: in response to detecting an interaction with the set of virtual objects, cause a mesh of virtual objects included in the set of virtual objects to be presented.
20. A non-transitory computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method, the method comprises: Display a plurality of virtual objects in an immersive environment; Generate a set of virtual objects including a first virtual object and a second virtual object, the set of virtual objects including a handle interface element for interacting with the set of virtual objects and an optional indicator associated with the first virtual object and the second virtual object; Display the set of virtual objects, the handle interface element, and the optional indicator in the immersive environment, wherein a first user interaction using a first gesture performed in the immersive environment received using the handle interface element causes the set of virtual objects to be manipulated using the first gesture in the immersive environment; and In response to detecting a selection of the optional indicator using a second gesture performed in the immersive environment, display a user interface for interacting with at least one of the first virtual object and the second virtual object included in the set of virtual objects in the immersive environment.
21. A method for interacting with a thumbnail representation, comprising: displaying one or more virtual objects in an immersive environment; generating a set of video objects in an empty state; displaying the set of video objects in the empty state; in response to receiving a user interaction to move at least one of the one or more virtual objects onto the set of video objects in the empty state, displaying a second set of video objects in a thumbnail state; and in response to receiving a second user interaction with the second set of video objects, playing back the at least one virtual object included in the second set of video objects in the thumbnail state.
Citation Information
Patent Citations
Dynamic switching and merging of head, gesture and touch input in virtual reality
CN107533374A
Mobile terminal and method for controlling the same
KR1020120033659A