representation of media data
By receiving and analyzing metadata sets to represent media data in virtual reality space, the challenges of media data representation and interaction in virtual reality are solved, enabling an intuitive user experience and data editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2020-01-16
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to effectively represent and interact with media data, especially video data, in virtual reality spaces, resulting in an unintuitive user experience.
By receiving a set of metadata associated with media data, the representation of the media data in virtual reality space is determined based on this metadata, and sensor tracking of user input is used to achieve intuitive interaction, including mapping of location, motion, and time information.
It provides an intuitive representation and interaction of media data in virtual reality space, allowing users to easily select and edit media data through diving gestures and other methods, thus integrating informational and social network information.
Smart Images

Figure CN115244940B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of separately processing and preparing media data. Background Technology
[0002] In various computer systems, a user interface can be provided to control the playback of media data, allowing users to play the desired portions of the media. For example, it can enable users to fast forward or rewind video clips in real time. Virtual reality (VR) technology can provide immersive video experiences (e.g., through virtual reality headsets), allowing users to observe virtual content around them in a virtual reality space. Summary of the Invention
[0003] The present invention is provided to introduce, in a simplified form, some concepts further described in the following detailed description. The summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0004] The object of this invention is to determine the representation of media data in virtual reality space. The above and other objects are achieved through the features of the independent claims. Other implementations are apparent from the dependent claims, the specification, and the drawings.
[0005] According to a first aspect, an apparatus is provided, such as a media data preparation apparatus for receiving media data. The apparatus may include a memory storing instructions executable by a processor to cause the apparatus to receive a set of metadata based on at least one spatial coordinate. The metadata set may be associated with the media data. The apparatus may also be used to: determine a representation of the media data in a virtual reality space based on the metadata set. This approach enables the provision of an informative representation of media data in a virtual reality space, which a user might perceive as being presented in a particularly realistic manner.
[0006] In one implementation of the first aspect, determining the representation of the media data may include mapping the metadata set to at least one location in the virtual reality space. This approach provides a media representation that reflects the location of the user or capture device during the capture of the media data. Therefore, the user can observe the location associated with the media data in the virtual reality space.
[0007] In another implementation of the first aspect, the media data may include multiple video frames. A subset of the metadata set may correspond to each of the multiple video frames. This scheme is capable of providing a video representation that reflects the position of the user or capturing device during the capture of each video frame.
[0008] In another implementation of the first aspect, the metadata set may include at least one of location information, motion information, and time information associated with the media data. This solution can provide a location-, motion-, and / or time-related media representation in the virtual reality space.
[0009] In another implementation of the first aspect, the location information may include multiple locations associated with the capture of the media data. This approach enables mapping portions of the media data corresponding to different capture locations to corresponding locations in the virtual reality space. Alternatively or additionally, the motion information may include multiple velocities associated with the capture of the media data. This approach enables determining the capture location based on the velocity and providing a location-related media representation in the virtual reality space.
[0010] In another implementation of the first aspect, the motion information may include gyroscope data and / or magnetometer data. This approach enables the motion information to be locally recorded in the capture device.
[0011] In another implementation of the first aspect, the representation of the media data may include a first representation of the first media data and a second representation of the second media data. The first and second representations may intersect in the virtual reality space. This approach allows the user to easily observe the spatial and temporal consistency of the media data in the virtual reality space. Furthermore, this approach can also determine the relationships between different media data and / or users associated with them.
[0012] In another implementation of the first aspect, the device can also be used to: detect user input at a first location of the representation of the media data in the virtual reality space. The device can also be used to: perform an operation associated with the media data based on the first location. This solution enables intuitive interaction between the user and the media data.
[0013] In another implementation of the first aspect, detecting the user input may include: detecting a user's body part or virtual reality controller to correspond with the representation of the media data in the virtual reality space at the first location. This approach is capable of detecting user interaction with the media representation.
[0014] In another implementation of the first aspect, detecting the user input may include detecting the user's head to correspond with the representation of the media data in the virtual reality space at the first location. For example, the position of the user's head may be tracked based on at least one sensor associated with the virtual reality headset. This approach allows the user to interact with the media representation through intuitive diving gestures, etc.
[0015] In another implementation of the first aspect, the operation may include: initiating playback of the media data from a portion of the media data associated with the first location. This approach allows the user to select a location in the virtual reality space for initiating media playback based on intuitive diving gestures, etc. Alternatively or additionally, the operation may include: providing a preview of a portion of the media data associated with the first location. This approach allows the user to select a desired portion of the media data. Alternatively or additionally, the operation may include: editing the representation of the media data. This approach enables an intuitive user interface for modifying the content and / or playback order of the media data in the virtual reality space.
[0016] In another implementation of the first aspect, the device may further be used to: determine social network information based on the representation of the media data and / or the metadata set. This scheme can determine the relationships between users associated with different media data based on the intersection of media representations, etc. Alternatively or additionally, the device may be used to: receive social network information associated with the media data. The device may also be used to: determine the representation of the media data based on the social network information and the metadata set. This scheme can provide the user with a media representation that takes into account the relationships between users associated with different media data.
[0017] In another implementation of the first aspect, the device may further be used to: receive a set of context metadata associated with the media data. The device may also be used to: determine a prediction of at least one event based on the set of context metadata and / or the set of metadata. This scheme can inform the user of future events or the nature of future events, such as encounters with other users.
[0018] According to a second aspect, a method for preparing media data is provided. The method may include: receiving a set of metadata based on at least one spatial coordinate. The metadata set may be associated with the media data. The method may further include: determining a representation of the media data in a virtual reality (VR) space based on the metadata set. This approach is capable of providing an informative representation of the media data in a virtual reality (VR) space.
[0019] In one implementation of the second aspect, the method may be executed in a device according to any implementation of the first aspect.
[0020] According to the third aspect, the computer program may include computer program code, which, when executed on a computer, is used to perform any implementation of the method according to the second aspect.
[0021] According to the fourth aspect, a computer program product may include a computer-readable storage medium storing program code, the program code including instructions for performing the method according to any implementation of the second aspect.
[0022] According to the fifth aspect, the device may include means for performing any implementation of the method of the second aspect.
[0023] Therefore, embodiments of the present invention can provide an apparatus, method, computer program, and computer program product for determining the representation of media data in a virtual reality space. These and other aspects of the invention will become apparent from the embodiments described below. Attached Figure Description
[0024] The accompanying drawings, which are included to further illustrate embodiments of the invention and form part of this specification, and together with the specification, facilitate understanding of the embodiments of the invention, wherein:
[0025] Figure 1 An example of a video system provided by an embodiment of the present invention is shown;
[0026] Figure 2 Examples of devices for implementing one or more embodiments of the present invention are shown;
[0027] Figure 3 An example of video representation in a virtual reality space provided by an embodiment of the present invention is shown;
[0028] Figure 4 This invention illustrates another example of video representation in virtual reality space provided by an embodiment of the invention;
[0029] Figure 5 Examples of determining and providing video representations provided by embodiments of the present invention are shown;
[0030] Figure 6 An example of an edited video representation provided by an embodiment of the present invention is shown;
[0031] Figure 7 An example of an edited video representation provided by an embodiment of the present invention is shown;
[0032] Figure 8 An example of the analytical video representation or metadata provided in an embodiment of the present invention is shown;
[0033] Figure 9 An example of analyzing video representations associated with different users, provided by an embodiment of the present invention, is shown;
[0034] Figure 10 An example of a video representation determined based on social network information provided in an embodiment of the present invention is shown;
[0035] Figure 11 An example of predicting future events based on video representation provided by an embodiment of the present invention is shown;
[0036] Figure 12 An example of an extended video representation based on the prediction of at least one future event, provided by an embodiment of the present invention, is shown;
[0037] Figure 13 An example of a video representation overlaid on geographic content provided by an embodiment of the present invention is shown;
[0038] Figure 14 This illustrates another example of a video representation overlaid on geographic content provided by an embodiment of the present invention;
[0039] Figure 15 An example of a method for preparing media data provided by an embodiment of the present invention is shown.
[0040] In the accompanying drawings, the same reference numerals are used to denote the same parts. Detailed Implementation
[0041] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The detailed description provided below, in conjunction with the drawings, is intended to illustrate examples of the invention and is not intended to represent the only ways in which the examples of the invention can be constructed or utilized. The detailed description illustrates the functionality of the examples of the invention and the order of steps in constructing and operating the examples of the invention. However, the same or equivalent functionality and order can be implemented through different examples.
[0042] Traditionally, it should be understood that there are three spatial dimensions (length, width, and depth) for locating an object's coordinates in space. Additionally, time can be considered a non-spatial fourth dimension. Conversely, this time dimension can be used to locate an object's position in time. Because the physical world is three-dimensional, time cannot be directly observed, as this transcends a single observable physical world. One can only perceive the direction of time's progression by observing changes in the physical world over time. However, according to another perspective, measured time is considered a purely mathematical value that itself does not exist physically. Accordingly, time is considered a fourth spatial dimension in an "eternal" universe.
[0043] Therefore, the exemplary embodiments disclosed herein are capable of representing and analyzing media data based on time, motion, and / or location information associated with media data (e.g., one or more video files). For example, an intuitive user interface for accessing or editing the media data can be provided in a virtual reality space, thereby implementing time and space factors in a form that allows the media data to be represented in an informative manner.
[0044] In one scenario, the media capture device can be implemented as an everyday wearable device or even a human implant. In this way, an entire life process, or a portion thereof, can be captured and visualized as a media representation within the virtual reality space. This allows the user to jump to a specific event or highlight in the blink of an eye. Furthermore, contextual metadata such as lifestyle data can be stored and associated with the media data. For example, the contextual metadata may include data received from health sensors and / or spectral cameras. The lifestyle data can then be analyzed to estimate its impact on lifespan. The result of the lifespan estimate can be reflected in the length of the media representation.
[0045] According to an exemplary embodiment, an apparatus for receiving media data may receive a set of metadata, such as at least one spatial coordinate associated with the media data. The metadata set may be used to determine a representation of the media data (e.g., a tube) in the virtual reality space, such that capture position and / or time information associated with the media data is reflected in the shape and / or length of the media representation.
[0046] Figure 1An example of a video system 100 provided in an embodiment of the present invention is shown. The video system 100 may include a video device, represented by a virtual reality (VR) headset 102. However, the video device may generally include any device suitable for providing or assisting in providing virtual reality content to a user. For example, the video device may include a standalone VR headset or other device (e.g., a mobile phone for coupling to the VR headset) so that the virtual reality content can be experienced on the display of the other device.
[0047] Virtual reality space can refer to the space observable by a user when consuming virtual reality content using a virtual reality device (e.g., the VR headset 102). For example, the virtual reality content may include omnidirectional video content, such that a portion of the content is displayed to the user based on their current gaze direction. The virtual reality content may be three-dimensional (e.g., stereoscopic), such that objects contained within the virtual reality content are displayed at specific locations within the virtual reality space. For example, a specific object may be displayed at a certain depth and in a certain direction within the virtual reality space. Virtual reality space can also blend real and virtual content, such that virtual content can be enhanced on top of a real-world view using transparent glasses or the like. Therefore, it should be understood that augmented reality (AR) or mixed reality (MR) can be considered different forms of virtual reality. Virtual reality space can also be referred to as virtual space, three-dimensional (3D) space, or 3D virtual space.
[0048] The VR system 100 may further include a video controller 104 for performing various control tasks, such as retrieving, decoding, and / or processing video content for display on the VR headset 102. The video controller 104 may include a video player. The video controller 104 may be implemented as a standalone device or integrated into the VR headset 102, for example, as a software and / or hardware component. Therefore, the video controller 104 can communicate with the VR headset 102 via one or more internal or external communication interfaces (e.g., data bus, wired connection, or wireless connection).
[0049] The video system 100 may further include a video server 108, which may be used to store video data and / or metadata associated with different users and provide the data to the video controller 104 upon request. The video server 108 may be a local server, or it may be accessible via a network 106 (e.g., the Internet). However, the video data and / or associated metadata may be stored locally in the VR headset 102 or the video controller 104. Therefore, exemplary embodiments can also be implemented without network access.
[0050] The video content can be captured by a mobile recording device, which may include a mobile phone or wearable device. For example, the mobile recording device can be worn around the user's wrist or neck, such as as a watch or bracelet. The mobile recording device can also be coupled to the user's head, such as as a VR headset or smart glasses. The mobile recording device can be used to record short and / or long events. The mobile recording device can provide the video content to the video server 108, or the video content can be captured locally by the mobile recording device, making the video content locally available to the video controller 104 and / or the VR headset 102. According to one embodiment, the VR headset 102 can capture the video content. For example, the recorded video content can be stored locally or remotely in a cloud accessible via a network.
[0051] Figure 2 Embodiments of a device 200 (e.g., a VR headset 102, a video controller 104, or a video server 108, etc.) are illustrated, which is used to implement one or more exemplary embodiments. According to some embodiments, the device 200 may be configured as a media data preparation device 200. The device 200 may include at least one processor 202. For example, the at least one processor 202 may include one or more of various processing devices (e.g., coprocessors, microprocessors, controllers, digital signal processors (DSPs), processing circuitry with or without an accompanying DSP) or various other processing devices including integrated circuits (e.g., application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontroller units (MCUs), hardware accelerators, dedicated computer chips, etc.).
[0052] The device 200 may further include at least one memory 204. The at least one memory 204 may be used to store computer program code, such as operating system software and application software. The at least one memory 204 may include one or more volatile memory devices, one or more non-volatile memory devices, and / or combinations thereof. For example, the at least one memory 204 may be implemented as a magnetic storage device (e.g., hard disk drive, floppy disk, magnetic tape, etc.), an optical-magnetic storage device, or a semiconductor memory (e.g., mask ROM, programmable ROM (PROM), erasable PROM (EPROM), flash memory ROM, random access memory (RAM), etc.).
[0053] The device 200 may further include a communication interface 208 for enabling the device 200 to send and / or receive information. The communication interface 208 may be used to provide at least one wireless connection, such as a 3GPP mobile broadband connection (e.g., 3G, 4G, 5G). Alternatively or additionally, the communication interface 208 may be used to provide one or more other types of connections, such as a wireless local area network (WLAN) connection, for example, a connection standardized by the IEEE 802.11 series or the Wi-Fi Alliance; a short-range wireless network connection, such as Bluetooth, near-field communication (NFC), or RFID; a wired connection, such as a local area network (LAN) connection, a universal serial bus (USB) connection, a high-definition multimedia interface (HDMI), or an optical network connection; or a wired internet connection. The communication interface 208 may include or be coupled to at least one antenna to send and / or receive radio frequency signals. One or more of the various types of connections can also be implemented as a separate communication interface, which can be coupled to or used to couple to multiple antennas.
[0054] The device 200 may further include a user interface 210, which includes or is used to couple to input devices and / or output devices. The input devices may take various forms, such as a keyboard, a touchscreen, and / or one or more embedded control buttons. The input devices may also include wireless control devices, such as a virtual reality manual controller. For example, the output devices may include at least one display, a speaker, a vibration motor, an olfactory device, etc.
[0055] When the device 200 is used to implement a certain function, one and / or some components of the device 200 (e.g., the at least one processor and / or the memory) can be used to implement that function. Furthermore, when the at least one processor is used to implement some functions, those functions can be implemented using program code 206, such as that included in the at least one memory 204.
[0056] The functions described herein may be performed at least in part by one or more computer program product components (e.g., software components). According to one embodiment, the device 200 includes a processor or processor circuitry (e.g., a microcontroller) configured by the program code to perform embodiments of the operations and functions described herein when the program code is executed. Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. Exemplary types of hardware logic components that may be used, such as but not limited to, include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), and graphics processing units (GPUs).
[0057] The device 200 includes means for performing at least one method described herein. In one example, the means includes at least one processor and at least one memory including computer program code, the at least one memory and the computer code being used by the at least one processor to cause the device to perform at least the method.
[0058] Although the device 200 is shown as a single device, it should be understood that, where applicable, the functionality of the device 200 can be distributed to multiple devices, for example, to implement the exemplary embodiment as a cloud computing service.
[0059] Figure 3An example of a video representation 302 in a virtual reality (VR) space provided by an embodiment of the present invention is illustrated. The VR space can be associated with a coordinate system, allowing a position within the VR space to be identified, for example, relative to three orthogonal axes 306 (x, y, z). Alternatively, other coordinate systems can be used. The VR space can provide six degrees of freedom (6DOF) capability, allowing a user 308 to move relative to the video representation 302 within the VR space. Alternatively, the representation can be provided in other virtual reality formats, such as 3DOF (where the user 308 can view the VR space from one position) or 3DOF+ (where the user can also move their head).
[0060] The video representation 302 can be generated based on video content (e.g., multiple video frames 304). A video frame can include a snapshot of the video data at a specific moment. Depending on the temporal resolution, the video data can include different numbers of frames per second (fps), such as 50 frames per second. Portions of the video data can be associated with a set of metadata. The metadata set can be based on at least one spatial coordinate. For example, the metadata set can include at least one coordinate associated with the capture of the video data. For example, the metadata set can be stored at the time the video data is captured. According to one example embodiment, a subset of the metadata can be associated with each video frame. For example, the metadata set can include the capture location of each video frame. Although embodiments have been described using video data as an example, it should be understood that these embodiments can be applied to other types of media data, such as audio data and / or olfactory data with or without associated video or image data.
[0061] The metadata set may include location information (e.g., coordinates) that may be associated with at least a portion of the captured video data. The location information may include positioning system coordinates. The location information can be determined based on any suitable positioning system, such as Global Positioning System (GPS), Wi-Fi positioning, etc. Typically, the location information may include multiple locations associated with the capture of the media data.
[0062] Alternatively or additionally, the metadata set may include motion information, such as multiple velocities associated with the capture of the media data (e.g., the video data). These multiple velocities may be associated with different portions of the video data (e.g., different video frames). The multiple velocities may include multiple directions of motion and corresponding multiple motion speeds. The motion information may be determined based on sensor data provided by one or more accelerometers, gyroscopes, and / or magnetometers embedded in or associated with the capture device. Therefore, the motion information may include gyroscope data and / or magnetometer data. The motion information may also be determined based on the tracking position of the capture device.
[0063] Alternatively or additionally, the metadata set may include time information associated with the media data, such as timestamps for each video frame or a subset of video frames. The time information may include absolute time (e.g., date and time in Coordinated Universal Time (UTC) format) or relative time (e.g., relative to the start position of the video data).
[0064] The metadata set associated with the video data can be used to provide the user 308 with a video representation 302 reflecting location information associated with the video data. For example, the video representation 302 may include a tube that propagates in the VR space according to the capture position of the video data. The length of the tube may reflect the time dimension. For example, the start position of the tube in the VR space may correspond to the start position of the video data, and the end position of the tube may correspond to the end position of the video data. Therefore, the four dimensions of the "eternal" universe can be advantageously presented to the user 308 in the VR space.
[0065] According to one embodiment, the video data may include a plurality of video frames 304. The metadata set may include a plurality of metadata subsets. Each video frame or subset of video frames may be associated with a subset of metadata. For example, a video capture device may be used to store a subset of metadata (e.g., spatial location) from time to time (e.g., periodically at certain time intervals), such that some video frames are associated with a subset of metadata while some video frames are not associated with a subset of metadata. For example, the metadata set may be stored as a metadata track in a video file; or, the metadata set may be stored separately.
[0066] The metadata set associated with the video frame can be mapped to the coordinate system of the virtual space, thereby determining the shape and / or position of the video representation 302 in the VR space. For example, video frame 310 can be associated with specific coordinates in the real world. These real-world coordinates can be mapped to coordinates (x0, y0, z0) in the VR space. Therefore, it can be determined that the representation 302 passes through a point (x0, y0, z0) in the VR space.
[0067] Figure 4 Another example of a video representation 402 in a virtual reality space provided by an embodiment of the present invention is shown. The video representation 402 can represent video data of a roller coaster (e.g., multiple video frames 404) such that the multiple wheels of the roller coaster form a closed tube in the VR space. Therefore, the location-based video representation 402 provides an informative representation of the video data to the user 308.
[0068] Combination Figure 3 and Figure 4 It should be noted that the user 308 can access the video data (e.g., multiple video frames 304, 404), or interact with the video representations 302, 402 in the VR space using various user inputs such as diving gestures 312 or pointing gestures 410, as will be further described below.
[0069] Figure 5 An example of determining and providing a video representation provided by an embodiment of the present invention is illustrated. The figure shows different operations performed by the VR headset 102, the video controller 104, and the video server 108. However, it should be understood that these operations can be performed by any suitable device or function of the video system. Furthermore, similar representations can be provided for other types of media (e.g., audio or image data) or multimedia content.
[0070] In step 501, the video controller 104 can retrieve a set of video data and / or metadata to provide a representation of video data in a VR space. The video data and / or metadata can be received from local memory (e.g., the internal memory of the video controller 104) or an external local memory device (e.g., an external hard disk drive). The video data may include multiple video frames. The metadata may be associated with the video data. The metadata may include spatial metadata, such as spatial information like location and / or motion associated with at least a portion of the video data. The metadata may also include temporal metadata, such as time information associated with at least a portion of the video data, like a timestamp. The portion of the video data may include one or more video frames.
[0071] Alternatively, the video data and / or metadata can be received from the video server 108, for example, via the network 106. For example, in 502, the video controller 104 can send a data request to the video server 108. The request may include one or more identifiers of the data and / or one or more conditions for providing the data. For example, the data request may indicate one or more users, one or more geographical regions, one or more time periods, etc., as conditions for providing the data. The data request may indicate whether the video controller 104 requests video data, metadata, or both.
[0072] The data request may indicate a request for video data for the purpose of providing a video representation. This allows for a reduction in the amount of data transfer between the video controller 104 and the video server 108. For example, in response to receiving such an instruction, the video server 108 may provide information necessary for generating the video representation. For example, the video server 108 may provide metadata without video data or metadata with a subset of video data. According to one embodiment, the video controller 104 may explicitly request the metadata and / or a subset of the video data. Where applicable, similar conditions or requests may also be applied when retrieving the video data from local storage.
[0073] In step 503, the video controller 104 can receive the video data from the video server 108.
[0074] In step 504, the video controller 104 may receive the metadata set from the video server 108. The video data and the metadata may be received as separate files or streams. Alternatively, the metadata may be provided along with the video data or a subset thereof, for example, as a metadata track. Typically, the video controller 104 may receive a metadata set based on at least one spatial coordinate, wherein the metadata set is associated with the media data. The media data may include multiple video frames, wherein a subset of the metadata set may correspond to each of the multiple video frames.
[0075] In step 505, the video controller 104 may determine at least one representation of the video data in the VR space based on the metadata set. The video representation may be determined such that the shape and / or position of the video representation in the VR space reflects positional information associated with the video data. Typically, the video controller 104 may determine the representation of the media data in the virtual reality space based on the metadata set. Determining the representation of the media data may include mapping the metadata set to at least one location in the virtual reality space. The metadata set may include at least one of positional information, motion information, and temporal information associated with the media data.
[0076] The location information (e.g., at least one real-world coordinate) can be directly mapped to the VR space; or, the location of the video representation in the VR space can be determined based on a transformation of the location information. For example, the location of the video representation in the VR space can be scaled to amplify or attenuate movement during video capture, etc. The video representation can also be determined based on time information associated with the capture of the media data. For example, the time information can be used to modify the video representation so that different locations in the VR space represent events associated with different times or periods of time. According to one embodiment, the length of the video representation in the VR space can be determined based on the duration of the video data. The video representation can include various geometries, such as geometric lines, splines, or tubes and / or branches of said geometric lines, splines, or tubes. The video controller 104 can also generate the video representation data in a format suitable for display on the VR headset 102.
[0077] In step 506, the video controller 104 can send the video representation data to the VR headset 102.
[0078] In step 507, the VR headset 102 can display the video representation. The user can experience the video representation in the VR space. For example, the user can observe how the video content propagates in the VR space over time.
[0079] In step 508, the VR headset 102 can send sensor data to the video controller 104. The VR headset 102 may embed sensors capable of tracking the user's current viewport. The viewport may include a portion of the VR space visible to the user at one time through the VR headset 102. The current viewport may be tracked based on the position and / or orientation of the VR headset 102, allowing the portion of the VR space corresponding to the user's current gaze to be presented. Furthermore, it may also allow the user to move relative to the video representation, thereby viewing the video representation from different viewpoints.
[0080] The video controller 104 can receive the sensor data from the VR headset 102. Based on the received sensor data, the video controller 104 can send video, image, and / or audio data corresponding to the current viewport to the VR headset 102. The sensor data may also include user input-related data, such as the position of the user's body parts (e.g., hands) in the VR space, the position of the VR controller, and the activation of buttons embedded in the VR headset 102 or associated VR controller.
[0081] In step 509, the video controller 104 can detect user input at a specific location in the video representation within the VR space. Typically, detecting user input at the video representation can include detecting body parts or virtual reality (VR) controllers to correspond with the video representation in the VR space. The video controller 104 can be used to track the user's position based on sensors attached to different body parts of the user, based on the VR controller worn by the user, and / or based on the position of the VR headset 102. For example, the video controller 104 can be used to track the position of the user's head, hand, or fingers to detect gestures at a specific location in the video representation. The video controller 104 can be used to perform at least one operation on the video representation or the video data based on the detected user input at the specific location. The VR controller can include a control device worn by the user. For example, the VR controller can include a manual controller. For example, the VR controller can be communicatively coupled to the VR headset 102 and / or the video controller 104 to transmit sensor data associated with the VR controller.
[0082] Typically, the video controller 104 can be used to: detect user input at a first location of the representation of the media data in the virtual reality space; and perform an operation associated with the media data based on the first location. Detecting the user input may include: detecting a body part or virtual reality controller of the user to correspond to the representation of the media data in the virtual reality space at the first location. Detecting the user input may include: detecting the user's head to correspond to the representation of the media data in the virtual reality space at the first location, wherein the position of the user's head may be tracked based on at least one sensor associated with the virtual reality headset. The operation includes at least one of the following: initiating playback of the media data from a portion of the media data associated with the first location; providing a preview of the portion of the media data associated with the first location; and editing the representation of the media data.
[0083] refer to Figure 3 For example, the user may be detected performing a diving gesture 312 on the video representation 302 at the first location (x1, y1, z1). Detecting the diving gesture may include detecting the user's head or the VR headset 102 to correspond to the video representation 302 at the first location (x1, y1, z1) in the VR space. For example, the position of the user's head may be tracked based on at least one sensor associated with the VR headset 102. The video controller 104 may be configured to perform an operation in response to detecting the diving gesture at the first location (x1, y1, z1). The operation may include initiating video playback from a portion of the video data corresponding to the first location (x1, y1, z1). For example, the video controller 104 may be configured to determine which of the plurality of video frames 304 corresponds to the first location (x1, y1, z1). In response to detecting the diving gesture or another predetermined gesture, the video controller 104 may determine to initiate video playback from that frame.
[0084] In step 510, for example, in response to the detected user input at the first position (x1, y1, z1) in the VR space, the video controller 104 can retrieve the video data. The video controller 104 can retrieve the video data from its internal memory, for example, from a video frame starting at the position corresponding to the first location (x1, y1, z1). The video controller 104 may have already retrieved or received necessary video content for initiating video playback in operation 503. Alternatively or additionally, the video controller 104 can retrieve the video data from the video server 108. For example, in step 511, the video controller 104 can send a video data request to the video server 108. The video data request may include an indication of the video frame corresponding to the first position (x1, y1, z1) in the VR space. In step 512, the video controller 104 can receive the requested video data from the video server 108.
[0085] In step 513, the video controller 104 can send the video data to the VR headset 102. In step 514, the VR headset 102 can display the video data. In step 515, the VR headset 102 can send sensor data to the video controller 104. It should be noted that even... Figure 5 The diagram illustrates a single operation of retrieving the video data 510, sending the video data 513, displaying the video data 514, and sending the sensor data 515. Providing the video data to the user may also include a continuous data stream from the video controller 104 and / or the video server 108 to the VR headset 102, wherein portions of the video data are provided based on sensor data (e.g., gaze direction and / or position) reported by the VR headset 102.
[0086] The video data may include three-dimensional video data or two-dimensional video data. Regardless of the video type, the video data can be presented to the user using the VR headset 102. For example, refer to... Figure 3 If the video data is three-dimensional (3D) video data, the video controller 104 can use the 3D video data to replace the 3D video representation 302. If the video data is two-dimensional (2D) video data, the video controller 104 can use the projection of the 2D video in the VR space to replace the 3D video representation 302, or use the projection of the 2D video data in the VR space to supplement the 3D video representation 302. For example, the 2D video can be displayed on a 2D surface in the VR space.
[0087] As described above, when the video data is displayed in 514, the VR headset 102 can report the sensor data to the video controller 104. In addition to the sensor data indicating the user's gaze and / or position, the sensor data may also include user input data indicating a request to terminate video playback. For example, the video controller 104 can terminate video playback in response to receiving such a request or upon reaching the end of the video data.
[0088] In step 516, the video controller 104 can be used to update the video representation. For example, when video playback is terminated, the video controller 104 can be used to determine a second position in the VR space corresponding to the last played video frame. The video representation can then be presented based on the position corresponding to the last played video frame in the VR space.
[0089] In step 517, the video controller 104 can send the updated video representation data to the VR headset 102.
[0090] In 518, the VR headset 102 can display the updated video representation.
[0091] For example, refer to Figure 3 Sensor data associated with a user input requesting termination of video playback can be received at a video frame corresponding to the second position (x2, y2, z2) in the VR space. In response to the user input, the video controller 104 can present the video representation 302 based on the second position (x2, y2, z2), for example, such that the user is located at or near the second position (x2, y2, z2) in the VR space. If the user does not terminate video playback, it can be determined that the last played video frame corresponds to the last video frame of the video data, and the video representation can be presented based on the end position of the video representation. Therefore, this embodiment provides an intuitive user interface for accessing media data in VR space. For example, the user can be allowed to enter a video tube in the VR space. At the end of video playback, the user exits the tube at the end of the tube or when video playback is terminated.
[0092] refer to Figure 4This allows the user to perform gestures 410 using body parts (e.g., hand 406) or VR controllers. The gestures can include any suitable gesture associated with the video representation 402, such as pointing gestures, drawing gestures, grabbing gestures, dragging gestures, wiping gestures, etc. Different gestures 410 can be associated with different operations. For example, a pointing gesture can be associated with a preview 408 that provides the video data in the VR space.
[0093] According to one embodiment, the video controller 104 may be used to: receive sensor data in 508; detect a gesture indicating a position (x3, y3, z3) in 509; and / or update the video representation using the preview 408 in 516. For example, the video controller 104 may determine which of the plurality of video frames 404 corresponds to the detected gesture indicating a position (x3, y3, z3). The video controller 104 may provide the preview 408 based on that video frame. The preview 408 may include a preview frame or a preview video. For example, the preview 408 may be provided on a two-dimensional surface in the VR space. Optionally, in 510, video data related to the update (e.g., the preview) may be retrieved. The video controller 104 may determine to initiate video playback automatically or based on subsequent user input. For example, the gesture indicating a position 410 may be followed by a diving gesture 312 to initiate video playback based on the gesture indicating a position or a subsequent diving position.
[0094] Figure 6 An example of an editable video representation 600 provided by an embodiment of the present invention is shown. The video representation 600 may include video representations 601 (including portions 601-1 and 601-2), 602, 603 (603-1 and 603-2), and 604 (604-1 and 604-2). However, it should be understood that, generally, the video representation 600 may include representations of portions of one or more video clips, files, or video data. Figure 6 As shown, video representations 601 and 602 can intersect in the VR space. Typically, the representation of media data can include a first representation of first media data and a second representation of second media data. The first representation and the second representation can intersect in the VR space.
[0095] Such as combination Figure 5 As discussed, user input can be detected in operation 509. In response to detecting the user input, an operation associated with the user input can be performed. For example, the operation may include updating the video representation in 516. For example, it may enable the user to edit the video representation.
[0096] exist Figure 6In one example, the video controller 104 may receive sensor data from the VR headset 102, at least one sensor associated with the VR headset 102, or the user, or a VR controller. Based on the received sensor data, the video controller 104 may detect user input at location 611, for example, the user input may be located at the intersection of video representations 601 and 602. The user input may be associated with an operation of cutting the video representation at location 611. According to one embodiment, the user input may include a predetermined gesture, such as a line-drawing gesture. In response to detecting the gesture, the video controller 104 may divide the video representation 601 into a first portion 601-1 and a second portion 601-2. Subsequently, the video controller 104 may detect user input at location 621, the user input may be located in the second portion 601-2. The user input may be associated with an operation of deleting an associated portion of the video representation. According to one embodiment, the user input may include a predetermined gesture, such as a wiping gesture.
[0097] The video controller 104 can also detect user input associated with cut operations at positions 613 and 614 corresponding to the video representations 603 and 604. Accordingly, the video representations 603 and 604 can be divided into a first portion (603-1, 604-1) and a second portion (603-2, 604-2). The video controller 104 can also detect one or more user inputs, such as wiping gestures, associated with deleting the first portion (603-1, 604-1) of the video representation.
[0098] The video controller 104 can also detect one or more user inputs associated with the operation of repositioning the second portion 603-2. According to one embodiment, the user input may include a predetermined gesture, such as a dragging gesture along trajectory 605. For example, the dragging gesture can be distinguished from the wiping gesture based on a movement speed below a threshold, finger position, and / or finger movement. According to one embodiment, the user input may also include a predetermined gesture, such as a drop gesture at position 632. For example, the drop gesture can be detected based on finger movement during the dragging gesture. Therefore, the second portion 603-2 can be repositioned at position 632. A similar repositioning operation can be performed on the second portion 604-2 of the video representation 604, which can be repositioned to position 642.
[0099] Figure 7An example of an edited video representation 700 provided by an embodiment of the present invention is shown. The video representation 700 can be obtained based on the user input and operations associated with the video representation 600. In step 516, in response to detecting the user input, the video controller 104 can update the video representation accordingly. The edited video representation 700 may include the first portion 601-1 of the video representation 601, the second video representation 602, and the repositioned second portions 603-2 and 604-2 of the video representations 603 and 604.
[0100] Figure 6 and Figure 7 Examples are shown that enable a user to interact with the video representation. For example, the user can combine different video clips to display desired video content to the user during video playback. Various embodiments also enable the creation of different storytelling options. For example, when playing a video clip associated with the video representation 602 and reaching the video frame corresponding to position 642, the user can be provided with the option to continue with the video representation 602 or proceed to the second portion 604-2 of the video representation 604.
[0101] Figure 8 An example of analyzing video representation or metadata provided by an embodiment of the present invention is illustrated. The figure shows different operations performed by the VR headset 102, the video controller 104, and the video server 108. However, it should be understood that these operations can be performed by any suitable device or function of the video system. Furthermore, similar representations can be provided for other types of media (e.g., audio or image data) or multimedia content.
[0102] In step 801, similar to operation 501, the video controller 104 can be used to retrieve a set of video data and / or metadata. For example, retrieving the video data and / or metadata may include receiving the data from local storage or requesting the video data and / or metadata from the video server 108. The video controller 104 can be used to: in step 802, generate a video representation; and in step 803, analyze the video representation. Alternatively or additionally, the video controller 104 can be used to: in step 803, analyze the metadata; and in step 806, generate the video representation. Therefore, the analysis in step 803 may be based on the metadata and / or the video representation.
[0103] In 807, similar to operation 506, the video controller 104 can send the video representation data to the VR headset 102.
[0104] In 808, similar to operation 507, the VR headset 102 can display the video representation.
[0105] Figure 9 An example of analyzing video representations associated with different users, provided by an embodiment of the present invention, is illustrated. The video representation 900 may include video representations 901, 902, and 903, respectively, associated with users 911, 912, and 913. For example, video representation 901 may have been captured by user 911, or user 911 may be associated with a video representation in other ways, such as by existing within the video data represented by video representation 901. The starting position of the video data may correspond to a point in time, such as 1990. Similarly, users 912 and 913 may be associated with video representations 902 and 903, respectively, whose starting positions correspond to 1993 and 1994.
[0106] In step 803, the video controller 104 can analyze the video representation 900. For example, the video controller 104 can determine that video representations 901 and 902 intersect at position 921 in the VR space. In step 804, the video controller 104 can determine social network information based on the video representation, such as based on the detected intersection of video representations 901 and 902. The video controller 104 can use metadata associated with the underlying video data to provide social network connection data. For example, the social network connection data may include the association between users 911 and 912 and / or contextual metadata associated with an encounter. For example, the contextual metadata may include the time of the encounter (e.g., 1995), the location of the encounter, or other information related to the encounter.
[0107] Alternatively or additionally, in 804, the video controller 104 can determine social network information based on location information associated with the video data. In 804, this can be achieved without generating the video representation. For example, the video controller 104 can compare the location and / or time information associated with the video data represented by video representations 901 and 903 and determine that these video representations will intersect at location 922 in the VR space. The video controller 104 can also determine that the encounter occurred in 1998. Based on the detected intersection, the video controller 104 can determine the social network information as described above. Subsequently, in 806, a video representation illustrating the encounter can be generated.
[0108] Typically, the video controller 104 can be used to: determine social network information based on the representation of the media data and / or the metadata set; or receive social network information associated with the media data and determine the representation of the media data based on the social network information and the metadata set.
[0109] Figure 10 An example of a video representation 1000 determined based on social network information according to an embodiment of the present invention is shown. The video representation 1000 may include video representations 1002 and 1003 associated with user 1001, video representations 1012 and 1013 associated with user 1011, and / or video representation 1022 associated with user 1021.
[0110] As described above, in operation 801, the video controller 104 can retrieve video data and / or a set of metadata associated with the video data, such as location information. The video controller 104 can also receive social network information associated with the video data or at least a portion of the video data. For example, the social network information can be requested from the video server 108 or other servers. Alternatively, the social network information can be retrieved from local storage. The social network information may include associations between users and / or the types of relationships between users.
[0111] According to one embodiment, the video controller 104 can determine the representation of the video data based on the social network information and the metadata set. For example, the video controller 104 can receive social network information indicating that user 1001 and user 1011 have a son 1021. Based on this social network information, the video controller 104 can determine a video representation of the lives of users 1001, 1011, and 1021 as visually represented in the VR space. For example, the users may have met in 1995, and their son 1021 may have been born in the same year. Accordingly, the video representations of users 1001 and 1011 may intersect at a point in time corresponding to 1005, and the video representation of the son 1021 may begin from that point in time. It should be noted that even though the video representation can be generated based on location information included in the metadata set, various methods can be applied to avoid excessive overlap in the video representations. For example, spatial offsets or spatial difference amplification between locations indicated in the metadata set can be used to provide the user with a clearer overall video representation in the VR space.
[0112] Figure 11An example of predicting future events based on video representation provided by an embodiment of the present invention is illustrated. In 801, the video controller 104 may receive a set of contextual metadata associated with at least a portion of video data. For example, the contextual metadata may include personal information such as a user's smoking habits, dietary habits, driving habits, hobbies, or other activities, and / or speed, location, or weather associated with the video data. The contextual metadata may include information about objects associated with the video data, such as vehicle specifications. Therefore, the contextual metadata may include information about the environment in which the video data has been captured or is being captured. In 801, the video controller 104 may also receive location information, as described above.
[0113] In operation 802, the video controller 104 may generate video representations 1101 and 1111. For example, these video representations may be associated with truck 1102 and car 1112, respectively. In operation 805, the video controller 104 may determine a prediction of at least one event 1120 based on the context metadata set and / or the metadata set. In this example, the context metadata may include indications that the user is driving a car, weather conditions, or traffic conditions. The metadata set may include spatial information, such as the vehicle's speed, direction, and / or location. In this example, the predicted event may include a collision between truck 1102 and car 1112. In response to determining the prediction of the event, the video controller 104 may send a notification of the predicted event to user 1103 and / or 1113, for example, to prevent adverse events such as collisions.
[0114] The prediction of event 1120 can be determined based on the intersection of the estimated extensions (1104, 1114) of the detected video representations (1101, 1111). For example, spatial information associated with each video data can be extrapolated to determine the positions of the extensions 1104, 1114, thereby determining their intersection.
[0115] In another example, the predicted event 1120 may include an unpleasant encounter or relationship between users 1103 and 1113 (excluding vehicles 1102 and 1112). The contextual metadata may include information about users 1103 and 1113, such as their lifestyle habits, values, or demographic information (e.g., their ages). Based on the contextual metadata, the video controller 104 can determine whether the encounter or relationship between users 1103 and 1113 is likely negative or positive. The video controller 104 may then provide notifications to users 1103 and / or 1113 accordingly.
[0116] Typically, the video controller 104 can be used to: receive a set of context metadata associated with the media data; and determine a prediction of at least one event based on the set of context metadata and / or the set of metadata.
[0117] The video controller 104 or other devices or functions may include artificial intelligence (AI) for determining the prediction of event 1120. For example, the AI may be implemented as a machine learning model such as a neural network. During training, one or more video representations, along with associated spatial information and contextual metadata, may be provided to the neural network. The contextual metadata may be automatically stored during video capture; alternatively, the user may be able to manually store any relevant information. Furthermore, the user may be requested to mark certain events (e.g., encounters with specific people) as favorable or unfavorable events.
[0118] Based on the labeled training data, the neural network can be trained to predict favorable or unfavorable events. During the inference phase, one or more video representations can be provided to the network, and the network can output an estimate of the expected favorable or unfavorable event based on available video data, contextual metadata, and / or spatial information. The video controller 104 can be configured to: update the video representation using notifications of the predicted events; and send the updated video representation to the VR headset 102 for display to the user.
[0119] Based on the analysis of artificial intelligence such as neural networks, data collected from millions of users can be used to predict future events and provide life guidance to specific users in areas such as health, finance, and / or social aspects. As mentioned above, the predictions can be based on statistical data of similar users. For example, the user could be given guidance such as: "Eating this food and ensuring this amount of sleep will have such a significant impact on your lifespan" or "Marrying this kind of woman (with certain lifelong behaviors) is / is not a wise choice."
[0120] Figure 12An example of providing an extended video representation based on predictions of at least one future event, as provided by an embodiment of the present invention, is illustrated. The video representation may include a first portion 1201 determined based on available video data, a metadata set, and / or a context metadata set. A second portion 1202 may be generated based on at least one predicted event. The shape of the second portion 1202 may be determined based on an estimate of future location data. The length of the second portion may be determined based on an estimated time of the last predicted event. Alternatively, for example, the shape of the video representation may be determined randomly without estimating future location data, while the length of the video representation may be determined based on the estimated time of the last predicted event.
[0121] According to one embodiment, the length of the second portion 1202 can be determined based on an estimate of the remaining lifespan of the user associated with the first portion 1201. As described above, the remaining lifespan can be estimated by artificial intelligence, such as a neural network, based on various contextual metadata. The contextual metadata can include lifestyle data. For example, the lifestyle data can include nutrition, stress, blood analysis, exercise, illness, smoking, drinking, and / or medical-related data. By using the contextual metadata as input, the neural network can perform a life expectancy estimate. Therefore, the lifespan estimate, and thus the length of the second portion 1202, may depend on the beneficial or harmful lifestyle factors analyzed.
[0122] Figure 13 An example of a video representation 1301 overlaid on geographic content provided by an embodiment of the present invention is illustrated. The geographic content may include maps, images, and / or videos of geographic locations. The geographic content may be two-dimensional or three-dimensional. For example, the geographic content may include a perspective view comprising two-dimensional and / or three-dimensional objects. For example, the geographic content may include a three-dimensional map view.
[0123] exist Figure 13 In the example, the geographic content includes a landscape view 1302. For example, the video representation 1301 can be overlaid on or enhanced on the landscape view 1202, allowing a user to observe the capture location of the video content relative to the landscape view 1302 in VR space.
[0124] According to one exemplary embodiment, a video representation can be overlaid on a real-world view using augmented reality headphones or the like. For example, the video controller 104 can use the metadata set (e.g., real-world coordinates) to determine the position and shape of the video representation in augmented reality space. For instance, the position where the video representation is presented at the augmented reality headphones can be determined, allowing the video representation to be augmented over the real-world capture position.
[0125] Figure 14 Another example of a video representation overlaid on geographic content, as provided in an embodiment of the present invention, is illustrated. In this example, the geographic content includes a three-dimensional illustration of a globe 1401. Based on location information and video data, video representations 1402 and 1403 can be overlaid on the globe 1401. For example, the location information included in the metadata set can be mapped to corresponding coordinates in VR space, such that the video representation reflects the location where the video was captured. For example, video representations 1402 and 1403 can be associated with two intercontinental flights.
[0126] Figure 15 An example of method 1500 for preparing media data is shown.
[0127] In 1500, the method may include: receiving a set of metadata based on at least one spatial coordinate, wherein the set of metadata is associated with the media data.
[0128] In 1501, the method may include: determining a representation of the media data in the virtual reality space based on the metadata set.
[0129] Other features of the method are derived directly from the functions and parameters of the video controller 104, the VR headset 102, or (typically) the media data preparation device 200, as described in the appended claims and throughout the specification, and therefore will not be repeated here.
[0130] Various exemplary embodiments disclose methods, computer programs, and apparatuses for generating video representations in a virtual reality space and interacting with said video representations in said virtual reality space. The exemplary embodiments improve the user experience, such as when accessing or editing videos. Taking into account location information and / or contextual metadata also enables the visualization of predictions of future events in an informative manner.
[0131] A device (e.g., a mobile phone, virtual reality headset, video player, or other virtual reality-enabled device) can be used to perform or cause any aspect of one or more methods described herein to be performed. Further, a computer program may include instructions for causing the device to perform any aspect of one or more methods described herein when executed. Further, the device may include means for performing any aspect of one or more methods described herein. According to an exemplary embodiment, the means includes at least one processor and a memory including program code, which, when executed by the at least one processor, are used to cause any aspect of the one or more methods to be performed.
[0132] Any ranges or device values given herein can be extended or changed without loss of the desired effect. Furthermore, unless expressly prohibited, any embodiment can be combined with other embodiments.
[0133] Although the subject matter of the invention has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as examples of implementing the claims, and other equivalent features and actions are intended to be included within the scope of the claims.
[0134] It should be understood that the above advantages and benefits may relate to one embodiment or several embodiments. The embodiments are not limited to embodiments that solve any or all of the described problems, nor are they limited to embodiments that have any or all of the described advantages and benefits. Furthermore, it should be understood that a reference to "one" item may refer to one or more of these items.
[0135] The steps or operations of the methods described herein can be performed in any suitable order, or simultaneously where appropriate. Additionally, individual blocks can be removed from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the above embodiments can be combined with aspects of any other described embodiments to form further embodiments without loss of the desired effects.
[0136] The term “comprising” is used herein to mean including identified methods, blocks or elements, but such blocks or elements do not include an exclusive list, and methods or devices may include additional blocks or elements.
[0137] Although a topic may be referred to as the "first" or "second" topic, this does not necessarily indicate the order or importance of the topics. Rather, such attributes can be used simply to distinguish topics.
[0138] It should be understood that the above description is provided by way of example only, and various modifications can be made by those skilled in the art. The foregoing specification, examples, and data provide a complete description of the structure and application of exemplary embodiments. Although various embodiments have been described above by way of a degree of specificity or in combination with one or more individual embodiments, those skilled in the art can make various modifications to the disclosed embodiments without departing from the scope of this specification.
Claims
1. A media data preparation device (102, 104, 200) for receiving media data, characterized in that, include: At least one processor (202); At least one memory (204) includes computer program code (206). The at least one memory (204) and the computer code (206) are used to cause the media data preparation device (102, 104, 200) to perform at least the following operations via the at least one processor (202): Receive a set of metadata associated with at least one spatial coordinate associated with media data capture, the set of metadata being associated with the media data; the media data comprising a plurality of video frames, a subset of the metadata set corresponding to each of the plurality of video frames; Based on the metadata set, determine the shape and / or position of the representation (302, 402, 600, 700, 900, 1000, 1101, 1111, 1201, 1301, 1402, 1403) of the media data displayed to the user in the virtual reality space, such that the capture position associated with the media data is reflected in the shape and / or position of the representation of the media data; The media data preparation devices (102, 104, 200) are also used for: Detect user input (312, 410) at a first location of the representation (302, 402) of the media data in the virtual reality space. Perform operations associated with the media data based on the first location.
2. The media data preparation device (102, 104, 200) according to claim 1, characterized in that, Determining the representation (302, 402, 600, 700, 900, 1000, 1101, 1111, 1201, 1301, 1402, 1403) of the media data includes mapping the metadata set to at least one location in the virtual reality space.
3. The media data preparation device (102, 104, 200) according to claim 1, characterized in that, The metadata set includes at least one of location information, motion information, and time information associated with the media data.
4. The media data preparation device (102, 104, 200) according to claim 3, characterized in that, The location information includes multiple locations associated with the capture of the media data, and / or the motion information includes multiple speeds associated with the capture of the media data.
5. The media data preparation device (102, 104, 200) according to claim 3, characterized in that, The motion information includes gyroscope data and / or magnetometer data.
6. The media data preparation device (102, 104, 200) according to any one of claims 1-5, characterized in that, The representations (600, 900) of the media data include a first representation (601, 901) of the first media data and a second representation (602, 902, 903) of the second media data, wherein the first representation (601, 901) and the second representation (602, 902, 903) intersect in the virtual reality space.
7. The media data preparation device (102, 104, 200) according to claim 1, characterized in that, Detecting the user input (312, 410) includes detecting a body part or virtual reality controller of the user (308) to be consistent with the representation (302, 402) of the media data in the virtual reality space at the first location.
8. The media data preparation apparatus (102, 104, 200) according to any one of claims 1-5, 7, characterized in that, Detecting the user input includes detecting the user's (308) head to correspond to the representation (302) of the media data in the virtual reality space at the first location, wherein the position of the user's (308) head is tracked based on at least one sensor associated with the virtual reality headset.
9. The media data preparation device (102, 104, 200) according to any one of claims 1-5 and 7, characterized in that, The operation includes at least one of the following: Playback of the media data is initiated from a portion of the media data associated with the first location; Provide a preview of a portion (408) of the media data associated with the first location; Edit the representation of the media data (600).
10. The media data preparation device (102, 104, 200) according to any one of claims 1-5, 7, characterized in that, The media data preparation devices (102, 104, 200) are also used for: Based on the representation (900) of the media data and / or the metadata set, determine social network information; or Receive social network information associated with the media data, and determine the representation of the media data based on the social network information and the metadata set (1000).
11. The media data preparation device (102, 104, 200) according to any one of claims 1-5, 7, characterized in that, The media data preparation devices (102, 104, 200) are also used for: Receive a set of context metadata associated with the media data; Based on the context metadata set and / or the metadata set, a prediction for at least one event (1120) is determined.
12. A method for preparing media data, characterized in that, include: Receive a set of metadata associated with at least one spatial coordinate associated with media data capture, wherein the set of metadata is associated with the media data; the media data includes a plurality of video frames, and a subset of the set of metadata may correspond to each of the plurality of video frames; Based on the metadata set, determine the shape and / or position of the representation (302, 402, 600, 700, 900, 1000, 1101, 1111, 1201, 1301, 1402, 1403) of the media data displayed to the user in the virtual reality space, such that the capture position associated with the media data is reflected in the shape and / or position of the representation of the media data; Detect user input (312, 410) at a first position of the representation (302, 402) of the media data in the virtual reality space. Perform operations associated with the media data based on the first location.
13. The method according to claim 12, characterized in that, The method is performed in a media data preparation device (102, 104, 200) according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The program code (206) is stored therein, which, when run on a computer, causes the computer to perform the method as described in any one of claims 12 and 13.
15. A computer program product, characterized in that, Includes program code (206), which includes computer program instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 12 and 13.
Citation Information
Patent Citations
Selecting portions of vehicle-captured video to use for display
CN108702454A
Providing virtual content based on user context
WO2019141903A1