Method, device and memory for presenting media items
By integrating display devices and input devices in a computing system, and utilizing qualitative feedback classifiers and target metadata features, media items are dynamically or pseudo-randomly selected and presented. This solves the problems of device wear and power consumption caused by user operations, achieves sporadic and dynamic selection of media content, and improves user experience.
Patent Information
- Application Number
- CN202110624286.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-18
- Filing Date
- 2021-06-04
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-06-04
AI Technical Summary
In the prior art, users need to manually operate when selecting and viewing media content, which leads to increased device wear and power consumption, and the workflow lacks sporadic and dynamic nature.
By integrating display devices and input devices in a computing system, qualitative feedback classifiers and target metadata features are used to dynamically or pseudo-randomly select and present media items, dynamically deliver media items based on user response information, and detect user interests through animation of virtual objects to achieve occasional media item delivery.
It reduces the wear and tear on the equipment and the power consumption of user operations, improves the sporadic and dynamic nature of media content selection, and enhances the user experience.
Smart Images

Figure CN113852863B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to media item delivery, and in particular to systems, devices, and methods for dynamic and / or episodic media item delivery. Background Art
[0002] First, in some cases, users manually select between groupings of tagged images or media content based on geolocation, facial recognition, events, and so on. For example, a user might select a Hawaiian vacation album and then manually select a different album or photo that includes a specific family member. This process involves multiple user inputs, which increases wear and tear on the associated input device and also consumes power. Second, in some cases, users simply select an album or event associated with a set of pre-categorized images. However, this workflow for viewing media content lacks sporadic nature. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] So that the present disclosure may be understood by those skilled in the art, a more detailed description may be had with reference to aspects of certain exemplary implementations, some of which are illustrated in the accompanying drawings.
[0004] Figure 1 is a block diagram of an exemplary operational architecture according to some implementations.
[0005] Figure 2 is a block diagram of an example controller according to some implementations.
[0006] Figure 3 is a block diagram of an exemplary electronic device according to some implementations.
[0007] Figure 4 is a block diagram of an exemplary training architecture according to some specific implementations.
[0008] Figure 5 is a block diagram of an exemplary machine learning (ML) system according to some specific implementations.
[0009] Figure 6 is a block diagram of an exemplary input data processing architecture according to some specific implementations.
[0010] Figure 7A is a block diagram of an exemplary dynamic media item delivery architecture according to some implementations.
[0011] Figure 7B Exemplary data structures for a media item repository are shown, according to some implementations.
[0012] Figure 8A is a block diagram of another exemplary dynamic media item delivery architecture according to some implementations.
[0013] Figure 8B An exemplary data structure for a user reaction history data repository is shown according to some implementations.
[0014] Figure 9 is a flowchart representation of a method of dynamic media item delivery according to some specific implementations.
[0015] Figure 10 is a block diagram of another exemplary dynamic media item delivery architecture according to some implementations.
[0016] Figures 11A to 11C A series of examples of episodic media item delivery scenarios are shown according to some implementations.
[0017] Figure 12 is a flowchart representation of a method for episodic media item delivery according to some specific implementations.
[0018] As is common practice, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, the dimensions of various features may be arbitrarily expanded or reduced for clarity. Furthermore, some of the accompanying drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used to denote similar features throughout the specification and accompanying drawings. Summary of the Invention
[0019] Various embodiments disclosed herein include devices, systems, and methods for dynamic media item delivery. According to some embodiments, the method is performed at a computing system including non-volatile memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices. The method includes: presenting a first group of media items associated with first metadata via a display device; while presenting the first group of media items, obtaining user reaction information collected by the one or more input devices; obtaining an estimated user reaction state for the first group of media items based on the user reaction information via a qualitative feedback classifier; obtaining one or more target metadata features based on the estimated user reaction state and the first metadata; obtaining a second group of media items associated with second metadata corresponding to the one or more target metadata features; and presenting the second group of media items associated with the second metadata via the display device.
[0020] Various embodiments disclosed herein include devices, systems, and methods for episodic media item delivery. According to some embodiments, the method is performed at a computing system comprising non-volatile memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices. The method includes: presenting an animation comprising a first plurality of virtual objects via the display device, wherein the first plurality of virtual objects correspond to virtual representations of a first plurality of media items, and wherein the first plurality of media items are pseudo-randomly selected from a media item repository; detecting, via the one or more input devices, a user input indicating interest in a corresponding virtual object associated with a particular media item in the first plurality of media items; and, in response to detecting the user input: obtaining a target metadata feature associated with the particular media item; selecting, from the media item repository, a second plurality of media items associated with a corresponding metadata feature corresponding to the target metadata feature; and presenting, via the display device, an animation comprising a second plurality of virtual objects, wherein the second plurality of virtual objects correspond to virtual representations of the second plurality of media items from the media item repository.
[0021] According to some specific implementations, an electronic device includes one or more displays, one or more processors, non-volatile memory, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the execution of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium has instructions stored therein that, when executed by one or more processors of the device, cause the device to perform or cause the execution of any of the methods described herein. According to some specific implementations, a device includes: one or more displays, one or more processors, non-volatile memory, and means for performing or causing the execution of any of the methods described herein.
[0022] According to some specific implementations, a computing system includes one or more processors, non-volatile memory, an interface for communicating with a display device and one or more input devices, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the execution of the operations of any of the methods described herein. According to some embodiments, a non-volatile computer-readable storage medium has instructions stored therein that, when executed by one or more processors of a computing system having an interface for communicating with a display device and one or more input devices, cause the computing system to perform or cause the execution of the operations of any of the methods described herein. According to some specific implementations, a computing system includes one or more processors, non-volatile memory, an interface for communicating with a display device and one or more input devices, and a device for performing or causing the execution of the operations of any of the methods described herein. DETAILED DESCRIPTION
[0023] Numerous details are described to provide a thorough understanding of the example implementations shown in the accompanying drawings. However, the accompanying drawings illustrate only some example aspects of the present disclosure and, therefore, should not be considered limiting. One of ordinary skill in the art will appreciate that other effective aspects and / or variations do not include all of the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits are not described in detail in order to avoid obscuring more relevant aspects of the example implementations described herein.
[0024] A physical environment refers to the physical world that people can sense and / or interact with without the help of electronic devices. A physical environment may include physical features, such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and the like. In the case of an XR system, a subset of a person's physical movements, or a representation thereof, is tracked, and in response, one or more features of one or more virtual objects simulated in the XR system are adjusted in a manner that conforms to at least one law of physics. For example, an XR system may detect head movement and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. As another example, an XR system can detect movement of an electronic device (e.g., a mobile phone, tablet, laptop, etc.) presenting an XR environment and, in response, adjust the graphical content and sound field presented to a person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust features of the graphical content in an XR environment in response to representations of physical movement (e.g., voice commands).
[0025] There are many different types of electronic systems that enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Instead of an opaque display, a head-mounted system may have a transparent or translucent display. The transparent or translucent display may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some embodiments, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology that projects graphic images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as a hologram or onto a physical surface.
[0026] Figure 1 1 is a block diagram of an exemplary operating architecture 100 according to some implementations. Although relevant features are shown, those skilled in the art will appreciate from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring more relevant aspects of the exemplary implementations disclosed herein. To this end, as a non-limiting example, the operating architecture 100 includes an optional controller 110 and an electronic device 120 (e.g., a tablet computer, mobile phone, laptop computer, near-eye system, wearable computing device, etc.).
[0027] In some implementations, the controller 110 is configured to manage and coordinate the XR experience (sometimes referred to herein as an "XR environment" or "virtual environment" or "graphical environment") of the user 150 and zero or more other users. In some implementations, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Figure 2The controller 110 is described in more detail. In some implementations, the controller 110 is a computing device that is located locally or remotely relative to the physical environment associated with the user 150. For example, the controller 110 is a local server located within the physical environment. As another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the physical environment. In some implementations, the controller 110 is communicatively coupled to the electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE802.16x, IEEE 802.3x, etc.). In some implementations, the functionality of the controller 110 is provided by the electronic device 120. Thus, in some implementations, components of the controller 110 are integrated into the electronic device 120.
[0028] In some implementations, the electronic device 120 is configured to present audio and / or video content to the user 150. In some implementations, the electronic device 120 is configured to present a user interface (UI) and / or an XR environment 128 to the user 150. In some implementations, the electronic device 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 The electronic device 120 is described in more detail.
[0029] According to some implementations, the electronic device 120 presents an XR experience to the user 150 while the user 150 is physically present within the physical environment. Thus, in some implementations, the user 150 holds the electronic device 120 in one or both of his / her hands. In some implementations, when presenting the XR experience, the electronic device 120 is configured to present the XR content on the display 122 and implement video pass-through of the physical environment. For example, the XR environment 128 including the XR content is stereoscopic or three-dimensional (3D).
[0030] In one example, the XR content corresponds to display-locked content, such that the XR content remains displayed at the same location on the display 122 despite translational and / or rotational motion of the electronic device 120. In another example, the XR content corresponds to world-locked content, such that the XR content remains displayed at its original location when translational and / or rotational motion is detected by the electronic device 120. Thus, in this example, if the field of view (FOV) of the electronic device 120 does not include the original location, the XR environment 128 will not include the XR content.
[0031] In some implementations, the display 122 corresponds to an additive display that enables optical see-through of the physical environment. For example, the display 122 corresponds to a transparent lens, and the electronic device 120 corresponds to a pair of glasses worn by the user 150. Thus, in some implementations, the electronic device 120 presents a user interface by projecting XR content onto the additive display, which is then superimposed on the physical environment from the perspective of the user 150. In some implementations, the electronic device 120 presents a user interface by displaying XR content on the additive display, which is then superimposed on the physical environment 105 from the perspective of the user 150.
[0032] In some implementations, the user 150 wears the electronic device 120, such as a near-eye system. Thus, the electronic device 120 includes one or more displays (e.g., a single display or one display per eye) provided to display XR content. For example, the electronic device 120 encompasses the FOV of the user 150. In such implementations, the electronic device 120 presents the XR environment 128 by displaying data corresponding to the XR environment 128 on one or more displays or by projecting data corresponding to the XR environment 128 onto the retina of the user 150.
[0033] In some implementations, the electronic device 120 includes an integrated display (e.g., a built-in display) that displays the XR environment 128. In some implementations, the electronic device 120 includes a head-mounted housing. In various implementations, the head-mounted housing includes an attachment area to which another device with a display can be attached. For example, in some implementations, the electronic device 120 can be attached to the head-mounted housing. In various implementations, the head-mounted housing is shaped to form a receiver for receiving another device including a display (e.g., the electronic device 120). For example, in some implementations, the electronic device 120 slides / snaps into the head-mounted housing or is otherwise attached to the head-mounted housing. In some implementations, the display of the device attached to the head-mounted housing presents (e.g., displays) the XR environment 128. In some implementations, the electronic device 120 is replaced with an XR room, housing, or room configured to present XR content, in which the user 150 does not wear the electronic device 120.
[0034] In some implementations, the controller 110 and / or the electronic device 120 causes the XR representation of the user 150 to move within the XR environment 128 based on movement information (e.g., body pose data, eye tracking data, hand / limb tracking data, etc.) from the electronic device 120 and / or optional remote input devices within the physical environment. In some implementations, the optional remote input devices correspond to fixed or movable sensory equipment within the physical environment (e.g., image sensors, depth sensors, infrared (IR) sensors, event cameras, microphones, etc.). In some implementations, each remote input device is configured to collect / capture input data while the user 150 is physically within the physical environment and provide the input data to the controller 110 and / or the electronic device 120. In some implementations, the remote input device includes a microphone, and the input data includes audio data (e.g., a voice sample) associated with the user 150. In some implementations, the remote input device includes an image sensor (e.g., a camera), and the input data includes an image of the user 150. In some implementations, the input data represents the body pose of the user 150 at different times. In some implementations, the input data represents the head pose of the user 150 at different times. In some implementations, the input data represents hand tracking information associated with the hands of the user 150 at different times. In some implementations, the input data represents the velocity and / or acceleration of a body part of the user 150 (such as his / her hands). In some implementations, the input data indicates the joint position and / or joint orientation of the user 150. In some implementations, the remote input device includes a feedback device, such as a speaker, a light, etc.
[0035] Figure 2is a block diagram of an example of a controller 110 according to some implementations. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some implementations, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0036] In some implementations, the one or more communication buses 204 include circuits that interconnect system components and control communications between system components. In some implementations, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a trackpad, a touch screen, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0037] The memory 220 includes a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some implementations, the memory 220 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. The memory 220 includes a non-transitory computer-readable storage medium. In some implementations, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the data described below with reference to Figure 2 The following programs, modules and data structures or their subsets are described.
[0038] The operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks.
[0039] In some implementations, the data acquirer 242 is configured to acquire data (e.g., captured image frames of the physical environment, presentation data, input data, user interaction data, camera pose tracking information, eye tracking information, head / body pose tracking information, hand / limb tracking information, sensor data, position data, etc.) from at least one of the I / O device 206 of the controller 110, the electronic device 120, and an optional remote input device. To this end, in various implementations, the data acquirer 242 includes instructions and / or logic for these instructions, as well as heuristics and metadata for the heuristics.
[0040] In some implementations, the mapper and locator engine 244 is configured to map the physical environment and track the orientation / position of at least the electronic device 120 relative to the physical environment. To this end, in various implementations, the mapper and locator engine 244 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.
[0041] In some implementations, the data transmitter 246 is configured to transmit data (e.g., presentation data such as rendered image frames associated with the XR environment, position data, etc.) to at least the electronic device 120. To this end, in various implementations, the data transmitter 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.
[0042] In some implementations, the training architecture 400 is configured to train various parts of the qualitative feedback classifier 420. Figure 4 The training architecture 400 is described in more detail. To this end, in various implementations, the training architecture 400 includes instructions and / or logic components for instructions, as well as heuristics and metadata for the heuristics. In some implementations, the training architecture 400 includes a training engine 410, a qualitative feedback classifier 420, and a comparison engine 430.
[0043] In some implementations, the training engine 410 includes a training data set 412 and an adjustment engine 414. According to some implementations, the training data set 412 includes input representation vectors and known user reaction state pairs. For example, the corresponding input representation vector is associated with user reaction information, which includes crowdsourced, user-specific and / or system-generated intrinsic user feedback measurements. In this example, the intrinsic user feedback measurements may include at least one of body posture features, voice features, pupil dilation values, heart rate values, respiration rate values, blood glucose values, blood oxygen saturation values, etc. Continuing with this example, the known user reaction state corresponds to a possible user reaction (e.g., emotional state, mood, etc.) to the corresponding input representation vector.
[0044] Therefore, during training, the training engine 410 feeds the corresponding input representation vectors from the training dataset 412 to the qualitative feedback classifier 420. In some implementations, the qualitative feedback classifier 420 is configured to process the corresponding input representation vectors from the training dataset 412 and output an estimated user reaction state. In some implementations, the qualitative feedback classifier 420 corresponds to a search engine or a machine learning (ML) system, such as a neural network, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), a state vector machine (SVM), a random forest algorithm, etc.
[0045] In some implementations, the comparison engine 430 is configured to compare the estimated user reaction state with the known user reaction state and output an error delta value. To this end, in various implementations, the comparison engine 430 includes instructions and / or logic components for these instructions as well as a heuristic and metadata for the heuristic.
[0046] In some specific implementations, the adjustment engine 414 is configured to determine whether the error delta value satisfies a threshold convergence value. If the error delta value does not satisfy the threshold convergence value, the adjustment engine 414 is configured to adjust one or more operating parameters (e.g., filter weights, etc.) of the qualitative feedback classifier 420. If the error delta value satisfies the threshold convergence value, the qualitative feedback classifier 420 is considered to be trained and ready for use at runtime. In addition, if the error delta value satisfies the threshold convergence value, the adjustment engine 414 is configured to abandon adjusting one or more operating parameters of the qualitative feedback classifier 420. To this end, in various specific implementations, the adjustment engine 414 includes instructions and / or logical components for these instructions as well as heuristics and metadata for the heuristics.
[0047] Although the training engine 410, qualitative feedback classifier 420, and comparison engine 430 are shown as residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the training engine 410, qualitative feedback classifier 420, and comparison engine 430 may be located in separate computing devices.
[0048] In some implementations, the dynamic media item delivery architecture 700 / 800 / 1000 is configured to deliver the media item in a dynamic manner based on user reactions to the media item and / or user interest indications. Figure 7A 、 Figure 8A and Figure 10Describe exemplary dynamic media item delivery architecture 700,800 and 1000 in more detail.For this reason, in various specific implementations, dynamic media item delivery architecture 700 / 800 / 1000 comprises instruction and / or the logical component and heuristic method and the metadata for heuristic method for instruction.In some specific implementations, dynamic media item delivery architecture 700 / 800 / 1000 comprises content manager 710, media item repository 750, gesture determiner 722, renderer 724, synthesizer 726, audio / video (A / V) presenter 728, input data ingestor 615, qualitative feedback classifier 652 through training, optional user interest determiner 654 and optional user reaction history data repository 810.
[0049] In some specific implementations, such as Figure 7A and Figure 8A As shown, the content manager 710 is configured to select a first set of media items from the media item repository 750 based on an initial user selection, etc. In some implementations, such as Figure 7A and Figure 8A As shown, the content manager 710 is further configured to select a second set of media items from the media item repository 750 based on the estimated user reaction states and / or user interest indications for the first set of media items.
[0050] In some specific implementations, such as Figure 10 As shown, the content manager 710 is configured to randomly or pseudo-randomly select a first set of media items from the media item repository 750. In some implementations, as shown in FIG. Figure 10 As shown, the content manager 710 is further configured to select a second set of media items from the media item repository 750 based on the user interest indication.
[0051] References below Figure 7A 、 Figure 8A and Figure 10 The content manager 710 and the media item selection process are described in greater detail. To this end, in various implementations, the content manager 710 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.
[0052] In some implementations, the media item repository 750 includes a plurality of media items such as audio / video (A / V) content and / or a plurality of virtual / XR objects, projects, scenes, etc. In some implementations, the media item repository 750 is stored locally and / or remotely relative to the controller 110. In some implementations, the media item repository 750 is pre-populated or manually authored by the user 150. Figure 7B The media item repository 750 is described in more detail.
[0053] In some implementations, the pose determiner 722 is configured to determine the current camera pose of the electronic device 120 and / or user 150 relative to the A / V content and / or virtual / XR content. To this end, in various implementations, the pose determiner 722 includes instructions and / or logic for these instructions, as well as heuristics and metadata for the heuristics.
[0054] In some implementations, the renderer 724 is configured to render A / V content and / or virtual / XR content from the media item repository 750 according to the current camera pose relative to the content. To this end, in various implementations, the renderer 724 includes instructions and / or logic for such instructions, as well as heuristics and metadata for such heuristics.
[0055] In some implementations, the compositor 726 is configured to composite the rendered A / V content and / or virtual / XR content with an image of the physical environment to produce a rendered image frame. In some implementations, the compositor 726 acquires (e.g., receives, retrieves, determines / generates, or otherwise accesses) a scene (e.g., Figure 1 In one embodiment, the compositor 726 includes depth information (e.g., a point cloud, a mesh, etc.) associated with the physical environment in the image to maintain z-axis order between the rendered A / V content and / or virtual / XR content and the physical objects in the physical environment. To this end, in various implementations, the compositor 726 includes instructions and / or logic for these instructions, as well as heuristics and metadata for the heuristics.
[0056] In some implementations, the A / V renderer 728 is configured to present or cause the rendered image frame to be presented (e.g., via one or more displays 312, etc.). To this end, in various implementations, the A / V renderer 728 includes instructions and / or logic for such instructions, as well as heuristics and metadata for such heuristics.
[0057] In some implementations, the input data ingestor 615 is configured to ingest user input data, such as user response information and / or one or more positive user feedback inputs collected by one or more input devices. According to some implementations, the one or more input devices include at least one of an eye tracking engine, a body posture tracking engine, a heart rate monitor, a respiratory rate monitor, a blood glucose monitor, a blood oxygen saturation monitor, a microphone, an image sensor, a body posture tracking engine, a head posture tracking engine, a limb / hand tracking engine, etc. Figure 6 The input data ingestor 615 is described in greater detail. To this end, in various implementations, the input data ingestor 615 includes instructions and / or logic for such instructions as well as heuristics and metadata for such heuristics.
[0058] In some implementations, the trained qualitative feedback classifier 652 is configured to generate an estimated user reaction state (or a confidence score associated therewith) for the first set of media items or the second set of media items based on the user reaction information (or a user representation vector derived therefrom). Figure 6 、 Figure 7A and Figure 8A The trained qualitative feedback classifier 652 is described in greater detail. To this end, in various implementations, the trained qualitative feedback classifier 652 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.
[0059] In some implementations, the user interest determiner 654 is configured to generate a user interest indication based on one or more positive user feedback inputs. Figure 6 、 Figure 7A 、 Figure 8A and Figure 10 The user interest determiner 654 is described in greater detail. To this end, in various implementations, the user interest determiner 654 includes instructions and / or logic for such instructions as well as heuristics and metadata for the heuristics.
[0060] In some implementations, the optional user reaction history data store 810 includes a history of past media items presented to the user 150, associated with an estimated user reaction state of the user 150 with respect to those past media items. In some implementations, the optional user reaction history data store 810 is stored locally and / or remotely relative to the controller 110. In some implementations, the optional user reaction history data store 810 is populated over time by monitoring the reactions of the user 150. For example, the user reaction history data store 810 is populated after detecting an opt-in input from the user 150. Figure 8A and Figure 8B The optional user response history data store 810 is described in more detail.
[0061] Although the data collector 242, mapper and locator engine 244, data transmitter 246, training architecture 400, and dynamic media item delivery architecture 700 / 800 / 1000 are shown as residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data collector 242, mapper and locator engine 244, data transmitter 246, training architecture 400, and dynamic media item delivery architecture 700 / 800 / 1000 may be located in separate computing devices.
[0062] In some implementations, the functions and / or components of the controller 110 are similar to those described below. Figure 3The electronic device 120 shown in FIG. 1 is combined with or provided by the electronic device 120 shown in FIG. Figure 2 It serves more as a functional description of various features present in a particular implementation than as a schematic diagram of the architecture of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0063] Figure 3 is a block diagram of an example of an electronic device 120 (e.g., a mobile phone, a tablet computer, a laptop computer, a near-eye system, a wearable computing device, etc.) according to some implementations. While some specific features are shown, those skilled in the art will appreciate from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, IEEE802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more displays 312, an image capture device 370 (one or more optional internal-facing and / or external-facing image sensors), memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0064] In some implementations, the one or more communication buses 304 include circuits that interconnect and control communications between system components. In some implementations, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetometer, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, an oxygen saturation monitor, a glucose monitor, etc.), one or more microphones, one or more speakers, a haptic engine, a heating and / or cooling unit, a skin shear engine, one or more depth sensors (e.g., structured light, time of flight, LiDAR, etc.), a positioning and mapping engine, an eye tracking engine, a body / head pose tracking engine, a hand / limb tracking engine, a camera pose tracking engine, etc.
[0065] In some implementations, one or more displays 312 are configured to present an XR environment to the user. In some implementations, one or more displays 312 are also configured to present flat video content to the user (e.g., a two-dimensional or "flat" AVI, FLV, WMV, MOV, MP4, etc. file associated with a TV series or movie, or a real-time video pass-through of a physical environment). In some implementations, one or more displays 312 correspond to a touch screen display. In some implementations, one or more displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some implementations, one or more displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, holographic, etc. For example, the electronic device 120 includes a single display. As another example, the electronic device 120 includes a display for each eye of the user. In some implementations, the one or more displays 312 can present AR and VR content. In some implementations, the one or more displays 312 can present AR or VR content.
[0066] In some implementations, the image capture device 370 corresponds to one or more RGB cameras (e.g., with a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), an IR image sensor, an event-based camera, etc. In some implementations, the image capture device 370 includes a lens assembly, a photodiode, and a front-end architecture.
[0067] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some implementations, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located away from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage media. In some implementations, memory 320 or a non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR rendering engine 340.
[0068] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, the presentation engine 340 is configured to present media items and / or XR content to a user via one or more displays 312. To this end, in various implementations, the presentation engine 340 includes a data acquirer 342, a renderer 344, an interaction processor 346, and a data transmitter 350.
[0069] In some implementations, the data acquirer 342 is configured to acquire data (e.g., presentation data, such as rendered image frames associated with a user interface / XR environment, input data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, sensor data, position data, etc.) from at least one of the I / O devices and sensors 306 of the electronic device 120, the controller 110, and a remote input device. To this end, in various implementations, the data acquirer 342 includes instructions and / or logic for these instructions, as well as heuristics and metadata for the heuristics.
[0070] In some implementations, the renderer 344 is configured to present and update media items and / or XR content (e.g., rendered image frames associated with a user interface / XR environment) via the one or more displays 312. To this end, in various implementations, the renderer 344 includes instructions and / or logic for such instructions, as well as heuristics and metadata for such heuristics.
[0071] In some implementations, the interaction processor 346 is configured to detect user interactions with the rendered media items and / or XR content. To this end, in various implementations, the interaction processor 346 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.
[0072] In some implementations, the data transmitter 350 is configured to transmit data (e.g., presentation data, position data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, etc.) to at least the controller 110. To this end, in various implementations, the data transmitter 350 includes instructions and / or logic for these instructions, as well as heuristics and metadata for the heuristics.
[0073] Although the data acquirer 342, renderer 344, interaction processor 346, and data transmitter 350 are illustrated as residing on a single device (e.g., electronic device 120), it should be understood that in other embodiments, any combination of the data acquirer 342, renderer 344, interaction processor 346, and data transmitter 350 may be located in separate computing devices.
[0074] also, Figure 3 It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0075] Figure 4 is a block diagram of an exemplary training architecture 400 according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the exemplary implementations disclosed herein. To this end, as a non-limiting example, the training architecture 400 includes a computing system such as Figure 1 and Figure 2 Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof.
[0076] In some implementations, the training architecture 400 (e.g., a training implementation) includes a training engine 410, a qualitative feedback classifier 420, and a comparison engine 430. In some implementations, the training engine 210 includes at least a training data set 412 and an adjustment unit 414. In some implementations, the qualitative feedback classifier 420 includes at least a machine learning (ML) system, such as Figure 5To this end, in some implementations, the qualitative feedback classifier 420 corresponds to a neural network, a CNN, a RNN, a DNN, a SVM, a random forest algorithm, or the like.
[0077] In some implementations, in the training mode, the training architecture 400 is configured to train the qualitative feedback classifier 420 based at least in part on the training dataset 412. Figure 4 As shown, the training data set 412 includes input representation vectors and known user reaction state pairs. Figure 4 , input representation vector 442A corresponds to possible known user reaction state 444A, and input representation vector 442N corresponds to possible known user reaction state 444N. Those skilled in the art will appreciate that the structure of training dataset 412 and the components therein may differ in various other implementations.
[0078] According to some specific implementations, the input representation vector 442A includes crowdsourced, user-specific and / or system-generated intrinsic user feedback measurements. In this example, the intrinsic user feedback measurements may include at least one of body posture features, voice features, pupil dilation values, heart rate values, respiration rate values, blood glucose values, blood oxygen saturation values, and the like. In other words, the intrinsic user feedback measurements include sensor information such as audio data, physiological data, body posture data, eye tracking data, and the like. As a non-limiting example, a set of sensor information associated with a user's known reaction state corresponding to a happy state (e.g., intrinsic user feedback measurements) includes: audio data indicating voice features of a slow speech cadence, physiological data including a heart rate of 90 beats per minute (BPM), a pupil diameter of 3.0 mm, body posture data of the user's arms being outstretched, and / or eye tracking data of a gaze focused on a particular object. As another non-limiting example, a set of sensor information associated with a user's known state corresponding to a stress state (e.g., intrinsic user feedback measurements) includes: audio data indicating speech characteristics associated with a staccato speech pattern, physiological data including a heart rate of 120 BPM, a pupil dilation diameter of 7.00 mm, body posture data of the user's arms being crossed, and / or eye tracking data indicating a shifting gaze. For another example, a set of sensor information associated with a user's known state corresponding to a calm state (e.g., intrinsic user feedback measurements) includes: audio data including a transcript of the user saying "I'm relaxed," audio data indicating a slow speech pattern, physiological data including a heart rate of 80 BPM, a pupil dilation diameter of 4.0 mm, body posture data of the user's arms being crossed behind their head, and / or eye tracking data indicating a relaxed gaze.
[0079] Therefore, during training, the training engine 410 feeds the corresponding input representation vector 413 from the training dataset 412 to the qualitative feedback classifier 420. In some implementations, the qualitative feedback classifier 420 processes the corresponding input representation vector 413 from the training dataset 412 and outputs an estimated user reaction state 421.
[0080] In some embodiments, the comparison engine 430 compares the estimated user reaction state 421 with the known user reaction state 411 from the training data set 412 associated with the corresponding input representation vector 413 to generate an error delta value 431 between the estimated user reaction state 421 and the known user reaction state 411.
[0081] In some specific implementations, the adjustment engine 414 determines whether the error delta value 431 meets the threshold convergence value. If the error delta value 431 does not meet the threshold convergence value, the adjustment engine 414 adjusts one or more operating parameters 433 (e.g., filter weights, etc.) of the qualitative feedback classifier 420. If the error delta value 431 meets the threshold convergence value, the qualitative feedback classifier 420 is considered to be trained and ready for use at runtime. In addition, if the error delta value 431 meets the threshold convergence value, the adjustment engine 414 abandons adjusting one or more operating parameters 433 of the qualitative feedback classifier 420. In some specific implementations, the threshold convergence value corresponds to a predefined value. In some specific implementations, the threshold convergence value corresponds to a deterministic value.
[0082] Although the training engine 410, qualitative feedback classifier 420, and comparison engine 430 are shown as residing on a single device (e.g., training architecture 400), it should be understood that in other embodiments, any combination of the training engine 410, qualitative feedback classifier 420, and comparison engine 430 may be located in separate computing devices.
[0083] also, Figure 4 It serves more as a functional description of various features that may be present in a particular implementation rather than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 4 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0084] Figure 5is a block diagram of an exemplary machine learning (ML) system 500 according to some implementations. While certain specific features are shown, those skilled in the art will appreciate from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some implementations, the ML system 500 includes an input layer 520, a first hidden layer 522, a second hidden layer 524, and an output layer 526. While the ML system 500 is illustrated as including two hidden layers, those skilled in the art will appreciate from this disclosure that, in various embodiments, one or more additional hidden layers may be present. Adding additional hidden layers increases computational complexity and memory requirements, but may improve performance in certain applications.
[0085] In various implementations, the input layer 520 is coupled (eg, configured) to receive the input representation vector 502 (eg, Figure 4 The input representation vector 422A shown below refers to Figure 6 The features and components of the exemplary input representation vector 660 are described in more detail. For example, the input layer 520 receives the input representation engine (e.g., Figure 6 The input characterization engine 640 (or associated data buffer 644) shown receives the input characterization vector 502. In various implementations, the input layer 520 includes a plurality of long short-term memory (LSTM) logic units 520a, etc., which are also referred to as models of neurons by those skilled in the art. In some such implementations, the input matrix from the feature unit to the LSTM logic unit 520a comprises a rectangular matrix. For example, the size of this matrix is a function of the number of features included in the feature stream.
[0086] In some implementations, the first hidden layer 522 includes a plurality of LSTM logic units 522a, etc. Figure 5 As shown in the example of , the first hidden layer 522 receives its input from the input layer 520. For example, the first hidden layer 522 performs one or more of the following operations: a convolution operation, a nonlinear operation, a normalization operation, a pooling operation, etc.
[0087] In some implementations, the second hidden layer 524 includes a plurality of LSTM logic units 524a, etc. In some embodiments, the number of LSTM logic units 524a is the same as or similar to the number of LSTM logic units 520a in the input layer 320 or the number of LSTM logic units 522a in the first hidden layer 522. Figure 5As shown in the example of , the second hidden layer 524 receives its input from the first hidden layer 522. Additionally and / or alternatively, in some implementations, the second hidden layer 524 receives its input from the input layer 520. For example, the second hidden layer 524 performs one or more of the following operations: a convolution operation, a nonlinear operation, a normalization operation, a pooling operation, etc.
[0088] In some implementations, the output layer 526 includes a plurality of LSTM logic cells 526a, etc. In some implementations, the number of LSTM logic cells 526a is the same as or similar to the number of LSTM logic cells 520a in the input layer 520, the number of LSTM logic cells 522a in the first hidden layer 522, or the number of LSTM logic cells 524a in the second hidden layer 524. In some implementations, the output layer 526 is a task-related layer that performs computer vision-related tasks such as feature extraction, object recognition, object detection, pose estimation, etc. In some embodiments, the output layer 526 includes an implementation of a polynomial logic function (e.g., a softmax function) that generates an estimated user reaction state 530.
[0089] Those skilled in the art will understand that Figure 5 The LSTM logic unit shown can be replaced with various other ML components. In addition, one of ordinary skill in the art will appreciate that in other implementations, the ML system 500 can be structured or designed in a variety of ways to ingest the input representation vector 502 and output the estimated user reaction state 530.
[0090] also, Figure 5 It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 5 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0091] Figure 6 is a block diagram of an exemplary input data processing architecture 600 according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the exemplary implementations disclosed herein. To this end, as a non-limiting example, the input data processing architecture 600 is included in a computing system such as Figure 1 and Figure 2 Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof.
[0092] like Figure 6 As shown, after or simultaneously with presenting the first group of media items, the input data processing architecture 600 (e.g., a runtime implementation) obtains input data associated with multiple modalities (sometimes also referred to herein as "sensor data" or "sensor information"), including audio data 602A, physiological measurement values 602B (e.g., heart rate values, respiratory rate values, blood glucose values, blood oxygen saturation values, etc.), body posture data 602C (e.g., body language information, joint position information, hand / limb position information, head tilt information, etc.), and eye tracking data 602D (e.g., pupil dilation values, gaze direction, etc.).
[0093] For example, audio data 602A corresponds to audio signals captured by one or more microphones of controller 110, electronic device 120, and / or an optional remote input device. For example, physiological measurements 602B correspond to information captured by one or more sensors of electronic device 120 and / or one or more wearable sensors on the body of user 150 communicatively coupled to controller 110 and / or electronic device 120. For example, body gesture data 602C corresponds to data captured by one or more image sensors of controller 110, electronic device 120, and / or an optional remote input device. For another example, body gesture data 602C corresponds to data obtained from one or more wearable sensors on the body of user 150 communicatively coupled to controller 110 and / or electronic device 120. For example, eye tracking data 602D corresponds to images captured by one or more image sensors of controller 110, electronic device 120, and / or an optional remote input device.
[0094] According to some implementations, the audio data 602A corresponds to a continuous or sequential time series of values. Then, the time series converter 610 is configured to generate one or more time frames of audio data from the continuous audio data stream. Each time frame of audio data includes a time portion of the audio data 602A. In some implementations, the time series converter 610 includes a multi-window processing module 610A that is configured to generate a time series of values for each of the time frames T1, T2, ..., T N One or more time frames or portions of the audio data 602A are marked and separated.
[0095] In some implementations, each time frame of audio data 602A is conditioned by a prefilter (not shown). For example, in some implementations, prefiltering includes bandpass filtering to isolate and / or emphasize portions of the frequency spectrum typically associated with human speech. In some implementations, prefiltering includes pre-emphasizing portions of one or more time frames of audio data to adjust the spectral composition of one or more time frames of audio data 602A. Additionally and / or alternatively, in some implementations, multi-window processing module 610A is configured to retrieve audio data 602A from non-transitory memory. Additionally and / or alternatively, in some implementations, prefiltering includes filtering audio data 602A using a low-noise amplifier (LNA) to substantially set the noise floor for further processing. In some implementations, the prefiltering LNA is positioned before time-series converter 610. Those skilled in the art will appreciate that many other prefiltering techniques can be applied to audio data, and those highlighted herein are merely examples of the many prefiltering options available.
[0096] According to some implementations, the physiological measurement values 602B correspond to a continuous or continuous time series of values. Then, the time series converter 610 is configured to generate one or more time frames of physiological measurement data from the continuous physiological measurement data stream. Each time frame of physiological measurement data includes a time portion of the physiological measurement value 602B. In some implementations, the time series converter 410 includes a multi-window processing module 610A that is configured to generate one or more time frames of physiological measurement data for the time periods T1, T2, ..., T N One or more portions of the physiological measurement 602B are labeled and separated. In some implementations, each time frame of the physiological measurement 602B is conditioned by a pre-filter or otherwise pre-processed.
[0097] According to some specific implementations, the body posture data 602C corresponds to a continuous or continuous time series of images or values. Then, the time series converter 610 is configured to generate one or more time frames of body posture data from the continuous body posture data stream. Each time frame of body posture data includes a time portion of the body posture data 602C. In some specific implementations, the time series converter 610 includes a multi-window processing module 610A, which is configured to generate one or more time frames of body posture data for the time T1, T2, ..., T N One or more time frames or portions of body gesture data 602C are labeled and separated. In some implementations, each time frame of body gesture data 602C is conditioned by a pre-filter or otherwise pre-processed.
[0098] According to some implementations, the eye tracking data 602D corresponds to a continuous or sequential time series of images or values. Then, the time series converter 410 is configured to generate one or more time frames of eye tracking data from the continuous stream of eye tracking data. Each time frame of eye tracking data includes a time portion of the eye tracking data 602D. In some implementations, the time series converter 610 includes a multi-window processing module 610A that is configured to generate one or more time frames of eye tracking data for each time T1, T2, ..., T N One or more time frames or portions of the eye-tracking data 602D are labeled and separated. In some implementations, each time frame of the eye-tracking data 602D is conditioned by a pre-filter or otherwise pre-processed.
[0099] In various implementations, the input data processing architecture 600 includes a privacy subsystem 620 that includes one or more privacy filters associated with user information and / or identification information (e.g., at least some portion of the audio data 602A, physiological measurements 602B, body posture data 602C, and / or eye tracking data 602D). In some implementations, the privacy subsystem 620 includes an opt-in feature where the device notifies the user which user information and / or identification information is being monitored and how the user information and / or identification information will be used. In some implementations, the privacy subsystem 620 selectively prevents and / or restricts the input data processing architecture 600, or portions thereof, from acquiring and / or transmitting user information. To this end, the privacy subsystem 620 receives user preferences and / or selections from the user in response to prompting the user for the user preferences and / or selections. In some implementations, the privacy subsystem 620 prevents the data processing architecture 600 from acquiring and / or transmitting user information unless and until the privacy subsystem 620 obtains informed consent from the user. In some implementations, the privacy subsystem 620 anonymizes (e.g., scrambles, obfuscates, encrypts, etc.) certain types of user information. For example, the privacy subsystem 620 receives user input specifying which types of user information the privacy subsystem 620 anonymizes. For another example, the privacy subsystem 620 independently specifies (e.g., automatically) that certain types of user information, which may include sensitive and / or identifying information, be anonymized.
[0100] In some implementations, a natural language processor (NLP) 622 is configured to perform natural language processing (or another speech recognition technique) on the audio data 602A or one or more time frames thereof. For example, the NLP 622 includes a processing model (e.g., a hidden Markov model, a dynamic time warping algorithm, etc.) or a machine learning node (e.g., a CNN, RNN, DNN, SVM, a random forest algorithm, etc.) that performs speech-to-text (STT) processing. In some implementations, a trained qualitative feedback classifier 652 uses the text output from the NLP 622 to help determine an estimated user reaction state 672.
[0101] In some implementations, the speech evaluator 624 is configured to determine one or more speech features associated with the audio data 602A (or one or more time frames thereof). For example, the one or more speech features correspond to pitch, rhythm, accent, diction, articulation, pronunciation, etc. For example, the speech evaluator 624 performs speech segmentation on the audio data 602A to separate the audio data 602A into words, syllables, phonemes, etc., and then determines one or more speech features therefor. In some implementations, the trained qualitative feedback classifier 652 uses the one or more speech features output by the speech evaluator 624 to help determine the estimated user reaction state 672.
[0102] In some implementations, the biometric data evaluator 626 is configured to evaluate physiological and / or biometric data from the user to determine one or more physiological measurements associated with the user. For example, the one or more physiological measurements may correspond to heart rate information, respiratory rate information, blood pressure information, pupil dilation information, glucose levels, blood oxygen saturation levels, and the like. For example, the biometric data evaluator 626 performs segmentation on the physiological measurements 602B to decompose the physiological measurements 602B into pupil dilation values, heart rate values, respiratory rate values, blood glucose values, blood oxygen saturation values, and the like. In some implementations, the trained qualitative feedback classifier 652 uses the one or more physiological measurements output by the biometric data evaluator 626 to help determine the estimated user reaction state 672.
[0103] In some implementations, the body pose interpreter 628 is configured to determine one or more pose features associated with the body pose data 602C (or one or more time frames thereof). For example, the body pose interpreter 628 determines the user's overall pose (e.g., sitting, standing, crouching, etc.) for each sampling period (e.g., each image within the body pose data 602C) or a predefined set of sampling periods (e.g., every N images within the body pose data 602C). For example, the body pose interpreter 628 determines the rotational and / or translational coordinates of each joint, limb, and / or body part of the user for each sampling period (e.g., each image within the body pose data 602C) or a predefined set of sampling periods (e.g., every N images or every M seconds within the body pose data 602C). For example, the body pose interpreter 628 determines the rotational and / or translational coordinates of a particular body part (e.g., head, hand, etc.) for each sampling period (e.g., each image within the body pose data 602C) or a predefined set of sampling periods (e.g., every N images or every M seconds within the body pose data 602C). In some implementations, the trained qualitative feedback classifier 652 uses one or more pose features output by the body pose interpreter 628 to help determine the estimated user reaction state 672.
[0104] In some implementations, the gaze direction determiner 630 is configured to determine a directional vector associated with the eye-tracking data 602D (or one or more time frames thereof). For example, the gaze direction determiner 630 determines a directional vector (e.g., X, Y, and / or focus coordinates) for each sampling period (e.g., each image within the eye-tracking data 602D) or a predefined set of sampling periods (e.g., every N images or every M seconds within the eye-tracking data 602D). In some implementations, the user interest determiner 654 uses the directional vector output by the gaze direction determiner 630 to help determine the user interest indication 674.
[0105] In some implementations, the input characterization engine 640 is configured to generate an input characterization engine 640 based on the outputs from the NLP 622, the speech evaluator 624, the biometric data evaluator 626, the body pose interpreter 628, and the gaze direction determiner 630. Figure 6 The input representation vector 660 is shown. Figure 6 As shown, input representation vector 660 includes speech content portion 662 corresponding to output from NLP 622. For example, speech content portion 662 may correspond to a user saying “Wow, I’m stressed out,” which may indicate a stressful state.
[0106] In some implementations, the input representation vector 660 includes a speech feature portion 664 corresponding to the output from the speech evaluator 624. For example, speech features associated with a fast speech cadence may indicate a state of nervousness. For another example, speech features associated with a slow speech cadence may indicate a state of fatigue. For another example, speech features associated with a normal speech cadence may indicate a state of concentration.
[0107] In some implementations, the input representation vector 660 includes a physiological measurement portion 666 corresponding to the output from the biological data evaluator 626. For example, physiological measurement values associated with high respiratory rate and high pupil dilation values can correspond to an arousal state. For another example, physiological measurement values associated with high blood pressure and high heart rate values can correspond to a stress state.
[0108] In some implementations, the input representation vector 660 includes a body posture feature portion 668 corresponding to the output from the body posture interpreter 628. For example, a body posture feature corresponding to a user crossing their arms across their chest can indicate an anxious state. For another example, a body posture feature corresponding to a user dancing can indicate a happy state. For another example, a body posture feature corresponding to a user crossing their arms behind their head can indicate a relaxed state.
[0109] In some implementations, the input representation vector 660 includes a gaze direction portion 670 corresponding to the output from the gaze direction determiner 630. For example, the gaze direction portion 670 corresponds to a vector indicating where the user is looking. In some implementations, the input representation vector 660 also includes one or more miscellaneous information portions 672 associated with other input modalities.
[0110] In some implementations, the input data processing architecture 600 generates an input representation vector 660 and stores the input representation vector 660 in a data buffer 644 (e.g., a non-transitory memory), which is accessible by the trained qualitative feedback classifier 652 and the user interest determiner 654. In some implementations, each portion of the input representation vector 660 is associated with a different input modality, including a speech content portion 662, a speech feature portion 664, a physiological measurement portion 666, a body posture feature portion 668, a gaze direction portion 670, a miscellaneous information portion 672, etc. One of ordinary skill in the art will appreciate that in other implementations, the input data processing architecture 600 can be structured or designed in a variety of ways to generate the input representation vector 660.
[0111] In some implementations, the trained qualitative feedback classifier 652 is configured to output an estimated user reaction state 672 (or a confidence score associated therewith) based on an input representation vector 660 that includes information derived from input data (e.g., audio data 602A, physiological measurements 602B, body pose data 602C, and eye tracking data 602D). Similarly, in some implementations, the user interest determiner 654 is configured to output a user interest indication 674 based on an input representation vector 660 that includes information derived from input data (e.g., audio data 602A, physiological measurements 602B, body pose data 602C, and eye tracking data 602D).
[0112] Although various aspects of specific implementations within the scope of the appended claims have been described above, it should be apparent that the various features of the above-described specific implementations can be embodied in a variety of forms, and any specific structures and / or functions described above are merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement an apparatus and / or a method can be practiced. In addition, in addition to or different from one or more aspects set forth herein, other structures and / or functions can be used to implement such an apparatus and / or such a method can be practiced.
[0113] also, Figure 6 It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 6 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0114] Figure 7A is a block diagram of an exemplary dynamic media item delivery architecture 700 according to some implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, the dynamic media item delivery architecture 700 includes a computing system such as Figure 1 and Figure 2Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof.
[0115] According to some implementations, the content manager 710 includes a target metadata determiner 714 and a media item selector 712 with an accompanying media item buffer 713. During runtime, the media item selector 712 obtains (e.g., receives, retrieves, or detects) an initial user selection 702. For example, the initial user selection 702 can correspond to a selection of a collection of media items (e.g., an album of images from a vacation or other event), one or more individually selected media items, a keyword or search string (e.g., Paris, rain, forest, etc.), etc.
[0116] In some implementations, the media item selector 712 obtains (e.g., receives, retrieves, etc.) a first set of media items associated with the first metadata from the media item repository 750 based on the initial user selection 702. As described above, the media item repository 750 includes a plurality of media items such as A / V content and / or a plurality of virtual / XR objects, projects, scenes, etc. In some implementations, the media item repository 750 is stored locally and / or remotely relative to the dynamic media item delivery architecture 700. In some implementations, the media item repository 750 is pre-populated or manually authored by the user 150. Figure 7B The media item repository 750 is described in more detail.
[0117] In some implementations, when the first set of media items corresponds to virtual / XR content, the pose determiner 722 determines a current camera pose of the electronic device 120 and / or the user 150 relative to the position of the first set of media items and / or the physical environment. In some implementations, when the first set of media items corresponds to virtual / XR content, the renderer 724 renders the first set of media items according to the current camera pose relative to the first set of media items. In some implementations, the pose determiner 722 updates the current camera pose in response to detecting translational and / or rotational movement of the electronic device 120 and / or the user 150.
[0118] In some implementations, when the first set of media items corresponds to virtual / XR content, the compositor 726 obtains (e.g., receives, retrieves, etc.) one or more images of the physical environment captured by the image capture device 370. Furthermore, in some implementations, the compositor 726 composites the first set of rendered media items with the one or more images of the physical environment to produce one or more rendered image frames. In some implementations, the compositor 726 obtains (e.g., receives, retrieves, determines / generates, or otherwise accesses) depth information (e.g., a point cloud, a mesh, etc.) associated with the physical environment to maintain z-order and reduce occlusion between the first set of rendered media items and physical objects in the physical environment.
[0119] In some implementations, the A / V renderer 728 renders or causes the one or more rendered image frames to be rendered (e.g., via one or more displays 312, etc.). Those skilled in the art will appreciate that the above steps may not be performed when the first set of media items corresponds to flat A / V content.
[0120] According to some embodiments, the input data ingestor 615 ingests user input data, such as user response information and / or one or more affirmative user feedback inputs collected by one or more input devices. In some embodiments, the input data ingestor 615 also processes the user input data to generate a user characterization vector 660 derived therefrom. According to some embodiments, the one or more input devices include at least one of an eye tracking engine, a body posture tracking engine, a heart rate monitor, a respiratory rate monitor, a blood glucose monitor, a blood oxygen saturation monitor, a microphone, an image sensor, a body posture tracking engine, a head posture tracking engine, a limb / hand tracking engine, etc. Reference is made to Figure 6 The input data ingestor 615 is described in more detail.
[0121] In some implementations, the qualitative feedback classifier 652 generates an estimated user reaction state 672 (or a confidence score associated therewith) for the first set of media items based on the user representation vector 660. For example, the estimated user reaction state 672 can correspond to an emotional state or mood of the user 150 in response to the first set of media items, such as happiness, sadness, excitement, stress, fear, etc.
[0122] In some implementations, the user interest determiner 654 generates a user interest indication 674 based on one or more positive user feedback inputs within the user representation vector 660. For example, the user interest indication 674 may correspond to a specific person, object, landmark, etc. that is the subject of the gaze direction that the user 150 is looking at, a pointing gesture made by the user 150, or a voice request from the user 150. For example, while viewing the first set of media items, the computing system may detect that the user 150 is focused on a specific person within the first set of media items, such as his / her spouse or children, indicating that they are interested in the same. For another example, while viewing the first set of media items, the computing system may detect a pointing gesture from the user 150 pointing to a specific object within the first set of media items, indicating that they are interested in the same. For another example, while viewing the first set of media items, the computing system may detect a voice command from the user 150 that corresponds to a selection of or interest in a specific object, person, etc. within the first set of media items.
[0123] In some implementations, the target metadata determiner 714 determines one or more target metadata features based on the estimated user reaction state 672, the user interest indication 674, and / or the first metadata associated with the first set of media items cached in the media item buffer 713. For example, if the estimated user reaction state 672 corresponds to happiness and the user interest indication 674 corresponds to being interested in a particular person, the one or more target metadata features can correspond to having a good time with the particular person.
[0124] Thus, in various implementations, the media item selector 712 obtains a second set of items associated with the one or more target metadata features from the media item repository 750. For example, the media item selector 712 selects the second set of media items from the media item repository 750 that match the one or more target metadata features. For another example, the media item selector 712 selects the second set of media items from the media item repository 750 that match the one or more target metadata features within a predefined tolerance. Thereafter, when the second set of media items corresponds to virtual / XR content, the pose determiner 722, the renderer 724, the compositor 726, and the A / V renderer 728 repeat the operations mentioned above with respect to the first set of items.
[0125] In some implementations, the second set of media items is presented in a spatially meaningful manner that takes into account the spatial context of the current physical environment and / or past physical environments (or features associated therewith) associated with the second set of media items. For example, if the first set of media items corresponds to an album of images of a user's children playing in their home, and the user is gazing at a rug, sofa, or other furniture item within the first set of media items, the computing system may present the second set of media items relative to the rug, sofa, or other furniture item within the user's current physical environment (e.g., a continuation of the album of images of the user's children playing in their home) as spatial anchors. For another example, if the first set of media items corresponds to an album of images from a day at the beach, and the user is gazing at his / her children building sandcastles within the first set of media items, the computing system may present the second set of media items relative to locations in the user's current physical environment that match at least a portion of the size, perspective, light direction, spatial features, and / or other features associated with a past physical environment that is associated with the album of images of the day at the beach within a certain degree of tolerance or confidence.
[0126] also, Figure 7A It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 7A Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0127] Figure 7B 1 and 2. Example data structures for a media item repository 750 according to some implementations are shown. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, the media item repository 750 includes a first entry 760A associated with a first media item 762A and an Nth entry 760N associated with an Nth media item 762N.
[0128] like Figure 7BAs shown, the first entry 760A includes inherent metadata 764A for the first media item 762A, such as length / running time, size (e.g., in MB, GB, etc.), resolution, format, creation date, last modification date, etc. when the first media item 762A corresponds to video and / or audio content. Figure 7B , the first entry 760A also includes contextual metadata 766A for the first media item 762A, such as a place or location associated with the first media item 762A, an event associated with the first media item 762A, one or more objects and / or landmarks associated with the first media item 762A, one or more people and / or faces associated with the first media item 762A, and the like.
[0129] Similarly, if Figure 7B As shown, the Nth entry 760N includes intrinsic metadata 764N and contextual metadata 766N for the Nth media item 762N. Those skilled in the art will appreciate that in various other implementations, the structure of the media item repository 750 and its components may be different.
[0130] Figure 8A is a block diagram of another exemplary dynamic media item delivery architecture 800 according to some implementations. To this end, as a non-limiting example, the dynamic media item delivery architecture 800 is included in a computing system such as Figure 1 and Figure 2 Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof. Figure 8A The dynamic media item delivery architecture 800 in is similar to Figure 7A Therefore, similar reference numerals are used herein and, for the sake of brevity, only the differences will be described.
[0131] like Figure 8A As shown, the first set of media items and the estimated user reaction state 672 are stored in association with the user reaction history data store 810. Thus, in some implementations, the target metadata determiner 714 determines one or more target metadata features based on the estimated user reaction state 672, the user interest indication 674, the user reaction history data store 810, and / or the first metadata associated with the first set of media items cached in the media item buffer 713.
[0132] also, Figure 8A It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 8A Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0133] Figure 8B FIG. 8 shows an exemplary data structure for the user reaction history data repository 810 according to some implementations. Figure 8B , the user reaction history data store 810 includes a first entry 820A associated with a first media item 822A and an Nth entry 820N associated with an Nth media item 822N. Figure 8B As shown, the first entry 820A includes a first media item 822A, an estimated user reaction state 824A associated with the first media item 822A, user input data 862A from which the estimated user reaction state 824A is determined, and contextual information 828A, such as time, location, environmental measurements, etc., that characterize the context when the first media item 822A is presented.
[0134] Similarly, in Figure 8B , the Nth entry 820N includes the Nth media item 822N, an estimated user reaction state 824N associated with the second media item 822N, user input data 862N from which the estimated user reaction state 824N is determined, and contextual information 828N, such as time, location, environmental measurements, etc., that characterize the context in which the Nth media item 822N is presented. Those skilled in the art will appreciate that in various other embodiments, the structure of the user reaction history data repository 810 and its components may be different.
[0135] Figure 9 is a flowchart representation of a method 900 for dynamic media item delivery according to some implementations. In various implementations, the method 900 is performed at a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices (e.g., Figure 1 and Figure 3 The electronic device 120 shown in FIG; Figure 1 and Figure 210; and / or suitable combinations thereof). In some implementations, method 900 is performed by processing logic (including hardware, firmware, software, or a combination thereof). In some implementations, method 900 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory). In some implementations, the electronic device corresponds to one of a tablet computer, a laptop computer, a mobile phone, a near-eye system, a wearable computing device, and the like.
[0136] In some cases, users manually select between groups of tagged images or media content based on geographic location, facial recognition, events, and so on. For example, a user may select a Hawaiian vacation album and then manually select a different album or photo that includes specific family members. In contrast, method 900 describes a process by which a computing system dynamically updates an image or media content stream based on user reactions to the image or media content stream (such as gaze direction, body language, heart rate, breathing rate, voice cadence, voice pitch, and so on). As an example, while viewing a media content stream (e.g., images associated with an event), the computing system dynamically changes the media content stream based on the user's reactions to the media content stream. For example, while viewing images associated with a birthday party, if the user's gaze is focused on a specific person, the computing system transitions to displaying images associated with that person. As another example, while viewing images associated with a specific place or person, if the user exhibits elevated heart and breathing rates and dilated eyes, the system may infer that the user is excited or happy and continue displaying more images associated with that place or person.
[0137] As shown in block 9-1, method 900 includes presenting a first set of media items associated with first metadata. For example, the first set of media items corresponds to an album of images, a set of videos, etc. In some implementations, the first metadata is associated with a specific event, person, location, object, landmark, etc.
[0138] For example, reference Figure 7A , the computing system or a component thereof (e.g., media item selector 712) obtains (e.g., receives, retrieves, etc.) a first set of media items associated with the first metadata from the media item repository 750 based on the initial user selection 702. Continuing with this example, when the first set of media items corresponds to virtual / XR content, the computing system or a component thereof (e.g., pose determiner 722) determines a current camera pose of the electronic device 120 and / or user 150 relative to the position of the first set of media items and / or the physical environment.
[0139] Continuing with the example, when the first set of media items corresponds to virtual / XR content, the computing system or a component thereof (e.g., the renderer 724) renders the first set of media items according to the current camera pose relative to the first set of media items. According to some implementations, the pose determiner 722 updates the current camera pose in response to detecting translational and / or rotational movement of the electronic device 120 and / or the user 150. Continuing with the example, when the first set of media items corresponds to virtual / XR content, the computing system or a component thereof (e.g., the compositor 726) obtains (e.g., receives, retrieves, etc.) one or more images of the physical environment captured by the image capture device 370.
[0140] In addition, when the first group of media items corresponds to virtual / XR content, the computing system or a component thereof (e.g., compositor 726) composites the first group of rendered media items with one or more images of the physical environment to produce one or more rendered image frames. Finally, the computing system or a component thereof (e.g., A / V renderer 728) renders or causes the one or more rendered image frames to be presented (e.g., via one or more displays 312, etc.). Those of ordinary skill in the art will appreciate that when the first group of media items corresponds to flat A / V content, the above steps may not be performed.
[0141] As shown in block 9-2, method 900 includes obtaining (e.g., receiving, retrieving, collecting / collecting, etc.) user reaction information collected by one or more input devices when presenting the first set of media items. In some implementations, the user reaction information corresponds to a user representation vector derived therefrom, the user representation vector including one or more intrinsic user feedback measurements associated with a user of the computing system, the one or more intrinsic user feedback measurements including at least one of body posture features, voice features, pupil dilation values, heart rate values, respiration rate values, blood glucose values, blood oxygen saturation values, and the like. For example, body posture features include head / hand / limb posture information, such as joint positions, and the like. For example, voice features include cadence, words per minute, pitch, and the like.
[0142] For example, reference Figure 7A , the computing system or a component thereof (e.g., input data ingestor 615) ingests user input data, such as user response information and / or one or more affirmative user feedback inputs collected by one or more input devices. Continuing with the example, the computing system or a component thereof (e.g., input data ingestor 615) also processes the user input data to generate a user characterization vector 660 derived therefrom. According to some specific implementations, the one or more input devices include at least one of an eye tracking engine, a body posture tracking engine, a heart rate monitor, a respiratory rate monitor, a blood glucose monitor, a blood oxygen saturation monitor, a microphone, an image sensor, a body posture tracking engine, a head posture tracking engine, a limb / hand tracking engine, etc. Reference above Figure 6The input data ingestor 615 and the input representation vector 660 are described in more detail.
[0143] As shown in block 9-3, method 900 includes obtaining (e.g., receiving, retrieving, or generating / determining) estimated user reaction states for the first set of media items based on user reaction information via a qualitative feedback classifier. In some implementations, the qualitative feedback classifier corresponds to a trained ML system (e.g., a neural network, CNN, RNN, DNN, SVM, random forest algorithm, etc.) that ingests a user representation vector (e.g., one or more intrinsic user feedback measurements) and outputs a user reaction state (e.g., an emotional state, mood, etc.) or a confidence score associated therewith. In some implementations, the qualitative feedback classifier corresponds to a lookup engine that maps the user representation vector (e.g., one or more intrinsic user feedback measurements) to a reaction table / matrix.
[0144] For example, reference Figure 7A , the computing system or a component thereof (e.g., the trained qualitative feedback classifier 652) generates an estimated user reaction state 672 (or a confidence score associated therewith) to the first set of media items based on the user representation vector 660. For example, the estimated user reaction state 672 may correspond to an emotional state or mood of the user 150 in response to the first set of media items, such as happiness, sadness, excitement, stress, fear, etc.
[0145] As shown in block 9-4, method 900 includes acquiring (e.g., receiving, retrieving, or generating / determining) one or more target metadata features based on the estimated user reaction state and the first metadata. In some implementations, the one or more target metadata features include at least one of a specific person, a specific place, a specific event, a specific object, or a specific landmark.
[0146] For example, reference Figure 7A , the computing system or a component thereof (e.g., target metadata determiner 714) determines one or more target metadata features based on estimated user reaction state 672, user interest indication 674, and / or first metadata associated with the first set of media items cached in media item buffer 713. For example, if estimated user reaction state 672 corresponds to happiness and user interest indication 674 corresponds to being interested in a particular person, the one or more target metadata features may correspond to having fun with the particular person.
[0147] In some implementations, method 900 includes: obtaining sensor information associated with a user of a computing system, wherein the sensor information corresponds to one or more positive user feedback inputs; and generating a user interest indication based on the one or more positive user feedback inputs, wherein one or more target metadata features are determined based on an estimated user reaction state and the user interest indication. For example, the user interest indication corresponds to one of a gaze direction, a voice command, a pointing gesture, etc. In some implementations, the one or more positive user feedback inputs correspond to one of a gaze direction, a voice command, or a pointing gesture. For example, if the estimated user reaction state 672 corresponds to happiness and the user interest indication 674 corresponds to interest in a particular person, the one or more target metadata features may correspond to having a good time with the particular person.
[0148] For example, reference Figure 7A , the computing system or a component thereof (e.g., user interest determiner 654) generates a user interest indication 674 based on one or more positive user feedback inputs within the user characterization vector 660. Continuing with this example, referring to Figure 7A , the computing system or a component thereof (e.g., a target metadata determiner 714) determines one or more target metadata features based on the estimated user reaction state 672, the user interest indication 674 and / or the first metadata associated with the first set of media items cached in the media item buffer 713.
[0149] In some implementations, method 900 includes associating the estimated user reaction state with a first set of media items in a user reaction history data repository. In some implementations, the user reaction history data repository may also be used with the user interest indication and / or the user state indication to determine one or more target metadata features. Figure 8B The user response history data store 810 is described in more detail. For example, reference Figure 8A , the computing system or a component thereof (e.g., a target metadata determiner 714) determines one or more target metadata features based on the estimated user reaction state 672, the user interest indication 674, the user reaction history data repository 810 and / or the first metadata associated with the first set of media items cached in the media item buffer 713.
[0150] As shown in block 9-5, method 900 includes obtaining (e.g., receiving, retrieving, or generating) a second set of media items associated with second metadata corresponding to one or more target metadata characteristics. Figure 7A, the computing system or a component thereof (e.g., the media item selector 712) obtains a second set of items associated with the one or more target metadata characteristics from the media item repository 750. For example, the media item selector 712 selects media items from the media item repository 750 that match the one or more target metadata characteristics. For another example, the media item selector 712 selects media items from the media item repository 750 that match the one or more target metadata characteristics within a predefined tolerance.
[0151] As shown in block 9-6, method 900 includes presenting (or causing to be presented) via a display device a second set of media items associated with the second metadata. Figure 7A , when the second set of media items corresponds to virtual / XR content, the computing system or components thereof (e.g., pose determiner 722, renderer 724, compositor 726, and A / V renderer 728) repeat the operations mentioned above with reference to box 9-1 to present or cause the presentation of the second set of media items.
[0152] In some implementations, the second set of media items is presented in a spatially meaningful manner that takes into account the spatial context of the current physical environment and / or past physical environments (or features associated therewith) associated with the second set of media items. For example, if the first set of media items corresponds to an album of images of a user's children playing in their home, and the user is gazing at a rug, sofa, or other furniture item within the first set of media items, the computing system may present the second set of media items relative to the rug, sofa, or other furniture item within the user's current physical environment (e.g., a continuation of the album of images of the user's children playing in their home) as spatial anchors. For another example, if the first set of media items corresponds to an album of images from a day at the beach, and the user is gazing at his / her children building sandcastles within the first set of media items, the computing system may present the second set of media items relative to locations in the user's current physical environment that match at least a portion of the size, perspective, light direction, spatial features, and / or other features associated with a past physical environment that is associated with the album of images of the day at the beach within a certain degree of tolerance or confidence.
[0153] In some implementations, the first group of media items and the second group of media items correspond to at least one of audio or visual content (e.g., images, video, audio, etc.). In some implementations, the first group of media items and the second group of media items are mutually exclusive. In some implementations, the first group of media items and the second group of media items include at least one overlapping media item.
[0154] In some implementations, the display device corresponds to a transparent lens assembly, and wherein the first set of media items and the second set of media items are projected onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system, and wherein presenting the first set of media items and the second set of media items includes compositing the first set of media items or the second set of media items with one or more images of a physical environment captured by an outward-facing image sensor.
[0155] Figure 10 is a block diagram of another exemplary dynamic media item delivery architecture 1000 according to some implementations. To this end, as a non-limiting example, the dynamic media item delivery architecture 1000 is included in a computing system such as Figure 1 and Figure 2 Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof. Figure 10 The dynamic media item delivery architecture 1000 in is similar to Figure 7A Dynamic Media Item Delivery Architecture 700 and Figure 8A Therefore, similar reference numerals are used herein and, for the sake of brevity, only the differences will be described.
[0156] like Figure 10 As shown, the content manager 710 includes a randomizer 1010. For example, the randomizer 1010 may correspond to a randomization algorithm, a pseudo-randomization algorithm, a random number generator that utilizes a natural source of entropy (e.g., radioactive decay, thermal noise, radio noise, etc.), etc. To this end, in some specific implementations, the media item selector 712 obtains (e.g., receives, retrieves, etc.) a first set of media items associated with the first metadata from the media item repository 750 based on a random or pseudo-random seed provided by the randomizer 1010. Thus, the content manager 710 randomly selects the first set of media items to provide the following references. Figures 11A to 11C and Figure 12 A more detailed description of the episodic user experience.
[0157] In addition, Figure 10 In some implementations, the target metadata determiner 714 determines one or more target metadata features based on the user interest indication 674 and / or the first metadata associated with the first set of media items cached in the media item buffer 713. For example, if the user interest indication 674 corresponds to an interest in a particular person, the one or more target metadata features can correspond to the particular person. Accordingly, in various implementations, the media item selector 712 obtains a second set of items associated with the one or more target metadata features from the media item repository 750.
[0158] also, Figure 10 It serves more as a functional description of various features present in a particular implementation than as a structural representation of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 10 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation and, in some specific implementations, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0159] Figures 11A to 11C A series of instances 1110, 1120, and 1130 of episodic media item delivery scenarios according to some implementations are shown. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, the series of instances 1110, 1120, and 1130 are executed by a computing system such as Figure 1 and Figure 2 Controller 110 shown; Figure 1 and Figure 3 The electronic device 120 shown; and / or suitable combinations thereof.
[0160] like Figures 11A to 11C As shown, the episodic media item delivery scenario includes a physical environment 105 and an XR environment 128 displayed on a display 122 of an electronic device 120. When a user 150 is physically present within the physical environment 105, the electronic device 120 presents the XR environment 128 to the user 150, the physical environment including the table 107 within the field of view (FOV) 111 of the external-facing image sensor of the electronic device 120. Thus, in some implementations, the user 150 holds the electronic device 120 in his / her hand, similar to Figure 1 in the operating environment 100.
[0161] In other words, in some implementations, the electronic device 120 is configured to present virtual / XR content on the display 122 and implement optical or video pass-through of at least a portion of the physical environment 105. For example, the electronic device 120 corresponds to a mobile phone, a tablet computer, a laptop computer, a near-eye system, a wearable computing device, etc.
[0162] like Figure 11AAs shown, during an instance 1110 of an episodic media item delivery scenario (e.g., associated with time T1), the electronic device 120 presents an XR environment 128 including a first plurality of virtual objects 1115 with a falling animation according to a gravity indicator 1125. Figures 11A to 11C In FIG, the first plurality of virtual objects 1115 are shown with a landing animation centered about the representation of the table 107 within the XR environment 128, but one of ordinary skill in the art will understand that the landing animation can be centered about a different point within the physical environment 105, such as about the electronic device 120 or the user 150. Furthermore, although the first plurality of virtual objects 1115 are shown in FIG. Figures 11A to 11C It is shown with a falling animation, but a person skilled in the art will understand that the falling animation can be replaced with other animations, such as a rising animation, a particle flow toward the electronic device 120 or the user 150, a particle flow away from the electronic device 120 or the user 150, etc.
[0163] exist Figure 11A In the embodiment of the present invention, the electronic device 120 displays a first plurality of virtual objects 1115 relative to or overlaid on the physical environment 105. Thus, in one example, the first plurality of virtual objects 1115 are composited with an optical perspective or a video perspective of at least a portion of the physical environment 105.
[0164] In some implementations, the first plurality of virtual objects 1115 includes virtual representations of media items having different metadata characteristics. For example, virtual representation 1122A corresponds to one or more media items associated with a first metadata characteristic (e.g., one or more images including a particular person or at least his / her face). For example, virtual representation 1122B corresponds to one or more media items associated with a second metadata characteristic (e.g., one or more images including a particular object such as a dog, cat, tree, flower, etc.). For example, virtual representation 1122C corresponds to one or more media items associated with a third metadata characteristic (e.g., one or more images associated with a particular event such as a birthday party). For example, virtual representation 1122D corresponds to one or more media items associated with a fourth metadata characteristic (e.g., one or more images associated with a particular time period such as a particular day, week, etc.). For example, virtual representation 1122E corresponds to one or more media items associated with a fifth metadata characteristic (e.g., one or more images associated with a particular location such as a city, state, etc.). For example, virtual representation 1122F corresponds to one or more media items associated with a sixth metadata feature (e.g., one or more images associated with a particular file type or format, such as a still image, a real-time image, a video, etc.). For example, virtual representation 1122G corresponds to one or more media items associated with a seventh metadata feature (e.g., one or more images associated with a particular system- or user-specified label / tag, such as a mood tag, an importance tag, etc.).
[0165] In some implementations, the first plurality of virtual objects 1115 correspond to virtual representations of a first plurality of media items, wherein the first plurality of media items are pseudo-randomly selected from Figure 7B and Figure 10 A media item repository 750 is shown.
[0166] like Figure 11B As shown, during instance 1120 of the episodic media item delivery scenario (e.g., associated with time T2), the electronic device 120 continues to present the XR environment 128 including the first plurality of virtual objects 1115 with a falling animation according to the gravity indicator 1125. Figure 11B As shown, the first plurality of virtual objects 1115 continue to “rain” on the table 107 , and a portion 1116 of the first plurality of virtual objects 1115 has accumulated on the representation of the table 107 within the XR environment 128 .
[0167] like Figure 11B As shown, the user holds the electronic device 120 with his / her right hand 150A and performs a pointing gesture within the physical environment 105 with his / her left hand 150B. Figure 11B, the electronic device 120 or a component thereof (e.g., a hand / limb tracking engine) detects a pointing gesture with the user's left hand 150B within the physical environment 105. In response to detecting the pointing gesture with the user's left hand 150B within the physical environment 105, the electronic device 120 or a component thereof displays a representation 1135 of the user's left hand 150B within the XR environment 128 and also maps the tracked location of the pointing gesture with the user's left hand 150B within the physical environment 105 to a corresponding virtual object 1122D within the XR environment 128. In some implementations, the pointing gesture indicates that the user is interested in the corresponding virtual object 1122D.
[0168] In response to detecting a pointing gesture indicating user interest in a corresponding virtual object 1122D, the computing system obtains a target metadata feature associated with the corresponding virtual object 1122D. For example, the target metadata feature corresponds to one or more of a specific event, person, location / place, object, landmark, etc., of a media item associated with the corresponding virtual object 1122D. Therefore, according to some embodiments, the computing system selects a second plurality of media items from the media item repository that are associated with a corresponding metadata feature corresponding to the target metadata feature. For example, the corresponding metadata feature and the target metadata feature match. In another example, the corresponding metadata feature and the target metadata feature are similar within a predefined tolerance threshold.
[0169] like Figure 11C As shown, during instance 1130 of the episodic media item delivery scenario (e.g., associated with time T3), in response to Figure 11B When a pointing gesture indicating that the user is interested in the corresponding virtual object 1122D is detected, the electronic device 120 presents the XR environment 128 including the second plurality of virtual objects 1140 with a falling animation according to the gravity indicator 1125. In some specific implementations, the second plurality of virtual objects 1140 include virtual representations of media items, and the media items have corresponding metadata features corresponding to the target metadata features.
[0170] Figure 12 is a flowchart representation of a method 1200 for episodic media item delivery according to some implementations. In various implementations, the method 1200 is performed at a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices (e.g., Figure 1 and Figure 3 The electronic device 120 shown in FIG; Figure 1 and Figure 210; and / or suitable combinations thereof). In some implementations, method 1200 is performed by processing logic (including hardware, firmware, software, or a combination thereof). In some implementations, method 1200 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory). In some implementations, the electronic device corresponds to one of a tablet computer, a laptop computer, a mobile phone, a near-eye system, a wearable computing device, and the like.
[0171] In some cases, current media viewing applications lack a sporadic nature. Typically, a user simply selects an album or event associated with a set of pre-categorized images. In contrast, in method 1200 described below, virtual representations of images "rain" within an XR environment, where the images are pseudo-randomly selected from the user's camera roll, etc. However, if the device detects that the user is interested in one of the virtual representations, the "pseudo-random rain" effect is changed to a virtual representation of the image corresponding to the user's interest. Thus, in order to provide a sporadic effect when viewing media, virtual representations of pseudo-randomly selected media items "rain" within the XR environment.
[0172] As shown in block 12-1, method 1200 includes presenting (or causing to be presented) an animation comprising a first plurality of virtual objects via a display device, wherein the first plurality of virtual objects correspond to virtual representations of a first plurality of media items, and wherein the first plurality of media items are pseudo-randomly selected from a media item repository. In some implementations, the media item repository includes at least one of audio or visual content (e.g., images, video, audio, etc.). For example, referring to Figure 10 , the computing system or a component thereof (e.g., the media item selector 712) obtains (e.g., receives, retrieves, etc.) a first plurality of media items from the media item repository 750 based on the random or pseudo-random seed provided by the randomizer 1010. Thus, the content manager 710 randomly selects the first set of media items to provide the above-referenced Figures 11A to 11C A more detailed description of the episodic user experience.
[0173] like Figure 11A As shown, for example, the electronic device 120 presents an XR environment 128 including a first plurality of virtual objects 1115 with a falling animation according to a gravity indicator 1125. Continuing with this example, the first plurality of virtual objects 1115 include virtual representations of media items having different metadata characteristics. For example, virtual representation 1122A corresponds to one or more media items associated with a first metadata characteristic (e.g., one or more images including a specific person or at least his / her face). For example, virtual representation 1122B corresponds to one or more media items associated with a second metadata characteristic (e.g., one or more images including a specific object such as a dog, a cat, a tree, a flower, etc.).
[0174] In some implementations, the first plurality of virtual objects correspond to three-dimensional (3D) representations of the first plurality of media items. For example, the 3D representations correspond to 3D models, 3D reconstructions, etc. of the first plurality of media items. In some implementations, the first plurality of virtual objects correspond to two-dimensional (2D) representations of the first plurality of media items.
[0175] In some implementations, the animation corresponds to a falling animation that simulates a rainfall effect centered on the computing system (e.g., rain, snow, etc.). In some implementations, the animation corresponds to a falling animation that simulates a rainfall effect offset from the computing system by a threshold distance. In some implementations, the animation corresponds to a flow of particles toward a first plurality of virtual objects of the computing system. In some implementations, the animation corresponds to a flow of particles away from the first plurality of virtual objects of the computing system. Those skilled in the art will appreciate that the above-described animation types are non-limiting examples, and that a myriad of animation types may be used in various other implementations.
[0176] As shown in block 12-2, method 1200 includes detecting, via one or more input devices, a user input indicating interest in a corresponding virtual object associated with a particular media item in the first plurality of media items. For example, the user input corresponds to one of a gaze direction, a voice command, a pointing gesture, and the like. In some implementations, the user input indicating interest in the corresponding virtual object may also be referred to herein as positive user feedback input. For example, referring to Figure 10 , the computing system or a component thereof (e.g., input data ingestor 615) ingests user input data, such as user response information and / or one or more affirmative user feedback inputs collected by one or more input devices. According to some specific implementations, the one or more input devices include at least one of an eye tracking engine, a body posture tracking engine, a heart rate monitor, a respiratory rate monitor, a blood glucose monitor, a blood oxygen saturation monitor, a microphone, an image sensor, a body posture tracking engine, a head posture tracking engine, a limb / hand tracking engine, etc. The above reference Figure 6 The input data ingestor 615 is described in more detail.
[0177] like Figure 11B, for example, the electronic device 120 or a component thereof (e.g., a hand / limb tracking engine) detects a pointing gesture with the user's left hand 150B within the physical environment 105. Continuing with this example, in response to detecting the pointing gesture with the user's left hand 150B within the physical environment 105, the electronic device 120 or a component thereof displays a representation 1135 of the user's left hand 150B within the XR environment 128 and also maps the tracked location of the pointing gesture with the user's left hand 150B within the physical environment 105 to a corresponding virtual object 1122D within the XR environment 128. In some implementations, the pointing gesture indicates that the user is interested in the corresponding virtual object 1122D.
[0178] In response to detecting the user input, as shown in block 12-3, method 1200 includes obtaining (e.g., receiving, retrieving, collecting / collecting, etc.) target metadata features associated with a particular media item. In some implementations, the one or more target metadata features include at least one of a particular person, a particular place, a particular event, a particular object, or a particular landmark. For example, referring to Figure 10 , the computing system or a component thereof (e.g., target metadata determiner 714) determines one or more target metadata features based on the user interest indication 674 (e.g., associated with the user input) and / or metadata associated with the first plurality of media items cached in the media item buffer 713.
[0179] In response to detecting the user input, as shown in block 12-4, method 1200 includes selecting a second plurality of media items from the media item repository that are associated with respective metadata features corresponding to the target metadata feature. Figure 10 The computing system or a component thereof (eg, the media item selector 712 ) obtains a second plurality of media items associated with the one or more target metadata characteristics from the media item repository 750 .
[0180] In response to detecting the user input, as shown in block 12-5, method 1200 includes presenting (or causing to be presented) via the display device an animation comprising a second plurality of virtual objects, wherein the second plurality of virtual objects correspond to virtual representations of a second plurality of media items from the media item repository. Figure 11C As shown, for example, in response to Figure 11B When a pointing gesture indicating that the user is interested in the corresponding virtual object 1122D is detected, the electronic device 120 presents the XR environment 128 including the second plurality of virtual objects 1140 with a falling animation according to the gravity indicator 1125. In some specific implementations, the second plurality of virtual objects 1140 include virtual representations of media items, and the media items have corresponding metadata features corresponding to the target metadata features.
[0181] For example, the corresponding metadata feature matches the target metadata feature. In another example, the corresponding metadata feature and the target metadata feature are similar within a predefined tolerance threshold. In some implementations, the first plurality of virtual objects and the second plurality of virtual objects are mutually exclusive. In some implementations, the first plurality of virtual objects and the second plurality of virtual objects correspond to at least one overlapping media item.
[0182] In some implementations, the display device corresponds to a transparent lens assembly, and wherein presenting the animation includes projecting an animation including the first plurality of virtual objects or the second plurality of virtual objects onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system, and wherein presenting the animation includes compositing the first plurality of virtual objects or the second plurality of virtual objects with one or more images of a physical environment captured by an outward-facing image sensor.
[0183] Although various aspects of specific implementations within the scope of the appended claims have been described above, it should be apparent that the various features of the above-described specific implementations can be embodied in a variety of forms, and any specific structures and / or functions described above are merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement an apparatus and / or a method can be practiced. In addition, in addition to or different from one or more aspects set forth herein, other structures and / or functions can be used to implement such an apparatus and / or such a method can be practiced.
[0184] It will also be understood that, although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first media item may be referred to as a second media item, and similarly, a second media item may be referred to as a first media item, which changes the meaning of the description as long as occurrences of the "first media item" are consistently renamed and occurrences of the "second media item" are consistently renamed. The first media item and the second media item are both media items, but they are not the same media item.
[0185] The terms used herein are merely for describing specific implementations and are not intended to limit the claims. As used in the description of this specific implementation and in the appended claims, the singular forms "a", "an" and "the" are intended to also cover the plural forms, unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" used herein refer to and cover any and all possible combinations of one or more of the associated listed items. It will also be understood that the term "comprising" when used in this specification specifies the presence of stated features, integers, steps, operations, elements and / or parts, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their groupings.
[0186] As used herein, the term “if” may be interpreted to mean “when the precondition is true” or “when the precondition is true” or “in response to determining” or “upon determining” or “in response to detecting” that the precondition is true, depending on the context. Similarly, the phrase “if it is determined that [the precondition is true]” or “if [the precondition is true]” or “when [the precondition is true]” is to be interpreted to mean “upon determining that the precondition is true” or “in response to determining” or “upon determining” that the precondition is true or “when detecting that the precondition is true” or “in response to detecting” that the precondition is true, depending on the context.
Claims
1. A method of presenting a media item, comprising: At a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices: presenting, via the display device, a first set of media items associated with first metadata; while presenting the first set of media items, obtaining user reaction information captured by the one or more input devices, wherein the user reaction information is associated with a user of the computing system; obtaining, via a qualitative feedback classifier, an estimated user reaction state for the first set of media items based on the user reaction information; determining one or more target metadata features based on a combination of the estimated user reaction state, the user interest indication, the user reaction history, and the first metadata for the first set of media items, wherein the target metadata features indicate that the user is interested in viewing media items related to a particular object; selecting a second group of media items associated with second metadata matching the one or more target metadata characteristics from a media item repository, wherein the second group of media items is different from the first group of media items and the second group of media items depicts the particular object indicated by the target metadata characteristics; as well as The second set of media items associated with the second metadata is presented via the display device.
2. The method of claim 1 , wherein the user reaction information corresponds to a user characterization vector, the user characterization vector comprising one or more intrinsic user feedback measurements associated with a user of the computing system, the one or more intrinsic user feedback measurements comprising at least one of a body posture feature, a voice feature, a pupil dilation value, a heart rate value, a respiration rate value, a blood glucose value, and a blood oxygen saturation value.
3. The method of claim 1, wherein the qualitative feedback classifier corresponds to a search engine, a neural network, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), a state vector machine (SVM), or a random forest algorithm.
4. The method of claim 1 , wherein the one or more input devices comprise at least one of an eye tracking engine, a body pose tracking engine, a heart rate monitor, a respiratory rate monitor, a blood glucose monitor, a blood oxygen saturation monitor, a microphone, an image sensor, a body pose tracking engine, a head pose tracking engine, or a limb / hand tracking engine.
5. The method according to claim 1, further comprising: obtaining sensor information associated with the user of the computing system, wherein the sensor information corresponds to one or more affirmative user feedback inputs; as well as The user interest indication is generated based on the one or more positive user feedback inputs, wherein the one or more target metadata features are determined based on the estimated user reaction state and the user interest indication. The method of claim 5 , wherein the one or more affirmative user feedback inputs correspond to one of a gaze direction, a voice command, or a pointing gesture.
7. The method according to claim 1, further comprising: The estimated user reaction state is associated with the first set of media items in a user reaction history data store. 8 . The method of claim 7 , wherein determining the one or more target metadata features comprises determining the one or more target metadata features based on the estimated user reaction state and the user reaction history data repository. 9 . The method according to claim 1 , wherein the specific object comprises at least one of a specific person, a specific place, a specific event, a specific object, or a specific landmark.
10. A device for presenting a media item, comprising: one or more processors; non-transitory memory; an interface for communicating with a display device and one or more input devices; and one or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the apparatus to: presenting, via the display device, a first set of media items associated with first metadata; while presenting the first set of media items, obtaining user reaction information captured by the one or more input devices, wherein the user reaction information is associated with a user of the device; obtaining, via a qualitative feedback classifier, an estimated user reaction state for the first set of media items based on the user reaction information; determining one or more target metadata features based on a combination of the estimated user reaction state, the user interest indication, the user reaction history, and the first metadata for the first set of media items, wherein the target metadata features indicate that the user is interested in viewing media items related to a particular object; selecting a second group of media items associated with second metadata matching the one or more target metadata characteristics from a media item repository, wherein the second group of media items is different from the first group of media items and the second group of media items depicts the particular object indicated by the target metadata characteristics; as well as The second set of media items associated with the second metadata is presented via the display device.
11. The device of claim 10 , wherein the user reaction information corresponds to a user characterization vector, the user characterization vector comprising one or more intrinsic user feedback measurements associated with a user of the device, the one or more intrinsic user feedback measurements comprising at least one of a body posture feature, a voice feature, a pupil dilation value, a heart rate value, a respiration rate value, a blood glucose value, and a blood oxygen saturation value.
12. The device of claim 10, wherein the one or more programs further cause the device to: obtaining sensor information associated with the user of the device, wherein the sensor information corresponds to one or more affirmative user feedback inputs; and The user interest indication is generated based on the one or more positive user feedback inputs, wherein the one or more target metadata features are determined based on the estimated user reaction state and the user interest indication.
13. The apparatus of claim 12, wherein the one or more affirmative user feedback inputs correspond to one of a gaze direction, a voice command, or a pointing gesture.
14. The device of claim 10, wherein the one or more programs further cause the device to: The estimated user reaction state is associated with the first set of media items in a user reaction history data store. 15 . The apparatus of claim 14 , wherein determining the one or more target metadata features comprises determining the one or more target metadata features based on the estimated user reaction state and the user reaction history data repository. The apparatus according to claim 10 , wherein the specific object comprises at least one of a specific person, a specific place, a specific event, a specific object, or a specific landmark.
17. A non-transitory memory storing one or more programs that, when executed by one or more processors of a device having an interface for communicating with a display device and one or more input devices, cause the device to: presenting, via the display device, a first set of media items associated with first metadata; while presenting the first set of media items, obtaining user reaction information captured by the one or more input devices, wherein the user reaction information is associated with a user of the device; obtaining, via a qualitative feedback classifier, an estimated user reaction state for the first set of media items based on the user reaction information; determining one or more target metadata features based on a combination of the estimated user reaction state, the user interest indication, the user reaction history, and the first metadata for the first set of media items, wherein the target metadata features indicate that the user is interested in viewing media items related to a particular object; selecting a second group of media items associated with second metadata matching the one or more target metadata characteristics from a media item repository, wherein the second group of media items is different from the first group of media items and the second group of media items depicts the particular object indicated by the target metadata characteristics; as well as The second set of media items associated with the second metadata is presented via the display device.
18. A non-volatile memory according to claim 17, wherein the user reaction information corresponds to a user characterization vector, the user characterization vector comprising one or more intrinsic user feedback measurement values associated with the user of the device, the one or more intrinsic user feedback measurement values comprising at least one of body posture features, voice features, pupil dilation values, heart rate values, respiration rate values, blood glucose values and blood oxygen saturation values.
19. The non-transitory memory of claim 17, wherein the one or more programs further cause the device to: obtaining sensor information associated with the user of the device, wherein the sensor information corresponds to one or more affirmative user feedback inputs; and The user interest indication is generated based on the one or more positive user feedback inputs, wherein the one or more target metadata features are determined based on the estimated user reaction state and the user interest indication.
20. The non-transitory memory of claim 19, wherein the one or more affirmative user feedback inputs correspond to one of a gaze direction, a voice command, or a pointing gesture.
21. The non-transitory memory of claim 17, wherein the one or more programs further cause the device to: The estimated user reaction state is associated with the first set of media items in a user reaction history data store.
Citation Information
Patent Citations
Hadoop cloud platform-based Web resource personalized recommendation system and method
CN106503140A
Method and system for pushing multimedia contents according to user emotions
CN108304458A