Active actions based on audio and body movement
By combining audio and body movement sensor data to identify users' interest in audio content, the problem of resource waste in existing technologies is solved, and personalized user experience and resource optimization are achieved.
Patent Information
- Application Number
- CN202210230091.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-08
- Filing Date
- 2022-03-10
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing technologies have difficulty in effectively identifying a user's interest in audio content and taking corresponding actions, resulting in a waste of electronic device resources.
By combining audio sensor and body movement sensor data, the temporal relationship between audio content and body movement is identified, the user's interest in the audio content is determined, and corresponding actions are taken based on this, such as presenting content identification, providing optional options, or adjusting device resource status.
It improves the efficiency of electronic device resource utilization, provides a personalized user experience, and enhances user interaction and access to audio content.
Smart Images

Figure CN115079816B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to electronic devices that use sensors to obtain information to understand the physical environment and provide audio and / or visual content. Background Art
[0002] Many electronic devices include microphones that capture audio from the physical environment around the device. Such audio information can be analyzed to identify songs and other sounds in the device's surroundings and provide information about such sounds. Summary of the Invention
[0003] Various embodiments disclosed herein include apparatus, systems, and methods for determining a user's interest in audio content by determining that a movement (e.g., a user's head shaking) has a time-based relationship with detected audio content (e.g., the beat of music playing in the background). Some embodiments relate to methods performed by a device having a processor that executes instructions stored on a non-transitory computer-readable medium. The method involves obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to body movement in the physical environment. The method identifies a time-based relationship between one or more elements of the audio and one or more aspects of the body movement based on the first sensor data and the second sensor data. For example, this may involve determining that a user of the device is shaking their head to the beat of music playing in the physical environment. Such head shaking may be considered a passive indication of interest in the music. In another example, the user's movement is considered an indication of interest based on the type of user movement (e.g., corresponding to excited behavior) and / or the movement occurring shortly after a significant event (e.g., a touchdown) occurs in the audio. The method identifies interest in the content of the audio based on identifying a time-based relationship. For example, this may involve determining that a particular song is playing, and determining that the user is interested in the song based on his / her movements matching the beat of the song.
[0004] Various actions can be proactively performed based on identifying interest in the content. For example, the device can present an identification of the content (e.g., displaying the name of the song, artist, etc.), present text corresponding to words in the content (e.g., lyrics), and / or present selectable options for replaying the content, continuing to experience the content after leaving the physical environment, purchasing the content, downloading the content, and / or adding the content to a playlist. In another example, characteristics of the content (e.g., musical genre, tempo range, instrument type, emotional mood, category, etc.) are identified and used to identify additional content for the user.
[0005] Device resources can be used effectively to determine that the user is interested in audio content. This can involve moving through different power states based on different triggers at the device. For example, audio analysis can be selectively performed, for example, based on detecting body movement, such as head shaking, foot tapping, jumping for joy, waving fists, facial reactions or other movements indicating user interest. Similarly, determining that music is playing can be selectively performed, for example, based on audio analysis. Comparing audio elements using various aspects of body movement can be selectively performed, for example, based on determining that body movement and music occur simultaneously in a physical environment. Identifying the source attributes of audio, for example, song title, artist, content provider, etc., can be based on successfully identifying time-based relationships. Selectively performing analysis only in appropriate circumstances can significantly contribute to the efficient use of processing, storage, power and / or communication resources of electronic devices.
[0006] According to some specific implementations, a device includes one or more processors, non-volatile memory, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the performance of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium has instructions stored therein that, when executed by one or more processors of the device, cause the device to perform or cause the performance of any of the methods described herein. According to some specific implementations, a device includes one or more processors, non-volatile memory, and means for performing or causing the performance of any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] So that the present disclosure may be understood by those skilled in the art, a more detailed description may be had with reference to aspects of certain exemplary implementations, some of which are illustrated in the accompanying drawings.
[0008] Figure 1 Example electronic devices are shown operating in a physical environment according to some implementations.
[0009] Figure 2 Shown according to some specific implementations Figure 1 An exemplary electronic device that provides a user interface enhanced with additional content based on detected body movement and audio Figure 1 view of the physical environment.
[0010] Figure 3 The embodiment disclosed herein obtains the mobile data. Figure 1 An exemplary electronic device of the present invention.
[0011] Figure 4is a flow chart illustrating a method of identifying interest in audio content by determining that movement has a time-based relationship to detected audio content, according to some implementations.
[0012] Figure 5 Based on some specific implementation Figure 1-Figure 3 Block diagram of an electronic device.
[0013] As is common practice, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, the dimensions of various features may be arbitrarily expanded or reduced for clarity. Furthermore, some of the accompanying drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used to denote similar features throughout the specification and accompanying drawings. DETAILED DESCRIPTION
[0014] Numerous details are described to provide a thorough understanding of the example implementations shown in the accompanying drawings. However, the accompanying drawings illustrate only some example aspects of the present disclosure and, therefore, should not be considered limiting. One of ordinary skill in the art will appreciate that other effective aspects and / or variations do not include all of the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits are not described in detail in order to avoid obscuring more relevant aspects of the example implementations described herein.
[0015] Figure 1 An exemplary electronic device 105 is shown being operated by a user 110 in a physical environment 100. In this example, the physical environment 100 is a room that includes people 120a-120d, wall decorations 125, 130, a wall speaker 135, and a vase 145 with flowers on a table. The electronic device 105 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to assess the physical environment 100 and the people and objects therein. Figure 1 In the example shown in FIG, wall speakers 135 are playing audio including a song within physical environment 100. In some implementations, a camera on device 105 captures one or more images of the physical environment and detects body movement (e.g., movement of user 110 and / or users 120a-120d). In some implementations, a microphone on device 105 captures audio in the physical environment, including the song playing via wall speakers 135.
[0016] Figure 2 Shown Figure 1 An exemplary electronic device 105 that provides a user interface enhanced with additional content 265 based on detected body movement and audio Figure 12. A user displays a view 200 of the physical environment 100. The view 200 includes depictions 220a-220d of the people 120a-120d; depictions 225, 230 of the wall hangings 125, 130; a depiction 245 of the flower 145; and a depiction 235 of the wall speaker 135. The electronic device 105 provides the view 200, which includes a depiction of the physical environment 100 from a viewer's position, which in this example is determined based on the position of the electronic device 105 in the physical environment 100. Thus, as the user moves the electronic device 105 relative to the physical environment 100, the viewer's position corresponding to the position of the electronic device 105 moves relative to the physical environment.
[0017] In this example, view 200 includes augmented content 265 that includes an information bubble with information and features selected based on detecting body movement and audio within physical environment 100. This can be based on electronic device 105 obtaining first sensor data corresponding to audio in the physical environment (e.g., via a microphone) and obtaining second sensor data corresponding to body movement in the physical environment 100 (e.g., via an image sensor and / or a motion sensor). The electronic device can identify a time-based relationship between one or more elements of the audio (e.g., periodically repeating elements in the audio signal corresponding to the beat / rhythm / tempo of a song) and one or more aspects of the body movement (e.g., periodically repeating positions of a part of the body) based on the first and second sensor data. For example, this can involve determining that a user of the device is moving their head to the beat of music playing in the physical environment.
[0018] Figure 3 Shown Figure 1 3. An exemplary electronic device 105 of FIG. 3 is provided, which obtains movement data (e.g., image / motion sensor data) corresponding to user 110, which corresponds to the user shaking his head back and forth between position 310 and position 320. Such head shaking can be considered as a passive indication of interest in music. In other examples, tapping one's foot to the beat of the music, walking, dancing, jumping, waving one's fist, facial reactions, and / or other movements in real time with the music are identified as indications of interest in the audio. In some examples, based on the type of user movement (e.g., corresponding to excited behavior) and / or movement shortly after a significant event occurs in the audio (e.g., jumping and cheering after a touchdown), the user movement is considered to be an indication of interest. The electronic device 105 identifies interest in the content of the audio based on identifying a time-based relationship.
[0019] The electronic device may perform one or more actions in response to identifying interest in the content of the audio. Figure 2In the example of FIG, the electronic device 105 analyzes the content to determine the name of the song (e.g., "Song Z") and the artist who performed the song (e.g., "Band Y") and displays this information within the enhanced content 265. The enhanced content 265 also includes a selectable option 270 that, if selected, enables the user 105 to download the song to the electronic device 105. Additional and / or different information and / or selectable options may be presented depending on configuration parameters and / or user preferences.
[0020] exist Figure 2 In the example of FIG2 , augmented content 265 is positioned adjacent to the depiction 135 of speaker 135 in view 200. This positioning can be based on detecting the spatial location of one or more audio sources within physical environment 100 and associating the audio content with one or more of those sound sources. Providing augmented content 265 adjacent to the sound source of the audio content to which it relates can provide an intuitive and desirable user experience.
[0021] exist Figure 1-Figure 3 In the example of , the electronic device 105 is shown as a single handheld device. The electronic device 105 can be a mobile phone, a tablet computer, a laptop computer, and the like. In some specific implementations, the electronic device 105 is worn by a user. For example, the electronic device 105 can be a watch, a helmet-mounted device (HMD), a head-mounted device (glasses), headphones, an ear-hook device, and the like. In some specific implementations, the functions of the device 105 are implemented by two or more devices, such as a mobile device and a base station or a head-mounted device and an ear-hook device. Various functions can be distributed among multiple devices, including but not limited to power functions, CPU functions, GPU functions, storage functions, memory functions, visual content display functions, audio content production functions, and the like. Multiple devices that can be used to implement the functions of the electronic device 105 can communicate with each other via wired or wireless communications.
[0022] According to some embodiments, the electronic device 105 generates and presents an extended reality (XR) environment to one or more users. In contrast to a physical environment that people can sense and / or interact with without an electronic device, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic device. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and the like. In the case of an XR system, a subset of a person's physical movements, or a representation thereof, is tracked, and in response, one or more features of one or more virtual objects simulated in the XR system are adjusted in a manner that conforms to at least one law of physics. For example, the XR system may detect head movement and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In another example, the XR system may detect movement of the electronic device (e.g., a mobile phone, tablet, laptop, etc.) that presents the XR environment and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the XR system may adjust characteristics of graphical content in the XR environment in response to representations of physical motion (e.g., voice commands).
[0023] There are many different types of electronic systems that enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Instead of an opaque display, a head-mounted system may have a transparent or translucent display. The transparent or translucent display may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some embodiments, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology that projects graphic images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as a hologram or onto a physical surface.
[0024] Figure 4 4 is a flow chart illustrating a method 400 for identifying interest in audio content by determining that movement has a time-based relationship with detected audio content, according to some implementations. In some implementations, a device such as the electronic device 105 performs the method 400. In some implementations, the method 400 is performed on a mobile device, a desktop computer, a laptop computer, an HMD, an ear-hook device, or a server device. The method 400 is performed by processing logic (including hardware, firmware, software, or a combination thereof). In some embodiments, the method 400 is performed on a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory).
[0025] At block 410, method 400 obtains first sensor data corresponding to audio in a physical environment and second sensor data corresponding to body movement in the physical environment. Examples of audio in the physical environment include, but are not limited to, songs, instrumental music, poetry, live and recorded radio broadcasts and podcasts, and / or live and recorded television program sounds. The audio can be obtained using a microphone or microphone array.
[0026] The second sensor data corresponding to body movement can correspond to the movement of the user of the device and / or one or more other persons. The body movement data can be obtained using, for example, a frame-based camera and / or using a motion sensor such as an accelerometer or a gyroscope. In some implementations, the user wears one or more electronic devices (e.g., a head-mounted device, a watch, a bracelet, an anklet, a wristband, jewelry, clothing, etc.) that include motion and / or position sensors that track the movement of one or more parts of the user's body over time.
[0027] At box 420, the method 400 identifies a time-based relationship between one or more elements of the audio and one or more aspects of the body movement based on the first sensor data and the second sensor data. Identifying a time-based relationship can involve determining that the timing of the repetitive body movements of the body movement matches the timing of the beat or rhythm of the content. For example, this can involve determining that the user is tapping his feet and / or shaking his head to the beat of the music. Such movements with timing corresponding to elements of the audio can be a passive indication of interest in music. Identifying a time-based relationship can involve matching the beats per minute of the song with the motion per minute of the movement, for example, within an error threshold. There may be metadata for the song that specifies the location of the beats within the song (which can be obtained via a network connection).
[0028] In another example, identifying a time-based relationship involves determining that a user is moving their lips (with or without producing a sound) in a manner corresponding to words in the content of the audio, e.g., lip syncing with current lyrics of music. In another example, identifying a time-based relationship involves determining that body movement is a reaction to a significant event identified in the content. For example, this may involve determining that a user movement indicating excitement occurs shortly after a significant event occurs in the audio (e.g., a touchdown).
[0029] At block 430, the method 400 identifies interest in the audio content based on identifying a time-based relationship. For example, this may involve determining that a particular song is playing and determining that the user is interested in the song based on the user's movements matching the beat of the song. In another example, this may involve determining that the user is interested in a football game on television based on determining that the user's body movements correspond to celebratory actions following an event in the football game (e.g., an identified touchdown, an increase in crowd noise in the television audio content, etc.).
[0030] In some specific implementations, additional and / or alternative information is used to identify and / or confirm interest in the content of the audio. In one example, identifying interest in the content is further based on determining that sounds in the physical environment (e.g., the device user) are singing along with the content. In another example, identifying interest in the content is further based on identifying gaze direction. For example, this can involve determining that the user is looking at the speaker that is playing the content or determining that the user is looking at a specific direction (e.g., up and left) that the system uses to indicate interest. In another example, identifying interest in the content is further based on recognized facial expressions, such as an expression of concentration combined with head movement that matches the beat. In another example, the user's voice is detected and used to identify and / or confirm that the user is interested in the content, such as based on detecting an expression of interest or verbal content corresponding to the content of the audio.
[0031] At block 440, based on identifying interest in the audio content, method 400 performs an action, such as presenting additional content. This may involve presenting an identifier for the content on the electronic device (e.g., displaying the song's title, artist, etc.). This may involve presenting text corresponding to words in the content. For example, the current playback position of a song may be determined, and real-time lyrics for the currently playing song may be displayed, e.g., to provide a karaoke mode. In some implementations, performing the action may include presenting one or more selectable options for: replaying the content, continuing to experience the content after leaving the physical environment, purchasing the content, downloading the content, and / or adding the content to a playlist. In other implementations, performing the action may include automatically replaying the content, continuing to experience the content after leaving the physical environment, purchasing the content, downloading the content, and / or adding the content to a playlist without presenting the user with one or more selectable options. For example, based on identifying interest in the content, characteristics of the content (e.g., musical genre, tempo range, instrument type, emotional context, category, etc.) may be identified and used to identify additional content based on the identified characteristics. For example, based on detecting interest in a certain rhythm, other songs or audio content having a similar rhythm may be suggested or identified to the user.
[0032] The identified interests can be used to enhance the user experience. For example, a speaker on an electronic device can play a song for the user at a louder volume than would be possible in the physical environment. The identified interests can be used to add songs to the user's collection, add metadata to the user's recommendation profile, recommend playlists or other audio content to the user, generate new playlists for the user, and otherwise provide a more customized user experience for the user's current interests.
[0033] The identified interests can be used to promote an improved user experience outside of the physical environment, for example, after the user leaves the current physical environment. For example, the next time the user accesses his music library, one or more suggestions for adding new content may be provided based on the identified interests. In some cases, it may be desirable to wait until a condition is met before providing information and / or selectable options based on the identified interest in the content of the audio. Thus, for example, method 400 may determine to limit the provision of features associated with the content based on determining the user state. Method 400 may wait to provide features when the user is focused, "engaged", driving, or otherwise busy. Method 400 may wait until the next time the user opens his music to identify / recommend songs, and thus avoid disturbing the user with notifications at inopportune moments.
[0034] Device resources can be used effectively to determine that the user is interested in audio content. This can involve moving through different power states based on different triggers at the device. For example, audio analysis can be selectively performed, for example, based on detecting body movement, such as head shaking, foot tapping, jumping for joy, waving fists, facial reactions or other movements indicating user interest. Similarly, determining that music is playing can be selectively performed, for example, based on audio analysis. Utilizing various aspects of body movement to compare audio elements can be selectively performed, for example, based on determining that body movement and music occur simultaneously in a physical environment. Identifying the source attributes of audio, such as song title, artist, content provider, etc., can be based on successfully identifying time-based relationships. Selectively performing analysis only in appropriate circumstances can significantly contribute to the efficient use of processing, storage, power and / or communication resources of electronic devices.
[0035] In some implementations, indications of interest from people are combined, for example, to provide information about a group's collective interest, such as determining how much the audience enjoys the music being played by a DJ or band. Thus, method 400 can involve identifying that multiple people in a physical environment are interested in the content of an audio based on identifying time-based relationships using audio and body movement sensor data from the multiple people.
[0036] Figure 5is a block diagram of electronic device 500. Device 500 illustrates an exemplary device configuration for electronic device 105. While some specific features are shown, those skilled in the art will appreciate from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some implementations, the device 500 includes one or more processing units 502 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 506, one or more communication interfaces 508 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 510, one or more output devices 512, one or more internal-facing and / or external-facing image sensor systems 514, memory 520, and one or more communication buses 504 for interconnecting these and various other components.
[0037] In some implementations, the one or more communication buses 504 include circuits that interconnect and control communications between system components. In some implementations, the one or more I / O devices and sensors 506 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time of flight, etc.), and / or the like.
[0038] In some implementations, the one or more output devices 512 include one or more displays configured to present a view of a 3D environment to a user. In some implementations, the one or more displays 512 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some implementations, the one or more displays correspond to waveguide displays such as diffraction, reflective, polarization, holographic, etc. In one example, the device 500 includes a single display. As another example, the device 500 includes a display for each eye of the user.
[0039] In some embodiments, the one or more output devices 512 include one or more audio generating devices. In some embodiments, the one or more output devices 512 include one or more speakers, surround sound speakers, speaker arrays, or headphones for producing spatialized sound, such as 3D audio effects. Such devices can virtually place sound sources in a 3D environment, including behind, above, or below one or more listeners. Generating spatialized sound can involve transforming sound waves (e.g., using head-related transfer functions (HRTFs), reverberation, or cancellation techniques) to simulate natural sound waves (including reflections from walls and floors) that emanate from one or more points in the 3D environment. Spatialized sound can trick the listener's brain into interpreting the sound as if it were occurring at one or more points in the 3D environment (e.g., from one or more specific sound sources), even though the actual sound may be produced by speakers in other locations. The one or more output devices 512 can additionally or alternatively be configured to generate tactile sensations.
[0040] In some implementations, the one or more image sensor systems 514 are configured to acquire image data corresponding to at least a portion of the physical environment 100. For example, the one or more image sensor systems 514 can include one or more RGB cameras (e.g., having a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), a monochrome camera, an IR camera, a depth camera, an event-based camera, etc. In various implementations, the one or more image sensor systems 514 also include an illumination source that emits light, such as a flash. In various implementations, the one or more image sensor systems 514 also include an on-camera image signal processor (ISP) that is configured to perform a plurality of processing operations on the image data.
[0041] Memory 520 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some implementations, memory 520 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 520 optionally includes one or more storage devices remotely located from one or more processing units 502. Memory 520 includes non-transitory computer-readable storage media.
[0042] In some implementations, the memory 520 or a non-transitory computer-readable storage medium of the memory 520 stores an optional operating system 530 and one or more instruction sets 540. The operating system 530 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, the instruction set 540 includes executable software defined by binary information stored in the form of electrical charge. In some implementations, the instruction set 540 is software that can be executed by one or more processing units 502 to implement one or more of the techniques described herein.
[0043] Instruction set 540 includes: a movement detection instruction set 542, which is configured to detect body movement as described herein when executed; an audio analysis instruction set 544, which is configured to analyze audio as described herein when executed; and a rendering instruction set 546, which is configured to present information and / or selectable options as described herein when executed. The rendering instruction set 546 can be configured to provide content, such as a view and / or sound of an XR environment. In some embodiments, the rendering instruction set 546 is executed to determine how to present the content based on the viewer's position. In some embodiments, the augmentation is overlaid on a 2D view, such as a video pass-through or optical see-through of the physical environment. In some embodiments, the augmentation is an assigned 3D position that corresponds to a position adjacent to a corresponding object in the physical environment. The instruction set 540 can be embodied as a single software executable file or multiple software executable files.
[0044] Although the instruction set 540 is shown as residing on a single device, it should be understood that in other implementations, any combination of elements may be located on separate computing devices. Figure 5 It serves more as a functional description of various features present in a particular implementation that are different from the structural diagrams of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately can be combined, and some items can be separated. The actual number of instruction sets and how the features are distributed among them will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.
[0045] It will be understood that the specific implementations described above are cited by way of example, and that the present disclosure is not limited to what has been particularly shown and described above. Rather, the scope includes both combinations and subcombinations of the various features described above, as well as variations and modifications of the various features that would occur to a person skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.
[0046] As described above, one aspect of the present technology is to collect and use sensor data, which may include user data, to improve the user experience of electronic devices. The present disclosure contemplates that, in some cases, the collected data may include personal information data that uniquely identifies a particular person or can be used to identify the interests, characteristics, or tendencies of a particular person. Such personal information data may include movement data, physiological data, demographic data, location-based data, phone number, email address, home address, device characteristics of a personal device, or any other personal information.
[0047] This disclosure recognizes that the use of such personal information data within the present technology can be used to benefit users. For example, personal information data can be used to improve the content viewing experience. Thus, the use of such personal information data may enable planned control of electronic devices. Furthermore, this disclosure contemplates other uses of personal information data that benefit users.
[0048] The present disclosure also contemplates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other uses of such personal information and / or physiological data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for the legal and reasonable purposes of the entity and not shared or sold outside of these legal purposes. In addition, such collection should only be carried out after the user's informed consent. In addition, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others who can access personal information data comply with their privacy policies and procedures. In addition, such entities can subject themselves to third-party assessments to prove that they comply with widely accepted privacy policies and practices.
[0049] Regardless of the foregoing, the present disclosure also contemplates specific implementations in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates the provision of hardware or software components to prevent or block access to such personal information data. For example, with respect to a content delivery service customized for a user, the technology of the present invention may be configured to allow the user to opt-in or opt-out of participating in the collection of personal information data during service registration. In another example, the user may choose not to provide personal information data to the target content delivery service. In yet another example, the user may choose not to provide personal information but allow the transmission of anonymous information for use in improving the functionality of the device.
[0050] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may be implemented without access to such personal information data. That is, various embodiments of the present technology will not fail to function due to the absence of all or a portion of such personal information data. For example, content may be selected and delivered to a user by inferring preferences or settings based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with the user, other non-personal information available to a content delivery service, or publicly available information.
[0051] In some embodiments, data is stored using a public / private key system that only allows the owner of the data to decrypt the stored data. In some other implementations, data can be stored anonymously (e.g., without identifying and / or personal information about the user, such as legal name, username, time and location data, etc.). This makes it impossible for other users, hackers, or third parties to determine the identity of the user associated with the stored data. In some implementations, a user can access their stored data from a user device that is different from the user device used to upload the stored data. In these cases, the user may need to provide login credentials to access their stored data.
[0052] Numerous specific details are set forth herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter may be practiced without these specific details. In other instances, methods, devices, or systems known to those skilled in the art are not described in detail in order not to obscure the claimed subject matter.
[0053] Unless otherwise specifically noted, it should be understood that throughout this specification, discussions utilizing terms such as "process," "compute," "calculate," "determine," and "identify" refer to the actions or processes of a computing device, such as one or more computers or similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, register, or other information storage device, transmission device, or display device of a computing platform.
[0054] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device that implements one or more specific implementations of the subject matter of the present invention. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in software used to program or configure a computing device.
[0055] The specific implementation of the method disclosed herein can be performed in the operation of such a computing device. The order of the blocks presented in the above examples can be changed, for example, the blocks can be reordered, combined and / or divided into sub-blocks. Some blocks or processes can be executed in parallel.
[0056] The use of "suitable for" or "configured to" herein is intended to be open and inclusive language, and does not exclude devices that are adapted or configured to perform additional tasks or steps. Furthermore, the use of "based on" is intended to be open and inclusive, as a process, step, calculation, or other action that is "based on" one or more stated conditions or values may, in practice, be based on additional conditions or values beyond those stated. The headings, lists, and numbers included herein are for ease of explanation only and are not intended to be limiting.
[0057] It will also be understood that, although the terms "first," "second," etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are simply used to distinguish one element from another. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description as long as all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. A first node and a second node are both nodes, but they are not the same node.
[0058] The terms used herein are merely for describing specific implementations and are not intended to limit the claims. As used in the description of this specific implementation and the appended claims, the singular forms "a" and "the" are intended to also cover the plural forms, unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" used herein refer to and cover any and all possible combinations of one or more of the associated listed items. It will also be understood that the term "comprising" when used in this specification specifies the presence of stated features, integers, steps, operations, elements and / or parts, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their groupings.
[0059] As used herein, the term “if” may be interpreted to mean “when the precondition is true” or “when the precondition is true” or “in response to determining” or “upon determining” or “in response to detecting” that the precondition is true, depending on the context. Similarly, the phrase “if it is determined that [the precondition is true]” or “if [the precondition is true]” or “when [the precondition is true]” is to be interpreted to mean “upon determining that the precondition is true” or “in response to determining” or “upon determining” that the precondition is true or “when detecting that the precondition is true” or “in response to detecting” that the precondition is true, depending on the context.
[0060] The foregoing description and summary of the present invention should be understood to be illustrative and exemplary in every respect and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the exemplary embodiments, but by the full breadth allowed by the patent law. It should be understood that the embodiments shown and described herein are only illustrative of the principles of the invention, and that various modifications can be implemented by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. A method for identifying interest in audio content, comprising: At an electronic device having a processor: obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to the audio in the physical environment and the second sensor data corresponding to body movement in the physical environment; Selectively switching to a different power state on the electronic device based on a plurality of triggers identified at the electronic device, wherein selectively switching to the different power state comprises: detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music; selectively triggering, based on detecting that the body movement corresponds to the type of body movement indicating user interest in music, performing audio analysis to determine whether music is playing based on the first sensor data; determining that music is playing based on the audio analysis; selectively triggering, based on determining that music is playing, performing a comparison of an audio element with one or more aspects of the body movement based on the first sensor data and the second sensor data; identifying, via the comparing, a time-based relationship between one or more elements of the audio and one or more aspects of the body movement; selectively triggering execution of a source attribute identification process based on identifying the time-based relationship; and identifying interest in the content of the audio based on the source attribute identification process; determining to wait to provide one or more features based on a user status corresponding to the user being busy; and One or more features are provided based on identifying interest in the content, wherein timing of providing the one or more features includes waiting based on the user status corresponding to the user being busy. 2 . The method of claim 1 , further comprising presenting an identification of the content on the electronic device based on identifying interest in the content.
3. The method of claim 1 further comprising presenting text corresponding to words in the content based on identifying interest in the content.
4. The method of claim 1 , further comprising presenting selectable options for the following operations based on identifying interest in the content: replaying said content; continuing to experience the content after leaving the physical environment; Purchase said content; download said content; or Add the content to a playlist. 5 . The method of claim 1 , further comprising, based on identifying interest in the content, identifying characteristics of the content and identifying additional content based on the identified characteristics. The method of claim 1 , further comprising determining to limit provision of features associated with the content based on determining the user status. The method of claim 1 , wherein identifying interest in the content is further based on determining that sounds in the physical environment are singing along with the content.
8. The method of claim 1, wherein identifying interest in the content is further based on an identified gaze direction corresponding to a device generating the audio, or a direction relative to the user's head corresponding to a psychological state indicative of user interest.
9. The method of claim 1, wherein identifying interest in the content is further based on a recognized facial expression.
10. The method of claim 1, wherein identifying the time-based relationship further comprises determining that the timing of repetitive ones of the body movements matches the timing of a beat or rhythm of the content.
11. The method of claim 1 , wherein identifying the time-based relationship comprises determining that lip movements of the body movement match words of the content, and identifying interest in the content of the audio is based on identifying from the time-based relationship that the user is singing along or lip syncing with words of the content.
12. The method of claim 1, wherein identifying the time-based relationship further comprises determining that the body movement is a reaction to an event in the content.
13. The method of claim 1 , wherein the second sensor data corresponding to the body movement comprises: Image data in which a body part moves over time; or Motion sensor data from a motion sensor attached to the part of the body.
14. The method of claim 1, further comprising identifying that a plurality of persons in the physical environment are interested in the content of the audio based on identifying time-based relationships using audio and body movement sensor data from the plurality of persons in the physical environment.
15. A system for identifying interesting audio content, comprising: non-transitory computer-readable storage medium; as well as one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the system to perform operations including: obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to the audio in the physical environment and the second sensor data corresponding to body movement in the physical environment; Selectively switching to a different power state on the electronic device based on a plurality of triggers identified at the electronic device, wherein selectively switching to the different power state comprises: detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music; selectively triggering, based on detecting that the body movement corresponds to the type of body movement indicating user interest in music, performing audio analysis to determine whether music is playing based on the first sensor data; determining that music is playing based on the audio analysis; selectively triggering, based on determining that music is playing, performing a comparison of an audio element with one or more aspects of the body movement based on the first sensor data and the second sensor data; identifying, via the comparing, a time-based relationship between one or more elements of the audio and one or more aspects of the body movement; selectively triggering execution of a source attribute identification process based on identifying the time-based relationship; and identifying interest in the content of the audio based on the source attribute identification process; determining to wait to provide one or more features based on a user status corresponding to the user being busy; and One or more features are provided based on identifying interest in the content, wherein timing of providing the one or more features includes waiting based on the user status corresponding to the user being busy.
16. The system of claim 15, wherein the operations further comprise: Based on identifying interest in the content: presenting an identification of the content on the electronic device; presenting text corresponding to words in the content; replaying said content; continuing to experience the content after leaving the physical environment; Purchase said content; download said content; or Add the content to a playlist. 17 . The system of claim 15 , wherein identifying interest in the content is further based on an identified gaze direction detected using an image obtained via an image sensor.
18. A non-transitory computer-readable storage medium storing program instructions, the program instructions being executable via one or more processors to perform operations comprising: obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to body movement in the physical environment; Selectively switching to a different power state on the electronic device based on a plurality of triggers identified at the electronic device, wherein selectively switching to the different power state comprises: detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music; selectively triggering, based on detecting that the body movement corresponds to the type of body movement indicating user interest in music, performing audio analysis to determine whether music is playing based on the first sensor data; determining that music is playing based on the audio analysis; selectively triggering, based on determining that music is playing, performing a comparison of an audio element with one or more aspects of the body movement based on the first sensor data and the second sensor data; identifying, via the comparing, a time-based relationship between one or more elements of the audio and one or more aspects of the body movement; selectively triggering execution of a source attribute identification process based on identifying the time-based relationship; and identifying interest in content of the audio based on the source attribute identification process; determining to wait to provide one or more features based on a user status corresponding to the user being busy; and One or more features are provided based on identifying interest in the content, wherein timing of providing the one or more features includes waiting based on the user status corresponding to the user being busy.
Citation Information
Patent Citations
Systems and methods for creation of a listening log and music library
CN107111642A
Dynamically adapting provision of notification output to reduce user distraction and / or mitigate usage of computational resources
US10368333B2
Ascertaining audience reactions for a media item
US10897647B1
Systems and methods for camera and microphone-based device
US20200296521A1