Enhanced audio in live photos and video

The integration of a computer vision module and generative audio system using AI/ML in devices addresses the challenge of inaccurate audio representation by creating immersive and interactive sounds, improving user experience in video playback.

WO2025198853A1PCT designated stage Publication Date: 2025-09-25QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/018528
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-20
Filing Date
2025-03-05
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing audio capture technologies often fail to accurately represent the actual sounds experienced by a user during video recording, leading to a diminished user experience due to missing or drowned-out sounds, and lack of immersive playback.

Method used

Employing a device with a computer vision module and generative audio system using artificial intelligence and machine learning to create or enhance audio based on image data, allowing for interactive sound addition and refinement.

Benefits of technology

Enhances user experience by providing immersive and interactive audio that better matches the user's desired sounds, even when actual captured audio is inadequate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025018528_25092025_PF_FP_ABST
    Figure US2025018528_25092025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure provides systems, methods, and devices for audio signal processing that support enhancing images with generative audio. In a first aspect, a device for enhancing images with generative audio includes a memory storing processorreadable code and one or more processors coupled to the memory. The one or more processors configured to execute the processor-readable code to cause the one or more processors to: obtain image data for a plurality of image frames associated with a live image; perform computer vision for the plurality of image frames to generate visual perception information based on the image data; generate generative audio information based on the visual perception information; and output the image data with the generative audio information. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

ENHANCED AUDIO IN LIVE PHOTOS AND VIDEOCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of India Patent Application No. 202441021066, entitled, “ENHANCED AUDIO IN LIVE PHOTOS AND VIDEO,” filed on March 20, 2024, which is expressly incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] Aspects of the present disclosure relate generally to audio signal processing, and more particularly, to enhancing video with generative audio information to increase user experience. Some features may enable and provide improved audio signal processing, including on-demand generation of sounds which correspond to video or live photos.INTRODUCTION

[0003] Stereophonic sound, commonly called stereo, reproduces a sound using two or more independent audio channels. The playback of stereo sounds involves at least to channels of data corresponding to, in some systems, a left speaker and a right speaker. A user listening to sound generated by the left speaker and the right speaker can perceive a location of certain sounds based on how much of the sound is reproduced by the left speaker versus the right speaker. For example, a sound originating entirely from the left speaker and not the right speaker will sound to the user as a sound on their left side. The left and right channels of stereo sounds can be recorded using two or more microphones placed a distance from each other such that each microphone records propagating sounds at different times. A user listening to stereo sounds from two signals will hear similar spatial information of the stereo sounds as recorded by the two microphones when reproduced from speakers with a similar arrangement as the microphones.

[0004] Audio playback devices are devices that can reproduce one or more audio signals, whether digital or analog signals. Audio playback can be incorporated into a wide variety of devices. By way of example, audio playback devices may comprise stand-alone audio devices, mobile telephones, cellular or satellite radio telephones, personal digital assistants (PDAs), panels or tablets, gaming devices, or computing devices.BRIEF SUMMARY OF SOME EXAMPLES

[0005] The following summarizes some aspects of the present disclosure to provide a basic understanding of the discussed technology. This summary is not an extensive overview of all contemplated features of the disclosure and is intended neither to identify key or critical elements of all aspects of the disclosure nor to delineate the scope of any or all aspects of the disclosure. Its sole purpose is to present some concepts of one or more aspects of the disclosure in summary form as a prelude to the more detailed description that is presented later.

[0006] Currently, when a device captures image data for videos (including for live photos), the device also captures sounds from a microphone of the device. However, sometimes the audio captured while recording the video will include sounds that reduce the user experience. Also, the capture audio may miss or not capture some sounds due to be being drowned out by other sounds or removed by noise cancelling algorithms. Thus, the user experience may be reduced when playing back the video as the sound of the video may not correspond to the actual experience or a user’s desired experience. Additionally, when a device captures image data for an image or live photo, the device does not capture sounds from the microphone of the device. Thus, the user experience may not be as immersive or enjoyable when the image or live photo is output or played back as there may be no sound.

[0007] In some aspects described herein, a device includes a computer vision module and a generative audio system for generative sound creation (e.g., re-construction, alteration, or enhancement) by artificial intelligence (Al) or machine learning (ML) techniques to better match sounds output by the device to the sounds desired by the user. For example, the device may utilize computer vision and generative Al techniques to create sounds, such as rain, wind breeze, or a flowing river, for the captured video and / or photo content to provide an immersive experience to the user. In the various aspects described herein, the device can flexibly use image data, and optionally the captured audio data, to more accurately output the actual sounds experienced by the user (and possibly not captured) or to customize or refine the audio of the video or live photo to the taste of the user. Refining the audio may include removing captured input sounds, adding a generative sound, and / or replacing a captured sound with a generative sound.

[0008] Additionally, it allows users to add their own specific sounds to photos interactively. For example, when the user highlights a specific area in the photo, device may generate and render the corresponding sound in real-time, making the experience both more immersive and interactive.

[0009] In one aspect of the disclosure, a device for enhancing images with generative audio includes a memory storing processor-readable code and one or more processors coupled to the memory. The one or more processors configured to execute the processor-readable code to cause the one or more processors to: obtain image data for a plurality of image frames associated with a live image; perform computer vision for the plurality of image frames to generate visual perception information based on the image data; generate generative audio information based on the visual perception information; and output the image data with the generative audio information.

[0010] In an additional aspect of the disclosure, a method for enhancing images with generative audio includes: obtaining image data for a plurality of image frames associated with a live image; performing computer vision for the plurality of image frames to generate visual perception information based on the image data; generating generative audio information based on the visual perception information; and outputting the image data with the generative audio information.

[0011] In an additional aspect of the disclosure, an apparatus includes: means for obtaining image data for a plurality of image frames associated with a live image; means for performing computer vision for the plurality of image frames to generate visual perception information based on the image data; means for generating generative audio information based on the visual perception information; and means for outputting the image data with the generative audio information.

[0012] In an additional aspect of the disclosure, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor, cause the processor to perform operations. The operations include: obtaining image data for a plurality of image frames associated with a live image; performing computer vision for the plurality of image frames to generate visual perception information based on the image data; generating generative audio information based on the visual perception information; and outputting the image data with the generative audio information.

[0013] Methods of audio signal processing described herein may be performed by a signal processing device. The audio signal processing may be applied audio data captured by one or more microphones of the signal processing device. Audio signal processing devices, devices that can playback, record, and / or process one or more audio recordings can be incorporated into a wide variety of devices. By way of example, audio signal processing devices may comprise stand-alone audio devices, such as entertainment devices and personal media players, wireless communication device handsets such asmobile telephones, cellular or satellite radio telephones, personal digital assistants (PDAs), tablets, gaming devices, computing devices such as webcams, video surveillance cameras, or other devices with audio recording or audio capabilities.

[0014] The audio signal processing techniques described herein may involve devices having microphones and processing circuitry (e.g., application specific integrated circuits (ASICs), digital signal processors (DSP), graphics processing unit (GPU), or central processing units (CPU)).

[0015] In some aspects, a device may include a digital signal processor or a processor (e.g., an application processor) including specific functionality for audio processing. The methods and techniques described herein may be entirely performed by the digital signal processor or the processor, or various operations may be split between the digital signal processor and the processor, and in some aspects split across additional processors. In some embodiments, the methods and techniques disclosed herein may be adapted using input from a neural signal processor (NSP) in which one or more parameters of the signal processing are controlled based on output from a machine learning (ML) model executed by the NSP.

[0016] In an additional aspect of the disclosure, a device configured for audio signal processing and / or audio capture is disclosed. The apparatus includes means for recording audio. Example means may include a dynamic microphone, a condenser microphone, a ribbon microphone, a carbon microphone, or a crystal microphone. The microphone may be construed as a microelectromechanical system (MEMS). These components may be controlled to capture first and / or second sound recordings, which may correspond to left and right channels of a recording.

[0017] For any of these types of microphones, the microphones may include analog and / or digital microphones. Analog microphones provide a sensor signal, which is some embodiments is conditioned or filtered. Analog microphones in a digital system include an external analog-to-digital converter (ADC) to interface with digital circuitry. Digital microphones include the ADC and other digital elements to convert the sensor signal into a digital data stream, such as a pulse-density modulated (PDM) stream or a pulse-code modulated (PCM) stream.

[0018] Other aspects, features, and implementations will become apparent to those of ordinary skill in the art, upon reviewing the following description of specific, exemplary aspects in conjunction with the accompanying figures. While features may be discussed relative to certain aspects and figures below, various aspects may include one or more of theadvantageous features discussed herein. In other words, while one or more aspects may be discussed as having certain advantageous features, one or more of such features may also be used in accordance with the various aspects. In similar fashion, while exemplary aspects may be discussed below as device, system, or method aspects, the exemplary aspects may be implemented in various devices, systems, and methods.

[0019] The method may be embedded in a computer-readable medium as computer program code comprising instructions that cause a processor to perform the steps of the method. In some embodiments, the processor may be part of a mobile device including a first network adaptor configured to transmit data, such as images or videos (with associated or embedded sounds) in a recording or as streaming data, over a first network connection of a plurality of network connections; and a processor coupled to the first network adaptor and the memory. The processor may cause the transmission of output image frames described herein over a wireless communications network such as a 5G NR communication network.

[0020] The foregoing has outlined, rather broadly, the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims.

[0021] While aspects and implementations are described in this application by illustration to some examples, those skilled in the art will understand that additional implementations and use cases may come about in many different arrangements and scenarios. Innovations described herein may be implemented across many differing platform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects and / or uses may come about via integrated chip implementations and other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, artificial intelligence (Al)-enabled devices, etc.). While some examples may or may not bespecifically directed to use cases or applications, a wide assortment of applicability of described innovations may occur. Implementations may range in spectrum from chiplevel or modular components to non-modular, non-chip-level implementations and further to aggregate, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical settings, devices incorporating described aspects and features may also necessarily include additional components and features for implementation and practice of claimed and described aspects. It is intended that innovations described herein may be practiced in a wide variety of devices, chip-level components, systems, distributed arrangements, end-user devices, etc. of varying sizes, shapes, and constitution.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] A further understanding of the nature and advantages of the present disclosure may be realized by reference to the following drawings. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If just the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.

[0023] FIG. 1 shows a block diagram of a system-on-chip (SoC) configured for performing signal processing according to one or more aspects of this disclosure.

[0024] FIG. 2 is a block diagram illustrating an example data flow path for audio signal processing in a multimedia device according to one or more aspects of the disclosure.

[0025] FIG. 3 is a block diagram illustrating an example of a device that supports generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure.

[0026] FIG. 4 is a flow diagram illustrating an example of generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure.

[0027] FIG. 4 is a flow diagram illustrating another example of generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure.

[0028] FIG. 4 is a flow diagram illustrating another example of generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure.

[0029] FIG. 7 is a flow chart illustrating a method for generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure.

[0030] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0031] The present disclosure provides systems, apparatus, methods, and computer-readable media that support signal processing, including techniques for generative audio enhancement operations for videos and live photos. In some aspects, the present disclosure provides techniques for outputting generative audio with videos and live photos. The generative audio may be created by a device based on image data captured by the device, and optionally based on audio data captured by the device. The generative audio may include one or more generative sounds for objects detected in the image data or related to a detected scene for the image data.

[0032] When capturing audio and video data, sometimes the captured audio does not quite represent the actual sounds experienced by the user or person who captured the data and / or may not represent the sounds expected by the user or another viewer. Shortcomings mentioned here are only representative and are included to highlight problems that the inventors have identified with respect to existing devices and sought to improve upon. Aspects of devices described below may address some or all of the shortcomings as well as others known in the art. Aspects of the improved devices described herein may present other benefits than, and be used in other applications than, those described above.

[0033] Particular implementations of the subject matter described in this disclosure may be implemented to realize one or more of the following potential advantages or benefits. In some aspects, the present disclosure provides techniques for enhancing the user experience for videos and live photos. For example, the techniques described herein enable users to add their own specific sounds to a video or live photo interactively. In some aspects, the user may highlight a specific area in an image frame, and the device generates and renders the corresponding generative sound in real-time, making the experience both more immersive and interactive.

[0034] The detailed description set forth below, in connection with the appended drawings to which the text references, is intended as a description of various embodiments and is not intended to limit the scope of the disclosure. Rather, the detailed description includes specific details for the purpose of providing a thorough understanding of the subjectmatter of this disclosure. It will be apparent to those skilled in the art that these specific details are not required in every case and that, in some instances, well-known structures and components are shown in block diagram form for clarity of presentation.

[0035] In the description of embodiments herein, numerous specific details are set forth, such as examples of specific components, circuits, and processes to provide a thorough understanding of the present disclosure. The term “coupled” as used herein means connected directly to or connected through one or more intervening components or circuits. Also, in the following description and for purposes of explanation, specific nomenclature is set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that these specific details may not be required to practice the teachings disclosed herein. In other instances, well known circuits and devices are shown in block diagram form to avoid obscuring teachings of the present disclosure.

[0036] Some portions of the detailed descriptions which follow are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. In the present disclosure, a procedure, logic block, process, or the like, is conceived to be a self-consistent sequence of steps or instructions leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, although not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system.

[0037] An example device for recording sounds and / or processing sound signals using one or more microphones, such as a MEMS microphone, may include a configuration of one, two, three, four, or more microphones at different locations on the device. The example device may include one or more digital signal processors (DSPs), Al engines, or other suitable circuitry for processing signals captured by the microphones. The one or more digital signal processors (DSPs) may output signals representing sounds through a bus for storage in a memory, for reproduction by an audio system, and / or for further processing by other components (such as an applications processor). The processing circuitry may perform further processing, such as for encoding, storage, transmission, or other manipulation of the audio signals. In some embodiments, the example device may include audio circuitry including an audio amplifier (e.g., a class-D amplifier) for driving a transducer to reproduce the sounds represented by the audio signals. A speaker may be integrated with the device and coupled to the audio amplifier to be driven by the audioamplifier for reproducing the sounds. A connection may be provided by a jack or other connector on the device to couple an external transducer (e.g., an external speaker or headphones) to the audio amplifier to be driven by the audio circuitry to reproducing the sounds. In some embodiments, the jack may instead output a digital signal for conversion and amplification by an external device, such as when the jack is configured to be coupled to a digital device through a Universal Serial Bus (USB) Type-C (USB-C) connection and some or all of the audio circuitry is bypassed.

[0038] FIG. 1 shows a block diagram of a system-on-chip (SoC) configured for performing signal processing according to one or more aspects of this disclosure. The SoC 100 may include several components coupled together through a bus 102, which may be a network-on-a- chip (NoC) or a plurality of NOCs interconnecting various components. For example, although FIG. 1 illustrates several components coupled to the bus 102, the several components may be coupled to different busses with additional busses connecting the different busses to provide a path for communication between the components.

[0039] One example component in the SoC 100 is a digital signal processor 112 for signal processing. The DSP 112 may process audio signals received from microphones 130A, 130B, and 130C of microphone array 130. The DSP 112 may include hardware customized for performing a limited set of operations on specific kinds of data. For example, a DSP may include transistors coupled together to perform operations on streaming data and use memory architectures and / or access techniques to fetch multiple data or instructions concurrently. Such configurations may allow the DSP 112 to operate on real-time data, such as video data, audio data, or modem data, in a power-efficient manner.

[0040] Additionally, the DSP 112 may process image data received from cameras 160A and / or 160B of camera array 160. The DSP 112 may include hardware customized for performing a limited set of operations on specific kinds of data. For example, the DSP 112 (or a second DSP) may include transistors coupled together to perform operations on image data, including streaming image data, and use memory architectures and / or access techniques to fetch multiple data or instructions concurrently. Such configurations may allow the DSP 112 to operate on real-time video data in a power-efficient manner.

[0041] The SoC 100 also includes a central processing unit (CPU) 104 and a memory 106 storing instructions 108 (e.g., a memory storing processor-readable code or a non-transitory computer-readable medium storing instructions) that may be executed by a processor of the SoC 100. The CPU 104 may be a single central processing unit (CPU) or a CPUcluster comprising two or more cores such as core 104 A. The CPU 104 may include hardware capable of performing generic operations on many kinds of data, such as hardware capable of executing instructions from the Advanced RISC Machines (ARM®) instruction set, such as ARMv8 and ARMv9. For example, a CPU 104 may include transistors coupled together to perform operations for supporting executing an operating system and user applications (e.g., a camera application, a multimedia application, a gaming application, a productivity application, a messaging application, a videocall application, an audio recording application, a video recording application). The CPU 104 may execute instructions 108 retrieved from the memory 106. In some embodiments, the CPU 104 executing an operating system may coordinate execution of instructions by various components within the SoC 100. For example, the CPU 104 may retrieve instructions 108 from memory 106 and execute the instructions on the DSP 112.

[0042] The SoC 100 may further include a neural signal processor (NSP) 124 for executing machine learning (ML) models relating to multimedia applications. The NSP 124 may include hardware configured to perform and accelerate convolution operations involved in executing machine learning algorithms. For example, the NSP 124 may improve performance when executing predictive models such as artificial neural networks (ANNs) (including multilayer feedforward neural networks (MLFFNN), the recurrent neural networks (RNN), and / or the radial basis functions (RBF)). The ANN executed by the NSP 124 may access predefined training weights stored in the memory 106 for performing operations on user data.

[0043] The SoC 100 may be coupled to a display 114 for interacting with a user. The SoC 100 may also include a graphics processing unit (GPU) 126 for rendering images on the display 114. In some embodiments, the CPU 104 may perform rendering to the display 114 without a GPU 126. In some embodiments, the GPU 126 may be configured to execute instructions for performing operations unrelated to rendering images, such as for processing large volumes of datasets in parallel.

[0044] Processing algorithms, techniques, and methods that are described herein may be executed by at least one processor of the SoC 100, which may include execution by all steps on one of the processors (e.g., DSP 112, CPU 104, NSP 124, GPU 126) or may include execution of steps across a combination of one or more of the processors (e.g., DSP 112, CPU 104, NSP 124, GPU 126). In some embodiments, at least one of the DSP 112 or the CPU 104 executes instructions to perform various operations described herein, including generative audio enhancement operations for videos and live photos. Forexample, execution of the instructions by the CPU 104 as part of a multimedia application (e.g., a voice recorder, a sound recording, or a video recorder) may instruct the DSP 112 to begin or end capturing audio from one or more microphones 130A-C. The operations of the CPU 104 may be based on user input. For example, a voice recorder application executing on processor 104 may receive a user command to begin a voice recording upon which audio comprising one or more channels is captured and processed for playback and / or storage. Audio processing to determine “output” or “corrected” signals, such as according to techniques described herein, may be applied to one or more segments of audio in the recording sequence.

[0045] As another example, execution of the instructions by the CPU 104 as part of a multimedia application may instruct the DSP 112 to begin or end capturing image data from one or more cameras 160A-B. The operations of the CPU 104 may be based on user input. For example, a video recorder application executing on processor 104 may receive a user command to begin a video (e.g., live image) recording upon which image data comprising one or more channels is captured and processed for playback and / or storage. Image processing to determine “output” or “corrected” signals, such as according to techniques described herein, may be applied to one or more segments of image data in the recording sequence.

[0046] Input / output components may be coupled to the SoC 100 through an input / output (I / O) hub 116. An example of a hub 116 is an interconnect to a peripheral component interconnect express (PCIe) bus. Example components coupled to hub 116 may be components used for interacting with a user, such as a touch screen interface and / or physical buttons. Some components coupled to hub 116 may also include network interfaces for communicating with other devices, including a wide area network (WAN) adaptor (e.g., WAN adaptor 152), a local area network (LAN) adaptor (e.g., LAN adaptor 153), and / or a personal area network (PAN) adaptor (e.g., PAN adaptor 154). A WAN adaptor 152 may be a 4G LTE or a 5G NR wireless network adaptor. A LAN adaptor 153 may be an IEEE 802.11 WiFi wireless network adapter. A PAN adaptor 154 may be a Bluetooth wireless network adaptor. Each of the WAN adaptor 152, LAN adaptor 153, and / or PAN adaptor 154 may be coupled to an antenna that may be shared by each of the adaptors 152, 153, and 154, or coupled to multiple antennas configured for primary and diversity reception and / or configured for receiving specific frequency bands. In some embodiments, the WAN adaptor 152, LAN adaptor 153, and / or PAN adaptor 154 may share circuitry, such as portions of a radio frequency front end (RFFE).

[0047] Audio circuitry 156 may be integrated in SoC 100 as dedicated circuitry for coupling the SoC 100 to a speaker 120 external to the SoC 100, which may be a transducer such as a speaker (either internal to or external to a device incorporating the SoC 100) or headphones. The audio circuitry 156 may include coder / decoder (CODEC) functionality for processing digital audio signals. The audio circuitry 156 may further include one or more amplifiers (e.g., a class-D amplifier) for driving a transducer coupled to the SoC 100 for outputting sounds generated during execution of applications by the SoC 100. Functionality related to audio signals described herein may be performed by a combination of the audio circuitry 156 and / or other processors of the SoC (e.g., CPU 104, DSP 112, GPU 126, NSP 124).

[0048] The SoC 100 may couple to external devices outside the package of the SoC 100. For example, the SoC 100 may be coupled to a power supply 118, such as a battery or an adaptor to couple the SoC 100 to an energy source. The signal processing described herein may be adapted to and achieve power efficiency to support operation of the SoC 100 from a limited-capacity power supply 118 such as a battery. For example, operations may be performed on a portion of the SoC 100 configured for performing the operation at a lowest power consumption. As another example, operations themselves are performed in a manner that reduces an amount of computations to perform the operation, such that the algorithm is optimized for extending the operational time of a device while powered by a limited-capacity power supply 118. In some embodiments, the operations described herein may be configured based on a type of power supply 118 providing energy to the SoC 100. For example, a first set of operations may be executed to perform a function when the power supply 118 is a wall adaptor. As another example, a second set of operations may be executed to perform a function when the power supply 118 is a battery.

[0049] The SoC 100 may also include or be coupled to additional features or components that are not shown in FIG. 1. Although components are shown integrated as a single SoC 100, which may include all components built on a single semiconductor die with a common semiconductor substrate, other arrangements of the illustrated blocks different number of dies, substrates, and / or packages may be arranged to accomplish the same functionality described in this disclosure.

[0050] The memory 106 may include a non-transient or non-transitory computer readable medium storing computer-executable instructions as instructions 108 to perform all or a portion of one or more operations described in this disclosure. The instructions 108 may include a multimedia application (or other suitable application such as a messagingapplication) to be executed by the SoC 100 that records, processes, or outputs audio signals. The instructions 108 may also include other applications or programs executed by the SoC 100, such as an operating system and applications other than for multimedia processing.

[0051] In addition to instructions 108, the memory 106 may also store audio data. The SoC 100 may be coupled to an external memory and configured to access the memory for writing output audio files for later playback or long-term storage. For example, the SoC 100 may be coupled to a flash storage device comprising NAND memory for storing video files (e.g., MP4-container formatted files) including audio tracks and / or storing audio recordings (e.g., MPEG-1 Layer 3 files, also referred to as MP3 files). Portions of the video or audio files may be transferred to memory 106 for processing by the SoC 100, with the resulting signals after processing encoded as video or audio files in the memory 106 for transfer to the long-term storage.

[0052] While the SoC 100 is referred to in the examples herein for performing aspects of the present disclosure, some device components may not be shown in FIG. 1 to prevent obscuring aspects of the present disclosure. Additionally, other components, numbers of components, or combinations of components may be included in a suitable device for performing aspects of the present disclosure. As such, the present disclosure is not limited to a specific device or configuration of components, including the SoC 100.

[0053] The SoC of FIG. 1 may be operated to obtain improved audio recordings and / or improved user experience through higher quality audio playback by creating generative audio and outputting the generative audio with image data, such as for videos or live photos. One example method of performing multimedia operations is shown in FIG. 2 and described below.

[0054] FIG. 2 is a block diagram illustrating an example data flow path for audio signal processing in a multimedia device according to one or more aspects of the disclosure. SoC 100 of system 200 may execute multimedia control 210, such as part of an operating system or driver, to control the capture of sounds from microphones or other audio sources and / or to control the configuration of audio processing circuitry 156. The audio configuration applied by multimedia control 210 to either output devices (e.g., speakers) or input devices (e.g., microphones) may include parameters that specify, for example, a bit depth, a sampling rate, a data rate, a magnitude, or other parameters.

[0055] Multimedia control 210 may be managed by or provide services to a multimedia application 204. The multimedia application 204 may also execute on the SoC 100. Themultimedia application 204 provides settings accessible to a user such that a user can specify individual playback settings or select a profile with corresponding playback settings. The multimedia application 204 may be, for example, a video recording application, a screen sharing application, a virtual conferencing application, an audio playback application, a messaging application, a video communications application, or other application that processes audio data. The multimedia application 204 may include a generative audio system 206 to improve the quality of audio presented to the user during execution of multimedia application 204 by adding generative audio sounds to output audio data. The generative audio system 206 may perform one or more or a combination of the techniques described herein to add generative audio sounds when no audio is present, to enhance captured audio with generative audio sounds, and / or to replace captured audio with generative audio sounds. The system 200 of FIG. 2 may be configured to perform the generative audio operations described with reference to any of FIGS. 3-7 to out a generative audio signal.

[0056] FIG. 3 illustrates an example 300 of a device 301 that supports enhanced video and live images with generative audio in accordance with aspects of the present disclosure. In some examples, the device 301 may include or correspond to the SoC 100 of FIG. 1 or the system 200 (e.g., cellular phone) of FIG. 2. For example, the device 301 may include a processing system for enhancing video and live images with generative audio and a wireless communication system to interact with a wireless communication network. The wireless communication network may include or correspond to a cellular network or a Wi-Fi network, as illustrative, non-limiting examples. The wireless communication network may include multiple wireless devices, such as user equipments (UEs).

[0057] Device 301 may be configured to communicate via one or more portions of the electromagnetic spectrum. For example, a wireless interface 320 of the device 301 may be configured to communicate via one or more portions of the electromagnetic spectrum associated with Bluetooth transmissions, Wi-Fi transmissions, or cellular transmissions (including sub-6 GHz and 6 GHz).

[0058] Device 301 may be configured to communicate via one or more channels or component carriers (CCs). Each channel or CC may have a corresponding configuration, such as configuration param eters / settings. The configuration may include bandwidth, bandwidth part, control channel resources, data channel resources, or a combination thereof. Additionally, or alternatively, one or more channels or CCs may have or be assigned to a Cell ID, or a Bandwidth Part (BWP) ID. The Cell ID may include a unique cell ID forthe channel or CC, a virtual Cell ID, or a particular Cell ID of a particular channel or CC of the plurality of channels or CCs. Each channel or CC may also have corresponding management functionalities, such as, beam management or BWP switching functionality. In some implementations, two or more channels or CCs are quasi co-located, such that the channels or CCs have the same beam and / or same symbol.

[0059] In some implementations, control information may be communicated by network devices (e.g., a base station 105) to the device 301. For example, the control information may be communicated using Bluetooth transmissions, Wi-Fi transmission, MAC-CE transmissions, RRC transmissions, DCI (downlink control information) transmissions, UCI (uplink control information) transmissions, SCI (sidelink control information) transmissions, another transmission, or a combination thereof.

[0060] Device 301 can include a variety of components (e.g., structural, hardware components) used for carrying out one or more functions described herein. For example, these components can include a processing system and memory configured to perform operations for enhancing video and live photos with generative audio, along with wireless communication components, such as a transceiver, an encoder, a decoder, and one or more antennas (not shown in FIG. 3 for simplicity and generally referred to as or part of the wireless interface 320).

[0061] As illustrated in the example of FIG. 3, the device 301 (e.g., a UE) includes a processor 302 and a memory 304. Processor 302 may be configured to execute instructions stored at memory 304 to perform the operations described herein. In some implementations, processor 302 includes or corresponds to the processing system and / or the SoC 100, the CPU 104, the DSP 112, and / or the audio processor 156 of FIG. 1 or FIG. 2, and memory 304 includes or corresponds to the memory 106 of FIGS. 1 or 2. Memory 304 may also be configured to store information and data for enhancing video and live photos with generative audio, as further described herein. For example, the memory 304 may be configured to store one or more of audio input data 360, image data 362, user input data 364, user preference data 366, scene data 368, object data 370, and visual perception information 372.

[0062] The audio input data 360 (or referred to herein as input audio data) may include or correspond to audio signal data for multiple audio frames in an audio space. The audio input data 360 may be generated by the microphone 322 of the sensor system 306. In some implementations, the audio input data 360 corresponds to spatial audio, and include spatial information associated with captured audio data.

[0063] Similarly, the image data 362 may include or correspond to image sensor data for multiple image frames in an image space. The image data 362 may be generated by the camera 324 of the sensor system 306 in some implementations. Alternatively, the image data 362 may include or correspond to processed image data. The image data 362 is used to determine sounds, or potential or candidate sounds, for generative audio operations. To illustrate, sounds may be generated by Al and / or ML techniques for sounds identified and that correspond to a scene and / or objects identified in and / or associated with the image frames. In some implementations, the image data 362 includes extended reality information, such as virtual or augmented reality information. The virtual or augmented reality information may include or correspond to virtual assets that are included with or output with the captured image data, and may correspond to virtual assets that are overlaid on the captured image data.

[0064] In some implementations, the audio input data 360 and the image data 362 correspond to a live photo or a video, a set of image frames and with corresponding audio frames that was captured temporally with the image frames. For example, the audio frames of the audio input data 360 and the image frames of the image data 362 may correspond to an overlapping time period and may be synchronized or associated with one another. The audio input data 360 and the image data 362, such as the frames thereof, can be played back or output to generate a video or a live image (short video). In some other implementations, the audio input data 360 and the image data 362 may be captured by another device, and received via wireless communication by the wireless interface 320.

[0065] However, the audio input data 360, when captured, may not contain audio data for all objects in the image due to distortion, capture issues, weak signal issues, device processing (e.g., noise cancellation), etc., and playback of the input audio may not provide a satisfying user experiences. Also, in some implementations, the audio input data may not be captured or transmitted with the image data 362.

[0066] The user input data 364 may include or correspond to data indicating a user input or selection related to generative audio determination. For example, the user input data 364 may include inputs regarding selection of generative audio sounds to output or to include in the generative audio played back with the image data. As another example, the user input data 364 may include inputs regarding sound mixing information, such as a frequency, an intensity, motion, or a location of the audio output with the image data 362. In some such implementations, the user input data 364 may include inputs for modifyingor adjusting the generative sounds themselves. For example, the user may control a type, a frequency, an intensity, motion, or a location of the generative sound.

[0067] The user preference data 366 may include or correspond to data indicating stored user preferences from prior inputs or selections related to generative audio determination or settings information. For example, the user preference data 366 may include data indicating rules or conditions for generating generative audio. To illustrate, the user preference data 366 may indicate selection criteria for which generative audio sounds to output or include in the generative audio. As another example, the user preference data 366 may include inputs regarding sound mixing information, such as a frequency, an intensity (e.g., gain), or a location of the audio output with the image data 362. To illustrate, the user may indicate to only generate sounds which have certain properties or to adjust generated sounds to fit within certain properties.

[0068] The scene data 368 may include or correspond to scene data indicating or identifying a scene or parameters thereof, for the multiple image frames of the image data 362. The scene data 368 may be generated by the computer vision module 314, such as the object detection and tracking system 336 thereof, and based on the image data 362. Additionally, or alternatively, the scene data 368 may include or correspond to a scene identified based on the audio input data 360 by the audio perception module 310.

[0069] The object data 370 may include or correspond to object data indicating or identifying an object or parameters thereof, associated with the scene, and / or identified or detected in the multiple image frames of the image data 362. The object data 370 may be generated by the computer vision module 314, such as the object detection and tracking system 336 thereof, and based on the image data 362 and optionally the scene data 368. Additionally, or alternatively, the object data 370 may include or correspond to an object identified based on the audio input data 360 by the audio perception module 310.

[0070] The object data 370 may include or correspond to output data from performing object detection operations using the image data 362, or using processed image data. For example, the image data 362 may be converted to another space, such as a feature space by a feature coder and decoder, for object detection. The object data 370 may include object detection data which indicates a type or class of detected object (e.g., person, pedestrian, car, sign, lane, etc.), a status of the detected object (e.g., moving, stationary, fast, slow, driving, etc.), a location of the detected object (e.g., bounding box information).

[0071] The object data 370 may also include motion or tracking data. The tracking data includes or corresponds to output data from performing object tracking operations using a portion of the object data 370. The tracking data may include object tracking data which indicates a position of detected objects over multiple frames. The tracking data may further indicate a direction or heading of the detected objects and a speed of the detected objects. The object data 370 may further include one or more of event data, motion estimation data, pose estimation data, and depth estimation data. The event data indicate or identify events associated with the identified scene and / or objects, the motion estimation information may indicate predicted future motion for the detected objects, the pose estimation information may indicate a current or predicted future pose of the detected objects, and the depth estimation information may indicate a current or predicted future distance to the detected object.

[0072] The visual perception information 372 may include or correspond to output information from performance of computer vision processing on the image data 362 and which indicates, directly or indirectly, potential candidate sounds that may be associated with and / or included in the audio for the image data 362. The visual perception information 372 may include visual perception information for the multiple image frames of the image data 362. The visual perception information 372 may include the object data 370 and / or one or more of the intermediary outputs generated by sub-modules of the computer vision module 314. In some implementations, the visual perception information 372 may include sound information 374 for identifying a sound, such as sound tag, associated with detected objects, scenes, and / or events.

[0073] Additionally, or alternatively, the memory 304 may also be configured to store one or more of sound information 374, scene database information 376, sound database information 378, generative sound information 380, generative audio information 382, output data 384, and AI / ML model data 386.

[0074] The sound information 374 may include or correspond to information regarding the identified sounds associated with the scene and / or objects in the scene based on performing the computer vision on the image data 362. For example, the sound information 374 may include sound tags or other sound identifying information to enable generation of generative audio by the generative audio system 318. Additionally, the sound information 374 may include sound parameter information indicating one or more parameters for the sounds, such as frames information, intensity information, frequency information, motion information, location information, depth information, poseinformation, etc., or any combination thereof. The sound information 374 in some implementations may include spatial audio information for enabling generation of generative spatial sounds.

[0075] The sound information 374 may include candidate sound information for generating candidate sounds which may optionally be added to the generative audio output. The candidate sounds may be generated or not based on one or more of user preferences, user input / selection, or audio input data. For example, the device 301 may filter or adjust the candidate sounds based on user information and the input audio to confirm which sounds should be in scene, or should be added to the scene, based on which sounds are already present and the user’s conditions for inclusion of generative sounds.

[0076] The scene database information 376 may include or correspond to a database of scene related information which enables the device 301 to identify a scene or scenes in the image data based on the visual perception information 372, such as the determined or estimated scene information thereof. For example, the scene database information 376 may include a data structure with scene tags or identification information, scene parameter information, scene sound information, scene object information, scene event information, scene context information, or a combination thereof. Each piece of information or entry of the types of data in the data structure may be associated with a corresponding piece of information or entry. To illustrate, the scene database information 376 may include a first scene (e.g., beach scene) that is associated with a plurality of beach related objects that are often included in a beach scene (e.g., palm tree, seagull, sand, waves, etc.), a plurality of beach related sounds that one may be expected (e.g., palm leaves rustling, sand being blown, bird flapping wings, bird cawing, waves crashing, etc.). The scene may also be associated scene parameters for identification or confirmation with the visual perception information 372 determined during computer vision. The second scene (e.g., an urban scene) of the scene database information 376 may include or be associated with corresponding scene and sound related information that corresponds to the second scene.

[0077] The sound database information 378 may include or correspond to information regarding a database of sound related information which enables the device 301 to identify and / or generate a sound in the image data based on the visual perception information 372. For example, the sound database information 378 may include a data structure with sound tags or identification information, sound parameter information, object information, scene information, sound context information, or a combination thereof. Each piece ofinformation or entry of the types of data in the data structure may be associated with a corresponding piece of information or entry. To illustrate, the sound database information 378 may include a first sound (e.g., wind sound) that is associated with a plurality of wind related objects that are often make a wind sound (e.g., palm tree, seagull, sand, waves, etc.), a plurality of related scenes that one may be expected to have such as sound (e.g., beach scene, ocean scene, mountain scene, etc.). The sound may also be associated sound parameters for generation of generative sound information for the sound based on the visual perception information 372 determined during computer vision. Additionally, or alternatively, the sound may also be associated sound parameters for identification or confirmation of the sound based on the visual perception information 372 determined during computer vision. The second sound (e.g., a waves crashing sound) of the sound database information 378 may include or be associated with corresponding scene and sound related information that corresponds to the second scene.

[0078] The generative sound information 380 may include or correspond to generative sound data that has been generated by Al and / or ML techniques for one or more objects in a scene identified in or based on the input data, the image data 362 and optionally the audio input data 360. For example, the generative sound information 380 may include or correspond to data for reproducing or enhancing sound of an object in the input scene or related to the input scene that has been generated by Al and / or ML techniques. To illustrate, the generative sound information 380 may include, as an illustrative, nonlimiting example, first generative sound information for producing a seagull sound for a beach scene, second generative sound information for reproducing a wind sound for the beach scene , or third generative sound information for enhancing a wave sound for the beach scene. Producing the sound may correspond to playback of the sound and produce audio waves or signals from the speaker 328.

[0079] The generative audio information 382 may include or correspond to information generated by Al and / or ML techniques for one or more objects in a scene identified in or based on the input data, the image data 362 and optionally the audio input data 360. For example, the generative audio information 382 may include or correspond one or more generative sounds of the generative sound information 380. To illustrate, the generative audio information 382 may include or correspond to one or more selected candidate sounds of the generative sound information 380, modified sounds of the generative sound information 380, or a combination thereof.

[0080] The output data 384 may include or correspond to audio and video data output by the device 301. For example, the output data 384 may include image data that includes or correspond to the image data 362, or that is generated based on the image data 362. As another example, the output data 384 includes the generative audio information 382 and optionally includes audio data that includes or correspond to the audio input data 360, or that is generated based on the audio input data 360.

[0081] The AI / ML model data 386 includes or corresponds to Al or ML models for one or more systems or modules of the device 301. For example, the device 301 may include a single Al or ML model for generative audio enhancement operations, such as a single Al or ML model associated with one or more of the sensor system 306, the output system 308, the audio perception module 310, the computer vision module 314 (e.g., the object detection and tracking system 336 thereof), the sound candidate generator 316, and / or the generative audio system 318. To illustrate, the Al or ML model for the generative audio enhancement operations may receive sensor data as input and output generative audio information as output. As another illustration, the Al or ML model for the generative audio enhancement operations may computer vision information, audio perception information, and / or sound tag information, as input and output generative audio information as output.

[0082] As another example, the generative audio enhancement operations may include multiple discrete Al or ML modules for different portions of the generative audio enhancement operations. To illustrate, the device 301 may include an Al or ML module for audio perception by the audio perception module 310, an Al or ML module for computer vision by the computer vision module 314, an Al or ML module for sound candidate generation and management by the sound candidate generator 316, and an Al or ML module for generative audio generation by the generative audio system 318. Each Al or ML module may have corresponding inputs and outputs as described herein with respect to each of the corresponding modules or components.

[0083] In some such implementations, one or more of the above Al or ML modules may have one or more Al or ML sub-modules or may be further broken into multiple discrete Al or ML sub-modules for different portions. For example, the an Al or ML module for computer vision by the computer vision module 314 may include or correspond to an Al or ML module for scene detection by the scene detector 334, an Al or ML module for object detection by the object detector 338, an Al or ML module for object tracking by the object tracker 340, an Al or ML module for event detection by the event detector 342,an Al or ML module for motion estimation by the motion estimator 344, an Al or ML module for pose estimation by the pose estimator 346, and an Al or ML module for depth estimation by the depth estimator 348.

[0084] The device 301 further includes a sensor system 306 including one or more different sensors. The sensors of the sensor system 306 may include or correspond to the sensors described with reference to FIGS. 1 or 2. As illustrated in the example of FIG. 3, the device 301 includes a microphone 322 and a camera 324. The microphone 322 may include or correspond to a microphone sensor or system of microphone sensors (e.g., a microphone array or arrays) configured to receive audio signals and generate audio data, such as the audio input data 360. The camera 324 may include or correspond to an optical sensor or system of optical sensors configured to generate image data, such as the image data 362.

[0085] The device 301 further includes an output system 308 including one or more output devices configured to output enhanced video or live photos with generative audio. As illustrated in the example of FIG. 3, the output system 308 of the device 301 includes a display 326 and a speaker 328. The display 326 may include or correspond to a LCD, LED, or OLED device configured to output image or video data, such as image data of the output data 384. The speaker 328 may include or correspond to a speaker or system of speakers configured to output or reproduce audio signals based on the output data 384, such as the generative audio data thereof (e.g., at least a portion of the generative audio information 382).

[0086] The device 301 may optionally include an audio perception module 310 for identifying sounds in the audio input data 360. As illustrated in the example of FIG. 3, the audio perception module 310 includes a sound identifier 330 and a candidate sound manager 332. The sound identifier 330 is configured to generate identify one or more sounds in audio input data 360. For example, the sound identifier 330 may detect and identify one or more sounds in the audio input data 360 based on the audio input data 360 and the sound database information 378. To illustrate, the sound identifier 330 may perform sound segmentation and compare segmented sounds to sounds of a sound database, such as sound database information 378, for sound identification. The sound identifier 330 may encode the audio input data 360 to perform the sound identification operations. For example, the sound identifier 330 may perform feature encoding and / or decoding on the input audio data 360 to convert that audio to a feature space for sound feature extraction and sound detection operations.

[0087] In some implementations, the audio perception module 310 further identifies a scene, an event, and / or objects related with or associated with the identified sounds. For example, the audio perception module 310 may determine one or more objects related to the sounds or to a determined scene or event related to the one or more of the identified sounds.

[0088] The candidate sound manager 332 may be configured to manage candidate sounds identified by the sound identifier 330 and to select sounds to output for sound generation. For example, the sound identifier 330 may generate candidate sounds in some implementations which are then selected or confirmed by the candidate sound manager 332. To illustrate, the candidate sound manager 332 may filter out candidate sounds which do not match or correspond to an identified scene or event or which are not related to one another. As an illustrative, non-limiting examples, if the sound identifier 330 identified multiple candidate sounds for beach scene or environment and a single candidate sound (e.g., a train hornjfor an urban scene or environment, the candidate sound manager 332 filter out or remove the single candidate sound (e.g., a train horn) for the urban scene or environment.

[0089] The audio perception module 310 may be configured to output sound tag information indicating the identified sounds in the input audio data 360 to one or more other modules of the device 301. For example, the audio perception module 310 may provide sound tag information, such as the sound information 374, to the computer vision module 314 and / or the sound candidate generator 316 for identification and / or confirmation of sounds identified based on the image data 362. As another example, the audio perception module 310 provides the sound tag information to the generative audio system 318 for generative sound generation. In some such implementations, the audio perception module 310 may provide audio perception information including sound parameter information to the generative audio system 318 for generative sound generation.

[0090] Additionally, the audio perception module 310 may be configured to output scene, event, and / or object information indicating the identified scenes, events, and / or objects associated with the identified sounds in the input audio data 360 to one or more other modules of the device 301. For example, the audio perception module 310 may provide scene and object information to the computer vision module 314 and / or the sound candidate generator 316 for identification and / or confirmation of sounds identified based on the image data 362. As another example, the audio perception module 310 provides the scene and object information to the generative audio system 318 for generative soundgeneration. Additional details on audio perception operations are described further with reference to FIG. 6.

[0091] The device 301 further includes a computer vision module 314, which includes one or more sub-modules for identifying sounds which are associated with the image data 362. As illustrated in the example of FIG. 3, the computer vision module includes one or more of a scene detector 334, an object detection and tracking system 336, an event detector 342, a motion estimator 344, a pose estimator 346, and a depth estimator 348. The computer vision module 314 is configured to receive the image data 362 and generate the sound information 374 based on the received the image data 362. For example, the computer vision module 314 is configured to generate a plurality of candidate sounds for generative audio processing, where the candidate sounds correspond to identified objects in a scene corresponding to the image data 362. The computer vision module 314 may be configured to generally perform image segmentation and image classification. Image Segmentation may include processing image frames to divide the image thereof into different regions based on characteristics of pixels to identify the boundary of the image or regions thereof. Image Classification may include categorizing and labelling the pixels within the image frames using a predefined tags on which algorithms have been trained on.

[0092] The scene detector 334 may be configured to perform scene detection operations based on the image data 362. For example, the scene detector 334 generates the scene data 368 based on performing scene detection operations on the image data 362. The scene data 368 may include scene type information, and may be provided to one or more other modules of the device 301, including one or more sub-modules of the computer vision module 314. For example, the scene data 368 is provided to the object detection and tracking system 336 for performing object detection and / or tracking operations, such as object detection operations. To illustrate, the scene information may be used to identify objects in the scene or confirm detected objects.

[0093] The object detection and tracking system 336 includes an object detector 338 and an object tracker 340. The object detector 338 may be configured to detect objects in the image data 362, such as associated with detected scenes or event of the image frames. For example, the object detector 338 may determine to generate and place bounding boxes based on the image data 362 (such as decoded feature information thereof), and then may identify the objects in the bounding boxes based on conventional object recognition or identification methods.

[0094] The object detector 338 generates the object data 370 based on performing object detection operations. The object information may include object position and type information and may be provided to one or more other modules of the device 301, including one or more sub-modules of the computer vision module 314. For example, the object data 370 is provided to the object tracker 340 for performing object tracking operations to generate tracking data which indicates object motion information. As other examples, the object data 370 is provided to one or more of the event detector 342, the motion estimator 344, the pose estimator 346, or the depth estimator 348.

[0095] The object tracker 340 may be configured to track objects detected by the object detector 338. For example, the object tracker 340 may be configured to generate tracking data (e.g., object tracking information) based on the object data 370 (e.g., object detection information thereof) received from the object detector 338. To illustrate, the object tracker 340 may be configured to generate tracking data (e.g., obj ect tracking information) which accounts for object movement based on the image data 362 and based on the detected objects of the object data 370. The tracking data (e.g., object tracking information) may be included with the object data 370 and provided to one or more other modules of the device 301, including one or more sub-modules of the computer vision module 314.

[0096] The event detector 342 may be configured to perform event detection operations based on the image data 362. For example, the event detector 342 generates event data based on performing event detection operations on the image data 362. The event data may include event type information. The event data may be provided to one or more components of the computer vision module 314, and optionally the sound candidate generator 316. For example, the event data may be provided to the scene detector 334 for performing scene detection operations, the object detector 338 for performing object detection operations, or the object tracker 340 for performing object tracking operations, as illustrative, non-limiting examples.

[0097] The motion estimator 344 may be configured to perform motion estimation operations based on the image data 362. For example, the motion estimator 344 generates motion estimation data based on performing motion estimation operations on the image data 362 and the object data 370. The motion estimation data may include estimated movement data, such as speed or velocity data, for the detected objects. The motion estimation data may be provided to one or more components of the computer vision module 314, and optionally the sound candidate generator 316. For example, the motion estimation datamay be provided to the object tracker 340 for performing object tracking operations, as an illustrative, non-limiting example.

[0098] The pose estimator 346 may be configured to perform pose estimation operations based on the image data 362. For example, the pose estimator 346 generates pose estimation data based on performing pose estimation operations on the image data 362 and the object data 370. The pose estimation data may include estimated pose data, such as orientation or axis data, for the detected objects. The pose estimation data may be provided to one or more components of the computer vision module 314, and optionally the sound candidate generator 316. For example, the pose estimation data may be provided to the object detector 338 for performing object detection operations, the object tracker 340 for performing object tracking operations, or the motion estimator 344 for performing motion estimation operations, as illustrative, non-limiting examples.

[0099] The depth estimator 348 may be configured to perform pose estimation operations based on the image data 362. For example, the depth estimator 348 generates depth estimation data based on performing depth estimation operations on the image data 362 and the object data 370. The depth estimation data may include estimated depth data, such as distance or coordinate data, for the detected objects. The depth estimation data may be provided to one or more components of the computer vision module 314. For example, the pose estimation data may be provided to the object tracker 340 for performing object tracking operations or the motion estimator 344 for performing motion estimation operations, as illustrative, non-limiting examples. Additional details on computer vision are described further with reference to FIGS. 4-6.

[0100] The device 301 also includes a sound candidate generator 316 (e.g., a sound candidate module) including one or more sub-systems or modules. As illustrated in the example of FIG. 4, the sound candidate generator 316 includes a sound identifier 350 and a candidate sound manager 332. The sound candidate generator 316 may be configured to generate one or more candidate sounds for potential generation, reproduction, or enhancement by the generative audio system 318 based on received computer vision information from the computer vision module 314, and optionally based on one or more the input audio (e.g., the audio input data 360) or the audio perception information from the audio perception module 310. Additionally, or alternatively, the sound candidate generator 316 may be configured to select one or more of the generated candidate sounds for generation, reproduction, or enhancement by the generative audio system 318 based on one or more of the user input information, user preference information, the input audio, or the audioperception information from the audio perception module. Although the sound candidate generator 316 is illustrated as separate from the computer vision module 314, in some implementations, the sound candidate generator 316 may be included in or part of the computer vision module 314 or the generative audio system 318.

[0101] The sound identifier 350 of the sound candidate generator 316 may include or correspond to an image-based sound identifier, as compared to the audio-based the sound identifier 330 of the audio perception module 310. The sound identifier 350 may be configured to identify sounds based on object information, scene information, and / or event information from the computer vision module 314. For example, the sound identifier 350 may be configured to generate sound tag information for the scene itself, for object identified in the scene, for events identified in the scene or image data, user input, or a combination thereof. Additionally, or alternatively, the sound identifier 350 may identify sounds based on audio input data or sound tag information received from the audio perception module and determined based on the input audio data. The sound identifier 350 may use one or more databases to identify sounds associated with identified scenes, objects, and events, such as the scene database information 376 or the sound database information 378.

[0102] The candidate sound manager 332 of the sound candidate generator 316 may include or correspond to the candidate sound manager 332 of the audio perception module 310. The candidate sound manager 332 may be configured to select candidate sounds for output by the sound candidate generator 316 and for generation by the generative audio system 318. For example, the candidate sound manager 332 may be configured to select a subset of identified candidate sounds for generation based on user preferences and / or inputs. Additionally, or alternatively, the candidate sound manager 332 may be configured to select or add candidate sounds based on received audio sound tag information from the audio perception module 310 and / or based on the input audio information. To illustrate, the candidate sound manager 332 may be configured to only select certain sounds for generation based on user preference and / or may only select original sounds for generation based on the identified sounds in the input audio.

[0103] The sound candidate generator 316 may then provide the selected candidate sounds as sounds to be generated for the generative audio information to the generative audio system 318. For example, the sound candidate generator 316 may provide the selected sound tags to the generative audio system 318.

[0104] The device 301 further includes a generative audio system 318 including one or more sub-systems or modules. As illustrated in the example of FIG. 3, the generative audiosystem 318 includes a sound generator 352, a sound mixer 354, a sound selector 356, and an error correction module 358. The generative audio system 318 is configured to generate generative audio, such as one or more sounds created by artificial intelligence. The generative audio system 318 may utilize one or more neural networks to create sounds using Al and / or ML techniques. The generative audio system 318 may use a neural network and / or LLM to learn statistical properties of the input audio, and then reproduce generative audio for the device using a database of sounds. The generative audio system 318 may use a pre-configured or stored neural network and / or LLM, which it may update or revise over time during operation, and / or receive updated neural network and / or LLM information from time-to-time.

[0105] The sound generator 352 may be configured to create or generate the generative audio information based on one or more of the input audio data, the identified sounds tags, the sound parameter information, user input information, or user preference information. The identified sounds tags may include the sound tag information from computer vision module 314, the sound tag information from the audio perception module 310, or both. The sound parameter information may include the sound parameter information from computer vision module 314, the sound parameter information from the audio perception module 310, or both.

[0106] The sound mixer 354 may be configured to modify the generative audio information, the input audio, or both, based on user inputs and / or preferences. For example, the sound mixer 354 may perform sound mixing operations, equalization, balancing, sound location modification, etc. to combine the generated generative sound information 380 into generative audio information 382. The sound mixer 354 may be configured to perform sound mixing based on user inputs, and based on sound(s) selected by the sound selector 356.

[0107] The sound selector 356 may be configured to compare the generative audio information to the input audio information to adjust the generative audio information based on input audio information. For example, the error correction module 358 may remove generative audio information 382 that is not present in the actual scene, that is not associated with the scene, sounds, or objects in the scene and / or selected candidate generative sounds for inclusion in the output generative audio information 382.

[0108] The error correction module 358 may be configured to compare the generative audio information to the input audio information to adjust the generative audio information based on input audio information. For example, the error correction module 358 mayremove generative audio information 382 that is not present in the actual scene or that is not associated with the scene, sounds, or objects in the scene. To illustrate, the error correction module 358 may instruct the sound selector 356 to unselect certain generative sounds for sound mixing or output which are not identified in the input audio. As another example, the error correction module 358 may modify the generative audio information 382 based on the actual capture audio information and / or user preferences. To illustrate, the error correction module 358 may instruct the sound generator 352 to regenerate one or more generative sounds and / or instruct the sound mixer 354 to adjust or modify one or more properties of the generative sounds based on sound parameters of the corresponding sounds in the input audio.

[0109] The device 301 further includes a wireless interface 320 configured to perform wireless communications. For example, the wireless interface 320 may include one or more components, such as a transceiver, an encoder, a decoder, and one or more antennas. The wireless interface 320 may be configured to wirelessly transmit and receive communications. The communications may include audio data, image data, generative audio data, or a combination thereof. The wireless interface 320 may enable the device 301 to transmit its generative audio information 382 for its own captured videos and live photos, and enables the device 301 to receive image data captured by another device and to add generative audio information 382 for the received image data.

[0110] During operation, the device 301 may perform computer vision and optionally audio perception to engage in operations to enhance videos or live photos with generative audio. The device 301 captures image data 362 using the cameras 324, and optionally captures the audio input data 360. The device 301 may determine sounds associated with a scene and objects from the image data 362 from performing computer vision on image data 362, and the sounds are used as input for the generative audio creation operations by the generative audio system 318. The resulting output of the generative audio system 318 may then be provided to output system 308 for playback of enhanced videos or live photos with generative audio. Because the generative audio was determined based on the image data, and optionally the input audio data, the user has an enhanced and immersive experience for videos and live photos as compared to conventional systems which may not output any audio or only output input audio or processed input audio. Detailed operations of the device 301, are described further with reference to FIGS. 4-6.[OHl] In some implementations, the device 301 may be configured to transmit the captured image data 362 and the generative audio information 382 via wireless interface 320 toanother device, such as another UE. Additionally, or alternatively, the device 301 may be configured to receive image data and optionally audio data via wireless interface 320 from another device, such as another UE, and that was captured by the other device. The device 301 may be configured to generate and output generative audio for the received image data based on performing computer vision on the received image data. The device 301 may also be configured to transmit the generative audio back to the other device, and / or one or more other devices to enhance received image data with generative audio.

[0112] In the example of FIG. 3, the device 301 may be able to output captured image data with generative audio information to improve user immersion and user experience. Accordingly, device performance and user experience may be increased by enhancing videos and live photos with generative audio.

[0113] Referring to FIGS. 4-6, FIGS. 4-6 correspond to examples of different example operations for creation of generative audio for live photos and videos. The example operations for creation of generative audio for live photos and videos of FIGS. 4-6 may be performed by device 301 of FIG. 3. FIG. 4 corresponds to an example of adding generative audio to captured image data or replacing captured audio data with generative audio, based on sounds determined from the captured image data alone. FIG. 5 corresponds to an example of adding user selected or user modified generative audio to captured image data or replacing captured audio data with user selected or user modified generative audio, based on sounds determined from the captured image data and based on user inputs or preferences. FIG. 6 corresponds to an example of adding generative audio to captured image data or replacing captured audio data with generative audio, based on sounds determined from the captured image data and the audio input data.

[0114] Referring to FIG. 4, a block diagram 400 of one example of creation of generative audio for live photos and videos. The example operations of FIG. 4 may be performed by a device as described herein, such as the device 301 of FIG. 3 and one or more components thereof. In the example of FIG. 4, the example includes capturing input image data for multiple image frames at 410, performing computer vision on the input image data at 412, creating generative audio information at 414, and outputting image data 418 and the generative audio information 420 at 416.

[0115] At 410, the device captures input image data for multiple image frames. For example, the camera 324 captures image signals and generates the image data 362 for multiple image frames, as described with reference to FIG. 3.

[0116] At 412, the device performs computer vision on the input image data. For example, the computer vision module 314 receives the image data 362, or processed image data based thereon, and generates visual perception information 372 and sound information 374 for the multiple image frames based on the image data 362, as described with reference to FIG. 3. To illustrate, the computer vision module 314 and the sound candidate generator 316 generate sound tag and parameter information for identified objects in the scene for the multiple frames based on the image data 362. The computer vision module 314 performs computer vision on the image data 362 to determine a scene and object and corresponding visual perception information 372, and provides scene and object information to the sound candidate generator 316. The sound candidate generator 316 generates sound tag and parameter information based on the scene and object information, and optionally the visual perception information 372.

[0117] At 414, the device creates generative audio information. For example, the generative audio system 318 receives the visual perception information 372 and sound information 374 and generates the generative audio information 382 for the multiple image frames based on the image data 362, as described with reference to FIG. 3.

[0118] At 416, the device outputs the image data 418 and the generative audio information 420. For example, the output system 308 outputs the output data 384 based on the image data 362 and including the generative audio information 382, as described with reference to FIG. 3. To illustrate, the device outputs the generative audio information 382 via the speaker 328 and outputs image data based on the image data 362 via the display 326.

[0119] In a particular use case for the example operations for creation of generative audio for live photos and videos of FIG. 4 described with reference to FIG. 3, the device 301 may capture input image data 362 using the camera 324. The device 301 may optionally capture input audio data 360 using one or more microphones, such as microphone 322. The captured input audio data 360 may be used by the device 301 to generate generative audio information 382 for the captured input image data 362. In the particular use case, the device 301 captures input image data 362 of a beach scene including sand, multiple palm trees, and an ocean. In a few frames of the multiple image frames captured by the device 301, a seagull is captured flying overhead. Over the multiple frames, leaves of the palm trees sway in the wind and waves flow through the ocean and reach the shore.

[0120] The image data 362 from the multiple frames is provided to the computer vision module 314 for computer vision processing. For example, the computer vision module 314 (e.g., the scene detector 334 thereof) performs scene detection to identify the beach scenes inthe frames of the image data 362 and generates the scene data 368 identifying the beach scene. The computer vision module 314, such as the object detection and tracking system 336, may then perform object identification and tracking on the image data 362. For example, the object detector 338 may utilize the scene data 368, the beach scene, to identify objects in the scene, such as the palm trees, leaves, sea gull, etc. The object detector 338 may use image segmentation and object classification operations to identify the objects. Once the objects are identified and tagged, the computer vision module 314 may track the identified object from frame to frame and determine a motion estimation for the objects. To illustrate, the object tracker 340 may determine changes in positions for the objects and the motion estimator 344 may estimate the motion of the objects based on the position changes from frame-to-frame. As described with reference to FIG. 3, the computer vision module 314 may also employ event, depth, and pose estimation for identified object to further estimate the motion of the object, the orientation of the object, and the distance to the object.

[0121] The computer vision module 314 may determine one or more sounds for generation by the generative audio system 318, and output one or more indications of the sounds. For example, for the particular use case the computer vision module 314 may determine to generate a wave sound (or wave sounds), a leaves rustling sound (e.g., palm tree leave rustling sound), a wind sound (e.g., an ocean breeze sound), and a sea gull. The computer vision module 314 may output sound tag information identifying the sounds to the generative audio system 318.

[0122] Additionally, the computer vision module 314 may provide additional information for generating the sounds. For example, the computer vision module 314 may output sound parameter information for one or more of the identified sounds. For example, the computer vision module 314 may output frequency information, intensity information, location of the sound information, movement of the sound information, etc., so that the generative audio system 318 can generate an artificial sound that correspond to and is tailored for the input image data. To illustrate, the frequency and intensity of the wave sounds may correspond to the timing of the waves crashing, the size of the waves, and the depth of the waves (e.g., distance to the shore) in the scene.

[0123] The generative audio system 318 then generates one or more generative sounds based on the identified sounds and the sound parameter information. For example, in the use case for FIG. 4, the generative audio system 318 generates a wave sound including a series ofsounds of waves hitting the sand and receding into the ocean corresponding to the identified waves in the input image data 362.

[0124] The generative audio system 318 also generates wind sounds for the image frames to generate sounds for the ocean breeze. The wind sounds include an ocean breeze sound based on the movement of the waves and / or the movement of the palm trees to match the intensity and frequency of the actual wind captured in the input image data 362.

[0125] The generative audio system 318 also generates leaf rustling sounds for the image frames to generate sounds for the palm trees moving in the wind. The leaf rustling sounds may include palm tree leaves rustling on one another, on other branches, and on other palm trees to match the intensity and frequency of the actual leaves rustling captured in the input image data 362.

[0126] The generative audio system 318 also generates sea gulls sounds for select image frames in which the sea gull was identified. The sea gull sounds may include wing flapping sounds and sea gull squawking sounds to correspond to sea gull captured in the image data. The sea gull sounds may be generated even if the sea gull made no noises (e.g., no cawing or squawking noises) during the captured input image data 362, and / or even if the audio or specific sounds was not captured (e.g., too inaudible or drowned out by other sounds or input audio processing, such as noise cancellation).

[0127] The generative audio sounds may be combined or mixed together to generate generative audio information, and the generative audio information may be synchronized with the image data 362, such as the input image data or modified image data. The device 301 then outputs the image data 362 with the generative audio information 382 via the display 326 and speakers 328 to a user as a video or a live photo. To illustrate, the user may watch the video or live photo and hear the generative audio including one more generative sounds which enhances the user experience.

[0128] In some implementations, the original audio for the scene is not captured, or even if the audio is captured it may not be played back with the image data. That is the generative audio replaces the actual audio for the video or live photo. In other implementations, the original audio, or a portion thereof, may be supplemented with the generative audio, as further described with reference to FIGS. 5 and 6.

[0129] In a particular implementation, the generative audio output by the device 301 may include spatial audio or spatial audio information. For example, one or more of the generative sounds may correspond to spatial sound or spatial audio. The spatial sounds may be generated to correspond to movement of objects throughout the scene. For example, thesea gull sounds may be spatial audio and the generated audio sounds may be generated based on movement information for identified sea gull to simulate surround sounds.

[0130] Referring to FIG. 5, a block diagram 500 of one example of creation of generative audio for live photos and videos. The example operations of FIG. 5 may be performed by a device as described herein, such as the device 301 of FIG. 3. In the example of FIG. 5, the example includes capturing input image data for multiple image frames at 510, performing computer vision on the input image data at 512, creating generative audio information at 514 based on user preferences 522, and outputting image data 518 and the generative audio information 520 based on user input 524.

[0131] At 510, the device capture input image data for multiple image frames. For example, the camera 324 captures the image data 362 for multiple image frames, as described with reference to FIG. 3.

[0132] At 512, the device performs computer vision on the input image data. For example, the computer vision module 314 receives the image data 362 and generates visual perception information 372 and sound information 374 for the multiple image frames based on the image data 362, as described with reference to FIG. 3. To illustrate, the computer vision module 314 and the sound candidate generator 316 generate sound tag and parameter information for identified objects in the scene for the multiple frames based on the image data 362.

[0133] At 514, the device creates generative audio information based on user preferences 522. For example, the generative audio system 318 receives the visual perception information 372 and sound information 374 and generates the generative audio information 382 for the multiple image frames based on the image data 362, as described with reference to FIG. 3.

[0134] At 516, the device outputs the image data 518 and the generative audio information 520 based on user input 524. For example, the output system 308 outputs the output data 384 based on the image data 362 and including the generative audio information 382, as described with reference to FIG. 3. To illustrate, the device outputs the generative audio information 382 via the speaker 328 and outputs image data based on the image data 362 via the display 326.

[0135] In a particular use case for the operations for creation of generative audio for live photos and videos of FIG. 5 described with reference to FIG. 3, the device 301 may capture input image data 362 using the camera 324. The device 301 may optionally capture input audio data 360 using one or more microphones, such as microphone 322. The captured inputaudio data 360 may be used by the device 301 to generate generative audio information 382 for the captured input image data 362. In the particular use case, the device 301 captures input image data 362 of a beach scene including sand, multiple palm trees, and an ocean. In a few frames of the multiple image frames captured by the device 301, a seagull is captured flying overhead. Over the multiple frames, leaves of the palm trees sway in the wind and waves flow through the ocean and reach the shore.

[0136] The image data 362 from the multiple frames is provided to the computer vision module 314 for computer vision processing. For example, the computer vision module 314 (e.g., the scene detector 334 thereof) performs scene detection to identify the beach scenes in the frames of the image data 362 and generates the scene data 368 identifying the beach scene. The computer vision module 314, such as the object detection and tracking system 336, may then perform object identification and tracking on the image data 362. For example, an object detector 338 may utilize the scene data 368, a beach scene, to identify objects in the scene, such as the palm trees, leaves, sea gull, etc. The object detector 338 may use image segmentation and object classification operations to identify the objects. Once the objects are identified and tagged, the computer vision module 314 may track the identified object from frame to frame and determine a motion estimation for the objects. To illustrate, the object tracker 340 may determine changes in positions for the objects and the motion estimator 344 may estimate the motion of the objects based on the position changes from frame-to-frame. As described with reference to FIG. 3, the computer vision module 314 may also employ event, depth, and pose estimation for identified object to further estimate the motion of the object, the orientation of the object, and the distance to the object.

[0137] The computer vision module 314 may determine one or more sounds for generation by the generative audio system 318, and output one or more indications of the sounds. For example, for the particular use case the computer vision module 314 may determine to generate a wave sound (or wave sounds), a leaves rustling sound (e.g., palm tree leave rustling sound), a wind sound (e.g., an ocean breeze sound), and a sea gull. The computer vision module 314 may output sound tag information identifying the sounds to the generative audio system 318.

[0138] Additionally, the computer vision module 314 may provide additional information for generating the sounds. For example, the computer vision module 314 may output sound parameter information for one or more of the identified sounds. For example, the computer vision module 314 may output frequency information, intensity information,location of the sound information, movement of the sound information, etc., so that the generative audio system 318 can generate an artificial sound that correspond to and is tailored for the input image data. To illustrate, the frequency and intensity of the wave sounds may correspond to the timing of the waves crashing, the size of the waves, and the depth of the waves (e.g., distance to the shore) in the scene.

[0139] The generative audio system 318 then generates one or more generative sounds based on the identified sounds and the sound parameter information, and optionally based on one or more user preferences 522, indicated by the user preference data 366. The user preferences 522 may include or correspond to user preferences for which sound or sounds to generate, which sound or sounds not to generate, how to generate all sounds, how to generate certain sounds, which sounds to modify or enhance, which sounds to not modify or enhance, which original / input sounds to remove, which original / input sounds to keep, etc., or a combination thereof.

[0140] For example, in the use case for FIG. 5, the generative audio system 318 generates a wave sound including a series of sounds of waves hitting the sand and receding into the ocean corresponding to the identified waves in the input image data 362. The generative audio system 318 then modifies the generative wave sound based on the user preferences 522. To illustrate, the user preferences may indicate to emphasize wave sounds, and the generative audio system 318 increases an intensity or sound volume of the wave sounds. Although the above example is a two-stage example of generation followed by modification based on user preferences, the generative audio system 318 may instead generate a user modified wave sound based on the identified sounds, the sound parameter information, and the user preferences in one-step.

[0141] The generative audio system 318 also generates wind sounds for the image frames to generate sounds for the ocean breeze. The wind sounds include an ocean breeze sound based on the movement of the waves and / or the movement of the palm trees to match the intensity and frequency of the actual wind captured in the input image data 362. Similar to the example above in which the generative audio system 318 generates user modified generative wave sounds, the generative audio system 318 may generate user modified generative ocean breeze sounds based on a user preference of the user preferences 522. To illustrate, the generative audio system 318 may generate a user modified generative ocean breeze sound of low intensity ocean breeze sounds based on a user preference to add low intensity wind sounds.

[0142] The generative audio system 318 also generates leaf rustling sounds for the image frames to generate sounds for the palm trees moving in the wind. The leaf rustling sounds may include palm tree leaves rustling on one another, on other branches, and on other palm trees to match the intensity and frequency of the actual leaves rustling captured in the input image data 362. Similar to the example above in which the generative audio system 318 generates user modified generative wave or wind sounds, the generative audio system 318 may generate user modified generative leaf rustling sounds based on a user preference of the user preferences 522. To illustrate, the generative audio system 318 may not generate palm leaf rustling sounds based on a user preference to not generate leaf rustling sounds. As another illustration, the generative audio system 318 may generate a user modified generative palm leaf rustling sounds sound of low intensity palm leaf rustling sounds based on a user preference to add low intensity leaf rustling sounds.

[0143] The generative audio system 318 also generates sea gulls sounds for select image frames in which the sea gull was identified. The sea gull sounds may include wing flapping sounds and sea gull squawking sounds to correspond to sea gull captured in the image data 362. The sea gull sounds may be generated even if the sea gull made no noises (e.g., no cawing or squawking noises) during the captured input image data, and / or even if the sound was not captured (e.g., too inaudible or drowned out by other sounds or input audio processing, such as noise cancellation). Similar to the above examples, the generative audio system 318 may generate user modified sea gull sounds based on a user preference of the user preferences 522. To illustrate, the generative audio system 318 may not generate any sea gull squawking sounds based on a user preference to not generate bird chirping sounds and may generate sea gull wing flapping sounds based on a user preference to generate bird flying sounds. As another illustration, the generative audio system 318 may generate a user modified sea gull sound including spatial audio for both sea gull wing flapping and squawking based on a user preference to add spatial audio for object that move in and out of a scene, for animals, etc.

[0144] The generative audio sounds may be combined or mixed together to generate generative audio information 382, and the generative audio information 382 may be synchronized with the image data 362, such as the input image data or modified image data. The device may store or otherwise associate the image data with the generative audio information. The device 301 may then output the image data 362 with the generative audio information 382 via the display 326 and the speakers 328 to a user as a video or a live photo. Toillustrate, the user may watch the video or live photo and hear the generative audio including one more generative sounds which enhances the user experience.

[0145] In some implementations, the device 301 outputs the synchronized image and generative audio data based on one or more user inputs. For example, an indication of the generative sounds of the generative audio information 382 may be provide to a user prior to, concurrently with, or after output of the output data 384, that is the synchronized image and generative audio data. The user may then select one or more of the generative sounds of the list for output by the device 301. For example, the generative audio system 318 may generate one or more generative sounds, some of which may be modified or generated based on user preferences, and then the user may select one or more of the generative sounds for inclusion into the generative audio data for output by the device. To illustrate, the user may select to only have wave generative wave sounds output with the image data. As another illustration, the user may select to send the image data and certain generative sounds to another device, such as sea gull and wave sounds only.

[0146] As yet another example, the user input may include a spatial audio input selection and the device 301 may adjust a generative sound to be a generative spatial sound. To illustrate, the device 301 may generate the generative sea gull sound as a standard sound, and the device 301 may modify the generative sea gull sound to be a spatial sound or regenerate the generative sea gull sound as a spatial sound.

[0147] Additionally, or alternatively, in implementations where the original audio is captured, the user input may include user inputs relating to input audio data 360. For example, the user inputs may include or correspond to selections of which original audio sounds to mix with the generative audio sounds or to replace the generative audio sounds. For example, the device 301 may output indications of identified sounds in the input audio data 360, and the user may select which sound or sounds to keep, which to remove, which to replace or enhance with generative sound data, or a combination thereof. The device 301 may then output the image data 362 with a portion of the audio input data 360 and with generative audio information 382. To illustrate, the user may select to keep the original leaf rustling noises and not to use the generative sound data for the leaf rustling noises. Additional description of outputting a portion of the input audio data or modifying the generative sound based on the input audio data is further described with reference to FIG. 6.

[0148] Referring to FIG. 6, a block diagram 600 of one example of creation of generative audio for live photos and videos. The example operations of FIG. 6 may be performed by adevice as described herein, such as the device 301 of FIG. 3. In the example of FIG. 6, the example includes capturing input image data for multiple image frames at 610, performing computer vision on the input image data at 612, and creating generative audio information at 614.

[0149] Additionally, the example of FIG. 6 includes capturing input audio data for the multiple image frames at 630, performing audio perception on the input audio data at 632, and comparing the input audio data and the generative audio information at 634, outputting image data 618 and the input and generative audio information 620 based on the comparison and user input 624.

[0150] At 610, the device capture input image data for multiple image frames. For example, the camera 324 captures the image data 362 for multiple image frames, as described with reference to FIG. 3.

[0151] At 630, the device capture input audio data for the multiple image frames. For example, the microphone 322 captures the audio input data 360 for the multiple image frames, as described with reference to FIG. 3.

[0152] At 612, the device performs computer vision on the input image data. For example, the computer vision module 314 receives the image data 362 and generates visual perception information 372 and sound information 374 for the multiple image frames based on the image data 362, as described with reference to FIG. 3. To illustrate, the computer vision module 314 and the sound candidate generator 316 generate sound tag and parameter information for identified objects in the scene for the multiple frames based on the image data 362.

[0153] At 632, the device performs audio perception on the input audio data. For example, the audio perception module 310 receives the audio input data 360 and generates audio perception information, such as sound information 374 for the multiple image frames based on the audio input data 360, as described with reference to FIG. 3. To illustrate, the audio perception module 310 generates sound tag information for identified sounds in the audio input data 360.

[0154] At 614, the device creates generative audio information. For example, the generative audio system 318 receives the visual perception information 372 and sound information 374 from and generates the generative audio information 382 for the multiple image frames based on the image data 362, as described with reference to FIG. 3.

[0155] At 634, the device compares the generative audio information and input audio information. For example, the sound selector 356 and the error correction module 358receive the audio input data 360 and the candidate generative sound information and compare the audio input data 360 and the candidate generative sound information to determine which candidate generative sounds to select for inclusion into the generative audio information 382 and which candidate generative sound to modify for the generative audio information 382 based on the comparison, as described with reference to FIG. 3.

[0156] At 636, the device performs sound selection on the generative audio information and input audio information based on the comparison. For example, the sound selector 352 may determine which sounds or audio portions, from the generative audio information 382 and input audio information 360, to output or include in the output data 384, and the sound mixer 354 mixes the sound and generates the output based on the selections of the sound selector 352. To the device 301 determines to use generative audio for two sounds of the input audio data and to keep original or unmodified input audio data pertaining to one original sound. When using generative audio, the generative audio for a particular sound may be added to the original audio to add a new sound, may be used or substituted to replace an identified original sound in the original audio, or may be used to enhance or modify the original audio to generate a modified sound of original and generated audio. In some implementations, the sound selection is further based on user preferences or user inputs, such as user input 624.

[0157] At 616, the device outputs the image data 618 and the generative audio information 620 based on user input 624. For example, the output system 308 outputs the output data 384 based on the image data 362 and including the generative audio information 382, as described with reference to FIG. 3. To illustrate, the device 301 outputs the generative audio information 382, and optionally a portion of the input audio data 360, via the speaker 328 and outputs image data based on the image data 362 via the display 326. Because the output data, output data 384, including the sounds of the generative audio information 382, was determined based on the user selection and / or preference, the audio output of the output data 384 corresponds to sounds (generative and optionally original) which reflect the desired user experience.

[0158] In a particular use case for the operations for creation of generative audio for live photos and videos of FIG. 6 described with reference to FIG. 3, the device 301 captures input image data 362 using the camera 324 and captures input audio data 360 using one or more microphones 322. The captured input audio data 360 is used by the device to generate and / or refine the generative audio information 382 for the captured input image data 362. In the particular use case, the device 301 captures input image data 362 of a beach sceneincluding sand, multiple palm trees, and an ocean and captures corresponding audio input data 360 for the image data 362. In a few frames of the multiple image frames captured by the device 301, a seagull is captured flying overhead. Over the multiple frames, leaves of the palm trees sway in the wind and waves flow through the ocean and reach the shore.

[0159] The audio input data 360 may include corresponding wave sounds, wind sounds, tree sounds, and bird sounds. Additionally, the audio input data 360 may include other sounds not capturable in the image data 362, such as sea gull sounds when the sea gull is not in the image frame, human talking sounds, other animal sounds, vehicle noise, etc. In some implementations, the one or more of the sounds may not be captured by the microphones 322. For example, the sea gull sounds may not be captured if they are too low or far away, or may be obscured by other more intense and closer sounds. Additionally, or alternatively, the device 301 may process the input audio data 360 and remove one or more sounds or type of noise from the phone to generate processed input audio. For example, the device 301 may use noise cancelling techniques to remove wind noise captured by the microphone 322.

[0160] Similarly, the audio data for the multiple frames, such as the original input audio data 360 or the processed input audio data, is provided to the audio perception module 310 for audio perception processing. For example, the audio perception module 310 performs sound detection to identify one or more sounds captured during the frames of the image data. The audio perception module 310 may perform audio image segmentation and sound classification operations to identify the sounds in the input audio data 360. Once the sounds are identified and tagged, the audio perception module 310 may output the identified sounds to the computer vision module 314, the generative audio system 318, or both, such as sound information 374.

[0161] The image data 362 from the multiple frames is also provided to the computer vision module 314 for computer vision processing, as described with reference to any of FIGS. 3-5. For example, the computer vision module 314 performs scene detection to identify the beach scenes in the frames of the image data. However, in the example of FIG. 6, the computer vision module 314 may receive sound tag information from the audio perception module 310 and may generate its output, visual perception information 372 and / or sound information 374, based on the received sound tag information (e.g., visual perception information and / or sound information 374) from the audio perception module 310. For example, the computer vision module 314 generates or refines the output or candidate sounds for generative audio processing based on the actual identified sounds inthe scene. To illustrate, the computer vision module 314 may either add sounds identified in the audio data 360 to the sound identified based on the image data 362 and / or remove sounds not identified in the audio data 360 from the sound identified based on the image data 362.

[0162] Additionally, or alternatively, the computer vision module 314 may perform one or more of the intermediary processing steps based on the sounds identified in the audio data 360. For example, the computer vision module 314 may identify the scene and / or objects in the scene based on the sounds identified in the audio data 360. To illustrate, the scene detector 334 may detect a beach scene based on identified wave and sea gull sounds in the audio sounds tags. As another example, the event detector 342 may detect a particular event based on sounds identified from the audio information. The detected scenes, objects, and / or events, which were detected in part based on the audio input data 360, may then be used to determine the candidate sounds for generation by the generative audio system 318.

[0163] In the example of FIG. 6, for the particular use case the computer vision module 314 may determine to generate a wave sound (or wave sounds), a leaves rustling sound (e.g., palm tree leave rustling sound), a wind sound (e.g., an ocean breeze sound), and a sea gull. The computer vision module 314 may output sound tag information identifying the sounds to the generative audio system 318. The audio perception module 310 may also output sound tag information identifying the sounds to the generative audio system 318, in addition to or in the alternative of, providing the output sound tag information to the computer vision module. For example, the audio perception module 310 outputs sound tag information (e.g., audio sound tag information) identifying the sounds identified in the captured input or processed audio to the generative audio system 318. In the example of FIG. 6, for the particular use case the audio perception module 310 outputs sound tag information identifying wave sounds, wind sounds, bird sounds, and airplane sounds identified in the captured input audio data 360.

[0164] In some implementations, the computer vision module 314 may provide additional information for generating the sounds as described with reference to FIGS. 3-5. Additionally, or alternatively, the audio perception module 310 may provide additional information for generating the sounds. For example, the audio perception module 310 may output sound parameter information, sound information 374 and / or audio perception information, for one or more of the identified sounds, identified based on the audio input data 360. For example, the audio perception module 310 may output frequencyinformation, intensity information, location of the sound information, movement of the sound information, etc., so that the generative audio system 318 can generate an artificial sound that corresponds to and is tailored for the input image data 362 and that are based on the actual sounds (and corresponding features / parameters) of the captured audio input data 360. To illustrate, the frequency and intensity of the wave sounds may correspond to the timing of the waves crashing, the size of the waves, and the depth of the waves (e.g., distance to the shore) in the scene. This sound parameter information may be used by the generative audio system 318, such as the error correction module 358 thereof, to generate or refine the generative sounds.

[0165] The generative audio system 318 then generates one or more generative sounds based on the identified sounds, from the input image data 362 and optionally the input audio data 360, and the sound parameter information. The generative sounds may optionally be further determined based on one or more user preferences, such as user preferences 522, as described with reference to FIG. 5. The user preferences, user preference data 366, may include or correspond to user preferences for which sound or sounds to generate, which sound or sounds not to generate, how to generate all sounds, how to generate certain sounds, which sounds to modify or enhance, which sounds to not modify or enhance, which original / input sounds to remove, which original / input sounds to keep, etc., or a combination thereof.

[0166] In the use case for FIG. 6, the generative audio system 318 generates a wave sound including a series of sounds of waves hitting the sand and receding into the ocean corresponding to the identified waves in the input image data 362 and / or the identified wave sounds in the input audio data 360. The generative audio system 318 may generate the wave sounds based on the sound parameter information from the computer vision module 314, the sound parameter information from the audio perception module 310, or both. To illustrate, the user preferences may indicate to use sound parameter information from both sources, or to prioritize the sound parameters from the actual input audio data 360 (or processed input sound) to generate the generative wave sounds.

[0167] The generative audio system 318 also generates wind sounds for the image frames to generate sounds for the ocean breeze. The wind sounds include an ocean breeze sound based on the movement of the waves and / or the movement of the palm trees to match the intensity and frequency of the actual wind captured in the input image data and / or the identified wind sounds in the input audio data. Similar to the example above, the generative audio system 318 may generate generative ocean breeze sounds based on thesound parameter information from the computer vision module 314, the sound parameter information from the audio perception module 310, or both, and optionally based on user preferences.

[0168] The generative audio system 318 also generates leaf rustling sounds for the image frames to generate sounds for the palm trees moving in the wind. The leaf rustling sounds may include palm tree leaves rustling on one another, on other branches, and on other palm trees to match the intensity and frequency of the actual leaves rustling captured in the input image data and / or the identified leaf rustling sounds in the input audio data. Similar to the example above, the generative audio system 318 may generate generative leaf rustling sounds based on the sound parameter information from the computer vision module 314, the sound parameter information from the audio perception module 310 or both, and optionally based on user preferences.

[0169] The generative audio system 318 also generates sea gulls sounds for select image frames in which the sea gull was identified and / or for select image frames in which the sea gull was identified in corresponding audio frames. The sea gull sounds may include wing flapping sounds and sea gull squawking sounds to correspond to the sea gull captured in the image data and / or the identified sea gull sounds in the input audio data. The sea gull sounds may be generated even if the sea gull made no noises (e.g., no cawing or squawking noises) during the captured input image data, and / or even if the sound was not captured (e.g., too inaudible or drowned out by other sounds or input audio processing, such as noise cancellation). Similar to the above examples, the generative audio system 318 may generate generative sea gull sounds based on the sound parameter information from the computer vision module 314, the sound parameter information from the audio perception module 310, or both, and optionally based on user preferences.

[0170] The generative audio sounds may be combined or mixed together to generate generative audio information 382, and the generative audio information 382 may be synchronized with the image data, such as the input image data 362 or modified image data. The device 301 may store or otherwise associate the image data with the generative audio information 382. The device 301 may then output the image data with the generative audio information 382, as output data 384, via the display 326 and the speakers 328 to a user as a video or a live photo. To illustrate, the user may watch the video or live photo and hear a portion of the captured audio mixed with the generative audio including one more generative sounds, which enhances the user experience.

[0171] In some implementations, as illustrated in the example of FIG. 6, the generative audio system 318 outputs the generated generative audio information 382, such as candidate generative audio information with candidate generative sounds, to the audio comparison module for modification or correction and / or for selection or refinement of sounds. The audio comparison module also receives the input audio data 360 (or processed input audio data), and compares the audio data to the generative audio information. The audio comparison module, such as the generative audio system 318 or the error correction module 358 thereof, generates output audio information for outputting with the image data 362 based on the comparison.

[0172] For example, the audio comparison module may output sound selection information indicating which original sounds to use or keep, and which candidate generative sounds to keep for sound selection and mixing for the output audio information. The sound selection or mixing may be further performed based on user inputs or preferences, such as user input 624. To illustrate, the audio comparison module of the device 301 may output a preliminary indication of which sounds to keep or include, and this selection or preliminary recommendation may be further refined or confirmed by the user. To illustrate, the user may indicate to keep certain original sounds or to add more additional generative sounds.

[0173] As another example, the audio comparison module may refine or instruct the generative audio system 318 to refine the generative audio information 382 based on the comparison. To illustrate, generative audio for a particular sound may be determined to be too different from the original audio, such as by a comparison score exceeding a threshold score, and the generative audio may be fed back to the generative audio system 318 for recreation or refinement. The refined or recreated generative audio, along with the generative audio for other sounds, may be provided from the generative audio system 318 to the sound selector for output sound generation. In some such implementations, the sound selection module outputs a combination of a portion of the input audio data 360 and the refined generative audio information 382. To illustrate, the sound selection module may remove particular input sounds that are duplicates of the generative sounds based on the comparison and output received from the audio comparison module.

[0174] In the use case for FIG. 6, the sound selection module removes the wind and waves sounds from input audio information and modifies the location, intensity, and / or frequency of the generative wind and waves sounds based on the location, intensity, and / or frequency of the removed wind and waves sounds based on the comparison result received from theaudio comparison module. The sound selection module generates and outputs output sound information including the revised generative sound (generative audio information) and optionally a portion of the input audio data 360.

[0175] The device outputs the output sound information and the image data via the display 326 and the speakers 328 as enhanced video or live photo data that includes generative audio. As described with reference to FIG. 5, the example operations in FIG. 6 to generate the generative audio information 382 may be further based on user input information, such as user input 624. In some implementations, the user input 624 may be used to select or refine the generative sounds after generation. For example, the user input 624 may be received prior to or after performance of the comparison operations by the comparison module. To illustrate, the candidate generative sounds generated by the generative audio system 318 or the sounds tags provided to the generative audio system 318 may be selected or refined based on user input, such as described with reference to FIGS. 3 and 5. As another illustration, the generative sounds output by the comparison module may be selected or refined based on user input, such as described with reference to FIGS. 3 and 5. In some implementations, the user inputs may include or correspond to sound mixing inputs for the sound mixer 354 to adjust and / or mix the generative sounds, adjust or combine generative sounds of the generative sound information 380.

[0176] In some implementations, the user input may include a spatial audio input selection and the device may adjust a generative sound to be a generative spatial sound. To illustrate, the device may generate the generative sea gull sound as a standard sound, and the device may modify the generative sea gull sound to be a spatial sound or regenerate the generative sea gull sound as a spatial sound.

[0177] Referring to FIG. 7, FIG. 7 is a flow chart illustrating a method 700 for generative audio enhancement operations for videos and live photos according to some embodiments of the disclosure. In some implementations, the method may be performed by the SoC 100 of FIGS. 1 or 2, the system 200 of FIG. 2, or the device 301 of FIG. 3. For example, the generative audio system 206 of FIG. 2, or one or more of the computer vision module 314, the sound candidate generator 316, and the generative audio system 318 may perform generative audio enhancement operations for a sequence of image frames corresponding to a video, including a live photo.

[0178] The method 700 includes, at block 702, obtaining image data for a plurality of image frames associated with a live image. For example, the camera 324 of the sensor system 306 captures image data 362 for a plurality of image frames associated with a video, asdescribed with reference to FIGS. 3-6. As another example, the wireless interface 320 receives image data 362 from and captured by another device. As yet another example, the computer vision module 314 receives the image data 362 for a plurality of image frames associated with a video.

[0179] At block 704, the method 700 includes performing computer vision for the plurality of image frames to generate visual perception information based on the image data. For example, the computer vision module 314 performs computer vision operations on the image data 362 for the plurality of image frames to generate visual perception information 372, as described with reference to FIGS. 3-6. The computer vision operations may include the computer vision operations described with reference to FIGS. 3-6, including one or more of scene detection, object detection, object tracking, event detection, motion estimation, pose estimation, or depth estimation. Additionally, the computer vision operations may include the sound or candidate sound generation operations of the sound candidate generator 316. The computer vision module 314 may optionally generate sound tags, sound information 374, indicative of or corresponding to detected objects or scenes of the image frames.

[0180] At block 706, the method 700 includes generating generative audio information based on the visual perception information. For example, the generative audio system 318 generates the generative audio information 382 based on at least the visual perception information 372, as described with reference to FIGS. 3-6. To illustrate, the generative audio system 318 generates the generative audio information 382 based on sounds identified by the sound candidate generator 316, that were identified based on at least the visual perception information 372, and optionally, input audio data 360.

[0181] At block 708, the method 700 includes outputting the image data with the generative audio information. For example, the display 326 and the speakers 328 of the output system 208 output the image data 362 with the generative audio information 382, as described with reference to FIGS. 3-6. In some implementations, the speakers 328 output audio data based on the audio input data 360 in addition to the generative audio information 382. For example, processed input audio data and / or audio data for select sounds of the audio input data 360 may be mixed with the generative sounds of the generative audio information 382 and output by the speakers 328 with the image data 362 output via (e.g., displayed on) the display 326. Accordingly, the operations of the method 700 of FIG. 7 may enable devices enhance videos and live photos with generative audio to enhance user immersion and user experience.

[0182] In the figures, a single block may be described as performing a function or functions. The function or functions performed by that block may be performed in a single component or across multiple components, and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are described below generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure. Also, the example devices may include components other than those shown, including well-known components such as a processor, memory, and the like.

[0183] Unless specifically stated otherwise as apparent from the following discussions, it is appreciated that throughout the present application, discussions using terms such as “accessing,” “receiving,” “sending,” “using,” “selecting,” “determining,” “normalizing,” “multiplying,” “averaging,” “monitoring,” “comparing,” “applying,” “updating,” “measuring,” “deriving,” “settling,” “generating,” or the like, refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system’s registers, memories, or other such information storage, transmission, or display devices. The use of different terms referring to actions or processes of a computer system does not necessarily indicate different operations. For example, “determining” data may refer to “generating” data. As another example, “determining” data may refer to “retrieving” data.

[0184] The terms “device” and “apparatus” are not limited to one or a specific number of physical objects (such as one smartphone, one camera controller, one processing system, and so on). As used herein, a device may be any electronic device with one or more parts that may implement at least some portions of the disclosure. While the description and examples herein use the term “device” to describe various aspects of the disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. As used herein, an apparatus may include a device or a portion of the device for performing the described operations.

[0185] Certain components in a device or apparatus described as “means for accessing,” “means for receiving,” “means for sending,” “means for using,” “means for selecting,” “means for determining,” “means for normalizing,” “means for multiplying,” or other similarly- named terms referring to one or more operations on data, such as image data, may refer to processing circuitry (e.g., application specific integrated circuits (ASICs), digital signal processors (DSP), graphics processing unit (GPU), central processing unit (CPU), computer vision processor (CVP), or neural signal processor (NSP)) configured to perform the recited function through hardware, software, or a combination of hardware configured by software.

[0186] Those of skill in the art would understand that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0187] Components, the functional blocks, and the modules described herein with respect to the Figures referenced above include processors, electronics devices, hardware devices, electronics components, logical circuits, memories, software codes, firmware codes, among other examples, or any combination thereof. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, application, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, and / or functions, among other examples, whether referred to as software, firmware, middleware, microcode, hardware description language or otherwise. In addition, features discussed herein may be implemented via specialized processor circuitry, via executable instructions, or combinations thereof.

[0188] Those of skill in the art that one or more blocks (or operations) described with reference to FIG. 3 may be combined with one or more blocks (or operations) described with reference to another of the figures. For example, one or more blocks (or operations) of FIG. 3 may be combined with one or more blocks (or operations) of FIG. 1 or FIG. 2. As another example, one or more blocks associated with FIGS. 4-8 may be combined with one or more blocks (or operations) associated with FIGS. 1-3.

[0189] In a first aspect, a device includes a memory storing processor-readable code; and one or more processors coupled to the memory. The one or more processors are configured toexecute the processor-readable code to cause the one or more processors to: obtain image data for a plurality of image frames associated with a live image; perform computer vision for the plurality of image frames to generate visual perception information based on the image data; generate generative audio information based on the visual perception information; and output the image data with the generative audio information.

[0190] In a second aspect, alone or in combination with one or more of the above aspects, the generative audio information corresponds to sound information associated with at least one obj ect identified in one or more image frames of the plurality of image frames or with at least one scene identified in the one or more image frames of the plurality of image frames.

[0191] In a third aspect, alone or in combination with one or more of the above aspects, the generative audio information corresponds to spatial sound information associated with at least one object identified in one or more image frames of the plurality of image frames and representing motion of the at least one object.

[0192] In a fourth aspect, alone or in combination with one or more of the above aspects, the visual perception information includes object identification information, scene identification information, or a combination thereof.

[0193] In a fifth aspect, alone or in combination with one or more of the above aspects, the visual perception information includes one or more of frame information, object information, event detection information, motion estimation information, three-dimensional pose information, object identification information, or object tracking information.

[0194] In a sixth aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to cause the device to: obtain audio input data for the plurality of image frames, wherein the visual perception information, the generative audio information, or both are based on the audio input data, and wherein the one or more processors are configured to cause the device to output the image data with the generative audio information includes to: output the image data with the generative audio information and with at least a portion of the audio input data.

[0195] In a seventh aspect, alone or in combination with one or more of the above aspects, one or more processors are further configured to cause the device to: obtain audio input data for the plurality of image frames, wherein the visual perception information, the generative audio information, or both are based on the audio input data; and refrain from outputting the audio input data with the image data.

[0196] In an eighth aspect, alone or in combination with one or more of the above aspects, to perform computer vision for the plurality of image frames to generate visual perception information based on the image data, the one or more processors are further configured to: extract scenes from multiple image frames of the plurality of image frames and generate scene information based on the extracted scenes; perform object detection on the extracted scenes; perform object tracking for the detected objects over the multiple frames; and perform photo segmentation and context detection on the detected and tracked objects to identify candidate sounds for the plurality of image frames, wherein the candidate sounds are indicated by the visual perception information.

[0197] In a ninth aspect, alone or in combination with one or more of the above aspects, to generate the generative audio information the one or more processors are configured to: obtain audio input data for the plurality of image frames; identify sounds in the audio input data; wherein to perform the computer vision to generate the visual perception information the one or more processors are configured to: generate the visual perception information based on the image data and based on identified sounds from the audio input data.

[0198] In a tenth aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to: obtain audio input data for the plurality of image frames, determine a plurality of candidate sounds for generative audio processing based on object information, scene information, or both of the visual perception information, wherein to generate the generative audio information the one or more processors are configured to: generate the generative audio information based on the plurality of candidate sounds and based on identified sounds from the audio input data.

[0199] In an eleventh aspect, alone or in combination with one or more of the above aspects, to generate the generative audio information the one or more processors are configured to:

[0200] identify a preset sound of a sound database based on a particular detected object of detected object information of the visual perception information; and modify sound information associated with the identified preset sound based on object information of the visual perception information and for the particular detected object, wherein the generative audio information includes the modified sound information.

[0201] In a twelfth aspect, alone or in combination with one or more of the above aspects, to modify the sound information the one or more processors are configured to: adjust a location, a frequency, an intensity, or a combination thereof, of the sound informationassociated with the identified preset sound based on the object information of the visual perception information to generate the modified sound information.

[0202] In a thirteenth aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to: receive a user input; remove one or more generative sounds of the generative audio information; generate one or more second generative sounds for the generative audio information; and output the image data with second generative audio information, wherein the second generative audio information includes the one or more second generative sounds and does not include the one or more removed generative sounds of the generative audio information.

[0203] In a fourteenth aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to: receive a user preference information, wherein the generative audio information is further generated based on the user preference information.

[0204] In a fifteenth aspect, alone or in combination with one or more of the above aspects, to generate the generative audio information the one or more processors are configured to: generate candidate generative sound information for one or more candidate generative sounds based on the visual perception information; select or modify the candidate generative sound information for at least one candidate generative sound of the one or more candidate generative sounds based on the user preference information, wherein the generative audio information includes the selected candidate generative sound information, the modified candidate generative sound information, or both.

[0205] In a sixteenth aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to: receive a user input with sound modification information for one or more generative sounds of the generative audio information; and adjust a frequency or intensity of the one or more generative sounds of the generative audio information based on the sound modification information.

[0206] In a seventeen aspect, alone or in combination with one or more of the above aspects, the one or more processors are further configured to: output a list of candidate generative sounds; and receive a user input with sound selection information indicating a selection of one or more of the candidate generative sounds of the list, wherein the generative audio information is further generated based on the sound selection information, and wherein the generative audio information includes one or more selected candidate generative sounds and does not include one or more unselected candidate generative sounds.

[0207] In an eighteenth aspect, alone or in combination with one or more of the above aspects, the generative audio information includes spatial audio information, and wherein one or more generative sounds of the generative audio information change locations in the frame over multiple image frames of the plurality of image frames.

[0208] In a nineteenth aspect, alone or in combination with one or more of the above aspects, the visual perception information includes depth information and object motion information, and wherein the one or more processors are further configured to: generate spatial audio information for one or more generative sounds of the generative audio information based on the depth information and the object motion information.

[0209] In a twentieth aspect, alone or in combination with one or more of the above aspects, the device further includes: a camera configured to generate the image data; a microphone array configured to generate audio input data associated with the image data; a display configured to output the image data; and a speaker configured to output the generative audio information with the image data.

[0210] Those of skill in the art would further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure. Skilled artisans will also readily recognize that the order or combination of components, methods, or interactions that are described herein are merely examples and that the components, methods, or interactions of the various aspects of the present disclosure may be combined or performed in ways other than those illustrated and described herein.

[0211] The various illustrative logics, logical blocks, modules, circuits and algorithm processes described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. The interchangeability of hardware and software has been described generally, in terms of functionality, and illustrated in the various illustrative components, blocks, modules,circuits, and processes described above. Whether such functionality is implemented in hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0212] In one or more aspects, the operations described may be implemented in hardware, digital electronic circuitry, computer software, firmware, including the structures disclosed in this specification and their structural equivalents thereof, or in any combination thereof. Implementations of the subject matter described in this specification also may be implemented as one or more computer programs, which is one or more modules of computer program instructions, encoded on a computer storage media for execution by, or to control the operation of, data processing apparatus.

[0213] The operations of a method or algorithm disclosed herein may be implemented in a processor-executable software module which may reside on a computer-readable medium and commercially made available as a computer program product as software. Computer- readable media includes both computer storage media and communication media including any medium that may be enabled to transfer a computer program from one place to another. A storage media may be any available media that may be accessed by a computer. By way of example, and not limitation, such computer-readable media may include random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer. Also, any connection may be properly termed a computer-readable medium. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc wherein disks usually reproduce data magnetically and discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0214] Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to some other implementations without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the implementations shown herein but are to be accorded the widest scope consistent with this disclosure, the principles and the novel features disclosed herein.

[0215] Additionally, a person having ordinary skill in the art will readily appreciate, opposing terms such as “upper” and “lower,” or “front” and back,” or “top” and “bottom,” or “forward” and “backward,” or “left” and “right” are sometimes used for ease of describing the figures, and indicate relative positions corresponding to the orientation of the figure on a properly oriented page, and may not reflect the proper orientation of any device as implemented.

[0216] Certain features that are described in this specification in the context of separate implementations also may be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation also may be implemented in multiple implementations separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0217] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed to achieve desirable results. Further, the drawings may schematically depict one or more example processes in the form of a flow diagram. However, other operations that are not depicted may be incorporated in the example processes that are schematically illustrated. For example, one or more additional operations may be performed before, after, simultaneously, or between any of the illustrated operations. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results.

[0218] As used herein, including in the claims, the term “or,” when used in a list of two or more items, means that any one of the listed items may be employed by itself, or any combination of two or more of the listed items may be employed. For example, if a composition is described as containing components A, B, or C, the composition maycontain A alone; B alone; C alone; A and B in combination; A and C in combination; B and C in combination; or A, B, and C in combination. Also, as used herein, including in the claims, “or” as used in a list of items prefaced by “at least one of’ indicates a disjunctive list such that, for example, a list of “at least one of A, B, or C” means A or B or C or AB or AC or BC or ABC (that is A and B and C) or any of these in any combination thereof.

[0219] The term “substantially” is defined as largely, but not necessarily wholly, what is specified (and includes what is specified; for example, substantially 90 degrees includes 90 degrees and substantially parallel includes parallel), as understood by a person of ordinary skill in the art. In any disclosed implementations, the term “substantially” may be substituted with “within [a percentage] of’ what is specified, where the percentage includes .1, 1, 5, or 10 percent.

[0220] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A device comprising: a memory storing processor-readable code; and one or more processors coupled to the memory, the one or more processors configured to execute the processor-readable code to cause the one or more processors to: obtain image data for a plurality of image frames associated with a live image; perform computer vision for the plurality of image frames to generate visual perception information based on the image data; generate generative audio information based on the visual perception information; and output the image data with the generative audio information.

2. The device of claim 1, wherein the generative audio information corresponds to sound information associated with at least one object identified in one or more image frames of the plurality of image frames or with at least one scene identified in the one or more image frames of the plurality of image frames.

3. The device of claim 1, wherein the generative audio information corresponds to spatial sound information associated with at least one object identified in one or more image frames of the plurality of image frames and representing motion of the at least one object.

4. The device of claim 1, wherein the visual perception information includes object identification information, scene identification information, or a combination thereof.

5. The device of claim 1, wherein the visual perception information includes one or more of frame information, object information, event detection information, motion estimation information, three-dimensional pose information, object identification information, or object tracking information.

6. The device of claim 1, wherein the one or more processors are further configured to cause the device to: obtain audio input data for the plurality of image frames, wherein the visual perception information, the generative audio information, or both are based on the audio input data, and wherein the one or more processors are configured to cause the device to output the image data with the generative audio information includes to: output the image data with the generative audio information and with at least a portion of the audio input data.

7. The device of claim 1, wherein the one or more processors are further configured to cause the device to: obtain audio input data for the plurality of image frames, wherein the visual perception information, the generative audio information, or both are based on the audio input data; and refrain from outputting the audio input data with the image data.

8. The device of claim 7, wherein to perform computer vision for the plurality of image frames to generate visual perception information based on the image data, the one or more processors are further configured to: extract scenes from multiple image frames of the plurality of image frames and generate scene information based on the extracted scenes; perform object detection on the extracted scenes; perform object tracking for the detected objects over the multiple image frames; and perform photo segmentation and context detection on the detected and tracked objects to identify candidate sounds for the plurality of image frames, wherein the candidate sounds are indicated by the visual perception information.

9. The device of claim 7, wherein to generate the generative audio information the one or more processors are configured to: obtain audio input data for the plurality of image frames; andidentify sounds in the audio input data; wherein to perform the computer vision to generate the visual perception information the one or more processors are configured to: generate the visual perception information based on the image data and based on identified sounds from the audio input data.

10. The device of claim 7, wherein the one or more processors are further configured to: obtain audio input data for the plurality of image frames; and determine a plurality of candidate sounds for generative audio processing based on object information, scene information, or both of the visual perception information, wherein to generate the generative audio information the one or more processors are configured to: generate the generative audio information based on the plurality of candidate sounds and based on identified sounds from the audio input data.

11. The device of claim 1, wherein to generate the generative audio information the one or more processors are configured to: identify a preset sound of a sound database based on a particular detected object of detected object information of the visual perception information; and modify sound information associated with the identified preset sound based on object information of the visual perception information and for the particular detected object, wherein the generative audio information includes the modified sound information.

12. The device of claim 11, wherein to modify the sound information the one or more processors are configured to: adjust a location, a frequency, an intensity, or a combination thereof, of the sound information associated with the identified preset sound based on the object information of the visual perception information to generate the modified sound information.

13. The device of claim 10, wherein the one or more processors are further configured to:receive a user input; remove one or more generative sounds of the generative audio information; generate one or more second generative sounds for the generative audio information; and output the image data with second generative audio information, wherein the second generative audio information includes the one or more second generative sounds and does not include the one or more removed generative sounds of the generative audio information.

14. The device of claim 10, wherein the one or more processors are further configured to: receive a user preference information, wherein the generative audio information is further generated based on the user preference information.

15. The device of claim 14, wherein to generate the generative audio information the one or more processors are configured to: generate candidate generative sound information for one or more candidate generative sounds based on the visual perception information; and select or modify the candidate generative sound information for at least one candidate generative sound of the one or more candidate generative sounds based on the user preference information, wherein the generative audio information includes the selected candidate generative sound information or the modified candidate generative sound information.

16. The device of claim 1, wherein the one or more processors are further configured to: receive a user input with sound modification information for one or more generative sounds of the generative audio information; and adjust a frequency or intensity of the one or more generative sounds of the generative audio information based on the sound modification information.

17. The device of claim 1, wherein the one or more processors are further configured to: output a list of candidate generative sounds; andreceive a user input with sound selection information indicating a selection of one or more of the candidate generative sounds of the list, wherein the generative audio information is further generated based on the sound selection information, and wherein the generative audio information includes one or more selected candidate generative sounds and does not include one or more unselected candidate generative sounds.

18. The device of claim 1, wherein the generative audio information includes spatial audio information, and wherein a perceived location of one or more generative sounds of the generative audio information changes over multiple image frames of the plurality of image frames.

19. The device of claim 18, wherein the visual perception information includes depth information and object motion information, and wherein the one or more processors are further configured to: generate spatial audio information for one or more generative sounds of the generative audio information based on the depth information and the object motion information.

20. The device of claim 1, further comprising: a camera configured to generate the image data; a microphone array configured to generate audio input data associated with the image data; a display configured to output the image data; and a speaker configured to output the generative audio information with the image data.

Citation Information

Patent Citations

  • Sound and video object tracking

    US20170364752A1