Systems and methods for performing acoustic zoom

By using the beamformer and the target enhancer to combine the signal during video playback, the problem of audio in the area of interest is submerged in noisy environments, and the enhancement of audio in the area of interest and the suppression of ambient noise is achieved, providing a clearer audio experience.

CN114727193BActive Publication Date: 2025-08-05SNAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210491087.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-03
Filing Date
2019-08-30
Publication Date
2025-08-05
Estimated Expiration
2039-08-30

AI Technical Summary

Technical Problem

During video playback, audio from areas of interest in noisy environments may be flooded, and prior art is difficult to effectively enhance audio related to areas of interest in video.

Method used

By using multiple beamformers to generate beamformer signals corresponding to multiple tiles of video content, and using the target enhancer to identify and combine the corresponding beamformer signals, a target enhancement signal associated with the zoom region is generated to enhance the audio of the region of interest.

Benefits of technology

During video playback, audio in the area of interest is enhanced, ambient noise is suppressed, and a clearer audio experience is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114727193B_ABST
    Figure CN114727193B_ABST
Patent Text Reader

Abstract

The method for performing acoustic zoom begins with a microphone capturing an acoustic signal associated with video content. A beamformer generates a beamformer signal using the acoustic signal. The beamformer signals correspond to tiles of the video content, respectively. Each beamformer is directed toward the center of each tile, respectively. A target enhancement signal is generated using the beamformer signal. The target enhancement signal is associated with a zoom area of the video content. The target enhancement signal is generated by: identifying tiles that each have at least a portion included in the zoom area, selecting beamformer signals corresponding to the identified tiles, and combining the selected beamformer signals to generate the target enhancement signal. Combining the selected beamformer signals may include: determining a proportion of each identified tile relative to the zoom area; and combining the selected beamformer signals based on the proportion to generate the target enhancement signal. Other embodiments are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application date of August 30, 2019, application number 201980056985.0, and invention name “Acoustic Zoom”.

[0002] priority

[0003] This application claims the benefit of priority to Indian patent application serial number 201811032980 filed on September 3, 2018, each of which is hereby claimed the benefit of priority and each of which is incorporated herein by reference in its entirety. Technical Field

[0004] Embodiments of the present application relate to acoustic zoom during video playback. Background Art

[0005] Currently, many consumer electronic devices are suitable for capturing audio and / or video content. For example, a user can quickly capture a video using his mobile device in a public place.

[0006] During playback of the video, the viewer can zoom in on the area of interest to see the selected area of interest in a larger format. However, if the environment in which the video is captured is noisy, the audio associated with the area of interest in the video may be drowned out. Summary of the Invention

[0007] According to one aspect of the present application, a system for performing acoustic zoom is provided, comprising: a plurality of beamformers that generate a plurality of beamformer signals corresponding to a plurality of tiles of video content associated with a plurality of acoustic signals, wherein each beamformer is directed to the center of each tile; and a target enhancer that identifies tiles having at least a portion included in a zoom region of the video content, selects beamformer signals corresponding to the identified tiles, and combines the selected beamformer signals to generate a target enhancement signal associated with the zoom region.

[0008] According to another aspect of the present application, a method for performing acoustic zoom is provided, comprising: causing a processor to cause a plurality of beamformers to generate a plurality of beamformer signals using a plurality of acoustic signals associated with video content, wherein the beamformer signals correspond to a plurality of tiles of the video content, wherein each beamformer is directed to the center of each tile; identifying a tile having at least a portion included in a zoom area of the video content, selecting the beamformer signals corresponding to the identified tiles, and combining the selected beamformer signals to generate a target enhancement signal associated with the zoom area.

[0009] According to another aspect of the present application, a computer-readable storage medium is provided, having instructions stored thereon, which, when executed by a processor, cause the processor to perform operations, the operations including: causing a plurality of beamformers to generate a plurality of beamformer signals using a plurality of acoustic signals associated with video content, wherein the beamformer signals correspond to a plurality of tiles of the video content, wherein each beamformer is directed to the center of each tile; identifying a tile having at least a portion included in a zoom region of the video content, selecting a beamformer signal corresponding to the identified tile, and combining the selected beamformer signals to generate a target enhancement signal associated with the zoom region.

[0010] According to another aspect of the present application, a system for performing acoustic zoom is provided, comprising: a plurality of beamformers for receiving a plurality of acoustic signals, the plurality of beamformers including a target beamformer and a noise beamformer, wherein the target beamformer points to a center of a field of view corresponding to a zoom area of video content and generates a target beamformer signal, and the noise beamformer has a zero value pointing to the center of the field of view and generates a noise beamformer signal; and a target enhancer for determining a field of view corresponding to the zoom area of the video content, and using the target beamformer signal and the noise beamformer signal to generate a target enhancement signal associated with the zoom area of the video content. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views. Like numerals with different letter suffixes may represent different instances of similar components. In the figures of the accompanying drawings, some embodiments are shown by way of example and not limitation, in which:

[0012] Figure 1 is an example of a system for performing acoustic zoom in use according to an example embodiment.

[0013] Figure 2 is a diagram showing a method according to an example embodiment Figure 1 A block diagram of the system with more details.

[0014] Figure 3 A system according to an example embodiment Figure 2 A block diagram of the details of the acoustic zoom controller 111 is shown in FIG.

[0015] Figure 4A -D shows the arrangement of tiles on video content according to an embodiment of the present invention ( Figure 4A ), the zoom area on the tile layout ( Figure 4B) and combining beamformer signals based on the tiles included in the zoom region ( Figures 4C-4D ).

[0016] Figure 5 A system according to an example embodiment Figure 2 A block diagram of the details of the acoustic zoom controller 111 is shown in FIG.

[0017] Figure 6 An example of a zoom area on video content and a field of view cone centered on the zoom area according to an embodiment of the present invention is shown.

[0018] Figure 7 is a flow chart of an example method for performing acoustic zoom according to one embodiment of the present invention.

[0019] Figure 8 is a flow chart of an example method for performing acoustic zoom according to one embodiment of the present invention.

[0020] Figure 9 is a block diagram illustrating a representative software architecture that can be used in conjunction with the various hardware architectures described herein.

[0021] Figure 10 is a block diagram illustrating components of a machine capable of reading instructions from a machine-readable medium (eg, a machine-readable storage medium) and performing any one or more of the methodologies discussed herein, according to some example embodiments. DETAILED DESCRIPTION

[0022] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody the illustrative embodiments of the present disclosure. In the following description, for the purpose of explanation, many specific details are set forth in order to provide an understanding of the various embodiments of the subject matter of the present invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the present invention may be practiced without these specific details. Typically, well-known instruction examples, protocols, structures, and techniques need not be shown in detail.

[0023] The embodiments described herein improve upon current systems by enabling acoustic zooming during video playback. Specifically, acoustic zooming refers to enhancing the audio associated with a region of interest in a video. For example, when a user visually zooms in on a region of interest in a video during playback, the region of interest can be visually enhanced (e.g., in a larger format) and the audio corresponding to the region of interest can be enhanced by increasing the volume of sounds originating from the region of interest, suppressing sounds originating from outside the region of interest (e.g., ambient noise, other speakers, etc.), or any combination thereof.

[0024] Figure 1is an example of a system for performing acoustic zoom in use according to an example embodiment. Figure 1 As shown, the system 100 may be an apparatus such as a client device (e.g., Figure 10 A machine 1000 in captures a video including multiple objects and an acoustic signal corresponding to the video.

[0025] As used herein, the term "client device" may refer to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smartphone, a tablet computer, an ultrabook, a netbook, a portable computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communication device that a user can use to access a network.

[0026] Some embodiments may include one or more wearable devices, such as a pendant with an integrated camera that is integrated with, in communication with, or coupled to a client device. Any desired wearable device may be used in conjunction with embodiments of the present disclosure, such as a watch, glasses, goggles, headphones, a wristband, earbuds, clothing (e.g., a hat or jacket with an integrated electronic device), a clip-on electronic device, or any other wearable device.

[0027] Figure 2 is a block diagram showing more details of system 100 according to an example embodiment. Figure 2 As shown, the system 100 includes microphones 113_1 to 113_N (N>1), a camera module 112, and an acoustic zoom controller 111. The microphones 113_1 to 113_N may be air interface pickup devices that convert sound into electronic signals. Figure 1 In the embodiment, the system 100 includes six microphones 113_1 to 113_6, but the number of microphones may vary. In one embodiment, the system 100 may include at least two microphones, and may form a microphone array.

[0028] Microphones 113_1 to 113_N can be used to create a microphone array beam (i.e., a beamformer), which can be steered toward a given direction by emphasizing or deemphasizing selected microphones 113_1 to 113_N. Similarly, the microphone array can also present or provide nulls in other given directions. Therefore, beamforming (also known as spatial filtering) can be a signal processing technique for directional sound reception using a microphone array.

[0029] The camera module 112 includes a camera lens and an image sensor. The camera lens can be a perspective camera lens or a non-perspective camera lens. The non-perspective camera lens can be, for example, a fisheye lens, a wide-angle lens, an omnidirectional lens, etc. The image sensor captures digital video through the camera lens. The image can also be a still image frame or a video including multiple still image frames. In one embodiment, the system 100 can be separate from the camera module 112 but coupled to a client device including the camera module 112. In this embodiment, the system 100 can be a housing or shell including microphones 113_1 to 113_N and a window that allows the camera lens to capture image or video content.

[0030] exist Figure 1 In an embodiment of the present invention, the system 100 uses the camera module 112 to capture a video including multiple objects and uses the microphones 113_1 to 113_N to capture an acoustic signal corresponding to the video. During playback, the acoustic signal is synchronized with the video in time. The acoustic signal may include a desired (or target) audio signal and peripheral or environmental noise. For example, in Figure 1 In the embodiment of the present invention, if the user of the system 100 intends to capture the audio signal from the source located in the center, the audio signals from the remaining sources (eg, the top and bottom sources) will also be captured as the ambient noise acoustic signal.

[0031] In one embodiment, when playing back the captured video and corresponding audio signal, the acoustic zoom controller 111 in the system 100 determines the field of view (or zoom area) of the video content and enhances the audio signal corresponding to the field of view. In another embodiment, the acoustic zoom controller 111 determines the field of view (or zoom area) of the video content in real time and enhances the audio signal corresponding to the field of view in real time.

[0032] Figure 3 A system according to an example embodiment Figure 2 A block diagram of the details of the acoustic zoom controller 111 is shown in FIG. Figure 3 , the acoustic zoom controller 111 includes a time-frequency converter 310 , a neural network 320 , a beamformer unit 330 including a plurality of beamformers, a target enhancer 340 and a frequency-time converter 350 .

[0033] The time-frequency converter 310 receives acoustic signals from the microphones 113_1 to 113_N and converts the acoustic signals from the time domain to the frequency domain. In one embodiment, the time-frequency converter 310 performs a short-time Fourier transform (STFT) on the acoustic signals in the time domain to obtain acoustic signals in the frequency domain.

[0034] Neural network 320 receives an acoustic signal in the frequency domain and generates a noise reference signal. Neural network 320 may be a deep neural network for generating a noise reference signal that estimates a noise covariance matrix that encodes the energy distribution of the noise in space. Neural network 320 may be trained offline to identify and encode the noise distribution in space.

[0035] In one embodiment, the neural network 320 is further used to mask noise in the acoustic signal in the frequency domain to generate an acoustic signal suppressed by noise in the frequency domain. The neural network 320 may also provide the acoustic signal suppressed by noise in the frequency domain to the beamformer unit 330 for further processing.

[0036] Figure 4A FIG. 4 shows an example of arranging tiles on video content according to one embodiment. The captured video content may be divided into a plurality of tiles 410_1 to 410_M (M>1). Figure 4A In an embodiment, the tiles of the video content are tiles of the same shape having an angular width of at least 10 degrees. For each tile 410_j (M≥j≥1), the beamformer unit 330 includes a beamformer directed to the center of the tile 410_j. Figure 4A In an embodiment of the present invention, the beamformer unit 330 includes nine (9) beamformers that are directed or steered toward the nine (9) centers of the nine (9) tiles, respectively. Thus, the beamformers each generate a beamformer signal that includes audio corresponding to a portion of the video content in each tile. The beamformers in the beamformer unit 330 can include a fixed beamformer directed toward the center of the tile 410_j, an adaptive beamformer such as a minimum variance distortionless response (MVDR) beamformer, or any combination thereof.

[0037] although Figure 4A The embodiment in FIG includes blocks 410_1 to 410_M of the same shape, but it should be understood that blocks 410_1 to 410_M can have any different shapes. Figure 4A The embodiment in includes tiles 410_1 to 410_M having an angular width of at least 10 degrees, but it is understood that the tiles 410_1 to 410_M may have different angular widths.

[0038] Figure 4B According to one embodiment, Figure 4A When a user selects an area of the video content to be displayed in a larger (zoomed) format, the user's field of view changes from including Figure 4A The first field of view of all tiles in becomes the corresponding portion including different tiles. Figure 4BThe second field of view of the zoom area 420 in FIG.

[0039] Figure 3 The target enhancer 340 in receives the plurality of beamformer signals from the beamformer unit 330 and generates a target enhancement signal associated with the zoom region 420 of the video content. In one embodiment, the target enhancer 340 generates the target enhancement signal by identifying tiles each having at least a portion included in the zoom region 420. Figure 4C , portions of four tiles 410_1 to 410_4 are identified as having at least a portion included in the zoom region 420. In this example, the entire tile 410_1 is included in the zoom region 420, and smaller portions of the tiles 410_2 to 410_4 are included in the zoom region 420. The target enhancer 340 selects beamformer signals corresponding to the identified tiles 410_2 to 410_4 and combines the selected beamformer signals to generate a target enhanced signal.

[0040] In one embodiment, the target enhancer 340 combines the selected beamformer signals in the same proportion as each identified tile contributes to the zoom area. Figure 4D 4 shows the combining performed by the target enhancer 340 according to one embodiment. In this embodiment, the target enhancer 340 determines the proportion of each identified tile relative to the zoom region 420 and combines the selected beamformer signals based on the proportion to generate a target enhanced signal. The target enhancer 340 can combine the selected beamformer signals by spectrally summing the selected beamformer signals based on the proportion.

[0041] The frequency-time converter 350 receives the target enhancement signal from the target enhancer 340 and converts the target enhancement signal from the frequency domain to the time domain. In one embodiment, the frequency-time converter 350 performs an inverse short-time Fourier transform (STFT) on the target enhancement signal in the frequency domain to obtain the target enhancement signal in the time domain.

[0042] Figure 5 A system according to an example embodiment Figure 2 A block diagram of the details of the acoustic zoom controller 111 in FIG. Figure 3 Details of the acoustic zoom controller 111, Figure 5The acoustic zoom controller 111 in the embodiment further includes a time-frequency converter 310, a neural network 320, and a frequency-time converter 350. However, in this embodiment, the acoustic zoom controller 111 includes a beamformer unit 530, which includes a target beamformer and a noise beamformer, and a target enhancer 540, which includes a feedback signal to the beamformer unit 530. The beamformer unit 530 receives the acoustic signal in the frequency domain from the time-frequency converter 310 and the noise reference signal from the neural network 320.

[0043] Figure 6 420 and a field of view circle 620 centered on the zoom area 420 in accordance with an embodiment of the present invention. Figure 6 The first field of view of the entire area 610 of the video content changes to Figure 6 The second field of view corresponding to the zoom region 420 in FIG. Figure 6 The second field of view is included as a circle 620, but the second field of view may be any shape.

[0044] In one embodiment, the beamformer unit 530 includes a target beamformer and a noise beamformer. The target beamformer is pointed at the center of the second field of view circle 620 corresponding to the zoom area 420 of the video content. In one embodiment, the second field of view circle 620 is an attempt to cover as much of the zoom area 420 as possible. In one embodiment, the target beamformer implements a steering vector that encodes the direction of the sound to be enhanced (e.g., the center of the second field of view circle 620). The noise beamformer is pointed at the first field of view 610 and has a zero value pointing at the center of the second field of view circle 620. The noise beamformer can be a cardioid or other beamforming pattern pointed away from the center of the second field of view circle 620 to capture ambient noise while contaminating the audio of interest (e.g., from the center of the second field of view circle 620) as little as possible. The noise beamformer generates a noise beamformer signal that captures acoustic signals that are not in the direction of the sound to be enhanced.

[0045] In one embodiment, the neural network 320 receives a plurality of acoustic signals to generate a noise reference signal. In this embodiment, the beamformer unit 530 receives the noise reference signal and uses the plurality of acoustic signals and the noise reference signal to generate a target beamformer signal and a noise beamformer signal.

[0046] The target enhancer 540 determines a second field of view circle 620 corresponding to the zoom region 420 of the video content. In one embodiment, the target enhancer 540 determines the position and orientation of the zoom region 420 relative to the first field of view 610. The target enhancer 540 may sequentially send data including the second field of view circle 620 to the beamformer unit 530, causing the beamformer unit 530 to steer the target beamformer and the noise beamformer accordingly. The target enhancer receives the target beamformer signal and the noise beamformer signal and uses the target beamformer signal and the noise beamformer signal to generate a target enhancement signal associated with the zoom region 420 of the video content. In one embodiment, the target enhancer 540 generates the target enhancement signal by spectrally subtracting the noise beamformer signal from the target enhancement signal.

[0047] The following embodiments of the present invention may be described as a process, which is often depicted as a flow chart, flowchart, structure diagram, or block diagram. Although a flow chart may depict operations as a sequential process, many operations can be performed in parallel or simultaneously. In addition, the order of operations can be rearranged. When the operations of a process are completed, the process terminates. A process may correspond to a method, procedure, etc.

[0048] Figure 7 7 is a flow chart of an example method for performing acoustic zoom according to one embodiment of the present invention. The method begins in block 701, with multiple microphones capturing multiple acoustic signals associated with video content. In block 702, multiple beamformers use the multiple acoustic signals to generate multiple beamformer signals. The beamformer signals may correspond to multiple tiles of the video content, respectively. Each beamformer may be directed toward the center of each tile. In block 703, a target enhancer uses the beamformer signals to generate a target enhancement signal. The target enhancement signal may be associated with a zoom region of the video content. In one embodiment, in block 703, the target enhancer generates the target enhancement signal by identifying tiles each having at least a portion included in the zoom region, selecting beamformer signals corresponding to the identified tiles, and combining the selected beamformer signals to generate the target enhancement signal. In one embodiment, combining the selected beamformer signals includes determining a ratio of each identified tile relative to the zoom region and combining the selected beamformer signals based on the ratio to generate the target enhancement signal.

[0049] Figure 8801 is a flow chart of an example method for performing acoustic zoom according to one embodiment of the present invention. The method begins, at block 802, with multiple microphones capturing multiple acoustic signals. A first field of view of video content may be associated with the multiple acoustic signals. At block 802, a target beamformer generates a target beamformer signal using the multiple acoustic signals. The target beamformer is directed toward the center of a second field of view corresponding to a zoomed area of the video content. At block 803, a noise beamformer generates a noise beamformer signal using the multiple acoustic signals. The noise beamformer is directed toward the first field of view and has a central null point directed toward the second field of view. At block 804, a target enhancer determines a second field of view corresponding to the zoomed area of the video content and, at block 805, generates a target enhancement signal associated with the zoomed area of the video content using the target beamformer signal and the noise beamformer. In one embodiment, generating the target enhancement signal by the target enhancer includes spectrally subtracting the noise beamformer signal from the target enhancement signal.

[0050] Software Architecture

[0051] Figure 9 is a block diagram illustrating an example software architecture 906 that can be used in conjunction with the various hardware architectures described herein. Figure 9 is only a non-limiting example of a software architecture, and it will be understood that many other architectures can be implemented to facilitate the functionality described herein. The software architecture 906 can be implemented in a system such as Figure 10 1000, including, among other things, a processor 1004, a memory 1014, and I / O components 1018. A representative hardware layer 952 is shown and may represent, for example, Figure 10 The machine 1000 of FIG. 1000 includes a representative hardware layer 952 including one or more processing units 954 having associated executable instructions 904. Executable instructions 904 represent executable instructions of a software architecture 906, including implementations of the methods, components, etc. described herein. The hardware layer 952 also includes a memory and / or storage module 956 also having executable instructions 904. The hardware layer 952 may also include other hardware 958.

[0052] As used herein, the term "component" may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, application program interfaces (APIs), or other techniques that provide partitioning or modularization for specific processing or control functions. Components can be combined through their interfaces with other components to perform machine processes. A component can be a packaged functional hardware unit designed to be used with other components and as part of a program that typically performs specific functions of related functions.

[0053] A component may constitute a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that is capable of performing certain operations and may be configured or arranged in some physical manner. In various example embodiments, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations described herein. A hardware component may also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations.

[0054] A hardware component may be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. After being configured by such software, the hardware component becomes a specific machine (or specific component of a machine) that is specifically customized to perform the configured function and is no longer a general-purpose processor. It will be understood that the decision to implement a hardware component mechanically in a dedicated and permanently configured circuit or in a temporarily configured circuit (e.g., configured by software) may be driven by cost and time considerations.

[0055] A processor may be or may include any circuit or virtual circuit (a physical circuit emulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "commands," "opcodes," "machine codes," etc.) and generates corresponding output signals suitable for operating a machine. A processor may be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may further be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.

[0056] Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, which is a physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) entity that operates in some manner or performs certain operations described herein. Considering embodiments in which hardware components are temporarily configured (e.g., programmed), each hardware component does not need to be configured or instantiated at any time. For example, in the case where the hardware component includes a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured as different special-purpose processors (e.g., including different hardware components) at different times. Therefore, the software configures a specific processor or processors accordingly, for example, to constitute a specific hardware component at one moment and another different hardware component at another different moment. The hardware components can provide information to other hardware components and receive information from other hardware components. Therefore, the described hardware components can be considered to be communicatively coupled. In the case where multiple hardware components are present at the same time, communication can be achieved by signal transmission between two or more hardware components (e.g., through appropriate circuits and buses). In embodiments where multiple hardware components are configured or instantiated at different times, communication between such hardware components may be achieved, for example, by storing and retrieving information in memory structures accessible to the multiple hardware components.

[0057] For example, a hardware component may perform an operation and store the output of the operation in a memory device to which it is communicatively coupled. Then, another hardware component may access the memory device at a later time to obtain and process the stored output. The hardware component may also initiate communication with an input or output device and may operate on a resource (e.g., a collection of information). The various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein may be implemented at least in part by a processor, wherein a specific one or more or processors are examples of hardware. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented components.

[0058] In addition, one or more processors may also be operable to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as an example of a machine including a processor), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of certain operations may be distributed among the processors, not only residing within a single machine, but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented component may be located in a single geographic location (e.g., in a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor-implemented component may be distributed across multiple geographic locations.

[0059] exist Figure 9 In the example architecture of , software architecture 906 can be conceptualized as a stack of layers, where each layer provides specific functionality. For example, software architecture 906 may include layers such as operating system 902, library 920, application 916, and presentation layer 914. In operation, applications 916 or other components within these layers can call application program interfaces (APIs) API calls 908 through the software stack and receive messages 912 in response to API calls 908. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile or dedicated operating systems may not provide framework / middleware 918, while other operating systems may provide such layers. Other software architectures may include additional or different layers.

[0060] The operating system 902 can manage hardware resources and provide public services. The operating system 902 may include, for example, a kernel 922, services 924, and drivers 926. The kernel 922 may act as an abstraction layer between the hardware and other software layers. For example, the kernel 922 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Services 924 may provide other public services to other software layers. Drivers 926 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 926 include display drivers, camera drivers, drives, flash drives, serial communication drivers (such as Universal Serial Bus (USB) drivers), drivers, audio drivers, power management drivers, etc., depending on the hardware configuration.

[0061] The libraries 920 may provide a common infrastructure that can be used by applications 916 or other components or layers. The libraries 920 generally provide functionality that allows other software components to perform tasks more easily than by directly interfacing with underlying operating system 902 functionality (e.g., kernel 922, services 924, and / or drivers 926). The libraries 920 may include system libraries 944 (e.g., the C standard library), which may provide functionality such as memory allocation, string manipulation, mathematical functions, and the like. Furthermore, the libraries 920 may include API libraries 946 such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats (e.g., MPEG4, H.264, MP3, AAC, AMR, JPG, PNG)), graphics libraries (e.g., the OpenGL framework for rendering 2D and 3D graphics content on a display), database libraries (e.g., SQLite, which may provide various relational database functions), networking libraries (e.g., WebKit, which may provide web browsing functionality), and the like. The libraries 920 may also include a variety of other libraries 948 to provide a number of other APIs to the applications 916 and other software components / modules.

[0062] The framework / middleware 918 (sometimes also referred to as middleware) provides a high-level, general-purpose infrastructure that can be used by applications 916 and / or other software components / modules. For example, the framework / middleware 918 can provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. The framework / middleware 918 can provide a wide range of other APIs that can be used by applications 916 or other software components / modules, some of which may be specific to a particular operating system 902 or platform.

[0063] Applications 916 include built-in applications 938 or third-party applications 940. Examples of representative built-in applications 938 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, or game applications. Third-party applications 940 may include applications that are created by entities other than the vendor of a particular platform using Android TM or iOS TM Applications developed with a software development kit (SDK) can be developed on mobile operating systems (such as iOS TM 、Android TM 、 Third-party applications 940 may invoke API calls 908 provided by a mobile operating system (such as operating system 902) to facilitate the functionality described herein.

[0064] Applications 916 can utilize built-in operating system functionality (e.g., kernel 922, services 924, and / or drivers 926), libraries 920, and framework / middleware 918 to create a user interface to interact with a user of the system. Alternatively or additionally, in some systems, interaction with the user can occur through a presentation layer, such as presentation layer 914. In these systems, the application / component "logic" can be separated from the aspects of the application / component that interact with the user.

[0065] Figure 10 A block diagram is shown of components (also referred to herein as “modules”) of a machine 1000 according to some example embodiments, which components are capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methodologies discussed herein. Specifically, Figure 10 A diagrammatic representation of a machine 1000 in the example form of a computer system is shown within which instructions 1010 (e.g., software, programs, applications, applet, application program, or other executable code) may be executed to cause the machine 1000 to perform any one or more of the methodologies discussed herein. In this manner, the instructions 1010 may be used to implement the modules or components described herein. The instructions 1010 transform a general-purpose, unprogrammed machine 1000 into a specific machine 1000 that is programmed to perform the functions described and illustrated in the manner described. In alternative embodiments, the machine 1000 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a network deployment, the machine 1000 may operate as a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a network appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 1010 that specify actions to be taken by the machine 1000. Furthermore, while only a single machine 1000 is shown, the term "machine" should also be construed to include a collection of machines that individually or collectively execute instructions 1010 to perform any one or more of the methodologies discussed herein.

[0066] The machine 1000 may include a processor 1004, a memory / storage 1006, and an I / O component 1018, which may be configured to communicate with each other, for example, via a bus 1002. The memory / storage 1006 may include a memory 1014 (such as a main memory or other memory storage device) and a storage unit 1016, both of which may be accessed by the processor 1004, such as via the bus 1002. The storage unit 1016 and the memory 1014 store instructions 1010 that embody any one or more of the methodologies or functions described herein. During execution by the machine 1000, the instructions 1010 may also reside, in whole or in part, within the memory 1014, within the storage unit 1016, within at least one of the processors 1004 (e.g., within a cache memory of the processor), or any combination thereof. Thus, the memory 1014, the storage unit 1016, and the memory of the processor 1004 are examples of machine-readable media.

[0067] As used herein, the terms "machine-readable medium," "computer-readable medium," and the like may refer to a component, device, or other tangible medium capable of temporarily or permanently storing instructions and data. Examples of such media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache, other types of storage devices (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) capable of storing instructions. The term "machine-readable medium" should also be understood to include any medium or combination of multiple media capable of storing instructions (e.g., code) executed by a machine such that the instructions, when executed by one or more processors of the machine, cause the machine to perform any one or more of the methods described herein. Thus, "machine-readable medium" refers to a single storage device or device, as well as a "cloud-based" storage system or storage network comprising multiple storage devices or devices. The term "machine-readable medium" itself does not include signals.

[0068] The I / O components 1018 may include a variety of components to provide a user interface for receiving input, providing output, generating output, sending information, exchanging information, collecting measurements, etc. The specific I / O components 1018 of the user interface included in a particular machine 1000 will depend on the type of machine. For example, a portable machine such as a mobile phone will likely include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It should be understood that the I / O components 1018 may be included in Figure 10The I / O components 1018 are grouped according to their functionality for the purpose of simplifying the following discussion only and are by no means limiting. In various exemplary embodiments, the I / O components 1018 may include an output component 1026 and an input component 1028. The output component 1026 may include a visual component (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, a resistive mechanism), other signal generators, and the like. The input component 1028 may include an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input component), a point-based input component (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing instrument), a tactile input component (e.g., a physical button, a touch screen that provides touch location and / or force or touch gestures, or other tactile input component), an audio input component (e.g., a microphone), and the like. Input components 1028 may also include one or more image capture devices, such as a digital camera for generating digital images or video.

[0069] In further example embodiments, the I / O components 1018 may include, among other components, a biometric component 1030, a motion component 1034, an environmental component 1036, or a position component 1038. One or more of these components (or a portion thereof) may be collectively referred to herein as a "sensor component" or "sensor" for collecting various data related to the machine 1000, the environment of the machine 1000, a user of the machine 1000, or a combination thereof.

[0070] For example, the biometric component 1030 may include components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 1034 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a velocity sensor component (e.g., a speedometer), a rotation sensor component (e.g., a gyroscope), etc. The environmental component 1036 may include, for example, an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of hazardous gases for safety purposes or measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The location component 1038 may include a location sensor component (e.g., a global positioning system (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects the altitude at which the air pressure is obtained), an orientation sensor component (e.g., a magnetometer), etc. For example, the location sensor component may provide location information associated with the system 1000, such as the GPS coordinates of the system 1000 or information regarding the current location of the system 1000 (e.g., the name of a restaurant or other business).

[0071] A variety of technologies can be used to implement communications. The I / O components 1018 may include a communications component 1040 operable to couple the machine 1000 to the network 1032 or the device 1020 via coupling 1024 and coupling 1022, respectively. For example, the communications component 1040 may include a network interface component or other suitable device for interfacing with the network 1032. In further examples, the communications component 1040 may include a wired communications component, a wireless communications component, a cellular communications component, a near field communications (NFC) component, a Components (e.g. Low energy consumption), Device 1020 may be another machine or any of a variety of peripheral devices, such as peripheral devices coupled via a universal serial bus (USB).

[0072] In addition, the communication component 1040 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 1040 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product codes (UPC) barcodes, multi-dimensional barcodes (e.g., Quick Response (QR) codes, Aztec codes, Data Matrix, Digital Graphics, Maximal codes, PDF417, Supercodes, UCC RSS-2D barcodes), and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). In addition, various information can be obtained via the communication component 1040, such as location via Internet Protocol (IP) geolocation, location information via Internet Protocol (IP), ... Signal triangulation to obtain location, obtaining location via detection of NFC beacon signals that can indicate a specific location, etc.

[0073] When phrases similar to "at least one of A, B, or C," "at least one of A, B, and C," "one or more of A, B, or C," or "one or more of A, B, and C" are used, the phrases are intended to be interpreted to mean that A can exist alone in an embodiment, B can exist alone in an embodiment, C can exist alone in an embodiment, or any combination of elements A, B, and C can exist in a single embodiment; for example, A and B, A and C, B and C, or A, B, and C.

[0074] Changes and modifications may be made to the disclosed embodiments without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

Claims

1. A system for performing acoustic zoom, comprising: Multiple beamformers that: generating a plurality of beamformer signals corresponding to a plurality of tiles of video content associated with the plurality of acoustic signals, wherein each beamformer is directed to a center of each tile; and Target Enhancer, which: identifying a tile having at least a portion included in a zoom region of the video content, selecting a beamformer signal corresponding to the identified tile, and The selected beamformer signals are combined to generate a target enhancement signal associated with the zoom region.

2. The system according to claim 1, wherein: The target enhancer is further configured to: determining a scale of each identified tile relative to the zoom region; and The selected beamformer signals are combined based on the ratio to generate the target enhancement signal.

3. The system according to claim 2, wherein: The target enhancer is further configured to: The selected beamformer signals are spectrally summed based on the ratio.

4. The system according to claim 1, further comprising: a neural network for receiving the plurality of acoustic signals to generate a noise reference signal, A plurality of beamformers receives the noise reference signal and generates the plurality of beamformer signals using the plurality of acoustic signals and the noise reference signal.

5. The system according to claim 1, further comprising: a time-frequency converter for receiving the plurality of acoustic signals and converting the plurality of acoustic signals from a time domain to a frequency domain; as well as A frequency-time converter is configured to receive the target enhanced signal and convert the target enhanced signal from the frequency domain to the time domain.

6. The system of claim 1 , further comprising: A camera is used to capture the video content.

7. The system according to claim 1, wherein: The tiles of the video content are equal-shaped tiles having an angular width of at least 10 degrees.

8. A method for performing acoustic zoom, comprising: causing, by a processor, a plurality of beamformers to generate a plurality of beamformer signals using a plurality of acoustic signals associated with video content, wherein the beamformer signals correspond to a plurality of tiles of the video content, wherein each beamformer is directed toward a center of each tile; identifying a tile having at least a portion included in a zoom region of the video content, selecting a beamformer signal corresponding to the identified tile, and The selected beamformer signals are combined to generate a target enhancement signal associated with the zoom region.

9. The method according to claim 8, further comprising: determining a scale of each identified tile relative to the zoom region; as well as The selected beamformer signals are combined based on the ratio to generate the target enhancement signal.

10. The method according to claim 9, further comprising: The selected beamformer signals are spectrally summed based on the ratio.

11. The method according to claim 8, further comprising: generating, by a neural network, a noise reference signal using the plurality of acoustic signals, The plurality of beamformer signals are generated using the beamformer using the plurality of acoustic signals and the noise reference signal.

12. The method according to claim 8, wherein The tiles of the video content are equal-shaped tiles having an angular width of at least 10 degrees.

13. A computer-readable storage medium having stored thereon instructions that, when executed by a processor, cause the processor to perform operations comprising: causing a plurality of beamformers to generate a plurality of beamformer signals using a plurality of acoustic signals associated with video content, wherein the beamformer signals correspond to a plurality of tiles of the video content, wherein each beamformer is directed to a center of each tile; identifying a tile having at least a portion included in a zoom region of the video content, selecting a beamformer signal corresponding to the identified tile, and The selected beamformer signals are combined to generate a target enhancement signal associated with the zoom region.

14. The computer-readable storage medium of claim 13, wherein: The processor performs operations further comprising: determining a scale of each identified tile relative to the zoom region; and The selected beamformer signals are combined based on the ratio to generate the target enhancement signal.

15. The computer-readable storage medium of claim 13, wherein: The processor performs operations further comprising: generating a noise reference signal based on the plurality of acoustic signals using a neural network, The plurality of beamformer signals are generated using the plurality of acoustic signals and the noise reference signal.

16. The computer-readable storage medium of claim 13, wherein: The processor performs operations further comprising: transforming the plurality of acoustic signals from the time domain to the frequency domain; and The target enhanced signal is transformed from the frequency domain to the time domain.

17. A system for performing acoustic zoom, comprising: A plurality of beamformers for receiving a plurality of acoustic signals, the plurality of beamformers comprising a target beamformer and a noise beamformer, wherein: The target beamformer is directed toward a center of a field of view corresponding to a zoomed area of video content and generates a target beamformer signal, and The noise beamformer has a null value directed toward a center of the field of view and generates a noise beamformer signal; and Target Enhancer, which is used to: determining a field of view corresponding to the zoom region of the video content, A target enhancement signal associated with the zoom region of the video content is generated using the target beamformer signal and the noise beamformer signal.

18. The system according to claim 17, wherein: Generating the target enhancement signal by the target enhancer includes spectrally subtracting the noise beamformer signal from the target enhancement signal.

19. The system of claim 17, further comprising: a neural network for receiving the plurality of acoustic signals to generate a noise reference signal, The plurality of beamformers receive the noise reference signal and use the plurality of acoustic signals and the noise reference signal to generate the target beamformer signal and the noise beamformer signal.

20. The system of claim 17, further comprising: a time-frequency converter for receiving the plurality of acoustic signals and converting the plurality of acoustic signals from a time domain to a frequency domain; as well as A frequency-time converter is configured to receive the target enhanced signal and convert the target enhanced signal from the frequency domain to the time domain.

Citation Information

Patent Citations

  • Mobile terminal and audio zooming method thereof

    CN103516894A

  • Technologies for localized audio enhancement of a three-dimensional video

    US20160381459A1