Network-based processing and distribution of multimedia content of live musical performance

JP2025063094A5Pending Publication Date: 2025-05-19DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024231532
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-05-04
Filing Date
2024-12-27
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

Delivering high-quality audio and video for live music demonstrations over the Internet is challenging due to poor sound and video quality, especially in unacoustic venues, and the lack of technical expertise in recording and editing.

Method used

A system that receives rehearsal data from microphones and cameras, analyzes it to derive performance levels and locations, and uses this data to compile and edit video during live demonstrations, applying rules to highlight prominent performers and improve editing.

Benefits of technology

The system enables the production of well-balanced audio and professional-quality video without relying on expert recording and editing, allowing bands to deliver high-quality content to multiple end-user devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To disclose methods, systems, and computer program products for network-based processing and distribution of multimedia contents of a live performance.SOLUTION: In some implementations, recording devices can be configured to record a multimedia event (e.g., a musical performance). The recording devices can provide the recordings to a server while the event is ongoing. The server automatically synchronizes, mixes, and masters the recordings. The server performs automatic mixing and mastering using reference audio data previously captured during a rehearsal. The server streams the mastered recording to multiple end users through the Internet or another public or private network. The streaming can be live streaming.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This disclosure relates generally to the capture, processing and delivery of multimedia content of live music performances. [Background technology]

[0002] Delivery of high quality audio and video of a live performance over the Internet can be difficult. Many amateur-made video recordings uploaded to the Internet have poor video and sound quality. When a band performs in an acoustically untreated venue, the sound quality can be poor if the recording is uploaded directly without further processing. For example, when a drum set is used, the drum set may be played too loudly and other instruments in the band cannot be clearly heard. Furthermore, if a band does not properly set up their recording equipment, including, for example, multiple microphones, preamps, and a mixing desk, the recording of the performance may have poor sound quality. Even if the recording equipment is properly set up, the band may lack the technical expertise to use the recording equipment effectively. Similarly, professional quality video recording and editing of a performance may require technical expertise beyond the skills of the performers. Summary of the Invention [Means for solving the problem]

[0003] A system, program product, and method for video editing based on rehearsal and live data are disclosed. The system receives rehearsal data for a rehearsal of a performance from one or more microphones and one or more video cameras. The system matches sounds and performers based on the rehearsal data. During a live performance, the system receives live audio and video of the performance. Based on an analysis of the rehearsal data, the system derives a level at which the performers perform relative to the rehearsal and a representative position of the performers during the rehearsal on the one or more video cameras. The system then edits the video data, e.g., highlighting prominent performers, based on rules utilizing the derived levels and positions. The system optionally uses an analysis of the performance to improve the edit. The analysis generates, for example, tempo or beat data and performer motion tracking data. The system then associates the audio data with the edited video data for storage and streaming to one or more user devices.

[0004] A system, program product, and method for video processing under limited network bandwidth is disclosed. A video camera can capture high definition video (e.g., 4K video) of a performance, which may be difficult to stream live (or even upload offline) over a communications network. The video camera can submit one or more frames of video, optionally at a lower resolution, optionally compressed using a lossy video codec, to a server system. Based on the one or more frames and audio data described in the previous paragraph, the server system can generate editing decisions for the video data. The server system can instruct the video camera to crop a portion of the high definition video corresponding to a performer or group of performers and submit that portion of the video to the server system as a medium definition or low definition video (e.g., 720p), optionally compressed using a lossless video codec. At any one time, the video camera device can store a long buffer (e.g., tens of seconds) of high definition video (e.g., 4K) corresponding to the last captured frames, so that commands received from the server system can be performed on frames captured several seconds ago. The server system can then store the medium or low definition video or stream the medium or low definition video to the user device.

[0005] Implementations for network-based processing and delivery of multimedia of live performances are disclosed. In some implementations, a recording device can be configured to record an event (e.g., a live musical performance). The recording device provides the recording to a server during the performance. The server automatically synchronizes, mixes, and masters the recording. In one implementation, the server performs the automated mixing and mastering using reference audio data captured during a rehearsal where the recording device and audio sources were placed in the same acoustic (and in the case of a video recording device, visual) configuration as at the event. The server provides the mastered recording to multiple end user devices, for example by live streaming.

[0006] In some implementations, a server streams video signals of a live event to multiple users. Using reference audio data (also referred to as rehearsal data) recorded during a rehearsal session, the server determines the locations of various instruments and vocalists (hereinafter also referred to as "sound sources") and the locations of the performers at the recording location. During the live performance, the server determines one or more dominant sound sources based on one or more parameters (e.g., volume). An image capture device (e.g., a video camera) can capture live video of the performance and send it to the server. Using the locations of the dominant sound sources, the server determines the portions of the video to which video editing operations (e.g., zooms, transitions, visual effects) should be applied. Applying the video editing operations can occur in real time on the live video or on previously recorded video data. The server streams the portions of the video that correspond to the dominant sound sources (e.g., a close-up of the lead vocalist or lead guitar player) to the end user device. For example, the server may provide a video overlay or graphical user interface on the end user device that allows the user to control audio mixing (e.g., increasing the volume of a vocalist or an instrument during a solo performance) or video editing (e.g., zooming in on a particular performer). In some implementations, the server may issue commands to an audio or video recording device to adjust one or more recording parameters, such as adjusting the recording level on a microphone preamplifier, the zoom level of a video recorder, turning a particular microphone or video recorder on or off, or any combination of the above.

[0007] The features described herein can achieve one or more advantages over conventional audio and video techniques that improve upon conventional manual audio and video processing techniques by automated mixing and mastering of audio tracks based at least in part on reference audio data derived from reference audio data. Using the automated mixing and mastering disclosed herein, a band can produce a well-balanced sound without resorting to using professional recording, mixing and mastering engineers. If a band desires a mixing style from a particular professional, the band can retain the professional to mix and master their recordings remotely using the network-based platform disclosed herein.

[0008] Similarly, the disclosed implementations improve upon conventional video processing techniques by replacing manual camera operations (e.g., panning and zooming) with automated camera operations based, at least in part, on audio and video rehearsal data, allowing bands to generate and edit professional-quality videos of their live performances without retaining a professional videographer.

[0009] A band can provide high quality audio and video to multiple end user devices using a variety of techniques (e.g., live streaming). To enhance the end user experience, the streaming can be made interactive, allowing the end user to control various aspects of the audio mixing and video editing. For convenience herein, the term band can refer to a musical band of one or more performers and instruments. The term can also refer to a group of one or more participants in a non-musical environment (e.g., performers in a drama, speakers in a conference, or loudspeakers in a public announcement system).

[0010] The features and processes disclosed herein improve upon conventional server computers by configuring the server computer to perform automated synchronization, mixing and mastering of audio tracks and editing of video data of a live performance. The server computer can stream the processed audio and video to end user devices and provide controls that allow the end user to further mix and edit the audio and video. In various implementations, the server computer can store raw data of the live performance for offline use, mixing, mastering, repurposing, segmenting, curation. The server computer can store processed data for later distribution. The server computer can store data that has been through various stages of processing anywhere between raw data and fully processed data, inclusive. The server can store the data on a storage device (e.g., hard disk, compact disc (CD), remote storage site (e.g., cloud-based audio and video service), or memory stick).

[0011] The features and processes described herein improve upon conventional server computers by allowing the server computer to automatically edit video data based on various rules. A server computer implementing the disclosed techniques can instruct a recording device, e.g., a video camera, to automatically focus on a performer, e.g., a soloist, when that performer is performing differently (e.g., louder) than other performers, or when the performer moves, or when the performer performs without musical accompaniment (e.g., a cappella). The server computer can cut and change scenes according to the tempo and beat of the music. The server computer instructs the recording device to track the movement of the sound source, e.g., switch from a first performer to a second performer, where the switch can be a hard cut or a slow pan from the first to the second performer. Tracking can be performed on the recorded data without physically moving the recording device. Thus, based on audio data analysis, the server computer can mimic what a human cameraman can do. In this manner, the disclosed techniques have the technical advantage of moving the view of an event without physically moving the recording device.

[0012] The features and processes disclosed herein improve upon conventional server computers by reducing the bandwidth requirements for transmitting high definition video data. High definition video data, e.g., 4K video, may require high bandwidth for transmission. The disclosed features can select a highlight of the video to be transmitted, e.g., a portion corresponding to the position of a soloist, and focus on that location. The system can transmit that portion of the video data at a lower resolution, e.g., 720p video. In this way, when the audience is only watching the soloist, the system does not need to transmit a video of the entire stage in 4K video. The system can still preserve the perceived definition and clarity of the soloist. Thus, the system achieves the technical advantage of transmitting high quality video at a reduced bandwidth.

[0013] In another embodiment, the camera device or devices capture high resolution (e.g., 4K) video and the methods described herein are used to stream an intermediate edited, lower resolution video (e.g., 1080) so that the server system can further decide to edit within the 1080 frames and provide 720p to the audience.

[0014] The details of one or more implementations of the disclosed subject matter are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the disclosed subject matter will become apparent from the description, drawings, and claims. [Brief description of the drawings]

[0015] [Figure 1] FIG. 1 illustrates a first exemplary placement of recording devices at an event.

[0016] [Diagram 2] FIG. 13 illustrates a second exemplary placement of recording devices at an event.

[0017] [Diagram 3] FIG. 2 is a block diagram illustrating an exemplary architecture of a recording device.

[0018] [Figure 4] FIG. 1 illustrates an exemplary audio and video system architecture for network-based audio processing.

[0019] [Diagram 5] FIG. 2 is a block diagram showing example signal paths for audio and video processing.

[0020] [Figure 6] 1 is a flow chart of an exemplary process for audio processing.

[0021] [Figure 7] FIG. 1 is a block diagram illustrating an exemplary automated mixing and mastering unit.

[0022] [Figure 8] 1 is a flowchart illustrating an exemplary process for automated leveling.

[0023] [Figure 9] 1 is a flow chart illustrating an example process for automated panning.

[0024] [Figure 10] FIG. 1 illustrates an example angle transformation at maximum distortion.

[0025] [Figure 11] 4 is a flowchart illustrating an example process for estimating energy level from a microphone signal.

[0026] [Figure 12] 1 is a flowchart illustrating an example process for estimating energy levels in a frequency band.

[0027] [Figure 13] 1 is a flowchart illustrating an exemplary process for automatically equalizing individual sound sources.

[0028] [Figure 14A] FIG. 2 illustrates an example three-instrument mix to be equalized.

[0029] [Figure 14B] FIG. 1 illustrates an example gain in automatic equalization.

[0030] [Figure 15] 1 is a flowchart illustrating an example process for segmenting a video based on novelty accumulation in audio data.

[0031] [Figure 16] FIG. 1 illustrates an exemplary novelty accumulation process.

[0032] [Figure 17] 4 is a flowchart illustrating an example process for synchronizing signals from multiple microphones.

[0033] [Figure 18] FIG. 1 illustrates an exemplary sequence for synchronizing five microphones.

[0034] [Figure 19] A and B show an exemplary user interface displaying the results of automated video editing.

[0035] [Figure 20] 1 is a flowchart of an exemplary process for automated video editing.

[0036] [Figure 24] 24 is a flow chart illustrating an example process 2400 for noise reduction.

[0037] [Diagram 25] FIG. 13 is a block diagram illustrating an exemplary technique for video editing based on rehearsal data.

[0038] [Figure 26] 1 is a flowchart illustrating an exemplary process of video editing based on rehearsal data.

[0039] [Figure 27] A block diagram illustrating an example technique for selecting sub-frame regions from full-frame video data.

[0040] [Figure 28]5 is a flowchart illustrating an exemplary process for selecting sub-frame regions from full-frame video data performed by a server system.

[0041] [Figure 29] 4 is a flowchart illustrating an exemplary process for selecting sub-frame regions from full-frame video data performed by a video capture device.

[0042] [Figure 21] FIG. 2 is a block diagram illustrating an example device architecture of a mobile device that implements the features and operations described with reference to FIGS. 1-20 and 24-29.

[0043] [Figure 22] FIG. 27 is a block diagram of an example network operating environment for the mobile devices of FIGS. 1-20 and 24-29.

[0044] [Figure 23] A block diagram of an exemplary system architecture for a server system implementing the features and operations described with reference to Figures 1-20 and 24-29.

[0045] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0046] Exemplary High-Level Architecture FIG. 1 illustrates a first exemplary arrangement of recording devices at a live performance event 100. Event 100 can be any event at which audio content (e.g., spoken, vocal, or instrumental) and optionally video content is produced. In particular, event 100 can be a live concert in which one or more musical instruments and / or one or more vocalists are performing. One or more sound sources can be present at event 100. Each sound source can be an instrument, a vocalist, a loudspeaker, or any item that produces sound. For simplicity, sound sources, including non-instrumental sound sources, are collectively referred to as instruments throughout this specification.

[0047] In some implementations, the devices 102, 104, 106 can be configured to record audio and video of the event 100. The devices 102 and 104 can be mobile devices (e.g., smartphones, wearable devices, or portable audio and video recorders). The devices 102 and 104 can include built-in microphones, be coupled to external microphones, or both. If an external microphone is used, the external microphone can be coupled to one or more microphone preamplifiers. The external microphones can be coupled to the devices 102 and 104 using wired or wireless connections. In some implementations, each of the devices 102 and 104 can be coupled to one or more external sound generating devices. Here, the sound generating devices generate audio signals directly in the form of analog electrical signals (e.g., keyboard output) or digital signals (e.g., digital sounds generated by a laptop computer). Such signals can be directly provided to the devices 102 and 104 via corresponding adapters.

[0048] Each of the devices 102 and 104 can execute an application program for recording audio content of the event 100. The application program can send the recorded audio tracks to a remote server computer over a communications network 110. The communications network 110 can be a personal area network (PAN, e.g., a Bluetooth network), a local area network (LAN), a cellular network (e.g., a 4G or 5G data network), a wide area network (WAN, e.g., the Internet), or an ad-hoc network. The communication can be through a gateway (e.g., wireless device 108) or individually. In some implementations, the server computer can be local to the event 100. For example, the server computer can be any of the devices 102 and 104.

[0049] In some implementations, each of the devices 102 and 104 can include a client application linked to an online user account for audio processing. The client application can perform user authentication and authorization prior to sending audio tracks to the online user account. The client application can include audio processing functions that respond to commands from a remote server, such as commands to adjust filters (e.g., low pass, high pass, shelf filters), gain or frequency bands of microphone preamplifiers incorporated or coupled to the device 102 or 104. Additionally or alternatively, the commands can control the bit depth and sample rate of the recording (e.g., 16 bits at 44.1 Hz).

[0050] Each of the devices 102 and 104 can submit the recorded audio tracks to a server through a wired device (or a wired or wireless router) or other wired or wireless devices 108. The wireless devices 108 can be wireless access points (APs) or cellular towers for a wireless local area network (WLAN). The wireless devices 108 can be connected to a communications network 110. The devices 102 and 104 can send the recorded live audio tracks to the server through the communications network 110. The data submission can be made in real time, e.g., while the performance is in progress, or offline after the performance has been partially or completely completed, e.g., by the devices 102 and 104, either simultaneously or sequentially. The devices 102 and 104 can store the recorded audio tracks for offline submission.

[0051] In some implementations, the device 106 is an image capture device configured to capture images and audio of the event 100. For example, the device 106 can be configured to capture high definition video (e.g., 4K resolution video). The device 106 can capture still images and video of the event 100. The device 106 can send the captured still images or video to a server through the wireless device 108 and the communication network 110. The device 106 can send the still images or video in real time or offline. In some implementations, the device 106 can perform the operations of a server computer.

[0052] In some implementations, one or two microphones of the devices 102, 104, 106 are designated as one or more main microphones (e.g., "room" microphones) that capture audio from all instruments and vocalists. The signals output from the one or two main microphones can be designated as main signals (e.g., main mono or main stereo signals) or main channel signals. Other microphones, sometimes placed at each individual sound source (e.g., vocal microphones) or individual sound sources (e.g., drum microphones), are designated as spot microphones, also referred to as satellite microphones. Spot microphones can augment the main microphones by providing more localized capture of sound sources (e.g., kick drum microphone, snare drum microphone, hi-hat microphone, overhead microphones for capturing cymbals, guitar and bass amplifier microphones, etc.).

[0053] In some implementations, each of devices 102, 104, 106 can be configured to perform the actions of a server by executing one or more computer programs. In such implementations, the processing of the audio signal can be performed on-site at the event 100. The device performing the action (device 106) can then upload the processed signal to a storage device or end-user device over the communications network 110.

[0054] 2 illustrates a second exemplary arrangement of recording devices at an event 100. The integrated recorder 200 can be configured to record audio and video signals of the event 100. The integrated recorder 200 can include microphones 202 and 204. Each of the microphones 202 and 204 can be an omnidirectional, directional, or bidirectional microphone or a microphone with any directional pattern. Each of the microphones 202 and 204 can be positioned to point in a given direction. The microphones 202 and 204 can be designated as primary microphones. In various implementations, the integrated recorder 200 can be coupled to one or more spot microphones for additional audio input.

[0055] The integrated recorder 200 may include an image capture device 206 to capture still images or video of the event 100. The integrated recorder 200 may include or be coupled to a user interface for specifying one or more attributes (e.g., target loudness level) of the sound sources of the event 100. For example, the integrated recorder 200 may be associated with a mobile application by a device identifier. The mobile application may have a graphical user interface (GUI) for display on a touch-sensitive surface of the mobile device 207. The GUI may include one or more user interface items configured to accept user input for specifying attributes of the sound sources (e.g., target volume or gain levels of a guitar, lead vocalist, bass, or drums). The attributes may include, for example, how many decibels (dB) apart two sound sources (e.g., between a lead vocalist and another sound source) should be, or how much (e.g., X dB) which sound source should be dominant (e.g., by playing louder than the volume level of the other sound source). The one or more user interface items can accept user input specifying a rehearsal session at which reference audio data from an audio source is to be collected for automated mixing.

[0056] The integrated recorder 200 can optionally perform one or more operations including, for example, synchronizing signals from a primary microphone and a spot microphone, separating audio sources from the recorded signal, mixing signals of different audio sources based on reference audio data, and mastering the mixed signal. The integrated recorder 200 can submit the mastered signal as a stereo or multi-channel signal to a server through a wireless device 208 connected to a communication network 210. Similarly, the integrated recorder 200 can provide a video signal to the server. The server can then deliver the stereo or multi-channel signal and the video signal to end-user devices in substantially real-time during the event 100. The communication network 210 can be a PAN, a LAN, a cellular data network (e.g., a 4G network or a 5G network), a WAN, or an ad-hoc network.

[0057] 3 is a block diagram illustrating an example architecture of a recording device 302. Recording device 302 can be device 102 or 104 of FIG. 1 or integrated recorder 200 of FIG.

[0058] Recording device 302 may include or be coupled to a primary microphone 304 and a video camera 306. Primary microphone 304 may be a built-in microphone or a dedicated microphone coupled to recording device 302. Primary microphone 304 may provide a baseline (also referred to as a bed) for audio signal processing, as described in more detail below. Video camera 306 may be a built-in camera or a dedicated camera coupled to recording device 302. Video camera 306 may be a Digital Cinema Initiative (CDI) 4K, DCI 2K, or Full HD video camera configured to capture video at a high enough resolution such that a portion of the captured video may be zoomed in and still utilize the full capacity of a typical monitor with a moderate (e.g., 1080p, 1080i, 720p, or 720i) resolution.

[0059] The recording device 302 may include an external microphone interface 308 for connecting to one or more spot microphones 310. The external microphone interface 308 is configured to receive signals from the one or more spot microphones 310. In some implementations, the external microphone interface 308 is configured to provide control signals to the one or more spot microphones 310. The recording device 302 may include an external camera interface 312 for connecting to one or more external cameras 314. The external camera interface 314 is configured to receive signals from the one or more external cameras 314 and provide control signals to the one or more external cameras 314.

[0060] The recording device 302 can include one or more processors 320. The one or more processors 320 can be configured to perform analog-to-digital conversion of audio signals from a microphone and digital compression of digital audio and video signals from a camera. In some implementations, the one or more processors 320 are further configured to synchronize audio signals from various channels, separate sound sources from those audio signals, automatically mix the separate sound sources, and master the mixed signal.

[0061] The recording device 302 can include a network interface 322 for submitting digital audio and visual signals to a server through a network device. In some implementations, the network interface 322 can submit mastered digital audio and video signals to the server. The network interface 322 can be configured to receive commands from the server to adjust one or more parameters of an audio or visual recording. For example, the network interface 322 can receive commands to pan and zoom in (or out) a video camera in a specified direction or to adjust recording levels for a particular microphone.

[0062] The recording device 302 may include a user interface 324 for receiving various user inputs that control attributes of the recording. The user interface 324 may include a GUI that is displayed on a touch-sensitive surface of the recording device 302. The user interface 324 may be displayed on a device separate from the recording device 302, such as a smart phone or tablet computer running a client application program.

[0063] 4 illustrates an example architecture of an audio and video system 400 for network-based audio and video processing. In network-based audio and video processing, a communication network 402 links an event to end-user devices, allowing end-users of the end-user devices to hear and see the live performance of an artist at the event 100 (of FIG. 1). The communication network 402 can be a PAN, a LAN, a cellular network, a WAN (e.g., the Internet), or an ad-hoc network. The audio and video system 400 can include one or more subsystems. Each subsystem is described below.

[0064] Studio-side system 404 is a subsystem of audio system 400 that includes equipment located and deployed at a location, for example, in a studio, concert hall, theater, stadium, living room, or other venue where an event takes place. Studio-side system 404 can include the architecture discussed with reference to FIG. 1, where multiple general-purpose devices (e.g., smart phones, tablet computers, laptop computers) each running an audio or video processing application program record and send the recorded signals to a server 408. Alternatively, studio-side system 404 can include the exemplary architecture discussed with reference to FIG. 2, where a dedicated integrated recorder records and sends the recorded signals to a server 408.

[0065] The server 408 is a subsystem of the audio system 400 that includes one or more computers or one or more discrete or integrated electronic circuits (e.g., one or more processors). The server 408 is configured to receive live audio and video content of the event 100 over the communication network 402, process the audio and video content, and provide the audio and video content to end user devices over the communication network 402. The server 408 can include one or more processors programmed to perform audio processing. In some implementations, the server 408 can control various aspects of the studio-side system 404. For example, the server 408 can increase or decrease microphone volume levels when clipping is detected, increase or decrease sample bit rate or bit depth, or select a compression type based on detected bandwidth limitations.

[0066] In some implementations, the server 408 automatically mixes and masters the audio signal. The server 408 can also automatically select specific portions of the video stream that correspond to the instruments being played. Further details about the components and operation of the server computer 408 are provided below with reference to FIG.

[0067] In some implementations, the server 408 allows equipment at an editor-side system 420 to perform mixing, mastering, and scene selection. The editor-side system 420 is a subsystem of the audio system 400 configured to allow a third-party editor to edit audio or video content during live content streaming. The editor-side system 420 can include one or more mixer devices 422. The mixer devices 422 can be operated by an end user, a player in a band or orchestra performing a live event, or a professional mixing engineer. The editor-side system 420 can include one or more video editing devices 424. The video editing devices 424 can be operated by an end user, a performer, or a professional videographer.

[0068] End users can listen to and view live content of the event 100 at a variety of end-user systems 410. In the various end-user systems 410, live or stored content can be played on a user audio device 412 (e.g., a stereo or multi-channel audio system with multiple loudspeakers), a user video device 414 (e.g., one or more computer monitors) or a combination of both (e.g., a television set, a smart phone, a desktop, laptop or tablet computer, or a wearable device).

[0069] In some implementations, the audio system 400 allows end users to provide feedback about and control various aspects of the live content using their end user devices, for example, the audio system 400 can allow real-time ratings of the live content based on votes or some type of authorized end user video panning.

[0070] 5 is a block diagram illustrating an example signal path for audio and video processing. The components of the signal path may be implemented on the server 408 (of FIG. 4). The components may include a synchronizer 502, a source separator 504, a mixing and mastering unit 506, a delivery front end 508, and an estimator 522. In some implementations, some or all of the components may be implemented in software on the server computer 408. In other implementations, some or all of the components may include one or more electronic circuits configured to perform various operations. Each electronic circuit may include one or more discrete components (e.g., resistors, transistors, or vacuum tubes) or integrated components (e.g., integrated circuits, microprocessors, or computers).

[0071] The synchronizer 502 can receive digital audio data of the event 100 from one or more recording devices. The digital audio data can be, for example, sampled audio data. Each recording device or each microphone coupled to a recording device can correspond to an audio channel or track of the musical performance. The signals from the recording devices are referred to as channel signals. Thus, the synchronizer 502 can receive Nm channel signals, where Nm is the total number of microphones recording the event 100, or more generally, all sound signals of the set captured in the event 100. For example, the Nm channel signals can include one or more channels from a direct output of a keyboard or from a line audio output of a computing device or portable music player. The Nm channel signals can include main channel signals from environmental microphones and spot channel signals (also referred to as beams) from spot microphones. The Nm channel signals can be recorded by microphones on the recording devices and sampled by analog-to-digital converters locally by the recording devices. The recording device can send the sampled audio data in a packetized audio format over the network to the synchronizer 502. Thus, the Nm channel signals can refer to digitized audio signals rather than analog signals directly from the microphones.

[0072] The Nm channel signals may become out of synchronization in time. For example, packets of the digital signal may not arrive at the server in the time order in which the corresponding captured sound signals were physically generated. The synchronizer 502 may generate an output including the Nm synchronized channel signals, for example based on timestamps associated with the packets. The synchronizer 502 may provide the Nm synchronized channel signals to a source separator 504. Further details of the operation of the synchronizer 502 are described below with reference to Figures 17 and 18.

[0073] The source separator 504 is a component of the server 408 configured to separate each audio source from the Nm synchronized signals. Each audio source can correspond to, for example, an instrument, a vocalist, a group of instruments, or a group of vocalists. The source separator 504 outputs Ns signals, each corresponding to an audio source. The number of audio sources (Ns) can be the same or different from the number of synchronized signals Nm. In some implementations, the source separator 504 can be bypassed.

[0074] The Ns signals output from the source separator 504 or the Nm synchronized signals output from the synchronizer 502 (if the source separator 504 is bypassed) may be input to one or more mixing and mastering units 506. The mixing and mastering units 506 may be software and / or hardware components of the server 408 configured to perform mixing operations on the channels of the individual audio sources based at least in part on the reference audio data and to perform mastering operations on the mixed audio signals to generate a final N-channel audio signal (e.g., stereo audio, surround sound). The mixing and mastering units 506 may output the N-channel audio signals to a distribution front end 508. In various implementations, the mixing and mastering units 506 may perform operations of applying mixing gain, equalizing each signal, performing dynamic range correction (DRC) on each signal, and performing noise reduction on each signal. The mixing and mastering unit 506 can perform these operations in various combinations, either on each signal individually or on multiple signals simultaneously.

[0075] The reference audio data may include audio content recorded by microphones at a rehearsal and processed by the estimator 522. At the rehearsal, microphones and sound sources are placed in the same acoustic arrangement as at the live event 100. The microphones then record audio signals when each sound source is played individually. Additionally, the microphones may record noise samples when the sound sources are not playing.

[0076] The estimator 522 is a component configured to collect and process audio data from a rehearsal session. The estimator 522 can instruct each source player at the performance position to play or sing their instrument individually. For example, the estimator 522 can instruct each performer to play their instrument at a low volume for X seconds and at a high volume for Y seconds (e.g., by prompting through the device user interface). Nm signals from rehearsal microphones can be recorded. The estimator 522 can process the Nm signals, determine a loudness matrix, derive source characteristics and positions, and provide the instrument characteristics and positions to the mixing and mastering unit 506 for mixing operations. The estimator 522 can receive additional inputs that constitute parameters for determining the instrument characteristics and positions. Further details of the components and operation of the estimator 522 are described below with reference to Figures 8, 9, 10, and 13.

[0077] The delivery front end 508 may include an interface (e.g., a streaming or web server) for providing the N-channel audio to an end user device for storage or for download, including live streaming (e.g., HyperText Transfer Protocol (HTTP) live streaming, Real Time Streaming Protocol (RTSP), Real Time Transport Protocol (RTP), RTP Control Protocol (RTCP)). Live streaming may occur substantially in real time during the event 100.

[0078] The server 408 may include a video editor 530. The video editor 530 is a component of the server configured to receive a video signal of the event 100 and automatically edit the video signal based at least in part on the audio content. Automatically editing the video may include, for example, zooming in (e.g., a close-up shot) on a particular instrument or player when the video editor 530 determines that the particular instrument is a dominant sound source. Further details of the operation of the video editor 530 are described below with reference to Figures 19A and 19B and 20.

[0079] 6 is a flow chart of an exemplary process 600 for audio processing. Process 600 may be performed, for example, by server 408 of FIG. 4. Process 600 improves upon conventional audio processing techniques by automating various mixing and mastering operations based, at least in part, on reference audio data recorded at a rehearsal. In this specification, the term rehearsal refers to a session.

[0080] The server 408 can receive (602) reference audio data from one or more channel signal sources. The reference audio data can include acoustic information of one or more sound sources playing individually in a rehearsal. The reference audio data can include acoustic information of the noise floor in a rehearsal, for example, when the sound sources are not playing. Each channel signal source can include a microphone or a line output. Each sound source can be, for example, an instrument, a vocalist, or a synthesizer. The server 408 can receive the reference audio data through a communication network (for example, communication network 402 of FIG. 4). The first channel signal can be captured by a first channel signal source (for example, device 102) recording the rehearsal at a first location (for example, front stage left or at a particular instrument). The second channel signal can be captured by a second channel signal source (for example, device 104) recording the rehearsal at a second location (for example, front stage right or at a particular instrument).

[0081] The server 408 can receive (604) one or more channel signals of a performance event, e.g., event 100, from one or more channel signal sources. Each channel signal can be a digital or analog signal from a respective channel signal source. Each channel signal can include audio signals from the one or more sound sources playing at the performance event. At the performance event, the sound sources and the channel signal source positions are located in the same acoustic arrangement (e.g., in the same position). In some implementations, the server 408 can automatically synchronize the first channel signal and the second channel signal in the time domain. After synchronization, the server 408 can determine the first sound source and the second sound source from the first channel signal and the second channel signal.

[0082] The server 408 can automatically mix (606) the one or more channel signals during the event 100 or after the conclusion of the event 100. The automated mixing operations can include adjusting one or more attributes of the sound effects from one or more sound sources of the event 100 based on reference audio data. For example, the automated mixing operations can include performing noise reduction on each sound source individually, balancing or leveling each sound source, and panning each sound source.

[0083] The mixing operations may also include automatically adjusting attributes of signals from one or more sound sources of the event 100 based at least in part on the reference audio data. Automatically adjusting the attributes of the one or more sound sources may include increasing or decreasing the gain of the one or more sound sources according to the respective volume levels of each sound source. Automatically adjusting the attributes of the one or more sound sources may include increasing or decreasing the gain of each channel signal from each sound source, or both, resulting in each of the one or more sound sources reaching or nearly reaching a target volume level. The server computer 408 may determine the respective volume levels at least in part from the reference audio data using the estimator 522. Other mixing operations may include, but are not limited to, applying compression, equalization, saturation or distortion, delay, reverberation, modulation, stereo, filtering and riding of vocal or instrument volumes.

[0084] The reference audio data may include audio signals recorded by a first recording device and a second recording device in a rehearsal session prior to the event 100. The reference audio data may be recorded individually for each sound source or group of sound sources in the rehearsal session. The reference audio data may include a first sound level signal (e.g., designated as a small or low volume) and a second sound level signal (e.g., designated as a large or high volume) for each sound source. The reference audio data may be recorded for background noise when the sound sources are not playing. In some implementations, the reference audio data may include a single sound level signal (e.g., when each sound source is playing at a medium volume).

[0085] The server 408 can determine, at least in part from the reference audio data, a respective gain for each sound source in the event 100. Determining the respective gains can include receiving an input specifying a target volume level for each sound source or group of sound sources (e.g., guitars, drums, background vocals). The server computer 408 can determine a respective volume level of the signals in the reference audio data using the estimator 502. The server 408 can determine the respective gains based on a difference between the volume level of the signal in the reference audio data and the target volume level.

[0086] In some implementations, the automated mixing (606) operations can include adjusting gain of signals from one or more audio sources according to input from a remote human mixing or mastering engineer logged onto the server system through a communications network, such that a remote mixing or mastering engineer who is not at the event 100 can mix or master the audio sources of the event 100 during live streaming.

[0087] The server 408 can provide 608 the downmix from the server system to a storage device or to an end user device as live content of the event 100, for example by live streaming. The end user device can play the content on one or more loudspeakers integrated or coupled to the end user device. In some implementations, the server 408 can automate video editing for the event 100. The video editing can be live editing while the event 100 is in progress or offline editing on previously recorded video of the event 100. Further details of automated video editing operations are described in Figures 19A and 19B and 20. In some implementations, a remote human video editor can use the platform to provide video editing during the event 100.

[0088] In some implementations, the server 408 can provide commands to the first and second recording devices based on the first channel signal or the second channel signal. The commands can adjust recording parameters of the recording devices. For example, the commands can instruct the recording devices to adjust gain, compression type, compression or sample rate (e.g., 44.1 Hz) or bit depth (e.g., 16 or 24 bits).

[0089] 7 is a block diagram illustrating components of an exemplary mixing and mastering unit 506. The mixing and mastering unit 506 may include various electronic circuits configured to perform mixing and mastering operations. The mixing and mastering unit 506 improves upon conventional mixing and mastering techniques by automating signal leveling and panning at the mixing stage, and by automating novelty-based signal segmentation when continuous, long crescendos are present.

[0090] The mixing and mastering unit 506 may include a mixing unit 702 and a mastering unit 704. The mixing unit 702 is a component of the mixing and mastering unit 704 configured to perform mixing operations on the Ns signals from the source separator 504 or the Nm synchronized signals from the synchronizer 502 automatically using reference audio data and inputs from one or more remote or local mixing consoles.

[0091] The mixing unit 702 may include, among other components, a leveling unit 706, a panner 708, a source equalizer 710, and a noise reduction unit 711. The leveling unit 706 is a component of the mixing unit 702 configured to adjust the respective gain for each source or microphone. The adjustments can be based at least in part on reference audio data, by input from a mixing desk, or a combination of both. Further details of the operation of the leveling unit 706 are described below with reference to FIG. 8.

[0092] Panner 708 is a component of mixing unit 702 that is configured to spatially position each sound source at a position in a virtual sound stage (e.g., left, right, center). Further details of the operation of Panner 708 are described below with reference to Figures 9 and 10.

[0093] The source equalizer 710 is a component of the mixing unit 702 that is configured to perform equalization (EQ) operations on individual sources rather than on the mixed audio signal as a whole. Further details of the operation of the source equalizer 710 are described below with reference to Figures 13 and 14A and B.

[0094] The noise reduction unit 711 is a component of the mixing unit 702 that is configured to perform noise reduction (NR) operations on individual signals rather than across the spectrum of all signals. Further details of the operation of the noise reduction unit 711 are described below with reference to FIG.

[0095] The mastering unit 704 may include, among other components, an equalizer 712 and a segmentation unit 714. The equalizer 712 is a module of the mastering unit 704 configured to smooth the sound levels across various frequencies for the mixed audio signals as a whole. The segmentation unit 714 is a module of the mastering unit 704 configured to split the video signal into multiple segments based on inherent characteristics of the audio signal. In some implementations, the segmentation unit 714 is a component of or coupled to the video editor 530 of FIG. 5. Further details of the operation of the segmentation unit 714 are described below with reference to FIG. 15 and FIG. 16.

[0096] FIG. 8 is a flow chart illustrating an exemplary process 800 for automatically leveling audio sources. Process 800 may be performed by the leveling unit 706 (of FIG. 7). In automatic leveling, the leveling unit 706 may automatically adjust the volume levels of each of the audio sources to a target level. Process 800 improves upon conventional mixing techniques by performing gain adjustments automatically based at least in part on reference audio data, rather than based on manual adjustments by a human. This allows for rapid processing of vast amounts of music content in real time.

[0097] The leveling unit 706 may receive (802) reference audio data (also referred to as rehearsal data). The reference audio data may include representations of channel signals from a channel signal source, e.g., a main microphone and a spot microphone of a multi-audio source. The representations may be direct channel signals from the channel signal source or partially processed signals, e.g., equalized or passed through dynamic range compensation.

[0098] The leveling unit 706 may determine 804 a respective correlation between each pair of channel signal sources, e.g., microphones. Details of determining the correlation are described below with reference to equation (3).

[0099] The leveling unit 706 may specify (806) the energy level of each of the primary microphones as a baseline associated with unity gain or some other reference level (eg, -18 dB).

[0100] In some implementations, the leveling unit 706 can determine 808 the respective contribution of each spot microphone to the baseline.

[0101] The leveling unit 706 may receive target level data specifying a target level for each sound source 810. The target level data may be received from a user interface.

[0102] The leveling unit 706 may determine a cost function for rescaling the audio signal to a target level according to the respective gains based on the respective contributions (812). The cost function may be a function of a variable (in this case the gains) that is to be solved so that the function has a minimum. Solving for the variables of the cost function is referred to as minimizing the cost function. Details and examples of solving for the variables of the cost function are provided below in the section entitled "Minimizing a Cost Function by Best Guess".

[0103] The leveling unit 706 may calculate a respective gain for each channel signal by minimizing a cost function (814). The leveling unit 706 may apply respective gains to the channel signals in the live audio data to achieve the target level for each sound source. The leveling unit 706 may provide the resulting signals to other components for further processing and playback on loudspeakers or headphones in an end user device. Further details and examples of the process 800 are described below.

[0104] A set of indexes i=1,……,N i can represent the sound source number. Here, N i is the total number of sound sources in event 100 (in Fig. 1). The set of indexes b=1,……,N b can represent the beam number, where each beam is a channel signal from a respective spot microphone, as mentioned earlier. N b is the total number of spot microphones. The set of indices M=L,R,1,……,N b can represent the combination of the dominant left microphone (L) and dominant right microphone (R) plus a beam index. If the dominant microphone is a mono microphone, the set of indices is M=Mono,1,……,N. bwhere the term Mono represents a mono microphone. The rest of the process is similar. Multiple sound sources may be assigned to the same beam. Thus, in some scenarios, N b <N i This may be the case, for example, when placing a spot microphone close to a guitar player who also sings. In this example, the vocals and guitar are assigned to the same spot microphone. Thus, the Leveling Unit 706 reduces the total number of signals that will be present in the final mix to N M can be specified as

[0105] One of the inputs to the algorithm performed by the leveling unit 706 is a loudness matrix L that quantifies the loudness level of each instrument i in each beam M (e.g., in dB). iM The estimator 522 estimates the loudness matrix L iM The leveling unit 706 calculates the loudness matrix L iM Therefore, the leveling unit 706 scales the energy of each instrument at each microphone by the energy matrix E iM can be expressed as:

number

[0106] The leveling unit 706 calculates the gain g MThe leveling unit 706 is configured to determine g, which is a vector representing the respective gain for each channel, including the main channel and the spot channel. Here, the gains for the two main channels are represented first. The leveling unit 706 can fix the absolute scale so that all energy is referenced to the energy in the main stereo channel. The leveling unit 706 can determine not to apply any gain to the energy in the main stereo channel. In this approach, the main stereo channel can be designated as a baseline with unity gain. The leveling unit 706 can calculate the contribution of each spot microphone above this baseline. Thus, the leveling unit 706 determines g M Change the first two entries of:

number

[0107] To estimate the energy after mixing the various signals, the leveling unit 706 first calculates, for each sound source i, the normalized correlation matrix (C i ) M,M' Each C i is obtained from the rehearsal of only the source i. The leveling unit 706 uses s to represent the signals captured by the M microphones (primary stereo plus beams) when the instrument i is rehearsed. iM The leveling unit 706 can calculate the normalized covariance matrix as follows:

number

[0108] Using this covariance matrix, the leveling unit 706 calculates the total energy E i can be expressed as:

number

number

[0109] The other input to the leveling unit 706 is the target loudness level (or target energy T on a linear scale) for each source i in the final mix. i ) In principle, only the relative target level matters. Since the Leveling Unit 706 has already fixed the global volume by fixing the gain of the main stereo channels to 1, this gives a physical meaning to the absolute target volume. The Leveling Unit 706 can determine one or more criteria for setting it to an appropriate level.

[0110] To do this, the leveling unit 706 calculates the desired relative target loudness level T i Of all the possible ways to arrive at , a specific item of data can be obtained to determine how the leveling unit 706 can specify the absolute scale, so that the leveling unit 706 can control the fraction of the total energy that results from the main stereo microphones versus the fraction that results from the spot microphones. In some implementations, the leveling unit 706 can set this fraction number as a user input parameter.

[0111] In some implementations, the leveling unit 706 can estimate this fraction by aiming for a given ratio between direct and reverberant energy, called the direct-to-reverberant ratio. For example, in an audio environment with strong reverberation (e.g., a church), the leveling unit 706 can apply a high level of relative spot microphone energy. Conversely, in an audio environment with low reverberation (e.g., an acoustically treated room), the leveling unit 706 can allow most of the energy to come from the dominant stereo microphone if it is in an optimal position. Thus, the leveling unit 706 can estimate the spot-to-dominant energy ratio R spotsAn input specifying θ can be obtained from a user or by automatic calculation. The leveling unit 706 can then determine the terms in the cost function using equation (6) below:

number

[0112] The leveling unit 706 can approximate this equation to simplify processing. The final source energy is correctly arrived at, i.e., in the following approximation:

number

number

[0113] In this approximation, these energies are g M , so the leveling unit 706 can apply few spot-to-full constraints before minimization. The leveling unit 706 sets the target energy

number

number

[0114] The leveling unit 706 then generates the appropriately scaled ^Ti Even if the leveling unit 706 uses R spots Even if you set =0, in that case

number

number

[0115] The leveling unit 706 is configured such that the cost function is (abbreviated as dB p [·]=20log 10 [·] and dB I [ ]=10log 10 (using [·])

number

[0116] N i -1The normalization factor in ensures that the absolute value of the first term can be compared across cases with different numbers of sources. The leveling unit 706 can use a cost function F to represent the mean squared error (e.g., in dB) that each source misses the target. In some implementations, the leveling unit 706 can derive the cost function to avoid approximations such as those above. The leveling unit 706 can include an additional cost term:

number

[0117] In various implementations, the Leveling Unit 706 can use algorithms that can provide better results if the Leveling Unit 706 has information about which sound sources are more important to reach a specified loudness target. For example, the input information can specify that the lead vocal should be 3 dB above the other instruments. This information can crucially determine the quality of the mix. If the other instruments miss the correct target by a few dB, the mix will not be judged as poor as if the lead vocal was below the target. To capture this aspect, the Leveling Unit 706 assigns for each sound source a set of importance weights w i imp The leveling unit 706 may define a cost function incorporating the importance weights as follows:

number

[0118] The leveling unit 706 is M To solve for , the cost function F can be minimized as described above. In some implementations, the leveling unit 706 assigns importance weights w according to whether the instrument is a lead instrument. i imp For example, the leveling unit 706 can set the importance weight w i imp can be set to 1 for non-lead instruments, and to a value between 2 and 5 for lead instruments.

[0119] Add a dedicated spot microphone In some circumstances, the algorithm tends to use very little energy from certain channel signal sources, e.g., dedicated spot microphones, because the level of the corresponding sound source can be achieved correctly using other microphones. This can occur in the case of leakage, such as when the spot microphone is omnidirectional (e.g., the built-in microphone of a smartphone). In general, when a dedicated spot microphone is used, the leveling unit 706 can be configured to obtain most of the energy of the corresponding instrument from such microphone.

[0120] From the rehearsal stage, the leveling unit 706 can define the degree of dedication that a given spot microphone has for a given sound source. The leveling unit 706 can set the degree of dedication to 1 if the spot microphone has little leakage from other sound sources. The leveling unit 706 can set the degree of dedication to 0 if the leakage from other sound sources is severe (e.g., above a threshold). Thus, for a sound source i with beam b(i), such a degree of dedication D(i) is

number

[0121] The leveling unit 706 can clamp D(i)∈[0,1]. In some implementations, the leveling unit 706 can set the values ​​for these parameters as follows: dBMaxRatio=3dB, dBMinRatio=-6dB. These settings imply that the dedication is 1 if the associated instrument is at least 3dB above the sum of all other instruments at that microphone, and 0 if it is -6dB or less.

[0122] The leveling unit 706 adds a new term N ded We can use D(i) to weight the:

number

number

[0123] The leveling unit 706 calculates the new term g by minimizing the cost function including this new term. M can be calculated.

[0124] Minimizing the best guess cost function In some implementations, these cost functions can be nonlinear. To minimize the nonlinear cost functions, the leveling unit 706 can take a guesswork approach. The leveling unit 706 can minimize all g M in 1 dB steps to find the combination that minimizes F. The leveling unit 706 can find the combination by starting from a best guess and taking steps away from the best guess through the range until the leveling unit 706 finds a minimum of the cost function.

[0125] To do so, the leveling unit 706 can perform a first guess. This can be obtained, for example, by ignoring leakage and diagonalizing E (assuming that the estimator 522 or the leveling unit 706 sorted the rows and columns, and the corresponding beams of the sources are on the diagonal). Then, for each i, only one beam contributes. That beam is labeled b(i). Thus,

number

number

number

[0126] If the leveling unit 706 is repeating the same beam for more than one instrument, the leveling unit 706 solves for g as follows:

number

[0127] A note on the sign of this numerator. The only fact that can be fully guaranteed is that the total target energy from all beams is equal to or greater than the energy from the primary microphone:

number

[0128] However, some individual terms in the sum may be negative for some i, i.e. for some sound sources i there is already enough loudness in the main stereo channels to reach the target. In such cases the leveling unit 706 may set the gain of the corresponding beam to 0. In some implementations the leveling unit 706 may search a range of possibilities, e.g. -15 dB.

[0129] Leveling Unit 706 is a system for setting various objectives. i Instead of expressing loudness in dB, the Leveling Unit 706 can express it in sones and convert it back to dB using the loudness model that the Leveling Unit 706 used for dB.

[0130] Automatic Panner Figure 9 is a flow chart illustrating an example process 900 for automatic panning. Process 900 may be performed by panner 708 of Figure 7. By performing process 900, panner 708 improves upon conventional panning techniques by automatically placing instruments in their correct positions on the soundstage.

[0131] The panner 708 may receive channel signals of the event 100 (902). The channel signals may be the output of the leveling unit 706. Each channel signal may correspond to a microphone. The panner 708 may receive reference audio data of sound sources in the event 100 (904). The reference audio data may be generated from a signal recorded in a rehearsal session. The panner 708 may calculate a total energy in the left channel and a total energy in the right channel as a contribution by each sound source based on the reference audio data (906). The panner 708 may calculate a left-right imbalance based on the total energies (908). The panner 708 may determine a cost function to minimize the imbalance. The panner 708 may calculate a natural pan of the sound source captured by the main microphone (910). The panner 708 may determine a cost function that maximizes the natural pan. The panner 708 may identify sound sources that cannot be panned (912). This may be based, for example, on an input that specifies the source as not pannable. The panner 708 may determine a cost function that respects sources that cannot be panned.

[0132] The panner 708 may determine a cost function with the pan angle as a variable for each channel signal (914). The cost function may have a first component corresponding to the imbalance, a second component corresponding to sound sources that can be panned, and a third component corresponding to sound sources that cannot be panned.

[0133] The Panner 708 can determine a pan position for each channel signal by minimizing a cost function (916). The pan position can be parameterized as a pan angle, a ratio between the left and right output channels, or a percentage for the left and right output channels. The Panner 708 can apply the pan positions to the channel signals to achieve the audio effect of placing a sound source between the left and right of a stereo sound stage for output to speakers.

[0134] In some implementations, the Panner 708 can perform audio panning based on the video data. The Panner 708 can determine the location of a particular audio source, such as a vocalist or instrument, using face tracking or instrument tracking in the video data. The Panner 708 can then determine a pan position for that audio source based on that location.

[0135] Once the leveling unit 706 determines the gain (g b ), the panner 708 can determine how to split each beam into left and right L / R in an energy-conserving manner. The panner 708 determines the pan angle θ b can be calculated:

number

[0136] Assuming that the panner 708 leaves the main stereo channels unchanged, the panner 708 may extend the index to M, where l L =r L = 1. The resulting mix as a function of the angle is:

number

[0137] Based on the reference audio data, the Panner 708 can calculate the total energy in the L / R channels due to each instrument:

number

[0138] One thing the panner 708 can impose is that the overall mix be balanced between L and R. Thus, the panner 708 imposes an LR imbalance cost function H LR-balance :

number

[0139] On the other hand, the panner 708 can be configured to respect the natural panning of the sound sources located in the event 100 from the point of view of the event 100. Natural panning is achieved by fully capturing the dominant stereo energy: iL , E iR Thus, the panner 708 can also impose:

number

[0140] In some implementations, the panner 708 can receive the desired position as an external input, rather than determining the position based on a natural pan obtained by analyzing the left and right channels. For example, the panner 708 can determine the natural position from an image or video. Additionally or alternatively, the panner 708 can determine the natural position from user input.

[0141] In addition, some sources should never be panned (e.g. lead vocals, bass, etc.). The Panner 708 can be configured to respect this as much as possible. These sources can be designated as non-pannable sources. The Panner 708 defines a set of pannable / unpannable sources, I P / I U Then, the panner 708 can generalize the above as follows:

number

number

[0142] The Panner 708 can control the pan amount, which indicates a tendency to pan the instrument wider as opposed to centering it in the soundstage. In some implementations, the Panner 708 can introduce another term. In some implementations, the Panner 708 can exaggerate the estimate from the dominant microphone. The Panner 708 can receive a parameter d∈[0,1], which indicates the divergence, as a preset or user input. The Panner 708 can perform the following transformation on the perceived dominant channel energy: The transformation introduces a transformation on the instrument angle:

number

[0143] Using d, the panner 708 uses the following pannable cost function:

number

[0144] Panner708's final cost function is:

number

[0145] FIG. 10 shows an example angle transformation for maximum distortion that can be performed by the panner 708 (d=1). The horizontal axis represents the original angle θ of one or more sound sources. The vertical axis represents the final angle θ of one or more sound sources. final The corner = 45 is the center pan.

[0146] <Congruent minimization> The Leveling Unit 706 may use fewer spot microphones than it can handle. For example, different configurations of input gains may compete, all leading to the same loudness. This may have a negative impact on the Panner 708, since the range of possibilities for the Panner 708 can be greatly reduced if only one or two spot microphones are used.

[0147] In some implementations, to reduce this uncertainty in the auto-level stage and favor configurations in which more spot microphones are used, the operation of the auto-level stage of the leveling unit 706 can be linked to the pan stage operation of the panner 708.

[0148] In such an implementation, the leveling and panning unit may combine the circuitry and functionality of the leveling unit 706 and the panner 708. The leveling and panning unit may receive reference audio data. The reference audio data may include representations of channel signals from a multiple channel signal source recorded in rehearsal of one or more sound sources. The leveling and panning unit may receive target level data. The target level data specifies a target level for each sound source. The leveling and panning unit may receive live audio data. The live audio data may include recorded or real-time signals from the one or more sound sources playing in the live event 100. The leveling unit may determine a joint cost function for leveling the live audio data and panning the live audio data based on the reference audio data. The joint cost function may have a first component for leveling the live audio data and a second component for panning the live audio data. The first component may be based on the target level data. The second component may be based on a first representation of imbalance between the left and right channels, a second representation of sources that can be panned among the sound sources, and a third representation of sources that cannot be panned among the sound sources. The leveling and panning unit may calculate a respective gain to be applied to each channel signal and a respective pan position for each channel signal by minimizing a joint cost function. The leveling and panning unit may apply the gain and pan position to a signal of the live audio data of the event to achieve an audio effect of leveling the sound sources in the live audio data and positioning the sound sources in the live audio data between the left and right of a stereo sound stage for output to a storage device or to a stereo sound reproduction system.

[0149] The joint cost function is shown below in equation (29), where some of the terms that appeared above have been renamed.

number

[0150] The cost function in the joint cost function is defined as follows:

number

[0151] Here, the automatic level process of the Leveling Unit 706 does not depend on the pan angle. It measures the overall loudness of the mono downmix. The automatic pan process of the Panner 708 depends on the beam gain g in addition to the pan angle. b Depends on.

[0152] Estimating Instrument RMS from Microphone Signals 11 is a flow chart illustrating an example process 1100 for estimating energy levels from microphone signals. The estimator 522 of FIGs. 5 and 7 may perform the process 1100 to measure the instrument RMS. The instrument RMS may be a root mean square representation of the energy levels of various sound sources.

[0153] The estimator 522 may receive 1102 reference audio data. The reference audio data may be audio data for i=1,...,N recorded during rehearsals. i The signal may include channel signals from m=1,...,M microphones at a sound source.

[0154] The estimator 522 may calculate 1104 a respective level (eg, loudness level, energy level, or both) for each instrument at each microphone based on the reference audio data.

[0155] The estimator 522 may determine a cost function based on the respective gains of each sound source (1108). In the cost function, the estimator 522 may give less weight to signals from key microphones than spot microphones. In the cost function, the estimator 522 may penalize estimating instrument loudness in the live data that is significantly higher than that represented in the reference audio data. In the cost function, the estimator 522 may scale the cost function by the cross-microphone average of the difference between the measured levels between the performance and the rehearsal.

[0156] The estimator 522 can determine 1110 a respective gain for each sound source by minimizing a cost function. The estimator 522 can provide the respective gains in the energy or loudness matrix to a processor (e.g., video editor 530) for processing the video signal, for example to identify which instruments are playing at a level above the threshold of other instruments and focus on that instrument or that instrument player. Further details and examples of the process 1100 are described below.

[0157] The audio scene of event 100 is recorded by m=1,……,M microphones and i=1,……,N iIn the rehearsal stage, each instrument is played separately. The estimator 522 estimates the loudness E of each instrument at each microphone. i,m , and converting the value into energy. In some implementations, the loudness index can be based on, for example, the R128 standard of the European Broadcasting Union (EBU), where we L / 10 Thus, in the rehearsal, the estimator 522 calculates the matrix e i,m can be calculated:

number

[0158] When the whole band plays together, the estimator 522 estimates the total loudness E m performance If the transfer functions from the instruments and microphones remain constant and equal to the rehearsal stage, and all instrument signals are statistically independent of each other, the following relationship holds:

number

number

[0159] In some implementations, the estimator 522 calculates the gain g iA cost function C can be used, which can be a function of .gamma. to ensure that the estimator 522 estimates the performance level such that the model is best satisfied in the least squares sense.

number

[0160] In some implementations, the estimator 522 can improve the results by giving less importance to the dominant stereo microphones, which are much less discriminatory than the spot microphones.

number

number

[0161] The problem of estimating energy levels can be underdetermined when there are fewer microphones than instruments. The uncertainty can correspond to obtaining the same overall loudness at each microphone by boosting the estimates of some instruments while attenuating others. To reduce this uncertainty, in some implementations, the estimator 522 can introduce a term that penalizes estimating instrument loudness significantly higher than that measured in rehearsal. One possible term is defined below.

number

[0162] For example, with α2=0.1 and n=6, there is essentially no penalty when the gain is lower than the rehearsal.

number

number

[0163] Thus, the estimator 522 may apply the following cost function:

number

[0164] In some implementations, the estimator 522 can measure in dB. Measuring everything in dB can provide better performance when estimating low levels.

number

[0165] In some implementations, the estimator 522 can apply an initial filtering stage before minimizing the cost function. In the initial stage, the estimator 522 can use the dedication degree D(i) of the source i to determine the sources whose given channel signal has small leakage from other instruments. For example, the estimator 522 can determine whether the dedication degree D(i) of the source i is higher than a given threshold. For each such source, the estimator 522 can obtain a corresponding gain by constraining the above cost function to include only the corresponding dedicated channel signal. For example, if the estimator 522 determines that the pair (^i,^m) for the instrument ^i and the dedicated microphone ^m satisfies the threshold, the estimator 522 can obtain the gain

number

number

[0166] Equation (40.1) allows the estimator 522 to perform the minimization using the following equation (40.2):

number

[0167] The estimator 522 can use the other cost functions above to determine each of the gain pairs, applying simplifications similar to those described with reference to equations (40.1) and (40.2) to include only one source-dedicated microphone pair at a time.

[0168] Having determined these gains in the early filtering stages, the estimator 522 reduces the problem of minimizing the cost function to determining only the gains of the instruments that do not have a dedicated channel signal. The estimator 522 can fix these gains to the gains found in the early filtering stages. The estimator 522 can then minimize the cost function with respect to the remaining gains.

[0169] Estimating Instrument RMS in Frequency Bands Estimating the signal RMS using the estimator 522 as described above can be extended in a frequency dependent manner to improve the estimation in cases where different instruments contribute to different parts of the overall frequency spectrum. FIG. 12 is a flow chart illustrating an example process 1200 for estimating the energy level in a frequency band. The estimator 522 can perform the operations of the process 1200.

[0170] The estimator 522 may receive 1202 reference audio data. The reference audio data may include audio signals recorded from a rehearsal in which sound sources and microphones are positioned in the same way as they will be during a live performance.

[0171] In the first stage, the estimator 522 estimates the rehearsal loudness E i,m,f rehearsal can be calculated (1204), where the frequency bands can be octave bands centered around frequencies f={32, 65, 125, 250, 500, 1000, 2000, 4000, 8000} obtained using standard filters following ANSI (American National Standards Institute) specifications.

[0172] In the next stage, the estimator 522 can calculate 1206 the total cost as a sum of costs per source using a cost function such as:

number

[0173] The estimator 522 may calculate (1208) a mass term in a frequency band:

number

[0174] The estimator 522 determines 1210 the gains for each frequency band by minimizing these costs. The estimator 522 can provide the gains to a processor for processing live data of the event 100. For example, the estimator 522 can provide the gains to the video editor 530 to identify a sound source that is playing above the level of other sound sources so that the video editor 530 can focus or zoom in on that sound source.

[0175] In some implementations, the estimator 522 can be biased towards estimating that the instrument is on or off. This can be done by modifying the second term of equation (43) to have a minimum at g=0,1. For example, the estimator 522 can be biased towards estimating that the instrument is on or off. i Into function f:

number

[0176] This function is symmetric under x←→-x in that it has only minima at x=0, x=±a, and a maximum with f=p only at x=a / √3. Thus, estimator 522 can use the value of p to control the value at the maximum (and therefore the size of the wall between the minimum at x=0,a). More generally,

number

number

[0177] Exemplary settings for the parameters in equation (45) are a=1, n=1, and p=5.

[0178] In some implementations, the estimator 522 may implement the following function:

number

[0179] In some implementations, the estimator 522 estimates the points (0,0), (x p ,y p ), (x a ,0), (x l ,y l ) symmetric sextic polynomial under x-parity:

number

[0180] Automated EQ in the loudness domain Figure 13 is a flow chart illustrating an example process 1300 for automatically equalizing individual sound sources. Process 1300 may be performed, for example, by source equalizer 710 of Figure 7. Source equalizer 710 is configured to apply equalization (EQ) to the levels of individual sound sources (rather than the overall stereo mix) for the specific purpose of cleaning, emphasizing, or de-emphasizing instruments in an automated manner.

[0181] The audio source equalizer 710 may receive 1302 audio data. The audio data may be live audio data or rehearsal data for the event 100. The audio data may include channel signals from an audio source.

[0182] The source equalizer 710 can map the respective signal for each sound source to an excitation in each frequency band (1304). The source equalizer 710 can sum the sounds from different sources in excitation space and sum the effects from different frequency bands in loudness space.

[0183] The source equalizer 710 then automatically equalizes one or more sources. The source equalizer 710 may generate a list of source-band pairs that map each source to each band. The source equalizer 710 may determine (1306) a respective need value for each source-band pair in the list. The need value may indicate the relative importance of the source represented in the pair to other sources and other frequency bands that are being equalized in that frequency band in the pair. The need value may be a product of the relative importance value and the masking level of the source by the other sources, or a mathematically expressible relationship that demonstrates that the need increases or decreases when either the relative importance or the masking level is increased or decreased, respectively.

[0184] The source equalizer 710 may compare the required values ​​to thresholds. If all required values ​​are less than the thresholds, the source equalizer 710 may terminate the process 1300.

[0185] Upon determining that the required value of a source-band pair exceeds the threshold, the source equalizer 710 may equalize 1308 for the source represented in that source-band pair to accentuate that pair. Equalization may include, for example, lowering other sources in that frequency band.

[0186] The source equalizer 710 may remove 1310 that source-band from the list of possible pairs to highlight and return to stage 1306 .

[0187] The source equalizer 710 may apply the process 1300 to the reference audio data only. The source equalizer 710 may then use fixed settings for the event 100. Alternatively or additionally, the source equalizer 710 may perform these operations and functions adaptively during the live event 100, optionally after using the reference audio data as a seed.

[0188] The source equalizer 710 uses an index i∈{1, . . . , N i} can be used (e.g., i=1 is bass, etc.). In some scenarios, all instruments to be mixed are well separated. The source equalizer 710 divides each instrument signal s i can be mapped to excitation in each frequency band b: s i →E(i,b) (49) where E(i,b) is the excitation for source i in band b. Frequency band b can be an ERB (equivalent rectangular bandwidth) frequency band.

[0189] This map can be expressed in terms of the Glasberg-Moore loudness model. By performing this mapping, the source equalizer 710 takes into account the head effects for frontal incidence (or in a diffuse field) and the inverse filter of the threshold-in quiet. Similarly, the source equalizer 710 can map from the excitation space to a specific loudness L[E(i,b)]: s i →E(i,b)→L[E(i,b)] (50) This models the compression applied by the basilar membrane. An example of such a function is L=α(E+1) above 1 kHz. 0.2 The term "specific loudness" can refer to loudness per frequency band, and "loudness" can refer to the sum over all frequency bands. The dependence on b as shown in these equations indicates specific loudness. Independence on b indicates summation.

[0190] The source equalizer 710 can sum the sounds from different sources in excitation space and sum the effects from different bands in loudness space:

number

[0191] The quantity of interest is the partial loudness of a signal in the presence of noise (or in the presence of some other signal). The source equalizer 710 is able to split all sources into one signal called pL with index i and all others with index i':

number

[0192] The maximum it can have is exactly the loudness of source i, pL(i,b) = L(i,b), where L(i,b) is the loudness of source i in band b. This value occurs when there is no masking between the sources, so the compressors act independently, i.e. L(ΣE) = ΣL(E). Masking can reduce this value.

[0193] The source equalizer 710 then automatically equalizes some sound sources. In some implementations, the equalization can avoid masking of some sound sources by other sound sources. Thus, the source equalizer 710 assumes that an earlier premix stage adjusted all sound sources to sound at a given target loudness. For example, all sound sources can have equal loudness except for the lead vocal, which the source equalizer 710 can leave at least 3 dB above all others.

[0194] The premix stage may only focus on the sound sources separately. However, when all sound sources play at the same time, masking may occur. The Source Equalizer 710 may perform an equalization operation to emphasize some sources, for example to help some sound sources stand out. A typical example is the bass. When the bass plays together with the organ or other wideband instruments, the Source Equalizer 710 may high pass such wideband instruments, leaving the bass more prominent at the low end. Conversely, in an organ solo audio event, the Source Equalizer 710 does not apply equalization for this reason. Thus, the problem is a cross-instrument problem.

[0195] One way for the Source Equalizer 710 to proceed is to detect which source or which band of which source has a greater need to be equalized. The Source Equalizer 710 can refer to this quantity as Need(i,b) for instrument i and band b. This need can depend on the following factors: i) how important that frequency band is to that source; and ii) how masked that source is by all other sources. The Source Equalizer 710 can use I(i,b), M(i,b) to quantify the importance and masking, respectively.

[0196] The importance of an instrument's frequency bands (unlike masking) depends only on that instrument. For example, bands at the low frequency end of a bass may be important to the bass, whereas bands at the low frequency end of an organ may be less important, since the organ is much more spread out in frequency. The source equalizer 710 can measure importance bounded by [0,1] as follows:

number

[0197] To measure masking by all other instruments, the Source Equalizer 710 designates all other sources as noise. The Source Equalizer 710 can use the fractional loudness for all others to get an indicator bounded in [0,1]:

number

[0198] Thus, the source equalizer 710 can implement the necessity function as follows:

number

[0199] The source equalizer 710 can simplify the notion of loudness for all sources other than i:

number

number

[0200] The source equalizer 710 may implement the following algorithm to achieve automatic equalization: We describe improvements in the next section: For convenience, the quantity Need(i,b) is simplified as N(i,b).

[0201] Stage 1: The Source Equalizer 710 can find the source and frequency band with the highest N(i,b). The Source Equalizer 710 can represent this as a source-band pair (^i,^b). For example (base, 3rd frequency band). The Source Equalizer 710 can compare this highest N(i,b) to a threshold t ∈ [0,1]. If N(i,b)>t, go to Stage 2, else stop (nothing else needs to be equalized).

[0202] Stage 2: The Source Equalizer 710 can equalize the remaining instruments to accentuate the selected pair. The Source Equalizer 710 can use i' to represent all sources other than ̂i. The Source Equalizer 710 can share responsibility between each source i' in a way proportional to the masking it causes for ̂i. The way to do that is by defining the gain reduction to each instrument band:

number

[0203] If g=1 then all gains are unity and each source-band pair reduces its excitation in proportion to how much it masks ^i. The masking that i' causes on ^i can be obtained from the same equation (54), but considering i' as noise and ^i as signal:

number

[0204] At the same time, the source equalizer 710 can also boost selected instrument-bands ^i, ^b:

number

[0205] Finally, the source equalizer 710 determines whether it is fully unmasked, i.e., M(^i,^b) for that source-band pair. <M threshold Solve for g by defining the target masking level to be, where M thresholdis the masking threshold. This is an implementation represented by an equation in one unknown (g). Equation (54) shows that the equation is nonlinear. The source equalizer 710 can solve for g by decreasing g in discrete steps, e.g., smaller than dB, until the bounds are met. The source equalizer 710 can invert the gain-to-loudness map, thus returning from the loudness domain to the linear domain.

[0206] In this context, the source equalizer 710 can impose a maximum level of acceptable equalization by setting a minimum value of g that is allowed, or for better control, by directly limiting the limits of the allowed values ​​of g(^i,^b) and g(i',^b).

[0207] Step 3: The source equalizer 710 can remove the pair (^i, ^b) from the list of possible pairs to be prominent. Return to step 1.

[0208] The above algorithm selects candidate pairs (^i,^b) by only examining the pair with the greatest need. This algorithm does not guarantee that the benefit is global across all instruments. One global approach would be to mimic spatial coding: find the pair that minimizes the global need when improved, and iterate.

[0209] For example, the source equalizer 710 may define the global need for the mix to be equalized as follows:

number

[0210] The source equalizer 710 then performs the following operations: First, the source equalizer 710 can calculate an initial global need Need(global). Second, the source equalizer 710 takes all possible pairs (i, b) or selects some pairs (e.g. 10 pairs) with higher N(i, b). For each pair, the source equalizer 710 designates it as a candidate to be improved and runs the above algorithm to find the gain g(i, b) to be applied to boost it and attenuate the others. For each pair thus considered, the source equalizer 710 can recalculate a new global need. Third, the source equalizer 710 selects the pair that minimizes the global need and applies that improvement gain. The source equalizer 710 then replaces Need(global) by its new value and returns to the first stage.

[0211] The source equalizer 710 may terminate the above iterative iterations if any of the following occurs: 1. Need(global) is already below a given threshold; 2. No choice of (i,b) led to a decrease in Need(global); or 3. The source equalizer 710 has iterated more than a given maximum number of times.

[0212] 14A is a diagram showing a three-instrument mix to be equalized. The horizontal axis represents frequency f. The three-instrument mix includes bass, organ, and other instruments. The vertical axis represents energy. As shown, the energy of the bass is concentrated in the lower ERB band. Thus, the source equalizer 710 can determine that the bass needs more equalization in the lower ERB band compared to the organ and other instruments.

[0213] 14B shows the gain in automatic equalization. Starting with g=1 and decreasing, we increase the gain of the bass in the lower ERB bands while attenuating all other instruments in the lower ERB bands. The organ is attenuated more because it masks the bass better than "another" instruments.

[0214] Segmentation based on novelty 15 is a flow chart illustrating an example process 1500 for segmenting a video based on novelty buildup in audio data. Process 1500 may be performed by segmentation unit 714 of FIG. 7. In some implementations, segmentation unit 714 may be implemented by video editor 530 of FIG.

[0215] The segment splitting unit 714 may receive an audio signal (1502). The segment splitting unit 714 may build a novelty index for the audio signal over time (1504). The segment splitting unit 714 may identify peaks in the audio signal that are above a threshold. The segment splitting unit 714 may determine a segment length based on an average cut length (1506). The cut length may be input, a preset value, or derived from past cuts (e.g., by averaging the past X cuts). The segment splitting unit 714 may determine a sum of novelty indices since the last cut (1508). The sum may be an integral of the novelty index over time.

[0216] Upon determining that the sum is higher than the novelty threshold, the segmentation unit 714 may determine (1510) a random time for the next cut, where the randomness of the time to the next cut averages out to the average segment length. The segmentation unit 714 may cut (1512) the audio signal, or a corresponding video signal synchronized with the audio signal, into a new segment at the random time and begin summing the novelty index for the next cut. The segmentation unit 714 may provide the new segment to a user device for streaming or downloading and for playback on a loudspeaker. Further details and examples of the process 1500 are described below.

[0217] Video editing can be based on novelty. Novelty can indicate points where the audio or video changes significantly. The segmentation unit 714 can build an index that measures novelty, called a novelty index. The segmentation unit 714 can build the novelty index by comparing a set of extracted features across different segments of the audio recording. The segmentation unit 714 can build a similarity index and convolve it with a checkerboard kernel to extract novelty.

[0218] The segmentation unit 714 can calculate a novelty index over time. It can first select peaks that are above a certain threshold. The size of the segments used to extract features can determine the scale at which novelty works. Short segments allow distinguishing individual notes. Long segments allow distinguishing coarser concepts, e.g. intro from chorus. The threshold at which a segment is considered novel can affect the frequency of cuts. The threshold can thus be set as a function of the desired average cut length, which itself can be set as a function of tempo. The segmentation unit 714 can thus perform the operations as follows: ·Get tempo → Set average cut length → Set threshold.

[0219] The segmentation unit 714 can properly handle song sections with crescendos. Such sections are characterized by a prolonged smooth increase in the novelty index. Such a smooth increase does not lead to a significant peak, and therefore to the absence of cuts for a very long period of time. The segmentation unit 714 can include a build-up processing module that quantifies the need to have cuts over a sustained period of time, independent of the fact that peaks may occur. The segmentation unit 714 can also ... last We can characterize this need by the integral of the novelty index:

number

[0220] N(t) is the threshold N thrIf it determines that the cut is above T seconds, the segmentation unit 714 starts a random drawing with an adjusted probability that, on average, there will be a cut within the next T seconds. thr can be assigned a value that is considered to be of great need, e.g., a sustained value of novelty=0.6 for at least 3 seconds. Similarly, the value of T can be linked to the required average cut length as described above.

[0221] FIG. 16 illustrates an exemplary novelty build-up process. The effect of novelty build-up processing on a song with a crescendo is shown in FIG. 16. The X-axis represents time in seconds. The Y-axis represents the value of the integral. There is a long crescendo between 150 and 180 seconds, as shown by curve 1602. Curve 1602 shows the novelty integral without build-up post-processing. If only the index peak were used, no event would be detected in this segment. Curve 1604 shows the integral after build-up post-processing, revealing both the presence of a new cut and the controlled randomness in its occurrence. The occurrence can be based on a hard threshold, or more preferably, on probability.

[0222] <Synchronization> Figure 17 shows an example process 1700 for synchronizing audio signals from multiple microphones. Process 1700 may be performed by synchronizer 502 of Figure 5. Synchronizer 502 performing process 1700 improves upon conventional synchronization techniques in which the synchronizer may synchronize audio signals based solely on the audio signals.

[0223] The synchronizer 502 can synchronize the microphones used in the audio scene simply by analyzing the audio. In some implementations, the synchronizer 502 can synchronize all microphones to a main stereo microphone using various correlation determination techniques, such as cross-correlation algorithms.

[0224] The synchronizer 502 may receive (1702) audio signals. The audio signals may be channel signals from a microphone. The synchronizer 502 may calculate (1704) respective quality values ​​of correlations between each pair of the audio signals. The synchronizer 502 may assign (1706) the quality values ​​in a map vector.

[0225] The synchronizer 502 may iteratively determine a set of delays and insert the delays into the map vector (1708) as follows: The synchronizer 502 may identify the signal pair in the map vector with the highest quality value. The synchronizer 502 may align the audio signals in the pair, downmix the pair to a mono signal, and add the alignment delay to the map vector. The synchronizer 502 may replace the first audio signal in the pair with the downmixed mono signal and remove the second audio signal from the list of indices for the maximum. The synchronizer 502 may keep the downmixed mono signal fixed and recalculate the quality value. The synchronizer 502 may iteratively iterate again from the identifying step above until only one signal is left.

[0226] Upon completing the successive iterations, the synchronizer 502 may synchronize 1710 the audio signal with each delay in the map vector according to the order of the delays inserted in the map vector. The synchronizer 502 may then submit the synchronized signal to other components (e.g., the source separator 504 of FIG. 5) for further processing and streaming. Further details and examples of the process 1700 are described below.

[0227] In some implementations, the synchronizer 502 runs a global synchronization algorithm, giving greater importance to delays calculated from cross-correlations with strong peaks. The synchronizer 502 can label the set of microphones in the scene as m=1,...,M. If one of the microphones is stereo, it has previously been polarity checked and downmixed to mono. The synchronizer 502 calculates their correlation over time C as follows: m,m' (t) can be determined.

number

[0228] For each pair, the synchronizer 502 respectively detects the higher and lower correlations C m,m' max / min Leading to max , t min The figure of merit (or quality of correlation, Q), which describes how good the correlation is, is

number

[0229] The synchronizer 502 can execute a recursive algorithm as follows: First, the synchronizer 502 can initialize an empty map vector Map. The vector will have (M-1) entries. The synchronizer 502 can then only iterate over the diagonal of Q (due to the symmetry of Q), and hence m1 <m2であるQ m1,m2 Think about it. 1. The Biggest Q m1,m2 Find the pair m1, m2 with 2.s m2s m1 align to (t m1,m2 ) to the Map. 3.s m1 Replace m with this downmix. Remove m2 from the list of indices to scan for the maximum of Q. 4. Fix m1 and Q for all m m,m1 Recalculate. 5. Repeat step 1 until only one microphone remains.

[0230] The synchronizer 502 has M-1 delays t m,m' where the second index appears only once for all microphones except the first one (which is typically the main stereo downmix). To reconstruct the delay required for microphone m to be synchronized with, say, the first microphone, the synchronizer 502 can follow the chain leading to the first microphone:

number

[0231] In some implementations, the synchronizer 502 can improve computation speed by avoiding recalculating all correlations after each downmix to mono. m,m' This means that the synchronizer 502 is calculating all c m,m' (T)= m (t)s m' After aligning and downmixing m' to m, the synchronizer 502 calculates the new signal:

number

number

number

number

[0232] Noise Reduction FIG. 24 is a flow chart illustrating an exemplary process 2400 of noise reduction. The process 2400 can be performed by the noise reduction unit 711 of FIG. 7. The noise reduction unit 711 can apply noise reduction to each channel signal. Advantages of the disclosed technique include, for example, that by applying different gains to each channel individually, the noise reduction unit 711 can reduce noise when a particular channel has a high enough audio level to mask the channel signal from other channels. Furthermore, the channel signals can come from different channel signal sources (e.g., microphones with different models, patterns) that may be located at separate points in the venue (e.g., more than two to three meters apart).

[0233] The noise reduction unit 711 may receive 2402 reference audio data. The reference audio data may include channel signals recorded during silent periods of a rehearsal session. The silent periods may be periods (e.g., X seconds) during which no musical instruments are playing.

[0234] The noise reduction unit 711 may include a noise estimator component. The noise estimator may estimate (2404) a respective noise level in each channel signal in the reference audio data. The noise estimator may designate the estimated noise level as a noise floor. The estimation of the respective noise level in each channel signal in the reference audio data may be performed across multiple frequency bands, referred to as frequency bins.

[0235] The noise reduction unit 711 may receive 2406 live performance data, which includes channel signals recorded during the event 100 in which one or more instruments are played that were silent during the rehearsal session.

[0236] The noise reduction unit 711 may include a noise reducer component. The noise reducer may individually reduce (2408) the respective suppression gains in each channel signal in the live performance data. The noise reducer may apply the respective suppression gains in each channel signal in the live performance data upon determining that, for each channel signal in the live performance data, a difference between the noise level in that channel signal in the live performance data and the estimated noise level meets a threshold. Reducing the respective noise levels in each channel signal in the live performance data may be performed in each frequency bin.

[0237] After reducing the noise level, the noise reduction unit 711 can provide 2410 the channel signal to a downstream device for further processing, storage, or delivery to one or more end user devices. The downstream device can be, for example, the delivery front end 508 of FIG. 5 or the mastering unit 704 of FIG. 7.

[0238] The estimation (2404) and reduction (2408) stages can be performed according to noise reduction parameters, including the threshold, slope, attack time, decay time, and octave size. Exemplary values ​​for the parameters are: threshold 10 dB; slope 20 dB per dB; attack time equal to decay time, 50 milliseconds (ms). Further details and examples of the noise reduction operation are provided below.

[0239] During the estimation (2404) stage, the noise estimator may perform the following operations individually for each channel signal in the reference audio data: The noise estimator may segment the channel signals into buffers of X samples (e.g., 2049 samples). The buffers may have half-length overlap. The noise estimator may apply a square root of a discrete window function (e.g., a Hann window) to each buffer. The noise estimator may apply a discrete Fourier transform. The noise estimator may calculate the noise level using equation (68) below: n(f)=10*log10(|·| 2 ) (68) where n(f) is the noise level for a particular frequency bin f.

[0240] The noise estimator can determine the noise floor by averaging the noise levels across the buffers using the following equation (69): n estimate (f)=<n(f)> buffers (69) where n estimate where (f) is the noise level designated as the noise floor for frequency bin f and < > is the average. As a result, the noise estimator estimates the number n for all frequency bins for all channel signals. estimate (f) can be determined.

[0241] During the noise reduction (2408) stage, the noise reducer may suppress the noise level in each channel signal in the live performance data of event 100 by performing the following operations on each channel signal individually: The noise reducer may segment the channel signals in the live performance data into buffers of X samples (e.g., 2049 samples). The buffers may have half-length overlap. The noise reducer may apply a square root of a discrete window function (e.g., a Hann window) to each buffer. The noise reducer may apply a discrete Fourier transform. The noise reducer may calculate the noise level using equation (68) above.

[0242] The noise reducer can calculate the difference between the noise level n(f) in the live performance data and the noise floor using the following equation (70): d(f)=n(f)-n estimate (f) (70) where d(f) is the difference.

[0243] The noise reducer can then apply a suppression gain to the channel signals in the live performance data in an expander mode. Applying the suppression gain in the expander mode can include determining whether the difference d(f) is less than a threshold. If determining that the difference d(f) is less than the threshold, the noise reducer can apply a gain that suppresses a number of dB per dB difference according to a slope parameter.

[0244] The noise reducer can smooth all the suppression gains over frequency bins or over a given bandwidth specified in the octave size parameter. The noise reducer can smooth all the suppression gains over time using the attack time and decay time parameters. The noise reducer can apply an inverse discrete Fourier transform and again apply the square root of a discrete window function. The noise reducer can then overlap and add the results.

[0245] 18 shows an exemplary sequence for synchronizing five microphones. First, synchronizer 502 aligns the signal from microphone 3 with the signal from microphone 2. Synchronizer 502 adds a delay t 23 and determining the delay t 23 Synchronizer 502 can add to the list. Synchronizer 502 downmixes the aligned signals to a mono signal. Synchronizer 502 then replaces the signal from microphone 2 with the mono signal. Synchronizer 502 can continue the process by aligning the mono signal with the signal from microphone 4, and then aligning the signal from microphone 1 with microphone 5. Finally, synchronizer 502 aligns the signal from microphone 1 with the signal from microphone 2. Synchronizer 502 eventually creates a list {t 23 ,t 24 ,t 15 ,t 12} can be obtained. In this case, t2=t 12 , t3=t 23 +t 12 , t4=t 24 +t 12 , t5=t 15 It is.

[0246] Video Editing 19A and 19B show an exemplary user interface for displaying the results of the automated video editing. The user interface can be presented on a display surface of a user device, such as video device 414. The features described in FIG. 19A and 19B can be implemented by video editor 530 (of FIG. 5).

[0247] FIG. 19A illustrates a user interface displaying a first video scene of the event 100 (of FIG. 1). In the illustrated example, a band is playing at the event 100. A video camera captures live video of the band playing. Each sound source in the band, e.g., vocalist 192 and guitar 194, as well as other sound sources, are playing at a similar level. A video editor 530 can receive the live video and audio data of the band playing. The live video can include real-time video of the event 100 or pre-stored video. The video editor 530 can determine from the audio data that the difference between the energy (or loudness) levels of each sound source is less than a threshold. In response, the video editor 530 can determine that the entire video scene 196 of the event 100 can be presented. The video editor 530 can then provide the entire video scene 196 in the live video for streaming. The video device 414 can receive the entire video scene 196 and present it for display.

[0248] FIG. 19B illustrates a user interface displaying a second video scene of the event 100 (of FIG. 1). During the live play described with reference to FIG. 19A, over a period of time, the video editor 530 can determine from the audio data that one or more sound sources are playing at a significantly higher level than the other instruments. For example, the video editor 530 can determine that the loudness or energy level of the vocalist 192 and the guitar 194 is higher than the loudness or energy level of the other instruments by more than a threshold level. In response, the video editor 530 can determine a pan angle of the one or more sound sources and focus or zoom in on a portion of the video data to obtain a partial video scene 198. In the illustrated example, the video editor 520 has focused and zoomed in on the position of the vocalist 192 and the guitar 194. The video editor 530 can then provide for streaming the partial video scene 198 including the vocalist 192 and the guitar 194 in the live video. The video device 414 can receive the partial video scene 198 including the vocalist 192 and the guitar 194. The video device 414 can present the partial video scene 198 for display.

[0249] 20 is a flow chart of an example process 200 for automated video editing. Process 2000 may be performed by video editor 530 (of FIG. 5), which is a component of a server configured to receive a live video recording of event 100.

[0250] The video editor 530 may receive (2002) video data of an event 100 (of FIG. 1) and audio data of the event 100. The video data and audio data may be live data. The live data may be real-time data or pre-stored data. The video data may include images of sound sources located at different positions in the event 100. The audio data may include the energy or loudness level of the sound sources and the pan angle of the sound sources.

[0251] The video editor 530 may determine from the audio data that a particular audio source is a dominant audio source 2004. For example, the video editor 530 may determine that a signal of a audio source represented in the audio data indicates that the audio source is playing at a volume level that is above a certain threshold amount relative to the volume levels of other audio sources represented in the audio data.

[0252] The video editor 530 may determine a location of the sound source in the video data (2006). In some implementations, the video editor 530 may determine the location based on a pan angle of the sound source in the audio data. For example, the video editor 530 may determine an angular width across a scene in the video data and determine a location across the scene that corresponds to an angle that corresponds to the pan angle of the sound source. The video editor 530 may determine a pan position of the sound source based on the audio data. The video editor 530 may designate the pan position of the sound source as the location of the sound source in the video data. In some implementations, the video editor 530 may determine the location based on the video data, for example by using face tracking or instrument tracking.

[0253] The video editor 530 may determine a portion of the live video data that corresponds to the location of the sound source 2008. For example, the video editor 530 may zoom in on a portion of the live video data according to a pan angle of the sound source.

[0254] The video editor 530 can synchronize and provide the audio data and the portion of the live video data for streaming to a storage device or end user device (2010). As a result, for example, when a vocalist or guitarist is performing a solo, the live video playback on the end user device can automatically zoom in on the vocalist or guitarist without the intervention and control of a camera operator.

[0255] Additionally, in some implementations, the video editor 530 can receive inputs identifying the locations of various sound sources in the event 100. For example, the video editor 530 can include or be coupled to a client-side application having a user interface. The user interface can receive one or more touch inputs on a still image or video of the event 100 input. Each touch input can associate a location with a sound source. For example, a user can designate a guitar player in a still image or video as "guitar" by touching the guitar player in the still image or video. During a rehearsal, a user can designate "guitar" and then record a section of the guitar playing. Thus, a sound labeled as "guitar" can be associated with a location in the still image or video.

[0256] When the event 100 is in progress, the video editor 530 may receive the live video recording as well as Ns signals for the Ns audio sources from the source separator 504. The video editor 530 may identify one or more dominant signals from the multiple signals. For example, the video editor 530 may determine that a signal from a particular audio source (e.g., a vocalist) is X dB louder than each of the other signals, where X is a threshold number. In response, the video editor 530 may identify a label (e.g., "vocalist") and identify a location in the live video recording that corresponds to the label. The video editor 530 may focus on the location, for example, by clipping a portion of the original video recording or zooming in on the portion of the original video recording that corresponds to the location. For example, if the original video recording is in 4K resolution, the video editor 530 may clip a 720p resolution video that corresponds to the location. The video editor 530 can provide the clipped video to the delivery front end 508 for streaming to end user devices.

[0257] Rehearsal-Based Video Processing 25 is a block diagram illustrating an example technique for video editing based on rehearsal data. An example server system 2502 is configured to provide editing decisions for live video data based on the rehearsal video data. The server system 2502 can include one or more processors.

[0258] The server system 2502 is configured to automatically edit live data 2504, e.g., live audio and video of a musical performance or live audio and video of any event, based on the rehearsal video data and rehearsal audio data. The live data includes M video signals 2506 of the performance captured by M video capturing devices, e.g., one or more video cameras. The audio data 2508 includes N audio signals from N audio capturing devices, e.g., one or more microphones. The number and location of the audio capturing devices can be arbitrary. Thus, the input gain of each of the audio capturing devices may be unknown. Due to the placement of the audio capturing devices, the level of the audio signal may not directly correlate to the natural or perceived level at which the performer is playing.

[0259] The server system 2502 can determine an approximation of which performers are playing at what level based on the live data 2504 and the rehearsal data 2510. Each performer can be an instrument, a person playing an instrument, a person performing as a vocalist, or a person otherwise operating a device that generates electronic or physical sound signals. As indicated above, the instruments, vocalists, and devices are referred to as sources. For example, in the live data 2504, the feed corresponding to a first source of a first performer (e.g., bass) may be lower than a second source of a second performer (e.g., guitar), even though in the actual performance the first source plays much louder than the second instrument. This discrepancy may be caused by the recording configuration. The various input stages involved in the chain of each source and the physical distance between the source and the audio capture device may be different.

[0260] Typically, a human operator (e.g., a sound engineer, cameraman, or video director) uses knowledge of who is performing at what level to determine how to edit the video. The server system 2502 can derive that knowledge from the rehearsal data 2510 and apply one or more editing rules that specify user preferences, e.g., artistic preferences, to perform an edit that simulates the editing of a human operator.

[0261] During the rehearsal phase, the server system 2502 can use the rehearsal data 2510 to determine where each performer is located in the camera feed. The server system 2502 then generates a map between the source sounds and the performers, without requiring the performers or operators to manually enter the mapping.

[0262] In the rehearsal phase, the band positions the sound sources at various locations on the stage in the same layout as in the live performance. One or more audio capture devices and one or more video capture devices are also positioned in the rehearsal in the same layout as in the live performance. Each audio capture device can be a pneumatic (air pressure gradient) microphone, a direct input feed (e.g., from an electronic keyboard) or a device that captures in the digital domain signal generated by a digital sound source (e.g., a laptop running music production software). At least one video capture device is a video camera positioned so that all sound sources and performers in the band can be captured in a single video frame. The server system 2502 can use the audio and video recordings of the rehearsal as rehearsal data 2510 to configure parameters for editing the live data 2504.

[0263] The rehearsal data 2510 includes rehearsal audio data 2512 and rehearsal video data 2514. The analysis module 2516 of the server system 2502 relates the loudness range of the sound sources to the digital loudness range that will be present in the final digital stream. In this way, the analysis module 2516 calibrates the multiple levels of steps involved between the capture of the signal and the final digital representation. In some implementations, the analysis module 2516 determines an average digital range for each of the sound sources captured by each of the audio capture devices. The average can be a weighted average between the EBU loudness levels between soft play at a low level and loud play at a high level.

[0264] The analysis module 2516 can analyze the rehearsal video data 2514 to determine where each performer is located in the video frame. The analysis module 2516 can do this using human detection, face detection algorithms, torso detection algorithms, pre-filtering with background subtraction, and any combination of the above and other object recognition algorithms. Some exemplary algorithms include principal component analysis (PCA), linear discriminant analysis (LDA), local binary patterns (LBP), facial trait code (FTC), ensemble voting algorithm (EVA), deep learning network (DLN), and the like.

[0265] In some implementations, the analysis module 2516 includes a sound source detector. The sound source detector is configured to analyze the rehearsal audio data 2512 to identify each individual sound source and apply media intelligence to determine what type of sound source it is at a low level (e.g., bass, pianistic, or vocal), at a high level (e.g., harmonic, percussion), and both. In some implementations, one or more musical instrument recognition (MIR) processes running in the analysis module 2516 can obtain a global descriptor of the event. The global descriptor indicates, for example, whether the genre of the musical piece being played is rock, classical, jazz, etc. The analysis module 2516 can provide the sound source type and the global descriptor to an automatic video editing engine (AVEE) 2518 for editing the live data 2504.

[0266] The analysis module 2516 associates each performer detected by the analysis module 2516 with a respective sound of the sound source. For example, the analysis module 2516 can map a face recognized by the analysis module 2516 with a particular sound, such as a guitar sound. In some implementations, the analysis module 2516 determines the mapping by ordering the sounds and faces. For example, the rehearsal audio data 2512 can include sound sources playing in sequence, such as from left to right as viewed from the video. The analysis module 2516 then associates the leftmost face detected with the first sound source that played in the rehearsal. In another example, the data is collected by direct human input via a customized graphical interface, such as by showing still frames of the band captured by one of the video devices and prompting the user to tap on each performer and select from a pre-populated menu which sound source he is playing.

[0267] After the band finishes rehearsing, the band may begin a live performance. The server system 2502 captures live data 2504 in the same manner as during the rehearsal. The sound sources, e.g., performers, can be located in approximately the same positions in the live performance and in the rehearsal. The audio and video capture devices are placed in the same positions as in the rehearsal. The server system 2502 provides the live audio data 2508 to the estimation module 2520 and the feature extraction module 2522. The estimation module 2520 is configured to determine the loudness of each sound source or each group of sound sources at a given moment. The output of the estimation module 2520 can include the sound level of each sound source or group of sound sources, e.g., in dB relative to the loudness played during the rehearsal, e.g., X dB from low, high, or average. Using the loudness as a reference during rehearsal eliminates ambiguities associated with potentially different level stages used during analog to digital conversion of each sound source.

[0268] The feature extraction module 2522 is configured to obtain time-varying features of the live audio data 2508, for example by using the MIR algorithm. The feature extraction module 2522 can perform operations including, for example, beat detection, including downbeat detection, calculation of novelty index, tempo, harmonicity, etc.

[0269] The server system 2502 can provide the live video data 2506 to an adaptive tracking module 2524. The adaptive tracking module 2524 is configured to perform adaptive face tracking, performer tracking, or other object tracking. In this manner, the adaptive tracking module 2524 takes into account performers that may leave the stage and therefore should not be in focus. The adaptive tracking module 2524 is also configured to track performers that move significantly from their original positions, for example, as a singer walks and dances on stage.

[0270] The analysis module 2516, the estimation module 2520, and the feature extraction module 2522 provide outputs to the AVEE 2518. The AVEE 2518 is a component of the system 2502 configured to perform operations including framing performers. A typical face detection algorithm may identify where each person's face is located, but not how to frame the face for zooming and cropping. The AVEE 2518 uses the respective size and respective position of each face to derive the respective size and respective position of a corresponding sub-frame of the original high definition video frame that provides a focused view of the corresponding performer. The sub-frames can be lower resolution (e.g., 720p) frames that the AVEE 2518 or a video capture device crops from a higher resolution (e.g., 4K) frame. The AVEE 2518 can present the sub-frames as images of the event. AVEE 2518 can determine frame size and location in size-based, position-based, and saliency-based cut decisions.

[0271] In size-based cut decisions, AVEE 2518 determines the size of the subframe using a facial proportion framing algorithm, where AVEE 2518 determines the size of the subframe to be proportional to the perceived face of the performer. For example, AVEE 2518 may determine that the height of the performer subframe is X (e.g., 5) times the diameter of the face. AVEE 2518 may determine that the width of the subframe is a multiple of the height that achieves a pre-specified aspect ratio. Similarly, AVEE 2518 may determine that the width of the performer subframe is Y (e.g., 8) times the diameter of the face. AVEE 2518 may determine that the height of the subframe is a multiple of the weight that achieves the aspect ratio. AVEE 2518 may determine that the face is positioned in the subframe horizontally centered and 1 / 3 down from the top of the subframe.

[0272] Alternatively or additionally, in some implementations, the AVEE 2518 determines the size of the sub-frames using a hand proportional algorithm, where the AVEE 2518 determines the size of the sub-frames to be proportional to the recognized hand or recognized hands of the performer. Alternatively or additionally, in some implementations, the AVEE 2518 determines the size of the sub-frames using a sound source proportional algorithm, where the AVEE 2518 determines the size of the sub-frames to be proportional to the recognized music source or recognized sound source or other area of ​​interest(s).

[0273] For position-based cut decisions, the AVEE 2518 can use motion tracking to determine the position of the sub-frame in the high resolution frame. For example, when the adaptive tracking module 2524 indicates that a performer is moving across the stage and provides a path of movement, the AVEE 2518 can follow the performer, identified by their face, and move the focus view sub-frame along the path.

[0274] In a saliency-based cut decision, AVEE 2518 places the subframe in the salient performer or in a group of salient performers. AVEE 2518 can determine the salience of the performer based on various conditions from the live audio data 2508. For example, from the output of the estimation module 2520 and the feature extraction module 2522, AVEE 2518 can determine the likelihood that the performer is salient at a particular moment in the performance. AVEE 2518 can select the performer in the next video cut based on the likelihood. The higher the likelihood, the higher the probability of being selected for the next cut. The likelihood that AVEE 2518 selects a subframe covering the performer is positively correlated to the likelihood that the performer is salient. For example, the higher the likelihood that the performer is salient, the more likely AVEE 2518 selects a subframe covering the performer. AVEE 2518 can determine the likelihood that the performer is salient based on audio traits. The audio features may include, for example, features such as the energy (e.g., RMS energy) of each audio signal of the corresponding performer, the RMS energy delta (increase or decrease) compared to the last N seconds of the performance, note onset frequency, tempo changes, etc. Additionally or alternatively, the AVEE 2518 may use various video features to determine salience. Video features may include, for example, movement within subframe boundaries.

[0275] The AVEE 2518 can generate a video edit that matches the pace and flow of the music. For example, the AVEE 2518 can determine cuts in such a way that the average frequency of cuts correlates with the tempo of the music as estimated by the feature extraction module 2522. The AVEE 2518 can determine the precise timing of each particular cut by aligning the cuts with changes in novelty of the live audio data 2508 that are above a given threshold. Optionally, the threshold is related to the tempo of the music. A faster tempo corresponds to a lower threshold and thus a higher frequency of cuts. The changes can include, for example, a change in overall loudness or timbre or one or more performers starting or stopping to play. The AVEE 2518 can time the cuts based on an evaluation of the musical structure of the performance. For example, the AVEE 2518 can time-align the cuts with bars or phrases of the music.

[0276] The AVEE 2518 may determine the selection of subframes to cut based on a performance metric that includes outputs from the analysis module 2516, the estimation module 2520, the feature extraction module 2522, and optionally the adaptive tracking module 2524. The performance metric may include a respective saliency metric for each performer, a respective designation of subframes for each performer, and size-based, position-based, and saliency-based cut decisions as described above. The AVEE 2518 may determine the selection of which subframes to cut using an exemplary process as follows:

[0277] The AVEE 2518 can detect the next peak in the novelty index. The AVEE can define a peak as a maximum loudness followed by a predefined and / or configurable threshold time, a decay above a threshold level, or both.

[0278] The AVEE 2518 may determine the time since the last cut and the time since a full frame shot showing all performers. If the AVEE 2518 determines that the time since the last cut is less than a predetermined and / or configurable minimum cut length, the AVEE 2518 may return to the first stage of detecting the next peak in the novelty index. If the AVEE 2518 determines that the time since a full frame shot exceeds a threshold, the AVEE 2518 may cut to full frames. The AVEE 2518 may define the threshold using a number of cuts or a duration derived from the tempo.

[0279] The AVEE 2518 may remove one or more performers from selection if their subframes are shown for a time that exceeds a threshold time. The AVEE 2518 may determine this threshold time based on tempo. For example, a faster tempo may correspond to a shorter threshold time.

[0280] The AVEE 2518 can boost the prominence of a performer designated as having a lead role to match the maximum prominence among all performers. The AVEE 2518 can designate one or more performers as having the lead role based on input received from a user interface.

[0281] The AVEE 2518 may build a list of performers whose salience values ​​are within X (e.g., 3) dB of the maximum salience among all performers. The AVEE 2518 may add an additional entry for the lead performer to this list. The AVEE 2518 may determine the number of additional entries to add to the list. For example, the number of additional entries may correlate to the total number of performers. The AVEE 2518 may randomly select a performer from the manner in which performers are selected as described above.

[0282] Based on the video editing decisions, the AVEE 2518 can edit the live video data 2506 in real-time, e.g., while the performance is in progress. The AVEE 2518 can provide the edited video data for storage or for streaming to one or more user devices. In the case of streaming, the AVEE 2518 can use a look-ahead time in the AVEE 2518 for buffering to perform the processing as described above. The look-ahead time can be pre-configured to be X seconds, e.g., more than 1 second, 5-10 seconds, etc. The AVEE 2518 can determine the look-ahead time based on the amount of buffering required at the cloud service application receiving the stream. In the offline case where the content is stored rather than streamed, the AVEE 2518 can set the look-ahead time to infinity or any period of time large enough to cover the entire performance or song.

[0283] For convenience, tracking is described with reference to a performer. In various implementations, tracking need not be limited to a performer. For example, it is possible to track an instrument (e.g., a guitar) or a part of an instrument (e.g., a guitar neck) or a part of a performer (a piano player's hands). AVEE 2518 can designate these areas as potential candidates that should be focused on and framed.

[0284] In Figure 25, the analysis module 2516, the AVEE 2518, the estimation module 2520, the feature extraction module 2522, and the adaptive tracking module 2524 are shown as separate modules for convenience. In various implementations, these modules may be combined or sub-divided. For example, in some implementations, the functionality of the analysis module 2516, the estimation module 2520, the feature extraction module 2522, and the adaptive tracking module 2524 may be implemented by the AVEE 2518. In some implementations, the AVEE makes video editing decisions and provides the decisions as instructions to one or more video capture devices, which then execute to implement the decisions.

[0285] 26 is a flow chart illustrating an example process 2600 for editing a video based on rehearsal data. The process 2600 can be performed by a server system, such as the server system 2502 of FIG.

[0286] The server system receives 2602 rehearsal data, including rehearsal video data and rehearsal audio data, from one or more recording devices. The rehearsal data represents a rehearsal of an event by one or more performers of the event. The one or more recording devices may include one or more microphones and one or more video cameras. The one or more video cameras may include at least one video camera designated as a high resolution video camera, e.g., a 4K capable video camera.

[0287] The server system recognizes (2604) a respective image of each of the one or more performers from the rehearsal video data. Recognizing the respective image can be based on video-based tracking of at least one of the performer or an instrument played by the performer. For example, the recognition can be based on facial recognition, instrument recognition, or other object recognition.

[0288] The server system determines (2606) corresponding sound attributes associated with each recognized image from the rehearsal audio data. The sound attributes may include sound type, sound level, or both. Sound type may indicate the type of instrument used by the performer, for example guitar, drums, or vocals.

[0289] The server system receives 2608 live data, including live video and audio data of the event, from the one or more recording devices. In some implementations, the live data can be buffered on the one or more recording devices, on the server system, or both, for a period of time depending on the time it takes to process the data and whether the results are stored or streamed to a user device.

[0290] The server system determines (2610) a respective prominence of each performer based on the recognized image and associated sound attributes. The system can derive a respective level at which each performer plays relative to the rehearsal as well as a respective position of each performer during the rehearsal captured by the one or more video cameras. The server system can determine the prominent performer using techniques for determining a dominant sound source as described above. In some implementations, determining that a first performer is a prominent performer can include the following operations: The server system normalizes respective loudness levels among the performers based on the live rehearsal audio data. The service system determines in the live audio data that at least one performer performs at a level above the normalized loudness levels of other performers by at least a threshold amount after normalization. The service system can then determine that the first performer is a prominent performer.

[0291] In some implementations, normalizing each loudness level may include the following operations: The server system may determine, from the rehearsal audio data, a first level of sound for each performer and a second level of sound for each performer, the first level being lower than the second level, and the server system may then normalize each loudness level by scaling and aligning the first level and scaling and aligning the second level.

[0292] In some implementations, determining that the first performer is a prominent performer may include the following operations: The server system determines, based on the live video data, that an amount of movement of the first performer exceeds an amount of movement of other performers by at least a threshold value. Then, the server system determines that the first performer is a prominent performer based on the amount of movement.

[0293] The server system edits (2612) the live video data and the live audio data according to one or more editing rules. In editing the live data, the server system emphasizes at least one performer based on their respective salience. For example, the edits may emphasize a vocalist, an instrument, or a section of a band or orchestra that includes multiple performers (e.g., a brass section or a woodwind section). The edits may be performed on the live video data by the server system. In some implementations, the edits may be performed by a recording device. For example, the server system may provide edit instructions to the recording device to cause the recording device to perform the edit operations.

[0294] Editing the live video data and the live audio data can include determining a pace and tempo of the event based on the live audio data. The server system can then cut the live video data according to the pace and tempo. Editing the live video data and the live audio data can include determining when a performer, e.g., a first performer, has started or stopped performing. The server system can then responsively cut the live video data at, e.g., the time when the performer started or stopped performing.

[0295] In some implementations, editing the live video data and the live audio data includes the following operations: The server system determines that the time elapsed since a full frame shot showing all performers exceeds a threshold time. The server system can cut the live video data in response. The server system can determine the threshold time based on a duration of time or a number of cuts derived from a tempo of the live audio data.

[0296] The server system then provides the edited data for playback (2614). The server system can store the association of the edited live video data and the edited live audio data in a storage device or stream the association of the edited live video data and the edited live audio data to a user device.

[0297] <Frame area selection> 27 is a block diagram illustrating an example technique for selecting sub-frame regions from full-frame video data. At a live event, e.g., a concert, at least one video capture device 2702 captures video of the event. At least one audio capture device 2704 captures audio of the event. The devices 2702 and 2704 submit live video and live audio data to a server system 2706 over a communications network 2708, e.g., the Internet. The video capture device 2702 may record video at high resolution, e.g., 4K. The high resolution video may consume too much bandwidth of the communications network 2708 to upload to the server system 2706 or to download from the server system 2706 to a user device.

[0298] The server system 2706 can be the same as or different from the server system 2502 described with reference to FIG. 25. The server system 2706 can store or stream medium resolution, e.g., 720p, video to a user device. The video capture device 2702 can be configured as a slave to the server system 2706. As a slave to the server system 2706, the video capture device 2702 follows commands from the server system 2706 to zoom, crop, and select a focal point of the video data. The server system 2706, receiving live audio data from the audio capture device 2704, makes a decision on selecting subframes and instructs the video capture device 2702 to submit the selected subframes to the server system 2706.

[0299] The server system 2706 receives at least one video frame of the full band at the event. The video frame includes all performers. The video frame does not have to be full resolution and can optionally be compressed using a lossy codec. The server system 2706 then determines where to focus and which sub-frame to select based on the live audio data. The server system 2706 instructs the video capture device 2702 to submit a medium resolution live video of only the selected sub-frame to the server system 2706.

[0300] The video capture device 2702 includes a video buffer 2710. The video buffer 2710 is a data store configured to store X seconds (e.g., 10 seconds) of video data at full resolution. The video data may include a series of full-band frames 2712 and associated time information. The video capture device 2702 includes a video converter 2714. The video converter 2714 converts the full-band frames 2712 from full resolution to a series of lower resolution (e.g., 720p or 640×480) images. The video converter 2714 submits the lower resolution images to the server system 2706 at a reduced frame rate (e.g., 1 fps). Meanwhile, the video capture device 2702 converts the video stream in the video buffer 2710 to a medium resolution video and submits the medium resolution video to the server system 2706 at a standard frame rate (e.g., 24 fps).

[0301] At an initial time t0, the submitted video may be a full-band coverage video with frames that match the lower resolution images. The video capture device 2702 then continues to submit the medium resolution video data and the images to the server system 2706 while waiting for instructions from the server system 2706 regarding editing decisions.

[0302] The server system 2706 includes an AVEE 2718, which may be the same as or different from the AVEE 2718 of FIG. 25. The AVEE 2718 receives the full frame image and live audio data. The AVEE 2718 is configured to determine which performer or instrument to focus on based on the full frame image and live audio data received from the audio capture device 2704. For example, the AVEE 2718 may determine that the singer is the prominent performer at time t1. The AVEE 2718 may then issue a command to zoom in on the prominent performer, in this example the singer, in the video from time t1. The command may be associated with sensor pixel coordinates, e.g., X pixels from the left, Y pixels from the bottom, a size, and a time t1.

[0303] In response to the command, the video capture device 2704 performs editing according to the command. The video capture device 2704 retrieves video data from the video buffer 2710 for the corresponding time t1. The video capture device 2704 crops the position according to the coordinates. The video capture device 2704 scales the cropped video data to the specified size and submits the scaled video data to the server system 2706 while continuing to submit the converted full-frame image to the server system 2706. The video capture device 2704 can submit the scaled video data at a standard frame rate, for example, 24 fps. Thus, the server system 2706 receives video that focuses on a prominent performer, for example, a singer, from time t1.

[0304] Based on the live audio data and images submitted to the server system 2706, the server system 2706 may determine that at a second time t2, a second performer, e.g., a violinist, will be the prominent performer. The server system 2706 then instructs the video capture device 2704 to focus on the portion of the live video that includes video of the violinist. Upon receiving the instruction specifying the location and size of the subframe and the time t2, the video capture device 2704 changes the cropping coordinates and submits the cropped and optionally resized video that includes the violinist to the server system 2706. Thus, the server system 2706 receives a medium resolution video of the violinist from time t2 onwards.

[0305] The server system 2706 includes an assembly unit 2720. The assembly unit 2720 is configured to assemble the medium resolution video data from the video capture devices 2702 and the live audio data from the audio capture devices 2704 for storage or streaming to a user device. For live streaming, the assembly unit 2720 can add a delay to the beginning of the assembled video stream. Both the video buffer 2710 and the delay can compensate for latency in the decision and data transmission. For example, the server system 2706 can determine to zoom on the drummer when he enters and instruct the video capture device 2702 to focus on the drummer at a location in the buffer held in memory that corresponds to the time when the drummer enters. This time may be X (e.g., 0.2) seconds before the video capture device 2702 receives the command. The server system 2706 then receives the new edited video and serves it to the audience, using the delay to hide the time it takes to make decisions and send commands.

[0306] 28 is a flow chart of an example process 2800 for selecting sub-frame regions from full-frame video data performed by a server system, which may be the server system 2706 of FIG.

[0307] The server system receives audio data of the event from one or more audio capture devices and at least one frame of video data of the event. The video data is captured by a video capture device configured to record video at a first resolution. The first resolution can be 4K or greater. The frame can have a resolution that is the same as or less than the first resolution. The frame of video data of the event can be a frame of a series of frames of video data. The series of frames can be received at the server system at a frame rate (e.g., 1 fps or less) lower than a frame capture rate (e.g., 24 fps or greater). In some implementations, a single frame capturing all performers of the event is sufficient. In some implementations, the video capture device submits multiple frames to the server system to cover performers who may have moved during the event.

[0308] The server system determines 2804 a respective location of each of the individual performers of the event based on images of the individual performers recognized from the frames of audio data and video data. The server system can make this determination based on the rehearsal data.

[0309] When the server system determines from the audio data that a first one of the individual performers is the prominent performer at the first time, the server system directs the video capture device to submit 2806 a first portion of the video data to the server system at a second resolution. The first portion of the video data is spatially oriented to a location of the first performer captured at the first time. The second resolution can be 1080p or less.

[0310] If the server system determines from the audio data that a second one of the individual performers is the prominent performer at the second time, the server system directs the video recorder to submit 2808 a second portion of the video data to the server system at the second resolution, the second portion of the video data being spatially oriented to a location of the second performer captured at the second time.

[0311] The server system designates (2810) the first and second portions of the video data as a video of the event at the second resolution. The server system then provides (2812) the association of the audio data and the video of the event at the second resolution to a storage device or user device as an audio and video recording of the event. For example, the server system may add a delay to the video of the event at the second resolution. The server system then streams the delayed video and associated audio data to one or more user devices.

[0312] In some implementations, the video capture device buffers a period of video data at a first resolution and, in response to a command from the server system, the video capture device selects a location in the buffered video data frames corresponding to the first performer and the second performer for submission to the server system.

[0313] FIG. 29 is a flow chart of an exemplary process 2900 performed by a video capture device for selecting sub-frame regions from full-frame video data.

[0314] The video capture device records (2902) video data at a first resolution. The first resolution can be 4K or greater. The video capture device stores (2904) the video data in a local buffer of the video capture device. The video capture device determines (2900) a sequence of one or more images from the recorded video data. The video capture device submits (2908) the sequence of one or more images to the server system at a first frame rate. The first frame rate can be one frame per second or less. The video capture device receives (2910) instructions from the server system to focus on a portion of the video data. The instructions indicate a temporal and spatial location of the portion of the recorded video data.

[0315] In response to the command, the video capture device converts (2912) the portion of the video data stored in the local buffer according to the indicated temporal and spatial locations into video data at a second resolution having a second frame rate higher than the first frame rate. The second frame rate can be 24 frames per second or greater. The video capture device then submits (2914) the converted video data at the second resolution to a server as live video data of the event.

[0316] Exemplary Recorder Architecture FIG. 21 is a block diagram illustrating an example device architecture 2100 of a device implementing the features and operations described with reference to FIGS. 1-20 and 24-29. The device can be, for example, the recording device 102 or 104 of FIG. 1 or the recording device 302 of FIG. 3. The device can include a memory interface 2102, one or more data processors, image processors and / or processors 2104 and / or a peripheral interface 2106. The memory interface 2102, the one or more processors 2104 and / or the peripheral interface 2106 can be separate components or can be integrated in one or more integrated circuits. The processor 2104 can include an application processor, a baseband processor and a radio processor. The various components in, for example, a mobile device can be coupled by one or more communication buses or signal lines.

[0317] Sensors, devices, and subsystems can be coupled to the peripheral interface 2106 to facilitate multiple functions. For example, a motion sensor 2110, a light sensor 2112, and a proximity sensor 2114 can be coupled to the peripheral interface 2106 to facilitate orientation, lighting, and proximity functions of the mobile device. A position processor 2115 can be coupled to the peripheral interface 2106 to provide geographic positioning. In some implementations, the position processor 2115 can be programmed to perform the operations of a GNSS receiver. An electronic magnetometer 2116 (e.g., an integrated circuit chip) can also be coupled to the peripheral interface 2106 to provide data that can be used to determine the direction of magnetic north. In this manner, the electronic magnetometer 2116 can be used as an electronic compass. The motion sensor 2110 can include one or more accelerometers configured to determine changes in speed and direction of motion of the mobile device. A barometer 2117 can be coupled to the peripheral interface 2106 and can include one or more devices configured to measure the pressure of the atmosphere around the mobile device.

[0318] A camera subsystem 2120 and an optical sensor 2122, such as a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) optical sensor, can be utilized to facilitate camera functions such as recording photographs and video clips.

[0319] The communication functions can be facilitated through one or more wireless communication subsystems 2124. The communication subsystem 2124 can include radio frequency receivers and transmitters and / or optical (e.g., infrared) receivers and transmitters. The specific design and implementation of the communication subsystem 2124 can depend on the communication network through which the mobile device is intended to operate. For example, the mobile device can include a communication subsystem 2124 designed to operate through a GSM network, a GPRS network, an EDGE network, a Wi-Fi™ or WiMax™ network, and a Bluetooth™ network. In particular, the wireless communication subsystem 2124 can include a host protocol such that the mobile device can be configured as a base station for other wireless devices.

[0320] The audio subsystem 2126 can be coupled to a speaker 2128 and a microphone 2130 to facilitate voice-enabled functions such as voice recognition, voice mimicry, digital recording, and telephone functions. The audio subsystem 2126 can be configured to receive voice commands from a user.

[0321] The I / O subsystem 2140 can include a touch surface controller 2142 and / or other input controllers 2144. The touch surface controller 2142 can be coupled to a touch surface 2146 or pad. The touch surface 2146 and touch surface controller 2142 can detect contact and movement or interruptions using, for example, any of a number of touch sensitive technologies. Touch sensitive technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch surface 2146. The touch surface 2146 can include, for example, a touch screen.

[0322] The other input controller 2144 can be coupled to other input / control devices 2148, such as one or more buttons, a rocker switch, a thumbwheel, an infrared port, a USB port, and / or a pointer device, such as a stylus. The one or more buttons (not shown) can include up / down buttons for volume control of the speaker 2128 and / or microphone 2130.

[0323] In some implementations, pressing the button for a first duration may unlock the touch surface 2146; pressing the button for a second duration longer than the first duration may turn power on or off to the mobile device. A user may be able to customize the functionality of one or more of the buttons. The touch surface 2146 may also be used to implement, for example, virtual or soft buttons and / or a keyboard.

[0324] In some implementations, the mobile device can present recorded audio and / or video files such as MP3, AAC, and MPEG files. In some implementations, the mobile device can include MP3 player functionality. Other input / output and control devices can also be used.

[0325] The memory interface 2102 can be coupled to a memory 2150. The memory 2150 can include high speed random access memory and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, and / or flash memory (e.g., NAND, NOR). The memory 2150 can store an operating system 2152, such as an embedded operating system, such as iOS, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks. The operating system 2152 can include instructions for handling basic system services and for performing hardware-dependent tasks. In some implementations, the operating system 2152 can include a kernel (e.g., a UNIX kernel).

[0326] The memory 2150 may also store communications instructions 2154 to facilitate communications with one or more additional devices, one or more computers, and / or one or more servers. The memory 2150 may include graphical user interface instructions 2156 for facilitating graphical user interface processing; sensor processing instructions 2158 for facilitating sensor-related processing and functions; telephony instructions 2160 for facilitating telephony-related processes and functions; electronic messaging instructions 2162 for facilitating electronic messaging-related processes and functions; web browsing instructions 2164 for facilitating web browsing-related processes and functions; media processing instructions 2166 for facilitating media processing-related processes and functions; GNSS / position instructions 2168 for facilitating general GNSS and position related processes and functions; camera instructions 2170 for facilitating camera-related processes and functions; magnetometer data 2172 and calibration instructions 2174 for facilitating magnetometer calibration. The memory 2150 may also store other software instructions (not shown), such as security instructions, web video instructions for facilitating web video related processes and functions, and / or web shopping instructions for facilitating web shopping related processes and functions. In some implementations, the media processing instructions 2166 are split into audio processing instructions and video processing instructions for facilitating audio processing related processes and functions and video processing related processes and functions, respectively. Activation records and an International Mobile Equipment Identity (IMEI) or similar hardware identifier may also be stored in the memory 2150. The memory 2150 may store audio processing instructions 2176 that, when executed by the processor 2104, cause the processor 2104 to perform various operations, including:For example, joining a group for recording services by logging into a user account, designating one or more microphones on the device as spot or primary microphones, recording audio signals for the group using one or more microphones, and submitting the recorded signals to a server. In some implementations, the audio processing instructions 2176 may cause the processor 2104 to perform operations of the server 408 described with reference to FIG. 4 and other figures. The memory 2150 may store video processing instructions that, when executed by the processor 2104, may cause the processor 2104 to perform various operations described with reference to FIGS. 25-29.

[0327] Each of the above-identified instructions and applications may correspond to a set of instructions for performing one or more functions described above. These instructions need not be implemented as separate software programs, procedures, or modules. The memory 2150 may include additional or fewer instructions. Additionally, various functions of the mobile device may be implemented in hardware and / or software, including in one or more signal processing and / or application specific integrated circuits.

[0328] FIG. 22 is a block diagram of an example network operating environment 2200 for the mobile devices of FIGS. 1-20 and 24-29. The devices 2202a and 2202b can communicate, for example, through one or more wired and / or wireless networks 2210 in data communication. For example, the wireless network 2212, e.g., a cellular network, can communicate with a wide area network (WAN) 2214, such as the Internet, by using a gateway 2216. Similarly, an access device 2218, such as an 802.11g wireless access point, can provide communication access to the wide area network 2214. Each of the devices 2202a and 2202b can be the device 102 or device 104 of FIG. 1, or the recording device 302 of FIG. 3.

[0329] In some implementations, both voice and data communications can be established through the wireless network 2212 and the access device 2218. For example, the device 2202a can make and receive telephone calls (e.g., using Voice over Internet Protocol (VoIP) protocol), send and receive electronic mail messages (e.g., using Post Office Protocol 3 (POP3)), and retrieve electronic documents and / or streams such as web pages, photos, and videos (e.g., using Transmission Control Protocol / Internet Protocol (TCP / IP) or User Datagram Protocol (UDP)) through the wireless network 2212, the gateway 2216, and the wide area network 2214. Similarly, in some implementations, the device 2202b can make and receive telephone calls, send and receive electronic mail messages, and retrieve electronic documents through the access device 2218 and the wide area network 2214. In some implementations, the device 2202a or 2202b can be physically connected to the access device 2218 using one or more cables, and the access device 2218 can be a personal computer. In this configuration, device 2202a or 2202b may be referred to as a "tethered" device.

[0330] The devices 2202a and 2202b may also establish communication by other means. For example, the wireless device 2202a may communicate with other wireless devices, e.g., other mobile devices, cell phones, etc., through the wireless network 2212. Similarly, the devices 2202a and 2202b may establish peer-to-peer communication 2220, e.g., a personal area network, through the use of one or more communication subsystems, such as a Bluetooth® communication device. Other communication protocols and topologies may also be implemented.

[0331] Device 2202a or 2202b can communicate with one or more services 2230, 2240, and 2250, for example, through one or more wired and / or wireless networks. For example, one or more audio and video processing services 2230 can provide audio processing services including automatic synchronization, automatic leveling, automatic panning, automatic source equalization, automatic segmentation, and streaming, as described above. Mixing service 2240 can provide a user interface that allows a mixing professional to log in through a remote console and perform mixing operations on live audio data. Visual effects service 2250 can provide a user interface that allows a visual effects professional to log in through a remote console and edit video data.

[0332] Device 2202a or 2202b may also access other data and content through one or more wired and / or wireless networks. For example, content publishers such as news sites, Really Simple Syndication (RSS) feeds, websites, blogs, social networking sites, developer networks, etc. may be accessed by device 2202a or 2202b. Such access may be provided, for example, by invoking a web browsing function or application (e.g., a browser) in response to a user touching a web object.

[0333] Example System Architecture FIG. 23 is a block diagram of a system architecture for an exemplary server system implementing the features and operations described with reference to FIGS. 1-20 and 24-29. Other architectures are possible, including architectures with more or fewer components. In some implementations, the architecture 2300 includes one or more processors 2302 (e.g., a dual-core Intel® Xeon® processor), one or more output devices 2304 (e.g., LCD), one or more network interfaces 2306, one or more input devices 2308 (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media 2312 (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can communicate and exchange data through one or more communication channels 2310 (e.g., a bus) that can utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0334] The term "computer-readable medium" refers to a medium that participates in providing instructions to the processor 2302 for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media, including, but not limited to, coaxial cables, copper wire, and fiber optics.

[0335] The computer readable medium 2312 may further include an operating system 2314 (e.g., a Linux operating system), a network communications module 2316, an audio processing manager 2320, a video processing manager 2330, and a live content distributor 2340. The operating system 2314 may be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system 2314 performs basic tasks including, but not limited to: recognizing input from and providing output to the network interface 2306 and / or device 2308; tracking and managing files and directories on the computer readable medium 2312 (e.g., memory or storage); controlling peripheral devices; and managing traffic on one or more communications channels 2310. The network communications module 2316 includes various components (e.g., software for implementing communications protocols such as TCP / IP, HTTP, etc.) for establishing and maintaining network connections.

[0336] The audio processing manager 2320 may include computer instructions that, when executed, cause the processor 2302 to perform various audio estimation and manipulation operations, such as those described above with reference to the server 408. The video processing manager 2330 may include computer instructions that, when executed, cause the processor 2302 to perform video editing and manipulation operations, such as those described above with reference to the video editor 530, the AVEE 2518, or the AVEE 2718. The live content distributor 2340 may include computer instructions that, when executed, cause the processor 2302 to perform operations of receiving reference audio data and live data of an audio event and, after the audio and visual data have been processed, streaming the processed live data to one or more user devices.

[0337] The architecture 2300 can be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software may include multiple software components or may be a single piece of code.

[0338] The described features may be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform a certain activity or bring about a certain result. Computer programs can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and can be written as stand-alone programs or as modules, components, subroutines, browser-based web applications, or other units suitable for use in a computing environment.

[0339] Suitable processors for the execution of a program of instructions include, by way of example, both general purpose and special purpose microprocessors, as well as the sole processor or one of multiple processors or cores of any kind of computer. Typically, a processor receives instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer will also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files. Such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Suitable storage devices for embodying computer program instructions and data include, by way of example, semiconductor memory devices, such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and all forms of non-volatile memory, including CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, ASICs (Application Specific Integrated Circuits).

[0340] To provide for interaction with a user, the functions may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retinal display, for displaying information to the user. The computer may have a touch surface input device (e.g., a touch screen) or a keyboard and pointing device, such as a mouse or trackball, by which the user may provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0341] The functionality may be implemented in a computer system that includes back-end components such as data servers, or that includes middleware components such as application servers or Internet servers, or that includes front-end components, e.g., a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system may be connected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include, e.g., LANs, WANs, and the computers and networks forming the Internet.

[0342] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server sends data (e.g., HTML pages) to a client device (e.g., for the purposes of displaying the data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received from the client device at the server.

[0343] One or more computer systems can be configured to perform certain actions by virtue of having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform the actions. One or more computer programs can be configured to perform certain actions by virtue of including instructions that, when executed by a data processing device, cause the device to perform the actions.

[0344] Although the present specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments herein can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and may even be initially claimed as such, one or more features from a claimed combination may in some cases be cut out from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0345] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or sequentially, or that all of the depicted operations be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0346] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

[0347] Several implementations of the present invention have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the invention.

[0348] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE).

[0349] [EEE1] 13. A method of leveling audio comprising: receiving, by a leveling unit including one or more electronic circuits, reference audio data, the reference audio data including representations of channel signals from a multiple channel signal source; receiving, by the leveling unit, target level data specifying a target level for each sound source; determining, by the leveling unit, a cost function for rescaling audio signals to the target levels according to respective gains based on the reference audio data; and calculating, by the leveling unit, a respective gain to be applied to each channel signal in the live audio data by minimizing the cost function. method. [EEE2] The method of EEE1, wherein the representation of the channel signal comprises an original channel signal or a processed channel signal, the processed channel signal comprising a channel signal processed by a noise reduction unit, an equalizer, a dynamic range compensation unit or an audio source separator. [EEE3] The method of any one of EEE1 and EEE2, further comprising determining, by said levelling unit, a respective correlation between each pair of said channel signal sources. [EEE4] 13. A method of panning audio comprising: receiving, by a panner including one or more electronic circuits, reference audio data for audio sources, the audio sources including one or more sources designated as one or more pannable sources and one or more sources designated as one or more non-pannable sources; receiving a channel signal of an event to be played by the audio source; determining a cost function based on the reference audio data, the cost function having a pan position for each channel signal as a variable, the cost function having a first component representing an imbalance between a left channel and a right channel, a second component representing the one or more pannable sources, and a third component representing the one or more non-pannable ones of the sound sources; determining a respective pan position for each channel signal by minimizing the cost function; applying the pan positions to the channel signals to achieve an audio effect of placing sound sources of the event between the left and right of a stereo sound stage for output to a stereo sound reproduction system. method. [EEE5] The method of EEE4, wherein the pan position includes at least one of a pan angle or a ratio between a left channel and a right channel, and the stereo sound reproduction system includes headphones or loudspeakers. [EEE6] 1. A method of leveling and panning audio, comprising: receiving, by a leveling and panning unit including one or more electronic circuits, reference audio data comprising representations of channel signals from a multi-channel signal source recorded in rehearsal of one or more sound sources; receiving, by said leveling and panning unit, target level data specifying a target level for each sound source; receiving, by said leveling and panning unit, live audio data, said live audio data including recorded or real-time signals from said one or more sound sources playing at a live event; determining, by the leveling unit, a joint cost function for leveling the live audio data and panning the live audio data based on the reference audio data, the joint cost function having a first component for leveling the live audio data and a second component for panning the live audio data, the first component being based on the target level data and the second component being based on a first representation of imbalance between a left channel and a right channel, a second representation of sources that can be panned among sound sources, and a third representation of sources that cannot be panned among sound sources; calculating a respective gain to apply to each channel signal and a respective pan position for each channel signal by minimizing the joint cost function; applying said gain and pan position to a signal of live audio data of the event to achieve an audio effect of leveling sound sources in the live audio data and positioning sound sources in the live audio data between the left and right of a stereo sound stage for output to a storage device or a stereo sound reproduction system; method. [EEE7] The method of EEE6, wherein each level is an energy level or a loudness level. [EEE8] 1. A method for determining an audio level, comprising: receiving, by an estimator including one or more electronic circuits, reference audio data, the reference audio data including channel signals respectively representative of one or more sound sources to be played during a rehearsal session; calculating, by said estimator, a respective level of each sound source at each microphone based on said reference audio data; determining a level difference between the live audio data and the reference audio data, the level difference comprising comparing each sound source represented in the live audio data with its respective level represented in the reference audio data; determining a cost function for each level of each sound source based on said difference; determining said respective levels by minimizing said cost function; and providing said level as an input to an audio or video processor. method. [EEE9] calculating, by said estimator, a respective level of each sound source in each frequency band of a plurality of frequency bands; the cost function comprises a respective sum of costs across frequency bands for each sound source; said respective levels being determined in each frequency band; The method described in EEE8. [EEE10] 1. A method of equalizing audio, comprising: receiving audio data, the audio data including signals from a plurality of sound sources, by an equalizer having one or more electronic circuits; mapping, by the equalizer, a respective signal for each sound source to an excitation in each frequency band; determining a requirement value for each source-band pair in the list of source-band pairs, each source-band pair representing a sound source and a frequency band, said requirement value indicating the relative importance of the sound source represented in that pair to other sound sources and other frequency bands being equalized in that frequency band in the pair, and the masking level of the sound source represented in that pair by one or more other sound sources; iteratively equalizing the signal of the sound source represented in the source-band pair in the list having the highest need value and removing the equalized source-band pair from the list until the highest need value of the remaining source-band pairs is less than a threshold value; and providing the equalized signal for playback on one or more loudspeakers. method. [EEE11] The method of claim EEE10, wherein the required value is a product of one or more values ​​representing the relative importance and one or more values ​​representing a masking level of the sound source. [EEE12] receiving an audio signal by a segmentation unit having one or more electronic circuits; The segmentation unit constructs a novelty index for the audio signal over time; determining a cut time for a next cut based on the peak of the novelty index; cutting the video content at said cut time; providing the cut video content as a new video segment to a storage device or to one or more end user devices; method. [EEE13] The cut time is determined by: determining a segment length based on the average cut length, the segment length corresponding to a length of the audio segment; determining the cut time based on a segment length; The method described in EEE12. [EEE14] Determining the cut time based on the segment length comprises: determining the sum of said novelty index over time since the last cut; and if it is determined that the sum is greater than the novelty threshold, determining the cut time as the time from when the sum of the novelty indexes satisfies the novelty threshold to the time of the next cut, wherein the randomness of the cut time averages out to the segment length. The method described in EEE13. [EEE15] 13. A method of synchronizing audio, comprising: receiving audio signals from a plurality of microphones; determining a respective quality value of correlation between each pair of said audio signals and assigning said quality values ​​in a map vector; iteratively determining a set of delays and inserting the delays into the map vector, where iteratively determining a set of delays includes iteratively aligning and downmixing a pair of audio signals having a highest quality value; and upon completion of the successive iterations, synchronizing the audio signal with each delay in the map vector in accordance with the order in which the delays were inserted into the map vector. method. [EEE16] 13. A method for reducing noise comprising: receiving, by a noise reduction unit including one or more electronic circuits, reference audio data, the reference audio data including channel signals recorded during a silence period rehearsal session; estimating, by a noise estimator of the noise reduction unit, a respective noise level in each channel signal in the reference audio data; receiving live performance data, the live performance data including channel signals recorded during an event in which one or more instruments that were silent in a rehearsal session are played; reducing, by a noise reducer of the noise reduction unit, a respective noise level in each channel signal in the live performance data individually, including applying a respective suppression gain to each channel signal in the live performance data when it is determined that a difference between a noise level in each channel signal in the live performance data and the estimated noise level satisfies a threshold; and after reducing the noise level, providing the channel signal to a downstream device for further processing, storage or delivery to one or more end user devices. method. [EEE17] estimating a respective noise level in each channel signal in the reference audio data is performed for a plurality of frequency bins; Reducing a noise level in each channel signal in the live performance data is performed in the frequency bins; said estimating and said reducing are performed in accordance with noise reduction parameters including a threshold, a slope, an attack time, a decay time and an octave size; The method described in EEE16. [EEE18] receiving, by a server system, reference audio data from one or more channel signal sources, the reference audio data including acoustic information of one or more sound sources to be played individually in a rehearsal; receiving, by the server system, one or more channel signals of a performance event from the one or more channel signal sources, each channel signal being from a respective channel signal source and including an audio signal from the one or more audio sources playing at the performance event; mixing, by the server system, the one or more channel signals, the mixing including automatically adjusting one or more audio attributes of one or more sound sources of the performance event based on the reference audio data; providing the mixed recording of the performance event from said server system to a storage device or to a plurality of end user devices; providing from the server system to a storage device the one or more channel signals of the performance event and a separate file describing the adjustment of at least one or more audio attributes. method. [EEE19] receiving, by a server system, reference audio data from one or more channel signal sources, the reference audio data including acoustic information of one or more sound sources to be played individually; receiving, by the server system, one or more channel signals of a performance event from the one or more channel signal sources, each channel signal being from a respective channel signal source and including an audio signal from the one or more audio sources playing at the performance event; mixing, by the server system, the one or more channel signals, the mixing including automatically adjusting one or more audio attributes of one or more sound sources of the performance event based on the reference audio data; providing the mixed recording of the performance event from the server system to a storage device or to a plurality of end user devices; method. [EEE20] Each channel signal source includes a microphone or a sound signal generator having a signal output; Each sound source is a vocalist, an instrument, or a synthesizer; the server system includes one or more computers connected to the one or more channel signal sources through a communications network; the one or more channel signal sources and the one or more sound sources have the same acoustic configuration during rehearsals and during the performance event; The method of any one of EEE1 to 19. [EEE21] the one or more channel signals include a first channel signal from a first channel signal source of the one or more channel signal sources and a second channel signal from a second channel signal source of the one or more channel signal sources; the method including synchronizing, by the server system, the first channel signal and the second channel signal in the time domain; The method of any one of EEE1 to 20. [EEE22] A method according to any one of EEE1 to 21, comprising separating a first sound source and a second sound source from the one or more channel signals, the separating step comprising separating the first sound source and the second sound source from a plurality of sound sources represented in the one or more channel signals, the one or more channel signals comprising a first signal representative of the first sound source and a second signal representative of the second sound source. [EEE23] The method of any one of EEE1 to EEE22, wherein the mixing includes leveling a first sound source and a second sound source and panning the first sound source and the second sound source by the server system. [EEE24] The method of claim EEE23, wherein leveling the first sound source and the second sound source comprises increasing or decreasing a gain of the one or more sound sources in response to a respective energy level of each sound source, the respective energy levels being determined by the server system from the reference audio data. [EEE25] Said reference audio data is: the signal of each sound source playing at a first level designated as a low level and a second level designated as a high level; or Each source signal plays at a single level The method of any one of EEE1 to 24, comprising at least one of: [EEE26] determining a respective gain for each sound source in the event from the reference audio data, wherein determining the respective gain comprises, for each sound source: receiving an input specifying a target level; determining a level of each of said signals in said reference audio data; determining respective gains based on a difference between a level of the signal in the reference audio data and the target level; The method of any one of EEE1 to 25. [EEE27] The method of any one of EEE1 to EEE26, wherein mixing the one or more channel signals includes adjusting gain of the one or more channel signals according to input from a mixer device logged onto the server system, the signals being from the one or more sound sources or both. [EEE28] performing a video edit on the event, the performing a video edit comprising: receiving, by a video editor of the server system, video and audio data of the event, the video data including video of sound sources positioned to appear at various locations in the event, and the audio data including energy levels of sound sources; determining from the audio data that a signal of a first sound source represented in the audio data indicates that the first sound source is playing at a level above a threshold amount relative to the levels of other sound sources represented in the audio data; determining a location of the first sound source in the video data; determining a portion of the video data that corresponds to a location of the first sound source; and providing said audio data and said portion of said video data synchronously to said storage device or said end user device. The method of any one of EEE1 to 27. [EEE29] Determining a location of a sound source in the video data comprises: determining a pan position of the first sound source based on the audio data; specifying the pan position of the first sound source as the position of a sound source in the video data. The method described in EEE28. [EEE30] The method of claim EEE28, wherein determining a location of a sound source in the video data comprises determining a location of a sound source using face tracking or instrument tracking. [EEE31] providing commands from the server system to the one or more channel signal sources based on the one or more channel signals, the commands configured to adjust recording parameters of the one or more channel signal sources, the recording parameters including at least one of a gain, a compression type, a bit depth, or a data transmission rate; The method of any one of EEE1 to 30.

[0350] Several aspects will be described. [Aspect 1] receiving, by a server system, reference audio data from one or more channel signal sources, the reference audio data including acoustic information of one or more sound sources to be played individually; receiving, by the server system, one or more channel signals of a performance event from the one or more channel signal sources, each channel signal being from a respective channel signal source and including an audio signal from the one or more audio sources playing at the performance event; mixing, by the server system, the one or more channel signals, the mixing including automatically adjusting one or more audio attributes of one or more sound sources of the performance event based on the reference audio data; providing the mixed recording of the performance event from the server system to a storage device or to a plurality of end user devices; method. [Aspect 2] Each channel signal source includes a microphone or a sound signal generator having a signal output; Each sound source is a vocalist, an instrument, or a synthesizer; the server system includes one or more computers connected to the one or more channel signal sources through a communications network; the one or more channel signal sources and the one or more sound sources have the same acoustic configuration during rehearsals and during the performance event; The method of embodiment 1. [Aspect 3] the one or more channel signals include a first channel signal from a first channel signal source of the one or more channel signal sources and a second channel signal from a second channel signal source of the one or more channel signal sources; the method including synchronizing, by the server system, the first channel signal and the second channel signal in the time domain; The method of embodiment 1. Aspect 4 4. The method of any one of aspects 1 to 3, further comprising separating a first sound source and a second sound source from the one or more channel signals, the separating comprising separating the first sound source and the second sound source from a plurality of sound sources represented in the one or more channel signals, the one or more channel signals comprising a first signal representing the first sound source and a second signal representing the second sound source. Aspect 5 5. The method of any one of aspects 1 to 4, wherein the mixing includes leveling a first sound source and a second sound source and panning the first sound source and the second sound source by the server system. Aspect 6 The method of embodiment 5, wherein leveling the first sound source and the second sound source includes increasing or decreasing a gain of the one or more sound sources in response to a respective energy level of each sound source, the respective energy levels being determined by the server system from the reference audio data. Aspect 7 Said reference audio data is: the signal of each sound source playing at a first level designated as a low level and a second level designated as a high level; or Each source signal plays at a single level The method of any one of embodiments 1 to 6, comprising at least one of: Aspect 8 determining a respective gain for each sound source in the event from the reference audio data, wherein determining the respective gain comprises, for each sound source: receiving an input specifying a target level; determining a level of each of said signals in said reference audio data; determining respective gains based on a difference between a level of the signal in the reference audio data and the target level; The method of any one of embodiments 1 to 7. Aspect 9 The method of any one of aspects 1 to 8, wherein mixing the one or more channel signals includes adjusting gain of the one or more channel signals according to input from a mixer device logged onto the server system, the signals being from the one or more sound sources or both. Aspect 10 performing a video edit on the event, the performing a video edit comprising: receiving, by a video editor of the server system, video and audio data of the event, the video data including video of sound sources positioned to appear at various locations in the event, and the audio data including energy levels of sound sources; determining from the audio data that a signal of a first sound source represented in the audio data indicates that the first sound source is playing at a level above a threshold amount relative to the levels of other sound sources represented in the audio data; determining a location of the first sound source in the video data; determining a portion of the video data that corresponds to a location of the first sound source; and providing said audio data and said portion of said video data synchronously to said storage device or said end user device. 10. The method of any one of embodiments 1 to 9. Aspect 11 Determining a location of a sound source in the video data comprises: determining a pan position of the first sound source based on the audio data; specifying the pan position of the first sound source as the position of a sound source in the video data. The method of embodiment 10. Aspect 12 11. The method of embodiment 10, wherein determining the location of a sound source in the video data includes determining the location of the sound source using face tracking or instrument tracking. Aspect 13 providing commands from the server system to the one or more channel signal sources based on the one or more channel signals, the commands configured to adjust recording parameters of the one or more channel signal sources, the recording parameters including at least one of a gain, a compression type, a bit depth, or a data transmission rate; 13. The method of any one of embodiments 1 to 12. Aspect 14 13. A method of leveling audio comprising: receiving, by a leveling unit including one or more electronic circuits, reference audio data, the reference audio data including representations of channel signals from a multiple channel signal source; receiving, by the leveling unit, target level data specifying a target level for each sound source; determining, by the leveling unit, a cost function for rescaling audio signals to the target levels according to respective gains based on the reference audio data; and calculating, by the leveling unit, a respective gain to be applied to each channel signal in the live audio data by minimizing the cost function. method. Aspect 15 15. The method of claim 14, wherein the representation of the channel signal comprises an original channel signal or a processed channel signal, the processed channel signal comprising a channel signal processed by a noise reduction unit, an equalizer, a dynamic range compensation unit or an audio source separator. Aspect 16 15. The method of claim 14, further comprising determining, by the leveling unit, a respective correlation between each pair of the channel signal sources. Aspect 17 13. A method of panning audio comprising: receiving, by a panner including one or more electronic circuits, reference audio data for audio sources, the audio sources including one or more sources designated as one or more pannable sources and one or more sources designated as one or more non-pannable sources; receiving a channel signal of an event to be played by the audio source; determining a cost function based on the reference audio data, the cost function having a pan position for each channel signal as a variable, the cost function having a first component representing an imbalance between a left channel and a right channel, a second component representing the one or more pannable sources, and a third component representing the one or more non-pannable ones of the sound sources; determining a respective pan position for each channel signal by minimizing the cost function; applying the pan positions to the channel signals to achieve an audio effect of placing sound sources of the event between the left and right of a stereo sound stage for output to a stereo sound reproduction system. method. Aspect 18 20. The method of embodiment 17, wherein the pan position comprises at least one of a pan angle or a ratio between a left channel and a right channel, and the stereo sound reproduction system comprises headphones or loudspeakers. Aspect 19 1. A method of leveling and panning audio, comprising: receiving, by a leveling and panning unit including one or more electronic circuits, reference audio data comprising representations of channel signals from a multi-channel signal source recorded in rehearsal of one or more sound sources; receiving, by said leveling and panning unit, target level data specifying a target level for each sound source; receiving, by said leveling and panning unit, live audio data, said live audio data including recorded or real-time signals from said one or more sound sources playing at a live event; determining, by the leveling unit, a joint cost function for leveling the live audio data and panning the live audio data based on the reference audio data, the joint cost function having a first component for leveling the live audio data and a second component for panning the live audio data, the first component being based on the target level data and the second component being based on a first representation of imbalance between a left channel and a right channel, a second representation of sources that can be panned among sound sources, and a third representation of sources that cannot be panned among sound sources; calculating a respective gain to apply to each channel signal and a respective pan position for each channel signal by minimizing the joint cost function; applying said gain and pan position to a signal of live audio data of the event to achieve an audio effect of leveling sound sources in the live audio data and positioning sound sources in the live audio data between the left and right of a stereo sound stage for output to a storage device or a stereo sound reproduction system; method. Aspect 20 20. The method of embodiment 19, wherein each level is an energy level or a loudness level. Aspect 21 1. A method for determining an audio level, comprising: receiving, by an estimator including one or more electronic circuits, reference audio data, the reference audio data including channel signals respectively representative of one or more sound sources to be played during a rehearsal session; calculating, by said estimator, a respective level of each sound source at each microphone based on said reference audio data; determining a level difference between the live audio data and the reference audio data, the level difference comprising comparing each sound source represented in the live audio data with its respective level represented in the reference audio data; determining a cost function for each level of each sound source based on said difference; determining said respective levels by minimizing said cost function; and providing said level as an input to an audio or video processor. method. Aspect 22 calculating, by said estimator, a respective level of each sound source in each frequency band of a plurality of frequency bands; the cost function comprises a respective sum of costs across frequency bands for each sound source; said respective levels being determined in each frequency band; The method of embodiment 21. Aspect 23 1. A method of equalizing audio, comprising: receiving audio data, the audio data including signals from a plurality of sound sources, by an equalizer having one or more electronic circuits; mapping, by the equalizer, a respective signal for each sound source to an excitation in each frequency band; determining a requirement value for each source-band pair in the list of source-band pairs, each source-band pair representing a sound source and a frequency band, said requirement value indicating the relative importance of the sound source represented in that pair to other sound sources and other frequency bands being equalized in that frequency band in the pair, and the masking level of the sound source represented in that pair by one or more other sound sources; iteratively equalizing the signal of the sound source represented in the source-band pair in the list having the highest need value and removing the equalized source-band pair from the list until the highest need value of the remaining source-band pairs is less than a threshold value; and providing the equalized signal for playback on one or more loudspeakers. method. Aspect 24 24. The method of claim 23, wherein the required value is a product of one or more values ​​representing the relative importance and one or more values ​​representing masking levels of the sound sources. Aspect 25 receiving an audio signal by a segmentation unit having one or more electronic circuits; The segmentation unit constructs a novelty index for the audio signal over time; determining a cut time for a next cut based on the peak of the novelty index; cutting the video content at said cut time; providing the cut video content as a new video segment to a storage device or to one or more end user devices; method. Aspect 26 The cut time is determined by: determining a segment length based on the average cut length, the segment length corresponding to a length of the audio segment; determining the cut time based on a segment length; The method of embodiment 25. Aspect 27 Determining the cut time based on the segment length comprises: determining the sum of said novelty index over time since the last cut; and if it is determined that the sum is greater than the novelty threshold, determining the cut time as the time from when the sum of the novelty indexes satisfies the novelty threshold to the time of the next cut, wherein the randomness of the cut time averages out to the segment length. The method of embodiment 26. Aspect 28 13. A method of synchronizing audio, comprising: receiving audio signals from a plurality of microphones; determining a respective quality value of correlation between each pair of said audio signals and assigning said quality values ​​in a map vector; iteratively determining a set of delays and inserting the delays into the map vector, where iteratively determining a set of delays includes iteratively aligning and downmixing a pair of audio signals having a highest quality value; and upon completion of the successive iterations, synchronizing the audio signal with each delay in the map vector in accordance with the order in which the delays were inserted into the map vector. method. Aspect 29 13. A method for reducing noise comprising: receiving, by a noise reduction unit including one or more electronic circuits, reference audio data, the reference audio data including channel signals recorded during a silence period rehearsal session; estimating, by a noise estimator of the noise reduction unit, a respective noise level in each channel signal in the reference audio data; receiving live performance data, the live performance data including channel signals recorded during an event in which one or more instruments that were silent in a rehearsal session are played; reducing, by a noise reducer of the noise reduction unit, a respective noise level in each channel signal in the live performance data individually, including applying a respective suppression gain to each channel signal in the live performance data when it is determined that a difference between a noise level in each channel signal in the live performance data and the estimated noise level satisfies a threshold; and after reducing the noise level, providing the channel signal to a downstream device for further processing, storage or delivery to one or more end user devices. method. Aspect 30 estimating a respective noise level in each channel signal in the reference audio data is performed for a plurality of frequency bins; Reducing a noise level in each channel signal in the live performance data is performed in the frequency bins; said estimating and said reducing are performed in accordance with noise reduction parameters including a threshold, a slope, an attack time, a decay time and an octave size; The method of embodiment 29. Aspect 31 receiving, by a server system, rehearsal data from one or more recording devices, the rehearsal data including rehearsal video data and rehearsal audio data, the rehearsal data representing a rehearsal of an event by one or more performers of the event; recognizing, by the server system, from the rehearsal video data, an image of each of the one or more performers; determining corresponding sound attributes associated with each recognized image from said rehearsal audio data; receiving, by the server system, live data from the one or more recording devices, the live data including live video data and live audio data of the event; determining, by said server system, a respective salience of each performer based on the recognized image and associated sound attributes; editing the live video data and the live audio data according to one or more editing rules, the editing rules being based on respective saliency and recognized images; providing the edited live video data and the edited live audio data for playback, the step including storing the association of the edited live video data and the edited live audio data in a storage device or streaming the edited live video data and the edited live audio data; the server system includes one or more computer processors; method. Aspect 32 the one or more recording devices include one or more microphones and one or more video cameras; recognizing each of the images is based on video-based tracking of at least one of a performer or a musical instrument played by the performer; The sound attributes include sound type and loudness level; The method of embodiment 31. Aspect 33 33. The method of embodiment 31 or 32, wherein determining a respective salience of each performer comprises: determining loudness levels of performers from said rehearsal audio data; determining, over time, a loudness level for each performer from said live audio data relative to a corresponding loudness level determined from said rehearsal audio data; Determine the relative salience of each performer over time by comparing their corresponding determined relative loudness levels; and editing the live video data based on the determined relative salience as an input to the editing rules. method. Aspect 34 34. The method of claim 33, wherein determining a loudness level from the rehearsal audio data comprises: determining from the rehearsal audio data a first level of sound for each performer and a second level of sound for each performer, the first level being lower than the second level; calculating a weighted average between the first level and the second level. method. Aspect 35 35. The method of any one of aspects 31 to 34, comprising determining that at least one performer is a prominent performer, wherein determining that the at least one performer is a prominent performer comprises: determining, based on the live video data, that an amount of movement of the first performer exceeds an amount of movement of other performers by at least a threshold value; determining that the first performer is a prominent performer based on the amount of movement. method. Aspect 36 The method of any one of aspects 31 to 34, wherein editing the live video data and the live audio data includes framing one or more prominent performers by zooming in or cropping onto one or more images of the one or more prominent performers. Aspect 37 37. The method of embodiment 36, wherein editing the live video data and the live audio data includes tracking movements of at least one performer. Aspect 38 Editing the live video data and the live audio data includes: determining a pace and tempo of the Event based on said live audio data; cutting said live video data according to said pace and tempo. The method of embodiment 37. Aspect 39 Editing the live video data and the live audio data includes: determining when at least one performer has started or stopped performing; cutting said live video data in response. The method of embodiment 38. Aspect 40 Editing the live video data and the live audio data includes: cutting the live video data in response to determining from the live video data that a time that has elapsed since a full frame shot showing all performers exceeds a threshold time; The threshold time is determined based on the number or duration of cuts. 40. The method of any one of embodiments 31 to 39. Aspect 41 Editing the live video data and the live audio data includes: determining that a time that has elapsed since at least one performer appeared in the edited live video data is less than a threshold time, the threshold time being determined based on a number or duration of cuts; and specifying that the at least one performer will not be shown in a next edited cut while the elapsed time remains below the threshold time. The method of any one of embodiments 31 to 40. Aspect 42 42. The method of any one of aspects 31 to 41, comprising providing edited live video data and edited live audio data for streaming to a user device. Aspect 43 43. The method of any one of aspects 31 to 42, wherein the editing is performed on the server system. Aspect 44 44. The method of any one of aspects 31-43, wherein the editing comprises providing instructions to a recording device, the instructions operable to cause the device to perform the editing operation. Aspect 45 receiving, by a server system, audio data for an event and frames of video data for the event from one or more audio capture devices, the video data being captured by video capture devices configured to record video at a first resolution; determining a respective location of each individual performer of the event based on the audio data and images of each performer appearing in the frames of the video data; instructing the video capture device to submit a first portion of the video data to the server system at a second resolution upon determining by the server system from the audio data that a first one of the individual performers is the prominent performer at a first time, the first portion of the video data being spatially oriented to a location of the first performer captured at the first time; instructing the video recorder to submit a second portion of the video data to the server system at the second resolution upon determining by the server system from the audio data that a second one of the individual performers is the prominent performer at a second time, the second portion of the video data being spatially oriented to a location of the second performer captured at the second time; designating the first portion and the second portion of the video data as a video of the event at a second resolution; providing the audio data and the associated video of the event at the second resolution to a storage device or a user device as an audio and video recording of the event; the server system includes one or more computer processors; method. Aspect 46 46. ​​The method of embodiment 45, wherein the first resolution is 4K or greater and the second resolution is 1080p or less. Aspect 47 The method of any one of aspects 45 or 46, wherein the frame of video data of the event is a frame of a series of frames of the video data, the series of frames being received at the server system at a frame rate lower than the frame capture rate of the video capture device. Aspect 48 48. The method of any one of aspects 45 to 47, wherein determining the respective positions of each individual performer of the event is based on rehearsal data. Aspect 49 A method as described in any one of aspects 45 to 48, wherein the video capture device buffers a period of the video data at the first resolution and, in response to a command from the server system, selects locations of frames of the buffered video data corresponding to the first performer and the second performer for submission to the server system. Aspect 50 applying a delay to video of the event at the second resolution; and streaming the delayed video and associated audio data to one or more user devices. 50. The method of any one of embodiments 45 to 49. Aspect 51 recording, by a video capture device, video data of the event at a first resolution; storing the video data in a local buffer of the video capture device; determining one or more sequences of images from the recorded video data; submitting said sequence of one or more images to a server system at a first frame rate; receiving an instruction from the server system to focus on a portion of the video data, the instruction indicating a temporal and spatial location of the portion within the recorded video data; in response to the command, converting the portion of the video data stored in the local buffer according to the indicated temporal and spatial locations into video data at a second resolution having a second frame rate higher than the first frame rate; and submitting the converted video data at the second resolution to the server as live video data of the event. method. Aspect 52 52. The method of embodiment 51, wherein the first resolution is 4K or greater, the second resolution is 1080p or less, the first frame rate is 1 frame per second or less, and the second frame rate is 24 frames per second or greater. Aspect 53 one or more processors; and a non-transitory computer-readable medium storing instructions, The instructions, when executed by the one or more processors, cause the one or more processors to perform operations including those described in any one of aspects 1 to 52. system. Aspect 54 A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including those described in any one of aspects 1 to 52.

Claims

1. A method of producing a rehearsal event comprising: receiving, by a server system, one or more rehearsal channel signals captured during a rehearsal event, the one or more rehearsal channel signals including one or more rehearsal source signals from one or more audio sources playing individually at the rehearsal event and video data capturing one or more performers; automatically estimating, by the server system, a volume level of each rehearsal source signal; receiving, by the server system, one or more performance channel signals captured during a performance event, the one or more performance channel signals including one or more performance source signals from the one or more sources playing together at the performance event; mixing, by the server system, the one or more performance channel signals, the mixing including automatically leveling the one or more performance source signals based on an estimated volume level of the rehearsal source signal; receiving, by the server system, from a video capture device, a cropped or cropped and resized portion of the video data indicative of a prominent performer, the prominent performer being determined at least in part based on the one or more rehearsal channel signals or one or more performance channel signals; providing said cropped or cropped and resized portions of the mixed one or more performance channel signals and video data from said server system to a storage device or a plurality of end user devices. method.

2. The method further includes obtaining, by the server system, information regarding which performance source signals are associated with prominent performers, and using the information to determine gains for adjusting volume levels of the performance source signals to respective target volume levels during a live performance event. The method of claim 1.

3. The method of claim 1 , wherein the first resolution is 4K or greater and the second resolution is 1080p or less.

4. 2. The method of claim 1, wherein the frame of video data of the event is a frame of a series of frames of the video data, the series of frames being received at the server system at a frame rate that is lower than a frame capture rate of the video capture device.

5. Determining that the first performer or the second performer is a prominent performer further comprises: determining a difference between an energy or loudness level associated with the first performer or the second performer and an energy or loudness level associated with other performers at the event; determining whether the first performer or the second performer is a prominent performer based on a difference between the energy or loudness levels. The method of claim 1.

6. Determining that the first performer or the second performer is a prominent performer further comprises: determining a difference between an energy or loudness level associated with the first performer or the second performer and an energy or loudness level of the first performer or the second performer captured during a rehearsal phase of the event; determining whether the first performer or the second performer is a prominent performer based on a difference between the energy or loudness levels. The method of claim 1.

7. The method of claim 1 , wherein the pan position is determined using face or instrument tracking.