Audio-visual analysis for object rendering in capture
By classifying and processing audio and visual frames independently, the method improves the quality of audiovisual content by enhancing clarity and immersion.
Patent Information
- Application Number
- JP2025515522
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-03
- Filing Date
- 2023-09-12
- Publication Date
- 2025-10-15
AI Technical Summary
Existing audiovisual content processing systems fail to effectively separate and process audio and visual frames independently, leading to suboptimal quality in the final output.
A method and system that analyzes and classifies audio and visual frames separately, applying distinct processing operations based on their classifications to enhance audio clarity, immersion, and spatiality, and then combines them for improved audiovisual output.
Enhances the quality of audiovisual content by optimizing audio and visual processing based on their respective classifications, resulting in improved clarity and immersion.
Smart Images

Figure 2025534236000001_ABST
Abstract
Description
[Technical Field]
[0001] 1. Related Applications This application claims the benefit of priority to PCT Patent Application Publication No. PCT / CN2022 / 118437, filed September 13, 2022, and U.S. Provisional Patent Application No. 63 / 449,726, filed March 3, 2023, which are incorporated herein by reference in their entireties.
[0002] 2. Related fields Various exemplary embodiments relate generally to media processing of multimedia content. In particular, exemplary embodiments relate to a system, method, or computer program product configured to generate automated audiovisual analytics for processing and rendering. Summary of the Invention
[0003] Various embodiments are disclosed herein for rendering objects captured within audiovisual data. Audiovisual content is divided into visual frames (e.g., image frames, video frames) and audio frames. The audio frames are analyzed to identify audio objects within the audio frames. The visual frames are analyzed to identify visual objects within the visual frames. The audio and visual objects are classified based on the detected objects and scenes. The classification may indicate, for example, whether the audiovisual content is indoors or outdoors, whether the captured content includes sports, people, landscapes, or may further indicate a particular object type, etc. The audio and visual frames are processed separately using the detected objects and classifications and then recombined to create a final audiovisual output.
[0004] Embodiments, examples, and aspects of the present disclosure provide a method for classifying and categorizing objects based on their contribution to overall audio clarity, immersion, and spatiality. By classifying and categorizing both audio and visual frames, the audio and visual aspects of the audiovisual content can be processed separately, improving the quality of the final audiovisual output.
[0005] According to an exemplary embodiment, a method for processing audiovisual content is provided. The method includes receiving content including a plurality of audio frames and a plurality of video frames, classifying each of the plurality of audio frames into a plurality of audio classifications, and classifying each of the plurality of video frames into a plurality of video classifications. The method includes processing the plurality of audio frames based on the respective audio classifications and processing the plurality of video frames based on the respective video classifications, where each audio classification is processed with a different audio processing operation and each video classification is processed with a different video processing operation. The method includes generating an audio / video representation of the content by merging the processed plurality of audio frames and the processed plurality of video frames.
[0006] According to another exemplary embodiment, a non-transitory computer-readable medium is provided that stores instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including: receiving content including a plurality of audio frames and a plurality of video frames; classifying each of the plurality of audio frames into a plurality of audio classifications; and classifying each of the plurality of video frames into a plurality of video classifications. The instructions include processing the plurality of audio frames based on the respective audio classifications and processing the plurality of video frames based on the respective video classifications. Each audio classification is processed with a different audio processing operation, and each video classification is processed with a different video processing operation. The instructions include generating an audio / video representation of the content by merging the processed plurality of audio frames and the processed plurality of video frames.
[0007] According to yet another exemplary embodiment, a video system for processing audiovisual content is provided. The system includes a processor for performing processing of the audiovisual content. The processor is configured to receive content including a plurality of audio frames and a plurality of video frames, and classify each of the plurality of audio frames into a plurality of audio classifications and classify each of the plurality of video frames into a plurality of video classifications. The processor is configured to process the plurality of audio frames based on the respective audio classifications and process the plurality of video frames based on the respective video classifications. Each audio classification is processed with a different audio processing operation, and each video classification is processed with a different video processing operation. The processor is configured to generate an audio / video representation of the content by merging the processed plurality of audio frames and the processed plurality of video frames. [Brief explanation of the drawings]
[0008] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings.
[0009] [Figure 1] 1 illustrates an exemplary process for a video / image delivery pipeline.
[0010] [Figure 2] 1 illustrates an example of an audiovisual analysis-based rendering system.
[0011] [Figure 3] 1 illustrates an example of a visual analytics system.
[0012] [Figure 4] 1 shows an exemplary video frame containing noise.
[0013] [Figure 5] 5 shows the example video frame of FIG. 4 after resizing.
[0014] [Figure 6] 1 shows a block diagram of an exemplary method for detecting a scene change.
[0015] [Figure 7A] An exemplary 3x3 pixel neighborhood is shown. [Figure 7B] An exemplary 3x3 pixel neighborhood is shown. [Figure 7C] An exemplary 3x3 pixel neighborhood is shown.
[0016] [Figure 8] 4 shows another block diagram of an exemplary method for detecting a scene change.
[0017] [Figure 9] 1 shows a block diagram of an exemplary method for processing audiovisual content. DETAILED DESCRIPTION OF THE INVENTION
[0018] The present disclosure and aspects thereof may be embodied in various forms, including hardware, devices, or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces, and application programming interfaces, as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The foregoing is intended merely to give an overall idea of various aspects of the present disclosure and is not intended to limit the scope of the present disclosure in any way.
[0019] <Example Video / Image Delivery Pipeline> 1 illustrates an exemplary process of a video delivery pipeline 100 showing various stages from video / image capture to video / image content display, according to an embodiment. A sequence of video / image frames 102 may be captured or generated using an image generation block 105. The video frames 102 may be captured digitally (e.g., by a digital camera, by a phone camera, etc.) or generated computer-generated (e.g., using computer animation) to provide video and / or image data 107. Audio data may be provided as part of the data 107 and associated with the video / image data. Alternatively, the frames 102 may be captured on film by a film camera. The film is then converted to a digital format to provide the video / image data 107.
[0020] Audio data may include channel-based audio (e.g., stereo, 5.1, 7.1, etc.), which assigns sound sources to specific channels, object-based audio, which allows sound sources to be assigned to specific channels, and any associated metadata. For example, an audio object may include one or more audio signals (e.g., a stream of data encoding audio data or audio essence, also referred to herein as an "audio object signal") and associated metadata. The metadata may describe one or more characteristics of the audio signal (e.g., information metadata) or indicate how the audio signal should be processed by downstream processes, such as rendering (e.g., control metadata).
[0021] In some embodiments, the metadata includes audio object position data, audio object size data, audio object gain data, audio object trajectory data, content type data (dialogue, effects, etc.), rendering constraint data, etc. Metadata corresponding to individual audio sources among the audio signals, also referred to as parametric source descriptions, includes a spatial audio description of each source (e.g., source position / 3D coordinates and source size).
[0022] Some audio objects may be static, while other audio objects may have time-varying metadata, which may move, change size, and / or have other time-varying characteristics.
[0023] When an audio object is monitored or played in a playback environment, the audio object may be rendered according to the positional metadata using playback speakers present in the playback environment, rather than being output to a predetermined physical channel, as is the case in conventional channel-based systems such as Dolby 5.1 and Dolby 7.1.
[0024] Some metadata associated with the audio objects may be received along with the audio data, but additional metadata may be generated during the audio-visual analysis process described herein.
[0025] In the production stage 110, the data 107 may be processed by a processor in the production stage 110 to provide a viewable video / image production stream 112. The data in the video / image production stream 112 may then be provided to a processor (or one or more processors, such as a central processing unit (CPU)) in a post-production block 115 for post-production editing. The post-production editing may be performed by a user of the video distribution pipeline 100, such as the creator who captured the frames 102. The post-production editing in block 115 may include, for example, intermediate adjustments or changes to the color or brightness of specific areas of the image to improve image quality or achieve a particular look according to the video producer's creative intent. This part of the post-production editing is sometimes referred to as "color timing" or "color grading." Other edits (e.g., scene selection and ordering, image cropping, adding computer-generated visual special effects, removing artifacts, etc.) may be performed in block 115 to generate a final version of the production 117 for distribution. In some examples, the operations performed in block 115 include detecting and classifying objects within data 107. During post-production editing 115, the video and / or images are viewed on a reference display 125.
[0026] Following post-production 115, the final version of data 117 may be delivered to a coding block 120 for further downstream delivery to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, coding block 120 may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate a coded bitstream 122. At the receiver, coded bitstream 122 is decoded by a decoding unit 130 to generate a corresponding decoded signal 132 that represents a copy or close approximation of signal 117. The receiver may be attached to a target display 140, which may have some or all different characteristics from the reference display 125. In that case, a display management (DM) block 135 may be used to map the decoded signal 132 to the characteristics of the target display 140 by generating a display-mapped signal 137. According to an embodiment, both the decoding unit 130 and the display management block 135 may include individual processors or may be integrated into a single processing unit.
[0027] The codecs used in coding block 120 and / or decoding block 130 enable video / image data processing and compression / decompression. Compression is used in coding block 120 to make the corresponding file or stream smaller. The decoding process performed by decoding block 130 typically involves decompressing the received video / image data file or stream into a usable format for playback and / or further editing. Examples of coding / decoding operations that can be used in coding block 120 and decoding unit 130 according to various embodiments are described in more detail below.
[0028] Audio-Visual Rendering System 2 shows a block diagram of an audiovisual analysis-based object rendering system 200. The described operations of the audiovisual analysis-based object rendering system 200 can be performed by an electronic processor in the post-production block 115. The audiovisual video input is divided into visual frames and audio frames by a visual frame extraction block 202 and an audio frame extraction block 204, respectively. A visual frame in this context includes an image frame or a video frame captured by a camera. An audio frame in this context includes audio data captured by a microphone associated with the visual frame (e.g., captured contemporaneously with the visual frame).
[0029] The visual frames are provided to a visual scene / object classifier block 206 for visual scene and / or object classification, and the audio frames are provided to an audio scene / object classifier block 208 for audio scene and / or object classification.
[0030] A scene class can include multiple different classifiers, each with a different purpose. For example, one scene classifier can be used to analyze a captured location, such as an outdoor location, an indoor location, or a type of transportation. Another scene classifier can be used to distinguish between captured content types, such as sports, food, landscapes, people, etc. A classifier can also perform object localization and segmentation. The output classes resulting from the classification of the audio and visual classifiers may be the same classifier, different classifiers, or related classifiers. For example, for a birdsong object, the audio class may be "birdsong," while the visual classes may be "tree," "bird," "birdcage," etc.
[0031] In some cases, in addition to classification, audio objects are assigned to four major categories based on their perceptual importance to the audio. Categories may include "essential objects" (e.g., Category 1), "high importance objects" (e.g., Category 2), "important objects" (e.g., Category 3), and "low importance objects" (e.g., Category 4). Objects assigned as "essential objects" represent objects that contribute to the intelligibility and spatiality of the audio data. For example, detected sounds may be classified as "essential objects," or any object that provides height information that may increase the sense of space in the audio. Objects assigned as "high importance objects" include non-sounds that contribute to audio immersion, such as animal sounds or directional objects with distinct time-frequency patterns. Objects assigned as "important objects" include background sounds that contain environmental information about the scene. These objects may be attenuated to increase the intelligibility of the audio, with only a minor impact on the immersion of the audio scene. Objects assigned as "low importance objects" include environmental sounds that may be highly attenuated to improve audio intelligibility. Thus, audio objects in one category (e.g., "important objects") may be attenuated at a first level and audio objects in a second category (e.g., "less important objects") may be attenuated at a second level greater than the first level, or vice versa. Although specific categories are provided, these categories are merely examples. Fewer or more categories may be provided to provide an order of importance for the classified audio objects.
[0032] In some embodiments, the classification of objects is performed automatically by an electronic processor. For example, the electronic processor may receive audio data and implement a machine learning model configured to classify the audio data. Each classification may be assigned a particular category level that the electronic processor automatically associates with the audio object. For example, if the electronic processor detects that an audio component of the audio data is "speech," it classifies the audio object as speech and automatically classifies the audio object as an "essential object." In another embodiment, the electronic processor receives input from an editor of audio frames indicating the classification and categorization of the audio frames.
[0033] The visual output classes and audio output classes are combined in an audio-visual combination block 210. The combination of the visual output classes and audio output classes may be deep learning based, rule-based, etc. The combined audio-visual information is used as context information for further audio-visual processing, which will be described in more detail below. The combined audio-visual information, like the original visual frames, is provided to a visual processing block 212 for visual processing.
[0034] The combined audio-visual information, as well as the original audio frames, is provided to an audio object separation block 214, which separates the objects classified by the audio scene / object classifier block 208 from the original audio frames. The separated audio objects, as well as the combined audio-visual information, are provided to an object selection / metadata generation block 216, which generates metadata describing the different audio objects. An audio rendering block 218 renders the audio objects.
[0035] The category of an audio object can influence the processing of audio frames in the audio object separation block 214, the object selection / metadata generation block 216, and / or the audio rendering block 218. For example, audio data may be assigned to a particular speaker based on its classification. Audio data associated with "high importance objects" may be expanded or contracted to increase the sense of space. Audio data associated with "low importance objects" may be suppressed to improve overall audio intelligibility while simultaneously preserving the audio objects of higher importance.
[0036] The processed video from the visual processing block 212 and the rendered audio from the audio rendering block 218 are fed to an audio-visual synthesis block 220. The audio-visual synthesis block 220 combines the rendered audio and the processed visuals to generate the video output.
[0037] 3 shows a block diagram of an example visual analytics system 300. The operations described with respect to the visual analytics system 300 may be performed by the visual scene / object classifier block 206, the audiovisual combining block 210, the visual processing block 212, or a combination thereof.
[0038] The visual frames from the visual frame extraction block 202 are received by the color diversity detection block 302. The color diversity detection block 302 is configured to filter the received video data and remove frames that contain only a single color or frames where the color diversity is below a set threshold. In this way, the color diversity detection block 302 removes frames that were captured when, for example, the camera lens was blocked or when the camera was too close to an object, resulting in loss of contour information about the environment or the object.
[0039] In one example of finding color distance within a visual frame, the color distance between two RGB pixels is calculated according to Equation 1:
number
number
[0040] Because humans have different sensitivities to different RGB color values, we can weight different color components differently. Thus, Equation 1 can be rewritten as:
number
[0041] In one implementation, the weights are w R =2, w G = 4, and w B = 3 because humans are more sensitive to changes in the green color channel and less sensitive to changes in the red color channel. R , w G , and w B will change based on different color types. For example, consider a situation where the weights are defined as follows:
number
number
[0042] In Equation 4, the green color channel has a fixed weight, and the red and blue color channels have dynamic weights based on the value of the red color channel. For example, if the value of the red color channel is close to 255, the red color channel gets the largest weight (w R =~3), and if the red color channel value is close to 0, the blue color channel gets the maximum weight (wB=~3).
[0043] In some cases, calculating the weighting function for all pixels in a visual frame can be very computationally demanding. Therefore, to reduce the computational demands, color distances may be calculated for a limited set of primary colors. A statistical analysis may be performed on the image to calculate the most represented colors in the image. In this example, the color diversity of a visual frame is given by Equation 4:
number
number
[0044] If a visual frame is characterized by noise, the most commonly represented color may be erroneously identified. For example, FIG. 4 provides an example of a video frame 400 having a scattered color (indicated by white dots). The video frame has a black background, which is selected as the most commonly represented color. The video frame 400 has a cr value of approximately 150 due to the white spots, even though the video frame 400 has no color diversity. To overcome this error, the video frame 400 may be resized to a smaller resolution, as shown as video frame 500 in FIG. 5. A smaller resolution reduces the impact of noise. The video frame 500 has a cr value of approximately 43, effectively discarding the noisy frame. The output of the color diversity detection block 302 may be an indication of whether the color diversity of the visual frame cr is equal to or exceeds a color diversity threshold.
[0045] Returning to FIG. 3 , the output of color diversity detection block 302 is provided to main object detection block 304 and scene change detection block 306. Main object detection block 304 is configured to identify the location of a main object of interest within the visual frame. For example, face and / or body detection methods may be performed to segment each person within the visual frame and estimate each person's position within the image coordinate system of the visual frame, the orientation of their face relative to the camera reference system, the distance of their face from the camera, etc. Main objects are not limited to just faces and bodies, but may also include objects such as animals, plants, buildings, or other subjects in the video frame. Main object detection block 304 outputs an indication of the main object and data associated with the main object.
[0046] In examples where the primary object is a person's head, the detected head position and head pose can be used as input for audio processing. For example, the size of a face in a visual frame may be used to control the volume of an audio object associated with that face in the audio mix (e.g., audio related to the person's speech). As the face moves closer to the camera in a second visual frame, the volume of the associated audio object increases for the second visual frame. As the face moves farther away from the camera in the second visual frame, the volume of the associated audio object decreases for the second visual frame. In some examples, the position of the face in the visual frame is referenced when spatializing the associated audio object. The face orientation is also referenced for speech spatialization and reverberation. This concept can be applied to other video objects. As a video object moves in the video from frame to frame, the image coordinates of the video object are used to spatialize the corresponding audio object.
[0047] The scene change detection block 306 determines a change in an object of interest (e.g., a scene change) between frames. For example, a first visual frame may have a first object of interest (e.g., a person), and a second visual frame may have a second object of interest (e.g., an animal). The scene change detection block 306 outputs an indication of the change in the object of interest.
[0048] There are three examples of scene change detection algorithms: gray value-based, edge contour-based, and motion-based. Gray value-based scene detection algorithms determine whether a scene change has occurred by comparing the difference between the gray values of a reference frame image and the current frame. For example, they compare the difference in the gray histograms of the two frames. Gray value-based scene detection algorithms have relatively low computational complexity, but can be prone to errors when video objects move between frames. Edge contour-based scene detection methods compare the edge contours of corresponding objects between the current and reference frames. This algorithm effectively detects soft handoffs such as ablation, fade-in, and fade-out. However, for detecting sudden scene handoffs, they offer no clear advantage over gray value-based detection methods and require increased computational complexity. Motion-based scene change detection algorithms detect discontinuities in the movement of video objects before and after a scene change. The average residual error of motion estimation is used as a decision criterion for scene changes. However, reliable motion-based scene change algorithms often have high computational complexity.
[0049] To mitigate these problems, the embodiments described herein provide a hybrid scene change detection algorithm that uses the luminance (Y), chrominance (U), and chroma (V) of a visual frame. Figure 6 provides an example method 600 for detecting scene changes performed in the scene change detection block 306. In step 602, the method 600 includes converting an RGB color frame to a YUV frame.
[0050] In step 604, method 600 includes transforming the Y component into a feature frame using binary weighting. For example, FIG. 7A shows an exemplary 3×3 pixel neighborhood centered on c(x,y). The position of each neighboring pixel is “p” and its corresponding value is denoted as g(p). In FIG. 7A, the position p of each pixel is labeled clockwise from 0 to 7, i.e., g(0)=7, g(1)=3, g(2)=1, g(3)=2, g(4)=4, g(5)=7, g(6)=9, and g(7)=8. FIG. 7B shows a binary value map obtained by comparing g(p) and g(c). FIG. 7C shows the decimal value of each neighboring pixel. The binary value map in FIG. 7B is obtained according to Equation 6.
number
[0051] In step 606, the method 600 includes generating a histogram of the feature frame. For example, Equation 7 provides for calculating the sum of absolute differences between the current frame t and the previous frame t-1.
number
[0052] In step 608, the method 600 includes determining whether a scene change occurs based on the histogram. For example, a threshold for detecting a scene change is set as provided by Equation 8.
number
[0053] For the first frame, C1(t) may be set to 1. The value C1(t) may be an output of the scene change detection block 306. Although the method 600 is described as calculating the difference between the current frame and a previous frame, the method 600 may instead calculate the difference between the current frame and a future frame.
[0054] 8 provides another exemplary method 800 for detecting a scene change. In step 802, the method 800 includes converting an RGB color frame to a YUV frame. In step 804, the method 800 includes calculating an average YUV value for the current frame. For example, the average YUV value for the current frame is determined according to Equations 9-11:
number
[0055] In step 806, the method 800 includes calculating the difference in average YUV values between the current frame and the subsequent frame. For example, the difference between frame t and frame t−1 can be calculated using Equation 12:
number
[0056] In step 808, the method 800 includes determining whether a scene change occurs based on the difference. For example, a threshold for detecting a scene change is set as provided by Equation 13.
number
[0057] The value C2(t) may be the output of the scene change detection block 306. Although the method 800 is described as calculating the difference between a current frame and a previous frame, the method 800 may instead calculate the difference between a current frame and a future frame.
[0058] In another example, multiple past or future frames can be analyzed to make scene change detection more robust. In this example, the average YUV values of k past frames are provided as follows:
number
number
[0059] In this example, the threshold can be set as provided by Equation 18.
number
[0060] In some implementations, a scene change is detected using a combination of parameters C1(t), C2(t), and C3(t) and their corresponding thresholds. For example, a scene change may be detected when C1(t) = 1, C2(t) = 1, and C3(t) = 1. In another example, a scene change is detected when at least two of parameters C1(t), C2(t), and C3(t) are 1. In this implementation, the scene change detection block 306 can effectively detect soft handoffs such as ablation, fade-in, and fade-out, as well as sudden scene handoff detection with minimal additional computational complexity. An example of potential thresholds for scene detection is th1 = 2500, th2 = 80, and th3 = 0.3.
[0061] The outputs of the main object detection block 304 and the scene change detection block 306 (e.g., whether a scene change has occurred) are provided to a scene / object classifier block 308. In some examples, the output of the color diversity detection block 302 is also provided to the scene / object classifier block 308. The scene / object classifier block 308 identifies and classifies the scene content, the type of content, and / or the primary video object in the scene captured in the visual frame. In some examples, the frequency at which the scene / object classifier block 308 operates depends on the frequency of scene changes. For example, the scene / object classifier block 308 may classify the content of the scene in response to the occurrence of a scene change. In another example, if a scene change is not detected, the scene / object classifier block 308 operates at set time intervals.
[0062] In some examples, the classification result for each visual frame or scene is weighted according to the analysis associated with the scene. For example, the color diversity values determined by the color diversity detection block 302 may be provided as weights to the scene / object classifier block 308. Scenes with high color diversity may be weighted more heavily than scenes with relatively low color diversity.
[0063] Thus, the audio-visual analysis based object rendering system 200 provides a system and process for identifying and classifying detected objects in both visual and audio frames. In some implementations, visual classes are defined to match audio scenes or objects. Visual classes are mapped to audio classes to create linked classes that define both visual and audio frames. For example, speech audio may be linked with detected faces.
[0064] The visual processing block 310 receives the classified visual frames and renders video based on the object classification. Different classification results may result in different processing strategies. For example, the video processing block 212 and the audio rendering block 218 process video and audio content for different content types, different scenes, and different objects independently. For example, automatic zooming may be performed on specific visual objects based on their classification. Different audio processing, such as leveling, equalization, and spatialization operations, may be performed on different audio objects based on their classification or categorization. Audio objects classified as "essential objects" may have their absolute or relative levels boosted (or increased), while audio objects classified as "low importance objects" may have their absolute or relative levels attenuated (or decreased). Boosting and attenuating audio objects includes adjusting the gain associated with the audio objects. In some embodiments, boosting and attenuating audio objects includes using other forms of audio processing (such as dialogue enhancement or application of filters) to increase or decrease intelligibility. Furthermore, audio and visual objects may be processed based on scene / object analysis performed in visual scene / object classifier block 206 and audio scene / object classifier block 208. For example, if a visual frame is classified as a forest scene, audio objects related to insect and bird sounds may be added or augmented to improve the sense of immersion.
[0065] 9 provides an example method 900 for processing audiovisual content. Various steps described herein with respect to method 900 may be performed simultaneously, in parallel, or in an order different from the illustrated sequential and iterative execution manner. Method 900 may be performed by post-production block 115. At block 902, method 900 includes receiving content including a plurality of audio frames and a plurality of video frames. For example, audio-visual analysis-based object rendering system 200 receives a video input. A plurality of audio frames are extracted from the video input by audio frame extraction block 204. A plurality of video frames are extracted by visual frame extraction block 202.
[0066] At block 904, the method 900 includes classifying each of a plurality of audio frames into a plurality of audio classes. For example, the audio scene / object classifier block 208 receives the audio frames from the audio frame extraction block 204 and classifies the audio frames as audio objects. The classification of the audio frames can be performed by the audio scene / object classifier block 208 in conjunction with the audio object separation block 214 and / or the object selection / metadata generation block 216. Each audio frame includes metadata indicating a classification of a detected audio object. At block 906, the method 900 includes classifying each of a plurality of video frames into a plurality of video classes. For example, the visual scene / object classifier block 206 receives the video frames from the visual frame extraction block 202 and classifies an object (e.g., a detected main object) in the video frame. Each video frame includes metadata indicating a classification of a detected visual object.
[0067] At block 908, the method 900 includes processing the audio frames based on their respective audio classifications. For example, the audio rendering block 218 processes the audio frames based on the classification and categorization of the audio frames. In some implementations, each audio frame is processed with a different audio processing operation. In other implementations, some audio frames may be processed with the same audio processing operation, while other audio frames are processed with different audio processing operations. For example, a first set of audio frames may be processed with a first audio processing operation, a second set of audio frames may be processed with a second audio processing operation, and a third set of audio frames may be processed with a third audio processing operation.
[0068] At block 910, the method 900 includes processing the plurality of video frames based on their respective video classifications. For example, the visual processing block 212 processes the video frames based on the classification of visual objects within the video frames. In some implementations, each video frame is processed with a different video processing operation. In other implementations, some video frames may be processed with the same video processing operation, while other video frames are processed with different video processing operations. For example, a first set of video frames is processed with a first video processing operation, a second set of video frames is processed with a second video processing operation, and a third set of video frames is processed with a third video processing operation.
[0069] At block 912, the method 900 includes generating an audio / video representation of the content. For example, the audio / visual synthesis block 220 generates a video output by merging the processed audio frames from the audio rendering block 218 and the processed video frames from the visual processing block 212.
[0070] Systems, methods and devices according to the present disclosure may incorporate one or more of the following features.
[0071] (1) A method for processing audiovisual content, comprising: receiving content including a plurality of audio frames and a plurality of video frames; classifying each of the plurality of audio frames into a plurality of audio classes; classifying each of the plurality of video frames into a plurality of video classifications; processing the plurality of audio frames based on respective audio classifications, wherein each audio classification is processed with a different audio processing operation; processing the plurality of video frames based on each video classification, wherein each video classification is processed with a different video processing operation; generating an audio / video representation of the content by merging the processed audio frames and the processed video frames; A method comprising:
[0072] (2) classifying each of the plurality of audio classifications into one of a plurality of priority categories; 10. The method of claim 1, wherein processing the plurality of audio frames includes processing the plurality of audio frames based on each priority category.
[0073] (3) the plurality of priority categories include a first category and a second category, and the first category indicates a higher priority than the second category; The step of processing the plurality of audio frames includes: boosting audio frames classified as said first category; attenuating audio frames classified as said second category; The method according to (2), comprising performing at least one selected from the group consisting of:
[0074] (4) The method of (3), wherein the first category includes conversation objects.
[0075] (5) The method according to (4), wherein the first category further includes objects having height information.
[0076] (6) The method according to any one of (4) to (5), wherein the second category does not include a conversation object.
[0077] (7) the plurality of priority categories includes a third category indicating a lower priority than the first category; A method according to any one of (3) to (6), wherein the step of processing the plurality of audio frames includes attenuating audio frames classified into the third category with a different level of attenuation than audio frames classified into the second category.
[0078] (8) extracting the audio frames from the content and separating the audio frames from the video frames; The method according to any one of (1) to (7), further comprising:
[0079] (9) for each video frame, determining the color diversity of the color frame; comparing, for each video frame, the color diversity to a color diversity threshold; for each video frame, in response to the color diversity being less than the color diversity threshold, discarding the video frame; The method according to any one of (1) to (8), further comprising:
[0080] (10) for each video frame, determining a weight value for the video frame based on the color diversity of the video frame; 10. The method of claim 9, wherein each of the plurality of video frames is classified based on the weight value.
[0081] (11) further comprising the step of determining whether a scene change occurs between the current video frame and the subsequent video frame; A method according to any one of (1) to (10), wherein the step of classifying each of the plurality of video frames into a plurality of video classifications includes a step of classifying the video frame in response to determining that a scene change has occurred for each video frame.
[0082] (12) The step of determining whether a scene change occurs comprises: converting the current frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent frame into a second YUV frame; generating a first histogram based on the first YUV frame; generating a second histogram based on the second YUV frame; determining whether the scene change occurs based on the first histogram and the second histogram; The method according to (11), comprising:
[0083] (13) The step of determining whether a scene change occurs based on the first histogram and the second histogram includes: calculating the sum of absolute differences between the first histogram and the second histogram; comparing the sum of absolute differences to a scene change threshold; The method according to (12), comprising:
[0084] (14) The step of determining whether a scene change occurs comprises: converting the current frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent frame into a second YUV frame; calculating a difference between a first average YUV of the first YUV frame and a second average YUV of the second YUV frame; determining whether the scene change occurs based on the difference between the first average YUV and the second average YUV; The method according to (11), comprising:
[0085] (15) The step of classifying each of the plurality of video frames into a plurality of video classifications includes: performing at least one of main object detection and scene detection to generate intermediate results; classifying the plurality of video frames based on the intermediate results; The method according to any one of (1) to (14), comprising:
[0086] (16) A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the method described in any one of (1) to (15).
[0087] (17) A system for processing audiovisual content, comprising: a processor for performing processing of audiovisual content, said processor comprising: receiving content including a plurality of audio frames and a plurality of video frames; classifying each of the plurality of audio frames into a plurality of audio classes; classifying each of the plurality of video frames into a plurality of video classifications; processing the plurality of audio frames based on each audio classification, each audio classification being processed with a different audio processing operation; processing the plurality of video frames based on each video classification, each video classification being processed with a different video processing operation; generating an audio / video representation of the content by merging the processed plurality of audio frames and the processed plurality of video frames; The system is configured as follows.
[0088] (18) The processor: further configured to classify each of the plurality of audio classifications into one of a plurality of priority categories; 18. The system of claim 17, wherein to process the plurality of audio frames, the processor is configured to process the plurality of audio frames based on each priority category.
[0089] (19) The plurality of priority categories include a first category and a second category, and the first category indicates a higher priority than the second category; To process the plurality of audio frames, the processor: boosting audio frames classified as the first category; attenuating audio frames classified as the second category; The system according to (18), configured as follows:
[0090] (20) A method according to any one of (17) to (19), wherein the processor is further configured to extract the plurality of audio frames from the content and separate the plurality of audio frames from the plurality of video frames.
[0091] (21) The processor: For each video frame, determine a color diversity of the color frame; for each video frame, comparing the color diversity to a color diversity threshold; for each video frame, in response to the color diversity being less than the color diversity threshold, discarding the video frame; The system according to any one of (17) to (20), further configured as follows:
[0092] (22) The processor is further configured to determine whether a scene change occurs between a current video frame and a subsequent video frame; The system of any one of (17) to (21), wherein the processor is configured to classify each of the plurality of video frames into a plurality of video classifications in response to determining that a scene change has occurred for each video frame.
[0093] (23) To determine whether a scene change occurs, the processor: converting the current frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent frame into a second YUV frame; generating a first histogram based on the first YUV frame; generating a second histogram based on the second YUV frame; determining whether the scene change occurs based on the first histogram and the second histogram; The system according to (22), configured as follows:
[0094] (24) To determine whether a scene change occurs based on the first histogram and the second histogram, the processor: calculating the sum of absolute differences between the first histogram and the second histogram; comparing the sum of absolute differences to a scene change threshold; The system according to (23), configured as follows:
[0095] (25) To determine whether a scene change occurs, the processor: converting the current frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent frame into a second YUV frame; calculating a difference between a first average YUV of the first YUV frame and a second average YUV of the second YUV frame; determining whether the scene change occurs based on the difference between the first average YUV and the second average YUV; The system according to (22), configured as follows:
[0096] While described herein with respect to processes, systems, methods, heuristics, etc., it should be understood that although the steps of such processes, etc. have been described as occurring according to a particular ordered sequence, such processes may be practiced with the described steps performed in an order different from that described herein. It should further be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the description of processes herein is provided for the purpose of describing particular embodiments and should not be considered as limiting the claims.
[0097] Therefore, it should be understood that the above description is intended to be illustrative, and not limiting. Many embodiments and applications other than the examples provided will become apparent upon reading the above description. The scope should be determined not with reference to the above description, but instead with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technology discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is capable of modification and alteration.
[0098] All terms used in the claims are intended to be given their broadest reasonable construction and ordinary meaning as understood by those skilled in the art described herein. In particular, the use of singular articles such as "a," "the," and "said" should be read to refer to one or more of the indicated elements, unless the claim expressly states a limitation to the contrary.
[0099] The Abstract of the Disclosure is provided to enable the reader to quickly assess the nature of the technical disclosure. It is understood that it is not used to interpret or limit the scope or meaning of the claims. Moreover, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as separately claimed subject matter.
[0100] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0101] Aspects of the systems described herein may be implemented in any suitable computer-based audio processing network environment that processes digital or digitized audio files. Portions of the adaptive audio system may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0102] Some embodiments may be implemented in the form of methods and apparatuses for practicing these methods. Some embodiments may also be implemented in the form of program code recorded on tangible media, such as magnetic recording media, optical recording media, solid-state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage media, where, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention. Some embodiments may also be implemented in the form of program code stored on, for example, a non-transitory machine-readable storage medium, where, when the program code is loaded into and executed by a machine, such as a computer or processor, the machine becomes an apparatus for practicing the patented invention. When implemented on a general-purpose processor, program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.
[0103] Unless expressly stated otherwise, each numerical value and range should be interpreted as approximate, as if the word "about" or "approximately" precedes the value or range.
[0104] Use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate claim interpretation, and such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0105] Where present, elements in the method claims that follow are recited in a particular order by corresponding labels, but these elements are not necessarily intended to be limited to being performed in that particular order, unless a recitation of a claim specifically indicates a particular order for implementing some or all of these elements.
[0106] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present disclosure. The appearances of the phrase "in one embodiment" in various places in this specification do not necessarily all refer to the same embodiment, nor do separate or alternative embodiments necessarily mutually exclude other embodiments. The same applies to the term "implementation."
[0107] Unless otherwise specified herein, the use of the ordinal adjectives "first," "second," "third," etc. to refer to multiple instances of similar objects merely indicates that different instances of such similar objects are being referenced and does not imply that the similar objects so referenced must be in a corresponding order or sequence, whether in time, space, ranking, or otherwise.
[0108] Unless otherwise specified herein, the conjunction "if," in addition to its plain meaning, can be interpreted to mean "when," "upon," "in response to determining," or "in response to detecting," depending on the particular context to which it corresponds. For example, the phrases "if it is determined" or "if [a stated condition] is detected" can be interpreted to mean "upon determining," "in response to determining [the stated condition or event]," or "in response to determining [the stated condition or event]."
[0109] Also, for purposes of this specification, the terms "couple," "coupling," "coupled," "connect," "connecting," or "connected" refer to any manner known in the art or later developed by which energy may be transferred between two or more elements, where the interposition of one or more additional elements is contemplated, but not required. Conversely, the terms "directly coupled," "directly connected," etc., imply the absence of such additional elements.
[0110] The term compatible, as used herein with respect to elements and standards, means that the element communicates with other elements in a manner defined, wholly or partially, by the standard and is recognized by other elements as being sufficiently capable of communicating with other elements in a manner specified by the standard. A compatible element need not operate internally in a manner defined by the standard.
[0111] The functionality of the various elements illustrated in the figures, including any functional blocks labeled as a "processor" and / or a "controller," may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functionality may be provided by a single dedicated processor, a single shared processor, or multiple individual processors, some of which may be shared. Furthermore, explicit use of the terms "processor" or "controller" should not be construed to refer solely to hardware capable of executing software, but may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM), random access memory (RAM), and non-volatile storage for storing software. Other hardware, conventional and / or custom, may also be included. Similarly, any switches illustrated in the figures are conceptual. Their functionality may be performed by the operation of program logic, dedicated logic, the interaction of program control and dedicated logic, or manually. The particular technique is selectable by the implementer, as more specifically understood from the context.
[0112] As used herein, the term "circuit" may refer to one or more or all of the following: (a) a hardware-only circuit implementation (e.g., an implementation with only analog and / or digital circuitry); (b) a combination of hardware circuitry and software (where appropriate), including: (i) a combination of analog and / or digital hardware circuitry and software / firmware; and (ii) a hardware processor with software (including a digital signal processor), software, and any portion of memory that work together to cause a device, such as a cell phone or server, to perform various functions; and (c) a hardware circuit and / or processor, e.g., a microprocessor or portion of a microprocessor, that requires software (e.g., firmware) to operate, but may not be present if software is not required for operation. This definition of circuitry applies to all uses of the term in this application, including the claims. As a further example, as used herein, the term "circuit" also includes a mere hardware circuit or processor (or processors), or a portion of a hardware circuit or processor, and its (or their) accompanying software and / or firmware implementation. The term "circuitry" also includes, for example, a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, where applicable to certain claim elements.
[0113] Those skilled in the art should appreciate that the block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that flowcharts, flow diagrams, state transition diagrams, pseudocode, and the like represent various processes substantially embodied on a computer-readable medium and executed by a computer or processor, whether or not a computer or processor is explicitly shown.
[0114] This Summary is intended to introduce some exemplary embodiments; additional embodiments are described in the Detailed Description and / or with reference to one or more drawings. This Summary is not intended to identify key elements or features of the claimed subject matter or to limit the scope of the claimed subject matter.
[0115] While the present disclosure includes reference to exemplary embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, and other embodiments within the scope of the disclosure that are obvious to those skilled in the art to which the disclosure pertains, are deemed to be within the principles and scope of the disclosure, as set forth, for example, in the following claims.
Claims
1. 1. A method for processing audiovisual content, comprising: receiving content including a plurality of audio frames and a plurality of video frames; classifying each of the plurality of audio frames into a plurality of audio classes; classifying each of the plurality of video frames into a plurality of video classifications; processing the plurality of audio frames based on respective audio classifications, wherein each audio classification is processed with a different audio processing operation; processing the plurality of video frames based on each video classification, wherein each video classification is processed with a different video processing operation; generating an audio / video representation of the content by merging the processed audio frames and the processed video frames; A method comprising:
2. classifying each of the plurality of audio classifications into one of a plurality of priority categories; The method of claim 1 , wherein processing the plurality of audio frames comprises processing the plurality of audio frames based on respective priority categories.
3. the plurality of priority categories include a first category and a second category, the first category indicating a higher priority than the second category; The step of processing the plurality of audio frames includes: boosting audio frames classified as said first category; attenuating audio frames classified as the second category; 3. The method of claim 2, comprising performing at least one selected from the group comprising:
4. The method of claim 3 , wherein the first category includes conversation objects.
5. The method of claim 4 , wherein the first category further includes objects having height information.
6. The method of claim 4 , wherein the second category does not include conversation objects.
7. the plurality of priority categories includes a third category indicating a lower priority than the first category; 4. The method of claim 3, wherein processing the plurality of audio frames includes attenuating audio frames classified into the third category with a different level of attenuation than audio frames classified into the second category.
8. extracting the audio frames from the content to separate the audio frames from the video frames; The method of claim 1 further comprising:
9. determining, for each video frame, the color diversity of the color frame; comparing, for each video frame, the color diversity to a color diversity threshold; for each video frame, in response to the color diversity being less than the color diversity threshold, discarding the video frame; The method of claim 1 further comprising:
10. for each video frame, determining a weight value for the video frame based on the color diversity of the video frame; The method of claim 9 , wherein each of the plurality of video frames is classified based on the weight value.
11. determining whether a scene change occurs between the current video frame and the subsequent video frame; 2. The method of claim 1, wherein classifying each of the plurality of video frames into a plurality of video classifications comprises classifying the video frame in response to determining, for each video frame, that a scene change has occurred.
12. The step of determining whether a scene change occurs comprises: converting the current video frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent video frame into a second YUV frame; generating a first histogram based on the first YUV frame; generating a second histogram based on the second YUV frame; determining whether the scene change occurs based on the first histogram and the second histogram; The method of claim 11 , comprising:
13. The step of determining whether a scene change occurs based on the first histogram and the second histogram includes: calculating the sum of absolute differences between the first histogram and the second histogram; comparing the sum of absolute differences to a scene change threshold; 13. The method of claim 12, comprising:
14. The step of determining whether a scene change occurs comprises: converting the current video frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent video frame into a second YUV frame; calculating a difference between a first average YUV of the first YUV frame and a second average YUV of the second YUV frame; determining whether the scene change occurs based on the difference between the first average YUV and the second average YUV; The method of claim 11 , comprising:
15. The step of classifying each of the plurality of video frames into a plurality of video classifications comprises: performing at least one of main object detection and scene detection to generate intermediate results; classifying the plurality of video frames based on the intermediate results; The method of claim 1 , comprising:
16. A non-transitory computer readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the method of any one of claims 1 to 15.
17. 1. A system for processing audiovisual content, comprising: a processor for performing processing of audiovisual content, said processor comprising: receiving content including a plurality of audio frames and a plurality of video frames; classifying each of the plurality of audio frames into a plurality of audio classes; classifying each of the plurality of video frames into a plurality of video classifications; processing the plurality of audio frames based on each audio classification, each audio classification being processed with a different audio processing operation; processing the plurality of video frames based on each video classification, each video classification being processed with a different video processing operation; generating an audio / video representation of the content by merging the processed audio frames with the processed video frames; The system is configured as follows.
18. The processor: further configured to classify each of the plurality of audio classifications into one of a plurality of priority categories; 20. The system of claim 17, wherein to process the plurality of audio frames, the processor is configured to process the plurality of audio frames based on respective priority categories.
19. the plurality of priority categories include a first category and a second category, the first category indicating a higher priority than the second category; To process the plurality of audio frames, the processor: boosting audio frames classified as the first category; attenuating audio frames classified as the second category; The system of claim 18 configured to:
20. 20. The system of claim 17, wherein the processor is further configured to extract the plurality of audio frames from the content and separate the plurality of audio frames from the plurality of video frames.
21. The processor: For each video frame, determine a color diversity of the color frame; for each video frame, comparing the color diversity to a color diversity threshold; for each video frame, in response to the color diversity being less than the color diversity threshold, discarding the video frame; 20. The system of claim 17, further configured to:
22. the processor is further configured to determine whether a scene change occurs between a current video frame and a subsequent video frame; 20. The system of claim 17, wherein the processor is configured to classify each of the plurality of video frames into a plurality of video classifications in response to determining, for each video frame, that a scene change has occurred.
23. To determine whether a scene change occurs, the processor: converting the current video frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent video frame into a second YUV frame; generating a first histogram based on the first YUV frame; generating a second histogram based on the second YUV frame; determining whether the scene change occurs based on the first histogram and the second histogram; 23. The system of claim 22, configured to:
24. To determine whether the scene change occurs based on the first histogram and the second histogram, the processor: calculating the sum of absolute differences between the first histogram and the second histogram; comparing the sum of absolute differences to a scene change threshold; 24. The system of claim 23, configured to:
25. To determine whether a scene change occurs, the processor: converting the current video frame into a first YUV (luminance-chrominance-chrome) frame; converting the subsequent video frame into a second YUV frame; calculating a difference between a first average YUV of the first YUV frame and a second average YUV of the second YUV frame; determining whether the scene change occurs based on the difference between the first average YUV and the second average YUV; 23. The system of claim 22, configured to: