Audio scene analysis based on audio content type recognition
By framing, feature extraction and machine learning classification of audio data, and dynamically adjusting audio settings, the problem that traditional audio systems cannot adapt to multiple audio types is solved, and a better listening experience and silent segment processing is achieved.
Patent Information
- Application Number
- CN202380088436.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-30
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional audio systems cannot dynamically identify and adapt to different audio content types, especially when multiple audio elements are interwoven, resulting in poor listening experience, improper processing of silent segments, and limitations in reliance on metadata tags or user manual adjustments.
By receiving audio data, segmenting it into multiple audio frames, extracting features and classifying them, using machine learning algorithms to identify audio types, dynamically adjust audio settings, smoothly transitioning silent segments, and automatically adapting to audio scenes.
In the absence of metadata marking, accurate classification and appropriate adjustment of complex audio scenes are achieved, the listening experience is enhanced, the sudden changes in audio settings are avoided, and immersive listening is provided.
Smart Images

Figure CN120359567A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to co - pending Indian Patent Application No. 202241077075, filed on December 30, 2022, entitled "AUDIO SCENE ANALYSIS". The subject matter of this related application is hereby incorporated herein by reference. Technical Field
[0003] Embodiments of the present disclosure generally relate to audio engineering, and more particularly, to dynamic audio scene analysis based on audio content type recognition. Background Art
[0004] Advances in audio technology have been marked by a series of incremental improvements that have gradually enhanced user interaction and the auditory experience. Traditional audio systems have been mainly static, providing a uniform audio setup regardless of the type or nature of the audio content. This approach has limited flexibility and cannot accommodate the wide variety of audio that listeners encounter, from the subtle bass and treble of classical symphonies to the diverse sound effects of live sports events. In multimedia entertainment, a single stream can combine dialogue, sound effects, and musical scores, and these static systems lack the ability to maintain the clarity and intended effects of each audio element, often resulting in a diluted or unbalanced overall sound. This limitation is particularly evident in environments where audio plays a crucial role in user engagement, such as in video games, virtual reality scenarios, or movies, where the inability to adjust audio settings in real - time can affect immersion and reduce enjoyment. The introduction of surround sound systems was a significant advancement as it began to place the user at the center of a more immersive sound field. However, even with these systems, due to the lack of dynamic adaptability, users typically have to manually adjust the settings for different types of content - a process that is both cumbersome and disruptive.
[0005] The diversification of audio and video content available to consumers, especially the proliferation of streaming platforms, has intensified the demand for audio systems that can intelligently adjust their output to suit the content type. It is expected that audio systems can not only identify different types of content during playback but also adjust the audio profile without latency. For example, when a user switches from a podcast to an action movie, the audio system should automatically shift from emphasizing vocal clarity to highlighting the dynamic range and spatial effects suitable for high-energy scenes. Similarly, when listening to a live concert recording, the sound system can enhance the sense of space and the atmosphere of the audience, thus transporting the listener to the scene. Ideally, these adjustments are made in real time, tracking content changes in a unified stream, such as a movie with interspersed dialogue, soundtrack, and sound effects, ensuring that each element is distinct audibly and contributes to a coherent auditory narrative. This intelligent responsiveness is particularly useful in interactive media such as video games, where the audio environment can change instantaneously based on the player's actions. In such applications, the rapid adaptability of the audio system can significantly impact the user's engagement and immersion, which determines whether it is a good or an outstanding auditory experience.
[0006] Traditional audio systems face significant challenges. One drawback is that traditional audio systems often rely on metadata marked by content creators to guide the adjustment of audio settings. This metadata can include information about the genre, the expected audio balance, and even specific cues for audio effects. However, the reliability of this approach depends on the integrity and accuracy of the metadata production, and the metadata can vary greatly between different content segments. In some cases, the metadata may be too general or not finely tuned to capture the nuances of the audio track, such as the low-frequency rumble of an approaching storm in a thriller or the subtle reverb that gives a live jazz recording a club-like feel.
[0007] Another drawback is the assumption that users will manually intervene to adjust audio settings based on their preferences, which overlooks the convenience of modern media consumption habits. Many listeners prefer a hands-free experience where they can immerse themselves in the content without being interrupted by fiddling with settings. This is especially important when listeners are engaged in other activities such as driving, exercising, or cooking, where adjustment actions are not only inconvenient but also unsafe or impossible. User intervention also requires a certain level of audio expertise, which the average listener often lacks. For example, when watching an action movie, users usually do not know how to adjust the audio system to ensure that the dialogue is not drowned out by the soundtrack or sound effects.
[0008] Another drawback is that the classification of audio content, especially in scenarios where multiple audio elements are present simultaneously, remains a challenge for current audio systems. Current audio systems may exhibit unstable dynamic responses when faced with such complex audio environments. For example, a conversation intertwined with the ambient noise of a sheet of music or a bustling cityscape. When these audio systems misidentify the main audio components, the adjustments made to the audio settings can be counterproductive. A system that wrongly identifies background music as the main element may enhance the music experience but at the cost of masking the voices of the actors, making it difficult for the audience to understand the dialogue.
[0009] Another drawback is that the silent segments within the audio content pose an additional challenge. Transitions into and out of silent periods are typically not handled well, and traditional audio systems are unable to smoothly adjust the audio settings. This can result in sudden changes in the audio output, which may draw attention and potentially disrupt the listening experience.
[0010] As previously mentioned, there is a need in the art for an advanced audio scene analysis system that can dynamically identify and adapt to different types of audio content, including the handling of silence. Summary of the Invention
[0011] In various embodiments, a computer-implemented method includes: receiving audio data; segmenting the audio data into a plurality of audio frames; extracting one or more audio features from the plurality of audio frames; classifying each of the plurality of audio frames based on the one or more features to generate audio classification predictions for the plurality of audio frames; determining the audio type of a current audio frame among the plurality of audio frames based on the audio classification predictions of the plurality of audio frames within a context window; determining one or more audio settings based on the audio type; processing the audio data of the current audio frame using the one or more audio settings; and outputting the processed audio data using one or more speakers.
[0012] Additional embodiments particularly provide a non-transitory computer-readable storage medium storing instructions for implementing the method set forth above, and a system configured to implement the method set forth above.
[0013] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, automatic adaptation of audio settings for dynamically changing an audio scene is possible without metadata tagging. The technology also allows for changing audio settings when it is not practical or safe for a user to adjust the audio settings in real time. The disclosed technology further accurately classifies complex audio scenes with multiple overlapping audio types, making appropriate adjustments based on context to enhance the overall listening experience without causing sudden changes in audio settings that would disrupt the listening experience. In addition, the disclosed technology can smoothly transition into and out of a silent period. These technical advantages provide one or more technical improvements over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To enable a more specific understanding of the manner in which the above-described features of the various embodiments can be obtained, the inventive concept briefly summarized above may be described in more detail by reference to the various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the inventive concept and are therefore in no way to be considered limiting in scope, and that there are other equally effective embodiments.
[0015] Figure 1 is a block diagram of a computing system configured to implement one or more aspects of the various embodiments;
[0016] Figure 2 is according to the various embodiments Figure 1 of a block diagram of an audio scene analyzer included in a computing device for processing audio data;
[0017] Figure 3 is according to the various embodiments Figure 2 of a block diagram of an audio type classifier included in the audio scene analyzer for processing the extracted features;
[0018] Figure 4 is according to the various embodiments Figure 2 of a block diagram of a context scene detector included in the audio scene analyzer;
[0019] Figure 5 shows an example of audio data being analyzed according to the various embodiments;
[0020] Figure 6 is a flowchart of method steps for audio scene analysis according to the various embodiments; and
[0021] Figure 7 is a flowchart of method steps for determining the context of an audio scene according to the various embodiments. DETAILED DESCRIPTION
[0022] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.
[0023] System Overview
[0024] Figure 1 FIG. is a block diagram of a computing system 100 configured to implement one or more aspects of the various embodiments. As shown, computing system 100 includes, but is not limited to, a computing device 102 and one or more speakers 104. Computing device 102 includes, but is not limited to, an I / O interface 106, a processor 108, a bus 110, and a memory 112. Memory 112 includes, but is not limited to, an audio scene analyzer 114, audio data 116, an audio settings application 118, and an audio processing application 120. In some embodiments, computing system 100 is incorporated into an audio device, such as an audio player, an audio / video player, a media player, a smartphone, a tablet, a laptop, a desktop computer, an in-vehicle system, and the like.
[0025] In operation, computing device 102 uses audio scene analyzer 114 to continuously evaluate and classify audio content. An audio input (not shown) may be received through I / O interface 106 and stored in memory 112 as audio data 116, which may include various types of audio scenes, such as conversations, music, ambient noise, and the like. Audio scene analyzer 114 processes audio data 116 to determine the audio scene type. Based on the determined audio scene type, audio settings application 118 configures parameters for processing the audio output based on the determined audio scene type. Then, audio processing application 120 applies the audio settings, which adjusts audio data 116 to enhance certain characteristics corresponding to the determined audio scene, such as the clarity of a conversation, the depth of music, and the like. The resulting audio is then transmitted to speaker 104, providing a listening experience that adapts in real time to the audio content being played.
[0026] Speaker 104 produces the audio output that the user is to hear. Speaker 104 is coupled to computing device 102 via I / O interface 106. For example, audio processing application 120 generates audio signals and sends these audio signals to speaker 104. In some embodiments, speaker 104 includes any type of speaker, such as a dynamic speaker, a planar magnetic speaker, an electrostatic speaker, and the like.
[0027] The I / O interface 106 facilitates communication between the computing device 102 and other external systems including the speaker 104. For example, audio input can be received through the I / O interface 106, such as from a separate media device (not shown) and / or a media streaming service. The I / O interface 106 can include any technically feasible interface type, such as a Universal Serial Bus (USB) interface, a Bluetooth interface, an optical audio interface, one or more network interfaces, etc.
[0028] The processor 108 performs various computing tasks during the operation of the computing device 120. The processor 108 is a processing device in any technically feasible form configured to process data and execute program code. The processor 108 can include, for example but not limited to, a System on Chip (SoC), a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a multi-core processor, or a combination of multiple processing units, such as a CPU configured to operate in conjunction with a GPU and / or a DSP.
[0029] The interconnect bus 110 connects the I / O interface 106, the processor 108, the memory 112, and any other components of the computing device 102. The interconnect bus 110 facilitates the flow of information and commands between the components of the computing device 102. The interconnect bus 110 can be any technically feasible type of bus system, including a serial bus, a parallel bus, etc.
[0030] The memory 112 includes one or more memory modules. In some embodiments, each of the one or more memory modules is a Random Access Memory (RAM) module, a flash memory cell, or any other type of memory cell or a combination thereof. The memory 112 stores information of various programs and applications executed by the processing unit 112, such as instructions and / or data. When executing programs and applications, the processing unit 112 is configured to read from and write to the memory 112. In various embodiments, the memory 112 includes non-volatile memory, such as Read Only Memory (ROM), such as flash memory, etc.
[0031] The audio scene analyzer 114 analyzes audio data 116 to assist in enhancing the playback of the audio data 116 to one or more users. The audio scene analyzer 114 segments the audio data 116 into frames. Then, the audio scene analyzer 114 processes each segment to extract one or more features of the audio data 116 in that segment. Then, the audio scene analyzer 114 classifies each frame to determine the audio type of that frame, such as dialogue, music, action sound effects, silence, etc. In some examples, the audio scene analyzer also includes separate techniques for detecting silence and detecting audio types. Then, the audio scene analyzer 114 uses the history of the audio type of the frames and the confidence scores of the audio types to determine the audio scene type. Then the audio scene type is provided to the audio settings application 118. The audio scene analyzer 114 will be described in further detail below in Figures 2 to 5 further detail.
[0032] In some embodiments, the audio scene analyzer 114 includes machine learning algorithms, such as a convolutional neural network (CNN) that identifies patterns within the spectrum of an audio frame and identifies the structured elements of music, a recurrent neural network (RNN) that processes the time-domain sequence of the audio data 116 to identify various audio scene types and that employs a long short-term memory (LSTM) architecture, etc.
[0033] The audio data 116 stored within the memory 112 includes recordings of audio content processed by the computing system 100. In various embodiments, the audio data 116 is either previously stored in the memory 112 or continuously streamed to the computing device 102 via the I / O interface 106. The audio data 116 includes a series of audio frames. In various embodiments, each audio frame stored in the audio data 116 includes the characteristics of the input audio within the corresponding audio frame. These characteristics can include any audio characteristic, such as frequency, amplitude, duration, etc.
[0034] The audio settings application 118 receives the audio scene type determined by the audio scene analyzer 114 and the confidence score corresponding to that audio scene type. Then, the audio settings application 118 uses the determined audio scene type and the confidence score to select or adjust audio parameters, such as equalizer settings, volume levels, and acoustic effects, that are customized to enhance the audio data 116 of the current audio segment. For example, for an audio scene of the dialogue type, the audio settings application 118 can select or adjust audio parameters to emphasize speech clarity. For a music audio scene, the audio settings application 118 enhances the richness and depth of the audio. Then the audio settings are provided to the audio processing application 120.
[0035] The audio processing application 120 applies the adjustments determined by the audio settings application 118 to the audio data 116. In various embodiments, the audio processing application 120 uses various signal processing methods, such as dynamic range compression (DRC) to maintain volume consistency between different audio scene types, equalization (EQ) to customize the frequency response according to the specific requirements of dialogue audio scene type or music audio scene type segments, and so on. In some embodiments, to obtain an immersive experience, the audio processing application 120 includes 3D audio processing technologies (such as binaural rendering, etc.) to create a three-dimensional sound field.
[0036] Determining an audio scene via audio scene analysis
[0037] Figure 2 is a block diagram of the audio scene analyzer 114 that processes the audio data 116 according to various embodiments. As shown, the audio scene analyzer 114 includes, but is not limited to, an audio segmentation module 200, a feature extraction module 202, a silence detector 206, an audio type classifier, and a context scene detector 214.
[0038] The audio segmentation module 200 divides the audio data 116 into multiple segments or frames for further analysis. The audio segmentation module 200 divides the audio data 116 into audio frames of a fixed length. In some examples, the duration of each audio frame is 0.48 seconds. In various embodiments, each frame overlaps a predetermined amount with the previous and the next audio frames. For example, the predetermined amount can be 50%, such that each part of the audio data 116 is included in two audio frames. By overlapping consecutive frames, the audio segmentation module 200 captures transitional elements in the audio data 116 that might be missed in non-overlapping segmentation. For example, in a scenario where the scene transitions from a quiet conversation to a sudden loud explosion, in the case of non-overlapping segmentation in the audio segmentation module 200, there is a risk that any segment might be misclassified, and the subtle changes in the audio data before the explosion or immediately after the conversation might be missed.
[0039] The feature extraction module 202 receives audio frames from the audio segmentation module 200 and extracts features 204 for each audio frame. In various embodiments, the feature extraction module 202 uses signal processing techniques to extract features, such as spectrograms, etc., which provide insights into the pitch and harmonic content of segments of the audio data 116. In some embodiments, the feature extraction module 202 extracts time-domain features, such as energy envelopes, zero-crossing rates, etc., which provide information about the rhythm and dynamics of the audio segments. For example, the feature extraction module 202 can extract Mel-frequency cepstral coefficients (MFCCs), which are very effective for capturing the timbre quality of each audio frame, and the feature extraction module 202 can extract features such as spectral flux or beat histograms to capture the varying energy and rhythm patterns of each audio frame. The extracted features 204 are then provided to the silence detector 206 and the audio type classifier 210.
[0040] The silence detector 206 analyzes the extracted features 204 of each audio frame to determine if there is no significant acoustic activity to detect audio frames corresponding to silence. In various embodiments, the silence detector 206 uses signal processing methods to detect silence scenarios, such as analyzing the energy envelope of the audio frame. For example, it is calculated by squaring the amplitude values of the spectrum of the audio signal, averaging the squared values over the audio frame, and then taking the square root of the average. Then, the silence detector 206 classifies audio frames with an energy envelope level below a certain threshold as silence. In some embodiments, the silence detector 206 includes an analysis of the zero-crossing rate, where the silence detector 206 evaluates the flatness of the audio spectrum. For silent or near-silent audio frames, the spectrum tends to be flat, indicating a lack of unique frequencies characterizing non-silent audio. Additionally, in at least one embodiment, the silence detector 206 uses statistical analysis, where the silence detector 206 analyzes the statistical distribution of the amplitudes within the frame. Silent frames typically have a very narrow distribution that approaches zero. In various embodiments, the silence detector 206 uses machine learning methods trained on various features (not limited to energy-based features) to distinguish between silent and non-silent audio frames.
[0041] In various embodiments, the silence detector 206 includes machine learning techniques to detect audio frames corresponding to silence. In some embodiments, the silence detector 206 uses a supervised learning algorithm, where a machine learning model is trained on a dataset of audio segments labeled as "silence" and "non - silence" to learn the characteristics of silence scenario types in various contexts. For example, the silence detector 206 can include a support vector machine (SVM) that is trained to identify a clear boundary between silence and sound by identifying the optimal hyperplane within the space of the extracted features 204. Alternatively, in some embodiments, the silence detector 206 uses a neural network (such as a recurrent neural network (RNN) employing long short - term memory (LSTM) cells) to predict silence in an audio sequence by learning the temporal patterns typically before and after silence intervals. Additionally, in various embodiments, the silence detector 206 uses an anomaly detection method, where any deviation from the norm (typical sound patterns) is detected as corresponding to silence.
[0042] Then, the silence detector 206 generates a silence prediction 208 for each audio frame based on the analysis of the extracted features 204. The silence prediction 208 is then provided to the context scenario detector 214.
[0043] The audio type classifier 210 processes the extracted features 204 to classify the audio scenario type within each audio frame, such as Figure 3As shown in more detail below. In various embodiments, the audio type classifier 210 uses various classification algorithms to classify various audio scene types, such as speech, music, conversation, silence, etc. In various embodiments, the audio type classifier 210 uses machine learning techniques, such as decision trees to make a preliminary judgment based on the extracted features 204, or uses a deep neural network (DNN), which learns from a large dataset and makes more subtle distinctions about the audio scene types. For example, machine learning techniques can be trained on the Google Audioset Strong dataset, which was created to solve the audio tagging task with noisy audio type labels, the human voices belonging to the conversation audio type, the SESA gunshot / gunshot audio dataset belonging to the action audio type, the condensed movie dataset containing real movie scenes with mixed audio type content can be added, the CommoVoice dataset belonging to the conversation audio type, the FMA-SMALL and GTZAN datasets belonging to the music audio type, the live music dataset belonging to the music audio type with crowd noise, the live sports commentary dataset belonging to the mixed action and conversation audio type, and the synthetically generated dataset in which the conversation, action, or music audio type is synthesized with different signal-to-noise ratios to generate different mixed audio scene scenarios. In some embodiments, the audio type classifier 210 uses a Gaussian mixture model (GMM) to identify the probabilistic features of speech. In some embodiments, when processing the music audio scene type, the audio type classifier 210 uses a convolutional neural network (CNN) to identify the patterns indicating the music structure. In some embodiments, the audio type classifier 210 uses a random forest, which can handle the variability of non-speech audio scene types by considering a large number of decision trees. The audio scene types classified by the audio type classifier 210 are provided as audio classification predictions 212 to the context scene detector 214.
[0044] The context scene detector 214 processes the history of the silence prediction 208 and the audio classification prediction 212 to determine the audio type 216 and the corresponding type confidence 218 for each audio frame. Figure 4 The context scene detector 214 is described in further detail below.
[0045] Figure 3 It is a block diagram of an audio type classifier 210 included in an audio scene analyzer 114 that processes the extracted features 204 according to various embodiments. As shown, the audio type classifier 210 includes, but is not limited to, a convolutional neural network (CNN) 300, a set of feature maps 302, and a fully connected classifier 304.
[0046] The CNN 300 analyzes the extracted features 204 of the audio clip, such as the Mel spectrogram, etc., and constructs a set of feature maps 302. The CNN 300 applies multi-layer filters to detect patterns in the features 204 extracted in the frequency domain and time domain. In various embodiments, the initial layer of the CNN 300 identifies basic patterns (such as the edges of sound bursts), while deeper layers detect more complex features (such as the texture of sound) to distinguish various audio scenarios, etc.
[0047] In some embodiments, the CNN 300 includes the MobileNetV1 architecture known for its computational efficiency on mobile devices. The MobileNetV1 processes the extracted features 204 to identify patterns that can be used to classify audio scene types. In some embodiments, the CNN 300 includes frozen layers whose weights remain constant during the retraining process. The frozen layers can accelerate training and utilize previously learned features that are still relevant for the classification of new audio frames.
[0048] The feature maps 302 are generated by the CNN 300. Each feature map 302 corresponds to a specific filter applied by the CNN 300, highlighting different aspects of the extracted features 204. For example, one feature map 302 can emphasize the boundaries of different sounds by capturing sharp changes in the spectrum, which are common at the start or end of a note or spoken word. Another feature map 302 can focus on the sustained frequency, identifying the steady hum of a machine or the continuous notes in a song. The feature maps 302 together represent a set of attributes extracted from the extracted features 204, such as texture, intensity, and time-domain variations, which are used by the fully connected classifier 304.
[0049] The fully connected classifier 304 receives the set of attributes from the feature maps 302 and classifies the audio data 116 into different audio scene types, such as music, action sound effects, dialogue, etc. In various embodiments, the fully connected classifier 304 includes one or more fully connected neural network layers that can distinguish different audio scene types. For example, if the features in the feature maps 302 indicate harmonic structure and rhythmic consistency, the fully connected classifier 304 can identify the audio scene type as music, or if the features in the feature maps 302 indicate clear patterns of human voice frequencies and speech rhythms, the fully connected classifier can identify the audio scene type as dialogue. The fully connected classifier 304 outputs an audio classification prediction 212.
[0050] Figure 4 It is a block diagram of the context scene detector 214 of the audio scene analyzer 114 according to various embodiments. As shown, the context scene detector 214 includes, but is not limited to, a context window 402, a background music detector 404, a smoother 406, a max pooling module 408, and decision logic 410.
[0051] The context window 402 includes the silence prediction 208 and the audio classification prediction 212 for a certain number of consecutive audio frames, the certain number of consecutive audio frames being Indicates, for example, Can be 2, 3, 4 or more, depending on the length of time the context is considered during audio scene analysis. For example, for a frame length of 0.5 seconds, if you divide the time The time index corresponding to the second The context window 402 at is set to include five audio frames , when using 50% audio frame overlap, the context window 402 includes five audio frames, corresponding to time periods (1.5s, 2.0s), (1.75s, 2.25s), (2.0s, 2.5s), (2.25s, 2.75s) and (2.5s, 3.0s), respectively.
[0052] Background music detector 404 determines the presence of background music in an audio frame. In various embodiments, background music detector 404 calculates the value of The sum of the absolute gradients of the confidence levels of the music audio scenes in the audio frames held by the context window 402.
[0053]
[0054] in is the sum of the absolute differences in confidence scores, It is The music audio scene detection confidence score for audio frames, is the index of the current audio frame, and is the size of the context window. If When the predefined background music threshold (ranging from 0.1 to 0.5) is exceeded, the background music detector 404 will classify the current audio frame as containing background music by setting the BGM flag to 1 to indicate that the underlying background music content should be considered when processing the audio scene. When the number of audio frames is 5 and the frame length is 0.5 seconds, the background music detector 404 receives the confidence score of the music audio scene detection as the confidence score. , then, the background music detector 404 calculates the sum of absolute gradients by summing the absolute differences between consecutive confidence scores :
[0055]
[0056]
[0057]
[0058] If the background music threshold is predefined as 0.2 and the calculated sum is 0.25 greater than 0.2, then the background music detector 404 will indicate the presence of background music by setting the BGM flag to 1. The gradient calculation allows the background music detector 404 to detect the occurrence or consistency of music elements over time, as opposed to transient music sounds that do not correspond to background music.
[0059] The smoother 406 uses a time-based decay function to moderate the fluctuations in the confidence scores within a specific window. The window is defined by the following parameters: , the duration of the smoothing window, e.g., can be 2, 3, 4, or more. In various embodiments, the smoother 406 includes a linear decay weight function with a weight of , which reduces the influence of the audio frame proportionally based on the index distance from the current frame . For example, the formula for calculating the weight of each audio frame can be as described in Equation 2.
[0060]
[0061] where is the weight applied to the confidence of the th audio frame, is the index of the audio frame within the context window, is the index of the current audio frame, and represents the number of audio frames used for smoothing. The linear decay weighting ensures that the most recent audio frames have a greater influence on the smoothed confidence score, while the influence of audio frames further in time gradually decreases. In various embodiments, the output audio frame confidence from the smoother 406 is normalized to avoid audio frame confidence greater than one. Alternatively, other non-linear decay functions can be used.
[0062] The max pooling module 408 extracts the most prominent audio type from the smoothed audio frame confidence sets provided by the smoother 406 for each audio frame. By identifying the audio frame type with the maximum confidence within each audio frame, the max pooling module 408 determines the most significant audio characteristic present in that audio frame. The max pooling operation in the max pooling module 408 can be represented mathematically as shown in Equation 3.
[0063]
[0064] Where Max_confidence[i] is the maximum confidence selected by the module at audio frame index i, and Smoothed_confidence is the confidence array that has been smoothed by the smoother 406. By using max pooling in the max pooling module 408, transient or less significant audio features are not allowed to have a disproportionate impact on the determination of the audio scene type.
[0065] The decision logic 410 determines the audio scene type and the corresponding confidence level. In various embodiments, the decision logic 410: calculates the percentage duration of each audio type represented by , where can be any audio type in the output of the max pooling module 408 within the pooling window represented by the number of audio frames, such as action sound effects, dialogue, music, silence, etc.; and calculates the confidence of each audio type as the average of the confidences . The decision logic 410 makes decisions based on several preset parameters . For example, , , , and . Then, for each audio type, the decision logic 410 calculates and identifies the dominant audio type as the audio type corresponding to the highest value of .
[0066] In various embodiments, after determining the dominant audio type, the decision logic 410 determines the audio type 216 and the type confidence 218 based on a set of heuristic rules. In at least one embodiment, the decision logic 410 uses the following heuristic rules to determine the audio type 216 and the type confidence 218. If the output of the background music detector 404 is BGM = 0, indicating no background music, the decision logic 410 outputs the dominant audio type as the audio type 216 and outputs the type confidence 218 as . If BGM = 1 indicates the presence of background music, and if or is greater than 0.4, the decision logic 410 outputs the audio type (action sound effect or dialogue) corresponding to the higher confidence value as the audio type 216 and has the corresponding type confidence 218; however, if the confidence value is below 0.1, the decision logic 410 outputs music as the audio type 216. If is silence, the decision logic 410 outputs the audio type 216 of the previous frame and sets the type confidence 218 to zero. If is not silent and If it is negative, the decision logic 410 outputs the as audio type 216 and sets the type confidence 218 to .
[0067] Figure 5 FIG. shows examples of audio data 502 being analyzed according to various embodiments. As shown, audio data 502 depicts the amplitude of an audio input over time. Spectrogram 504 depicts the spectrum of the audio input over time, where brighter regions represent silence or lower energy sounds, while darker regions indicate louder sounds. After being processed by audio segmentation module 200 and feature extraction module 202, audio data 502 is classified by audio type classifier 210 and silence detector 206 to determine the audio type of each audio frame in audio data 502, as shown by audio classification prediction 506. Audio classification prediction 506 shows the confidence level assigned to each classification of audio type determined by audio type classifier 210 and silence detector 206 on the same timeline as audio data 502 and spectrogram 504. Each line shows the audio classification prediction for a different audio type, such as conversation 508, music 512, action sound effect 514, and silence 510. The unsmoothed audio classification prediction 506 indicates a preliminary assessment of the likelihood that each audio frame contains audio of the indicated type when considered individually. Then, audio classification prediction 506 is further processed by context scene detector 214 to determine the audio type 216 and type confidence 218 of each audio frame by considering the audio classification probabilities over multiple audio frames.
[0068] Figure 6 is a flowchart of method steps for analyzing an audio scene according to various embodiments. Although the method steps are described in connection with Figures 1 to 4 the system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.
[0069] Method 600 begins at step 602, where computing device 102 receives audio data 116. In some examples, the audio input is received via I / O interface 106 as part of a streaming media or loaded from memory 112. For example, the audio input may correspond to the audio track of a movie that the user is watching. The audio input is stored in memory 112 as audio data 116.
[0070] At step 604, the audio segmentation module 200 segments the audio data 116 into audio frames. The audio data 116 is segmented into audio frames of a fixed length. In various embodiments, consecutive audio frames overlap, for example, by 50%. For example, if the fixed length of each audio frame is set to 0.5 seconds, the audio segmentation module 200 creates audio frames such that each audio frame starts halfway through the previous audio frame. Then, the first audio frame covers the audio data 116 from 0 to 0.5 seconds, the second audio frame covers the audio data from 0.25 to 0.75 seconds, and so on, thus creating a sequence of consecutive and overlapping audio frames. Once the audio segmentation module 200 receives enough audio data 116 for a new audio frame, the audio segmentation module 200 passes that audio frame to the feature extraction module 200.
[0071] At step 606, the feature extraction module 202 extracts audio features 204 from the audio frames received from the audio segmentation module 200. In various embodiments, the feature extraction module 202 analyzes the audio frames to identify and derive acoustic features for subsequent classification, such as spectral features (including but not limited to MFCC, spectral centroid, spectral flux, and spectral roll-off), time-domain features (which capture the temporal characteristics of the sound, such as the zero-crossing rate (the frequency at which the audio frame signal changes sign) or the time-domain envelope feature that tracks the change in the amplitude of the audio frame signal over time), and harmonic features (including but not limited to the harmonic-to-noise ratio or pitch).
[0072] At step 608, the silence detector 206 analyzes the features 204 extracted from the audio frames to determine whether the audio frame should be classified as silence. In various embodiments, the silence detector 206 calculates the energy envelope of the audio frame by taking the root mean square (RMS) of the amplitude values within the audio frame, thus providing a single value representing the power of the signal within the audio frame. If the energy envelope value is below a preset threshold, the silence detector 206 concludes that the audio frame should be classified as silence. In some embodiments, the silence detector 206 checks the zero-crossing rate, which indicates the frequency at which the audio frame signal changes sign or crosses the zero amplitude axis. If the zero-crossing rate is below the threshold, the silence detector 206 concludes that the audio frame should be classified as silence. The silence detector 206 outputs a silence prediction 208, which, in some embodiments, is a binary flag that indicates whether there is silence in the audio frame with values of 1 or 0, respectively. If silence is detected in the audio frame, method 600 proceeds to step 612. Otherwise, method 600 proceeds to step 610.
[0073] At step 610, the audio type classifier 210 classifies the audio frames. The audio type classifier 210 analyzes the features extracted from the audio frames to generate an audio type 212. The audio classification prediction 212 indicates the probability of the presence of each of one or more sound types (e.g., dialogue, music, silence, action sound effects, etc.) in the audio frame. In various embodiments, the audio type classifier 210 uses machine learning techniques to generate the audio classification prediction 212.
[0074] At step 612, the context scene detector 214 determines the context of the audio scene in the audio frame. The context scene detector 214 analyzes the silence prediction 208 from step 608 and the audio classification prediction 212 from step 610, and outputs the context of the audio scene using the audio type 216 and the type confidence 218. Step 612 is described in more detail in Figure 7 more detail.
[0075] At step 614, the audio processing application 120 processes the audio data 116 using audio settings based on the determined audio type 216. The audio settings application 118 selects and modifies the audio settings based on the audio type 216 and the type confidence 218 generated during step 612. The audio settings can include one or more of equalizer settings, volume levels, acoustic effects, etc., to enhance the determined audio type. For example, the audio settings application 118 can set the audio settings to increase the treble to enhance the speech clarity in an audio frame with dialogue, or enhance the bass to obtain a richer and deeper music experience in an audio frame with music. Then, the audio processing application 120 uses the audio settings to process the audio data 116. Then, the audio processing application 120 uses the audio settings to adjust the audio data 116 in the audio frame. This adjusts the audio data 116 in real time to enhance the audio data 116 based on the overall audio scene in the audio frame while taking into account the context of the previous audio frame, thereby improving the audio listening experience.
[0076] At step 616, the audio processing application 120 outputs the processed audio to the speaker 104 of the computing device 102. The speaker 106 converts the processed audio data 116 into sound waves customized based on the audio type.
[0077] After completing step 616, the method loops back to step 602 to analyze additional audio inputs by repeating method 600. By repeating method 600, the computing system 100 continues to dynamically update the audio settings for processing audio inputs in real time to improve the listening experience.
[0078] Figure 7 is a flowchart of the method steps corresponding to step 612 for determining the context of an audio scene according to various embodiments. Although combined withFigures 1 to 4 The system describes method steps, but those skilled in the art will understand that any system configured to execute the method steps in any order falls within the scope of the present disclosure.
[0079] At step 702, the context scene detector 214 receives a prediction history. The prediction history includes the silence prediction 208 and the audio classification prediction 212 within the context window 402. The context window 402 includes a fixed number of audio frame predictions, which allows for considering the temporal relationships and patterns that occur within a series of audio frames.
[0080] At step 704, the background music detector 404 determines whether there is background music by analyzing the extracted features 204 within the audio frames included in the context window 402. For example, the background music detector 404 can sum the absolute differences in the confidence levels of the music audio scenes and then compare the sum with a threshold. If background music is detected, the background music detector 404 sets the BGM flag to 1. Otherwise, the flag is set to 0.
[0081] At step 706, the smoother 406 applies decay smoothing to the confidence scores corresponding to the audio classification predictions 212 in the context window 402. The smoother 406 calculates and applies weights to the confidence scores of the current audio frame and the previous audio frame within the context window according to a time-based decay function, which provides a weight for the current audio frame starting from an initial value and then decreases the value for each previous frame. For example, the smoother 406 can generate weights according to Equation 2. In various embodiments, step 706 can be performed before step 704 or in parallel with step 704.
[0082] At step 708, the max pooling module 408 performs max pooling on the smoothed audio classification predictions 212 of the audio frames. Max pooling identifies the audio type corresponding to the highest smoothed classification prediction of the audio frames determined during step 706.
[0083] At step 710, the decision logic 410 analyzes the output of the max pooling within a fixed number of audio frames in the pool window to determine the dominant audio type. The decision logic 410 determines the dominant audio type by calculating the percentage duration and the average classification probability of each audio type within the pool window. Then, the decision logic 410 compares the average classification probability with a predefined threshold to identify the dominant audio type.
[0084] At step 712, decision logic 410 predicts the audio type 216 and type confidence 218 of an audio frame. Decision logic 410 applies a set of heuristic rules to predict the audio type 216 and type confidence 218. The heuristic rules consider multiple factors, including but not limited to the presence or absence of background music identified by the background music detector 404 in step 704, and the dominant audio type determined in step 710. Additionally, the heuristic rules take into account previous audio type predictions to more accurately predict the current audio type. For example, if silence is detected as the current dominant audio type, but previous audio type predictions indicate a repeating pattern of dialogue audio type or action audio type, then decision logic 410 adjusts the current audio type prediction to be consistent with the expected continuity of the audio type. The heuristic rules are designed to balance immediate audio characteristics and the broader context established by past and present audio frames to ensure that transient sounds or momentary changes in audio type do not disproportionately distort the audio type prediction.
[0085] At step 714, decision logic 410 outputs the predicted audio type 216 and type confidence 218. Then, the predicted audio type 216 (such as music, dialogue, action sound effects, silence, etc.) and type confidence 218 (i.e., the probability of an accurate audio type classification from 0 to 1) are provided to the audio settings application 118 to select and modify the audio settings in step 614.
[0086] In summary, the disclosed technology dynamically analyzes an audio scene based on content recognition. The incoming audio content is segmented into audio frames. The audio frames are then analyzed to extract audio features. The audio frames are then classified to determine which of one or more audio types are present in the audio frame. The classification can include not only the identification of the audio type, but also a confidence level indicating the certainty of the classification. The classification of the current audio frame and historical data from the previous audio frame are then analyzed to determine the current context of the audio scene. Once the audio type of the audio scene is determined, the audio settings for that audio type are selected. The audio content is then processed using the selected audio settings to generate audio output using one or more speakers.
[0087] At least one technical advantage of the disclosed technology over the prior art is that, with the disclosed technology, automatic adaptation of audio settings for dynamically changing the audio scene is possible without metadata tagging. The technology also allows for changing the audio settings when it is not practical or safe for the user to adjust the audio settings in real time. The disclosed technology further accurately classifies complex audio scenes with multiple overlapping audio types and makes appropriate adjustments based on the context to enhance the overall listening experience without causing sudden changes in the audio settings that would disrupt the listening experience. Additionally, the disclosed technology can smoothly transition into and out of silent periods. These technical advantages provide one or more technical improvements over prior art methods.
[0088] 1. In some embodiments, a computer-implemented method for adjusting audio includes: receiving audio data; splitting the audio data into a plurality of audio frames; extracting one or more audio features from the plurality of audio frames; classifying each of the plurality of audio frames based on the one or more features to generate an audio classification prediction for the plurality of audio frames; determining an audio type of a current audio frame of the plurality of audio frames based on the audio classification predictions within a context window of the plurality of audio frames; determining one or more audio settings based on the audio type; processing the audio data of the current audio frame using the one or more audio settings; and outputting the processed audio data using one or more speakers.
[0089] 2. The computer-implemented method according to clause 1, wherein the plurality of audio frames are overlapping audio frames.
[0090] 3. The computer-implemented method according to clause 1 or 2, wherein the one or more audio features include at least one of time domain features or Mel-frequency cepstral coefficients.
[0091] 4. The computer-implemented method according to any one of clauses 1 to 3, wherein the audio classification prediction includes a prediction of one or more audio types selected from the group consisting of silence, dialogue, action sound effects, and music.
[0092] 5. The computer-implemented method according to any one of clauses 1 to 4, wherein classifying each of the plurality of audio frames includes applying one or more machine learning models.
[0093] 6. The computer-implemented method according to any one of clauses 1 to 5, further comprising determining whether each of the plurality of audio frames includes silence based on an energy envelope or zero crossings of each audio frame of the plurality of audio frames; and determining the audio type of the current audio frame based on whether the current audio frame includes silence.
[0094] 7. The computer-implemented method according to any one of clauses 1 to 6, wherein determining the audio type of the current audio frame includes weighting the audio classification predictions of the plurality of audio frames within the context window according to a time-based decay function; determining the audio type with the highest prediction in each of the plurality of audio frames within the context window; and applying one or more heuristic rules to determine the audio type of the current audio frame.
[0095] 8. The computer-implemented method according to any one of clauses 1 to 7, wherein determining the audio type of the current audio frame further includes determining the percentage duration of each audio type within the context window; subtracting the corresponding confidence threshold of each audio type from the determined percentage duration; and applying the one or more heuristic rules based on the subtracted result.
[0096] 9. The computer-implemented method according to any one of clauses 1 to 8, wherein determining the audio type of the current audio frame further includes determining the average of the confidence scores of each audio type within the context window; and applying the one or more heuristic rules based on the average of the confidence scores.
[0097] 10. The computer-implemented method according to any one of clauses 1 to 9, wherein determining the audio type of the current audio frame further includes detecting whether there is background music in the context window; and applying the one or more heuristic rules based on the background music detection.
[0098] 11. The computer-implemented method according to any one of clauses 1 to 10, wherein detecting whether there is background music includes calculating the sum of the gradients of the confidence scores of the audio classification predictions corresponding to the audio type of music within the context window; and determining that there is background music when the sum is greater than a threshold.
[0099] 12. The computer-implemented method according to any one of clauses 1 to 11, further including determining the confidence of the prediction of the audio type; and further determining the one or more audio settings based on the confidence of the prediction.
[0100] 13. The computer-implemented method according to any one of clauses 1 to 12, wherein when the current audio frame is classified as silent, determining the audio type of the current audio frame includes setting the audio type to the previous audio type of the first audio frame before the current audio frame; and setting the confidence of the audio type to zero.
[0101] 14. In some embodiments, one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: receive audio data; segment the audio data into a plurality of audio frames; extract one or more audio features from the plurality of audio frames; classify each of the plurality of audio frames based on the one or more features to generate an audio classification prediction for the plurality of audio frames; determine an audio type of a current audio frame among the plurality of audio frames based on the audio classification predictions of the plurality of audio frames within a context window; determine one or more audio settings based on the audio type; process the audio data of the current audio frame using the one or more audio settings; and output the processed audio data using one or more speakers.
[0102] 15. The one or more non-transitory computer-readable media according to clause 14, wherein the plurality of audio frames are overlapping audio frames.
[0103] 16. The one or more non-transitory computer-readable media according to clause 14 or 15, wherein the audio classification prediction includes a prediction of one or more audio types selected from the group consisting of silence, dialogue, action sound effects, and music.
[0104] 17. The one or more non-transitory computer-readable media according to any one of clauses 14 to 16, wherein the steps further include determining whether each audio frame among the plurality of audio frames includes silence based on an energy envelope or zero crossing of each audio frame among the plurality of audio frames; and determining the audio type of the current audio frame based on whether the current audio frame includes silence.
[0105] 18. The one or more non-transitory computer-readable media according to any one of clauses 14 to 17, wherein determining the audio type of the current audio frame includes weighting the audio classification predictions of the plurality of audio frames within the context window according to a time-based decay function; determining the audio type with the highest prediction in each of the plurality of audio frames within the context window; subtracting a corresponding confidence threshold for each audio type from the determined percentage duration; and applying one or more heuristic rules.
[0106] 19. The one or more non-transitory computer-readable media according to any one of clauses 14 to 18, wherein determining the audio type of the current audio frame further includes detecting whether there is background music in the context window; and applying the one or more heuristic rules based on the background music detection.
[0107] 20. In some embodiments, a system includes a memory storing one or more instructions and one or more processors that, when executing the one or more instructions, are configured to perform steps including: receiving audio data; splitting the audio data into a plurality of audio frames; extracting one or more audio features from the plurality of audio frames; classifying each of the plurality of audio frames based on the one or more features to generate audio classification predictions for the plurality of audio frames; determining an audio type of a current audio frame among the plurality of audio frames based on the audio classification predictions within a context window; determining one or more audio settings based on the audio type; processing the audio data of the current audio frame using the one or more audio settings; and outputting the processed audio data using one or more speakers.
[0108] Any and all combinations of any claim elements recited in any of the claims of the claims and / or any elements described in this application are in the scope of the invention and protection in any way.
[0109] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0110] Aspects of the present embodiments may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects with hardware aspects, which embodiments are generally referred to herein as "modules", "systems", or "computers". Additionally, any hardware and / or software technologies, processes, functions, components, engines, modules, or systems described in the present disclosure may be implemented as a circuit or a collection of circuits. Further, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0111] Any combination of one or more computer-readable media can be utilized. A computer-readable medium can be either a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, by way of example and not limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following media: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing media. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0112] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine. These instructions, when executed by the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, by way of example and not limitation, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0113] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, depending upon the functionality involved, or may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0114] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which is determined by the claims that follow.
Claims
1. A computer-implemented method for adjusting audio, the method comprising: Receiving audio data; Dividing the audio data into a plurality of audio frames; Extracting one or more audio features from the plurality of audio frames; Classifying each of the plurality of audio frames based on the one or more features to generate an audio classification prediction for the plurality of audio frames; Determining an audio type of a current audio frame among the plurality of audio frames based on the audio classification predictions of the plurality of audio frames within a context window; Determining one or more audio settings based on the audio type; Processing the audio data of the current audio frame using the one or more audio settings; And Outputting the processed audio data using one or more speakers.
2. The computer-implemented method according to claim 1, wherein the plurality of audio frames are overlapping audio frames.
3. The computer-implemented method according to claim 1, wherein the one or more audio features include at least one of time-domain features or mel cepstral coefficients.
4. The computer-implemented method according to claim 1, wherein the audio classification prediction includes a prediction of one or more audio types selected from the group consisting of silence, dialogue, action sound effects, and music.
5. The computer-implemented method according to claim 1, wherein classifying each of the plurality of audio frames includes applying one or more machine learning models.
6. The computer-implemented method according to claim 1, further comprising: Determining whether each audio frame among the plurality of audio frames includes silence based on an energy envelope or zero crossings of each audio frame among the plurality of audio frames; And Determining the audio type of the current audio frame based on whether the current audio frame includes silence.
7. The computer-implemented method according to claim 1, wherein determining the audio type of the current audio frame includes: Weighting the audio classification predictions of the plurality of audio frames within the context window according to a time-based decay function; Determining the audio type with the highest prediction in each audio frame among the plurality of audio frames within the context window; And Applying one or more heuristic rules to determine the audio type of the current audio frame.
8. The computer-implemented method according to claim 7, wherein determining the audio type of the current audio frame further includes: Determining a percentage duration of each audio type within the context window; Subtracting a corresponding confidence threshold of each audio type from the determined percentage duration; And Applying one or more heuristic rules based on the subtracted result.
9. The computer-implemented method according to claim 8, wherein determining the audio type of the current audio frame further includes: Determining an average of confidence scores of each audio type within the context window; And Applying the one or more heuristic rules based on the average of the confidence scores.
10. The computer-implemented method according to claim 7, wherein determining the audio type of the current audio frame further includes: Detect whether there is background music in the context window; and Apply the one or more heuristic rules based on the background music detection.
11. The computer-implemented method according to claim 10, wherein detecting whether there is background music includes: Calculating the sum of the gradients of the confidence scores of the audio classification predictions corresponding to the audio types of the music within the context window; and Determining that there is background music when the sum is greater than a threshold.
12. The computer-implemented method according to claim 1, further comprising: Determining the confidence of the prediction of the audio type; and Further determining the one or more audio settings based on the confidence of the prediction.
13. The computer-implemented method according to claim 1, wherein when the current audio frame is classified as silent: Determining the audio type of the current audio frame includes setting the audio type to the previous audio type of the first audio frame before the current audio frame; and Setting the confidence of the audio type to zero. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: Receive audio data; Segment the audio data into a plurality of audio frames; Extract one or more audio features from the plurality of audio frames; Classify each of the plurality of audio frames based on the one or more features to generate audio classification predictions for the plurality of audio frames; Determine the audio type of the current audio frame among the plurality of audio frames based on the audio classification predictions of the plurality of audio frames within a context window; Determine one or more audio settings based on the audio type; Process the audio data of the current audio frame using the one or more audio settings; and Output the processed audio data using one or more speakers.
15. The one or more non-transitory computer-readable media according to claim 14, wherein the plurality of audio frames are overlapping audio frames.
16. The one or more non-transitory computer-readable media according to claim 14, wherein the audio classification predictions include predictions of one or more audio types selected from the group consisting of silence, dialogue, action sound effects, and music.
17. The one or more non-transitory computer-readable media according to claim 14, wherein the steps further include: Determining whether each audio frame among the plurality of audio frames includes silence based on the energy envelope or zero-crossing of each audio frame among the plurality of audio frames; and Determining the audio type of the current audio frame based on whether the current audio frame includes silence.
18. The one or more non-transitory computer-readable media according to claim 14, wherein determining the audio type of the current audio frame includes: Weighting the audio classification predictions of the plurality of audio frames within the context window according to a time-based decay function; Determining the audio type with the highest prediction in each audio frame among the plurality of audio frames within the context window; Subtract the corresponding confidence threshold for each audio type from the determined percentage duration; and Apply one or more heuristic rules.
19. The one or more non-transitory computer-readable media according to claim 18, wherein determining the audio type of the current audio frame further comprises: Detecting whether there is background music in the context window; and Applying the one or more heuristic rules based on the background music detection.
20. A system, comprising: A memory storing one or more instructions; and One or more processors configured to perform steps including the following when executing the one or more instructions: Receiving audio data; Segmenting the audio data into a plurality of audio frames; Extracting one or more audio features from the plurality of audio frames; Classifying each of the plurality of audio frames based on the one or more features to generate an audio classification prediction for the plurality of audio frames; Determining the audio type of a current audio frame among the plurality of audio frames based on the audio classification predictions of the plurality of audio frames within a context window; Determining one or more audio settings based on the audio type; Processing the audio data of the current audio frame using the one or more audio settings; and Outputting the processed audio data using one or more speakers.
Citation Information
Cited By
GIS disconnecting switch opening and closing state monitoring method based on multi-dimensional voiceprint features
CN120954449A