Method for dynamically adjusting subtitles and translation modes according to audio input type

The audio classification model is built through an intelligent audio signal classification algorithm, and the audio categories are detected in real time and the subtitles and translation modes are adjusted, which solves the problem of too many user interface options and realizes the concise and efficient operation of the user interface.

CN120581032APending Publication Date: 2025-09-02GUANGZHOU LANGO ELECTRONICS TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510717768.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In existing audio processing technologies, too many user interface options lead to high operational difficulty and poor user experience, making it difficult to automatically hide or display related options according to the audio type.

Method used

The audio classification model is constructed using intelligent audio signal classification algorithm, and through hidden Markov model and feature extraction technology, audio categories are detected in real time, and subtitles and translation modes are dynamically adjusted according to the categories, and related options are automatically hidden or displayed.

Benefits of technology

Simplify the user interface, improve operational efficiency, optimize user experience, reduce unnecessary setting options, and improve the intelligence of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581032A_ABST
    Figure CN120581032A_ABST
Patent Text Reader

Abstract

The invention provides a method for dynamically adjusting subtitles and translation modes according to audio input types. The method comprises the following steps: acquiring audio data from multiple sources and preprocessing the audio data; constructing an audio classification model based on an intelligent audio signal classification algorithm, training the audio classification model, and classifying the audio data; inputting the current audio into the audio classification model, and determining the category of the current audio; constructing a mapping relationship among the audio category, the subtitles and the translation modes, matching the corresponding subtitles and the translation modes according to the category of the current audio, and displaying the subtitles and translation contents in real time; automatically hiding related option settings based on subtitles and translation modes matched with the current audio; according to the invention, the category of the current audio can be detected in real time, the subtitle and translation mode can be dynamically adjusted according to the audio category, related option settings can be automatically hidden or displayed according to the subtitle and translation mode, unnecessary setting options are reduced, and the user experience is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio, and in particular to a method for dynamically adjusting subtitle and translation modes according to audio input types. Background Art

[0002] In current audio processing technology applications, user interfaces often present a multitude of option settings, which causes a lot of trouble for users. Take a general audio editing software as an example. Its interface is filled with various functional options, ranging from complex audio parameter adjustments, such as sampling rate and bit rate settings, to a variety of special function options for different usage scenarios, such as noise reduction mode selection in professional audio recording, and chord editing options in music creation. For ordinary users, the numerous options are not only dazzling, but also increase the difficulty of operation and the cost of learning. For example, a user who just wants to simply edit a voice audio has to struggle to find the required function in a screen full of options. Those complex special effects options for music audio, album cover related settings, etc., are useless when processing voice, but take up a lot of screen space, making the interface look extremely confusing.

[0003] How to intelligently and automatically hide or display some option settings based on audio from different sources is a problem that needs to be solved urgently. For example, for voice audio, only options related to voice processing are displayed on the interface, and options related to music are automatically hidden. For music audio, only options related to music are displayed on the interface, and options related to voice processing are automatically hidden. In this way, the user interface is greatly simplified, unnecessary interference items are removed, and users can find the required functions more quickly and accurately, thereby significantly optimizing the user experience and improving operational efficiency. Summary of the Invention

[0004] To solve the above problems, the present invention provides a method for dynamically adjusting subtitles and translation modes according to the audio input type. The method can detect the category of the current audio in real time, and dynamically adjust the subtitles and translation modes according to the audio category. It automatically hides or displays related option settings according to the subtitles and translation modes, reduces unnecessary setting options, and optimizes the user experience.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] The present invention provides a method for dynamically adjusting subtitles and translation modes according to audio input type, comprising the following steps:

[0007] Step S1: collecting audio data from multiple sources and preprocessing the collected audio data;

[0008] Step S2: Building and training an audio classification model based on an intelligent audio signal classification algorithm to classify the audio data;

[0009] Step S3: Input the current audio into the audio classification model to determine the category of the current audio;

[0010] Step S4: Constructing a mapping relationship between audio categories, subtitles, and translation modes, matching the corresponding subtitles and translation modes according to the current audio category, and displaying the subtitles and translation content in real time;

[0011] Step S5: Based on the subtitles and translation mode matched to the current audio, the relevant option settings are automatically hidden.

[0012] In step S1, a unique identification code is assigned to each audio acquisition source, and the unique identification code is associated with the relevant information of the audio acquisition device or data source and stored.

[0013] In step S2, a hidden Markov model is used to classify audio from different sources. The specific process includes:

[0014] Define the model structure: determine the state set based on the audio category, and determine the observation set based on the audio feature extraction method;

[0015] Initialize model parameters: set the probability of the audio data being in each state at the initial moment, define the probability of transition between states, and determine the probability of generating different observations in each state;

[0016] Label the preprocessed audio data, determine the category of each audio sample, and divide it into training set and validation set;

[0017] Perform feature extraction on the audio data in the training set to obtain an observation sequence;

[0018] The Baum-Welch algorithm is used to iteratively estimate the model parameters to maximize the log-likelihood function of the training data.

[0019] Use the validation set to evaluate the trained model.

[0020] In step S2, Mel-frequency cepstral coefficients are selected for feature extraction, the extracted features are normalized, and the feature vectors of each frame of audio after feature extraction and normalization are arranged in sequence to form an observation sequence.

[0021] In step S2, the model parameters are iteratively estimated using the Baum-Welch algorithm, and the specific process includes:

[0022] Initialize model parameters;

[0023] For each observation sequence in the training set, the forward-backward algorithm is used to assist in calculating its probability under the current model parameters, and the expected number of occurrences of each state and state transition is calculated;

[0024] Update model parameters based on expected number of occurrences;

[0025] Repeat the above steps and continuously update the model parameters until the change in the log-likelihood function is less than a preset threshold or the preset maximum number of iterations is reached.

[0026] In step S3, the category of the current audio is determined, and the specific process includes:

[0027] Extract features from the current audio data to obtain an observation sequence;

[0028] Use the Viterbi algorithm to infer the most likely hidden state sequence based on the obtained observation sequence and the trained audio classification model;

[0029] Determine the category to which the audio belongs based on the inferred hidden state sequence.

[0030] The step S4 further includes detecting whether the category of the current audio has changed at a preset time interval. If so, dynamically adjusting the subtitle and translation modes according to the audio category determined in real time, and displaying the subtitle and translation content after the adjustment mode.

[0031] In step S5, the front-end framework is used to implement the option setting hiding logic.

[0032] The step S5 further includes dynamically adjusting the subtitles and translation mode according to the audio category determined in real time when a change in the audio category is detected, and updating the display status of related option settings according to the re-matched subtitles and translation mode.

[0033] The beneficial effects of the present invention are:

[0034] The present invention provides a method for dynamically adjusting subtitles and translation modes according to audio input types. The method constructs an audio classification model based on an intelligent audio signal classification algorithm, matches corresponding subtitles and translation modes according to the category of the current audio, automatically hides related option settings according to the subtitles and translation modes, reduces unnecessary setting options, improves the intelligence of the system, simplifies the user interface, optimizes the user experience, and can detect whether the category of the current audio has changed at preset time intervals. If changed, the subtitles and translation modes are dynamically adjusted according to the audio category detected in real time, and the related option settings are hidden or displayed according to the new subtitles and translation modes. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0036] See also Figure 1 FIG. 1 is a flow chart of a method of the present invention, which relates to a method for dynamically adjusting subtitle and translation modes according to audio input type, comprising the following steps:

[0037] Step S1: collecting audio data from multiple sources and preprocessing the collected audio data;

[0038] Capture audio data from multiple sources. Audio data from different sources includes different types of audio, including voice recordings, music files, environmental sounds, etc. For example, for microphone input, DirectX audio capture technology is used in Windows systems, and an audio capture device is built with the help of the DirectSound API. By setting the appropriate audio format (such as a sampling rate of 44100Hz, a 16-bit quantization depth, and dual channels), high-quality audio is captured in real time. In Linux systems, the Advanced Linux Sound Architecture (ALSA) library is used to open the audio capture device through the alsa_pcm_open function and configure the corresponding audio parameters.

[0039] For network streaming media, use open source libraries based on the HTTP protocol, such as libcurl, to capture the audio stream. During the capture process, parse the audio stream segmentation information according to the streaming protocol specifications (such as the MPEG-DASH protocol) to achieve stable audio data download.

[0040] For local audio files, the FFmpeg open source multimedia framework is used to support reading more than 100 common audio formats. The audio file is opened through the avformat_open_input function, and then the av_read_frame function is used to read the audio data frame by frame.

[0041] In step S1, a unique universal unique identifier (UUID) is assigned to each audio acquisition source, and the UUID is associated with the relevant information of the audio acquisition device or data source (such as device name, network address, file name, etc.) and stored to facilitate subsequent tracing and management of the audio source.

[0042] The collected audio data is preprocessed, including denoising, sampling, quantization and framing.

[0043] Noise removal uses a denoising method based on wavelet transforms. First, an appropriate wavelet basis function (such as the DB4 wavelet basis) is selected to perform wavelet decomposition on the audio signal, obtaining wavelet coefficients at different frequency scales. Then, based on the characteristics of noise in the wavelet domain, a threshold is set to process the wavelet coefficients and remove the corresponding wavelet coefficients. Finally, wavelet reconstruction is performed to obtain the denoised audio signal.

[0044] The continuous audio signal is discretized according to a certain sampling frequency, such as 8kHz, 16kHz, 44.1kHz, etc. The higher the sampling frequency, the richer the audio information retained.

[0045] The sampled audio values ​​are quantized to a certain precision, that is, mapped to a finite number of discrete values. Quantization precision is usually expressed in bits, such as 8, 16, or 32 bits. The higher the bit count, the higher the quality of the quantized audio.

[0046] The audio signal is divided into several short frames, with a frame length generally between 20 and 40 milliseconds, such as a frame length of 30 milliseconds and a frame shift of 10 milliseconds. For example, the numpy library of Python can be used to implement the framing operation, and the audio data is sliced ​​according to the set frame length and frame shift to form individual audio frames. For each frame of audio data, appropriate zero-padding operations are performed at the start and end positions of the frame to avoid boundary effects, thereby ensuring the accuracy and continuity of feature extraction.

[0047] Step S2: Building and training an audio classification model based on an intelligent audio signal classification algorithm to classify the audio data;

[0048] The Hidden Markov Model (HMM) is used to classify audio from different sources. The specific process includes:

[0049] Define the model structure: determine the state set based on the audio category, and determine the observation set based on the audio feature extraction method;

[0050] Analyze the characteristics of audio from different sources and classify them into different states. For example, for audio data containing speech, music, and ambient sound, the state set S = {S1, S2, S3} can be set to correspond to the speech state, music state, and ambient sound state, respectively. Determine the observation set based on the method of extracting audio features. If Mel-frequency cepstral coefficients (MFCCs) are used as features, each Mel-frequency cepstral coefficient (MFCC) vector is an observation value, and the observation set O is the set of all possible Mel-frequency cepstral coefficient (MFCC) vectors.

[0051] Initialize model parameters: set the probability of the audio data being in each state at the initial moment, define the probability of transition between states, and determine the probability of generating different observations in each state;

[0052] Set the probability of the audio data being in each state at the initial moment. Assuming the prior probabilities of speech, music, and ambient sound are 0.4, 0.3, and 0.3, respectively, the initial state probability distribution π = [0.4, 0.3, 0.3].

[0053] Define the probability of transition between states. For example, the probability of speech remaining after speech is 0.7, the probability of speech transitioning to music is 0.2, and the probability of speech transitioning to ambient sound is 0.1; the probability of music transitioning to music after music is 0.6, and the probabilities of transitioning to speech and ambient sound are 0.2 and 0.2 respectively; the probability of ambient sound transitioning to ambient sound after ambient sound is 0.8, and the probabilities of transitioning to speech and music are 0.1 and 0.1 respectively. The state transition probability matrix A is:

[0054]

[0055] Determine the probability of generating different observations in each state. Assume that the probability of generating a specific MFCC vector is 0.6 in the speech state, 0.3 in the music state, and 0.1 in the ambient sound state. For different MFCC vectors, the values ​​in the observation probability matrix B will change accordingly.

[0056] Label the preprocessed audio data, determine the category of each audio sample, and divide it into training set and validation set;

[0057] The preprocessed audio data is labeled to determine the category of each audio sample. The labeling information includes audio type (speech, music, ambient sound, etc.), audio source (microphone, network, local file, etc.), speaker information (if it is speech type audio), music style (if it is music type audio), etc.

[0058] Perform feature extraction on the audio data in the training set to obtain an observation sequence;

[0059] Feature extraction is performed on the preprocessed audio data using Mel-frequency cepstral coefficients (MFCCs). First, a Fast Fourier Transform (FFT) is performed on each frame of audio to convert the time-domain signal into the frequency domain. The frequency-domain signal is then filtered using a Mel filter bank, converting it to a Mel-frequency scale to simulate the human ear's perception of sound frequency. Next, the filtered signal is logarithmized and subjected to a discrete cosine transform (DCT) to obtain the Mel-frequency cepstral coefficients. Typically, 12-40 Mel-frequency cepstral coefficients are extracted as feature vectors for the audio, effectively representing the audio's spectral characteristics.

[0060] In addition to Mel-frequency cepstral coefficients, other features can be extracted, such as linear predictive coding (LPC) coefficients (which describe the spectral envelope of the audio by predicting the linear relationship between the current audio sample and past samples), as well as spectral centroid, spectral bandwidth, and zero-crossing rate, which respectively represent the spectral center of gravity position, spectral width, and frequency of signal zero crossing of the audio signal. These features can be used alone or in combination with Mel-frequency cepstral coefficients (MFCCs) to more comprehensively describe the audio data.

[0061] The extracted features are normalized, mapping them to a specific range, such as normalizing feature values ​​to the interval [0, 1] or [-1, 1]. This eliminates dimensional differences between features and improves model training effectiveness and stability. Normalization methods used include min-max normalization and Z-score normalization. After feature extraction and normalization, the feature vectors of each frame of audio are arranged in sequence to form a feature sequence, which is the observation sequence in the hidden Markov model. This observation sequence serves as the model input for training and subsequent audio classification tasks.

[0062] Parameter estimation: Use the Baum-Welch algorithm to iteratively estimate the model parameters (initial state probability distribution π, state transition probability matrix A, observation probability matrix B) to maximize the log-likelihood function of the training data. The specific process includes:

[0063] Initialize model parameters, including initial state probability distribution π, state transition probability matrix A, and observation probability matrix B;

[0064] For each observation sequence in the training set, the forward-backward algorithm is used to assist in calculating its probability under the current model parameters, and the expected number of occurrences of each state and state transition is calculated;

[0065] Forward algorithm: For a given observation sequence O = (o1, o2, ..., o T ), T represents the length of the observation sequence, and the forward variable α1(i) is calculated, which means that at time t, the state is i and the sequence o1, o2, ..., o is observed. t probability.

[0066] Initialization: α1(i)=π(i)b i (o1), where π(i) is the probability of state i in the initial state probability distribution, b i (o1) is the probability that state i in the observation probability matrix B produces observation value o1.

[0067] Recursion: where a ij is the probability of transitioning from state i to state j in the state transition probability matrix A, and N is the total number of states.

[0068] termination: Where P(0|λ) is the probability of the observation sequence O under the model λ = (π, A, B).

[0069] Backward algorithm: Calculate the backward variable β t (i), means that at time t, it is in state i and observes sequence o t+1 , o t+2 ,…,o T probability.

[0070] Initialization: β T (i) = 1, for all i = 1, 2, ..., N.

[0071] Recursion:

[0072] Calculate the expected number of occurrences of states and state transitions

[0073] Calculate ξ t (i,j): It represents the probability of being in state i at time t and transitioning to state j at time t+1 given the observation sequence O and model λ.

[0074] Calculate γ t (i) It represents the probability of being in state i at time t given the observation sequence O and model λ.

[0075] Update model parameters

[0076] Update the initial state probability distribution π: π(i) = γ1(i): that is, the initial state probability distribution is updated to the probability of being in state i at time 1.

[0077] Update the state transition probability matrix A: That is, the state transition probability a ij Updated to the ratio of the expected number of transitions from state i to state j to the expected number of times in state i.

[0078] Update the observation probability matrix B: where v k is the kth element in the set of observations, i.e., the observation probability b j (k) Update to generate observation value v in state j k The ratio of the expected number of times to the expected number of times in state j.

[0079] Update model parameters according to the expected number of occurrences, including the initial state probability distribution π, the state transition probability matrix A, and the observation probability matrix B;

[0080] Repeat the above steps and continuously update the model parameters π, A, and B until the change in the log-likelihood function is less than a preset threshold or the preset maximum number of iterations is reached. The model parameters obtained at this time are the optimal parameters estimated by the Baum-Welch algorithm, which makes the model have the maximum log-likelihood under the given training data.

[0081] Use the validation set to evaluate the trained model and calculate metrics such as classification accuracy and recall. Based on the evaluation results, analyze whether the model performance meets the requirements. If the model performance does not meet the requirements, adjust the model structure or parameters and retrain and evaluate.

[0082] Step S3: Input the current audio into the audio classification model to determine the category of the current audio;

[0083] Extract features from the current audio data to obtain an observation sequence;

[0084] Use the Viterbi algorithm to infer the most likely hidden state sequence based on the obtained observation sequence and the trained audio classification model;

[0085] Assume that the observation sequence is o = {o1, o2, ..., o T}, where T is the length of the observation sequence, and the state set of the hidden Markov model (audio classification model) is S = {s1, s2, ..., s N}, where N is the number of states.

[0086] Initialize the probability matrix δ: δ1(i)=π(i)b i (o1), where i = 1, 2, ..., N, δ1(i) represents the maximum probability of being in state i at time 1 and observing the first observation value o1.

[0087] Initialize the path matrix ψ: ψ1(i) = 0, i = 1, 2, ..., N, ψ is used to record the previous optimal state of each state at each moment.

[0088] For t = 2, 3, ..., T and i = 1, 2, ..., N, perform the following operations:

[0089] Update probability matrix δ: δ t (i)=max 1≤j≤N [δ t-1 (j)a ji ]b i (o t ). This step calculates the value o at time t when the state is in state i and the observation value is observed. t The maximum probability of δ is obtained by selecting one of all possible states j at time t-1 that can make δ t-1 (j)a jiThe maximum value, multiplied by state i, generates the observation value o t The probability b i (o t ) obtained.

[0090] Update path matrix ψ: ψ t (i) = arg max 1≤j≤N [δ t-1 (j)a ji ]. t (i) Records the optimal state of the previous moment when the state is in state i at time t.

[0091] Determine the maximum probability state at the final moment: p * =max 1≤i≤N [δ T (i)], P * is the maximum probability of the entire observation sequence, It is the optimal state at the final moment.

[0092] Starting from the final moment T, the most likely hidden state sequence is obtained by backtracking according to the path matrix For t=T-1,T-2,…,1,

[0093] The audio is classified into the speech category based on the inferred hidden state sequence. For example, if most of the states in the hidden state sequence are speech state S1, the audio is classified as speech.

[0094] Step S4: Constructing a mapping relationship between audio categories, subtitles, and translation modes, matching the corresponding subtitles and translation modes according to the current audio category, and displaying the subtitles and translation content in real time;

[0095] A detailed mapping table is constructed. Once the audio type and source are identified, the corresponding subtitle generation and translation modules are automatically loaded by querying the mapping table. For example, if the audio is identified as a "lecture" type and the source is an "online education platform video," the corresponding subtitle generation function is called to generate subtitles according to the preset subtitle style, and a professional education translation engine is used for translation.

[0096] Matching subtitles and translation modes to the current audio category means using different display methods for different audio types, ensuring that the subtitles and translation modes are tailored to the audio content and more convenient for users. The following shows the subtitle and translation modes corresponding to three different audio sources and types.

[0097] For audio of the "lecture" type and sourced from "online education platform video", the mapped subtitle mode is "displaying the lecture content in chapters, highlighting key titles and key knowledge points, using Microsoft YaHei font, size 16, white color, and a translucent black background." The translation mode is "calling a professional education translation engine, prioritizing the translation of academic terms, using a sentence-by-sentence translation method, with the translation results displayed below the original subtitles in Arial font, size 14, and gray color."

[0098] For audio of the "Pop Music" type and sourced from the "Music Playback App", if it is a song, the subtitle mode is "dynamically displaying lyrics according to the song's beat, with the lyrics color changing with the song's emotions (e.g., bright colors in the climax and soft colors in the chorus), and the font is DynaFont Girls, size 18." The translation mode is "adopting a pop culture translation model, focusing on the translation of buzzwords and emotional expressions in the lyrics, and appropriately adding pop culture elements for explanation. The translated subtitles are displayed to the right of the original lyrics in Kaiti, size 16." If it is pure music, the subtitle mode is "displaying text information such as the music's creative background and theme introduction in bold, size 16, and blue." The translation mode is "accurately and fluently translating descriptive text, with the translated subtitles aligned above and below the original subtitles in Songti, size 14."

[0099] For audio of the "Multi-person Conference Discussion" type and sourced from the "Conference Room Microphone Array," the subtitle mode is "Real-time display of each speaker's name and speech content, arranged in order of speaking. The subtitle font is Founder Lanting Black, size 14, and green, with different subtitle background colors for different speakers." The translation mode is "Using the business conference translation model, supporting real-time translation in multiple languages, and displaying translation results by language. The font is Times New Roman, size 12."

[0100] The step S4 further includes detecting whether the category of the current audio has changed at a preset time interval. If so, dynamically adjusting the subtitle and translation modes according to the audio category determined in real time, and displaying the subtitle and translation content after the adjustment mode.

[0101] For example, use a Python timer module (such as threading.Timer) to set a preset interval of 3 seconds. Each time the timer triggers, a new thread is started to detect the current audio category. In the detection thread, a new segment of the current audio data (such as the last 1 second of audio data) is reframed, feature extracted, and classified.

[0102] If a change in the audio category is detected, the currently running subtitle generation and translation task is immediately stopped. According to the new audio category detected, the mapping table is re-queried to determine the corresponding subtitle and translation mode, and a new subtitle generation and translation task is started.

[0103] For example, if the audio is currently being processed according to the subtitle and translation mode of the "music" type, and it is detected that the audio is switched to the "speech" type, the operation of the music subtitle generation module and the music translation module is stopped by calling the corresponding stop function. Then, based on the new audio category identified in real time, the mapping table is re-queried to determine the corresponding subtitle and translation mode. Next, the new subtitle generation module and translation module are loaded, and the audio is processed according to the new mode. For example, for the new "speech" type audio, the speech subtitle generation module and the speech translation module are loaded, and subtitles are generated and translated according to the subtitle and translation mode of the speech.

[0104] In the process of performing corresponding subtitle generation and translation tasks according to the audio type, the audio recognition model and the machine translation system can be combined to match the extracted audio features with the pre-trained audio recognition model, map the audio features to the corresponding phonemes or syllables, and decode the recognition results in combination with the language model. The recognized phonemes or syllables are converted into text and arranged in chronological order to generate subtitle text.

[0105] The generated source language subtitles are input into a machine translation system. Based on neural networks or statistical models, the system establishes a mapping relationship between the languages ​​by learning from a large parallel corpus (i.e., text data corresponding to the source and target languages). The translated subtitles are formatted based on the characteristics of the target language and the requirements for subtitle display. This includes adjusting the font, size, color, and position of the subtitles to ensure that the subtitles are displayed clearly and aesthetically on the user interface, coherently with the audio and visuals, and optimizing the user experience.

[0106] Step S5: Based on the subtitles and translation mode matched to the current audio, the relevant option settings are automatically hidden.

[0107] In step S5, a front-end framework (such as Vue.js) is used to implement the option setting hiding logic.

[0108] In user interface development, front-end frameworks (such as Vue.js) are used to implement option hiding logic. When the audio category is "pure music," Vue.js directives (such as the v-if directive) are used to hide the "speaker" (speaker-related settings) and "mixed input" options during interface rendering. When the audio category is "ambient sound," the v-if directive is also used to hide the "speaker" option. When the audio category changes, Vue.js's responsive data mechanism is utilized to update the interface's option display status in real time. For example, when the audio type switches from "speech" to "pure music," Vue.js detects the change in audio category data and immediately hides the relevant options according to the option hiding logic. When the audio type switches back to "speech," the "speaker" and other related options are redisplayed. Furthermore, to enhance the user experience, transition animation effects (such as fade-in and fade-out effects) are added when options are hidden and displayed, making interface changes smoother and more natural.

[0109] The step S5 further includes dynamically adjusting the subtitles and translation mode according to the audio category determined in real time when a change in the audio category is detected, and updating the display status of related option settings according to the re-matched subtitles and translation mode.

[0110] For example, for audio in the music category, options include rhythm adjustment (used to adjust the speed of the music), music style selection (providing a variety of music style options, such as pop, rock, classical, etc., to facilitate users to categorize current music or find similar music based on style), album cover display, and singer information display options, as well as hidden speaker recognition and ambient noise suppression (music itself is the main subject that users want to listen to, and there is no need to suppress itself, so this option can be hidden in this scenario).

[0111] For audio in the voice category, the options of speech-to-text (converting voice content into text for user viewing, editing, or archiving, which is very practical in scenarios such as speeches, meeting minutes, and voice messages), speaker recognition (if there are multiple people speaking, different speakers can be identified, which helps to distinguish the speeches of different roles, such as in interviews, multi-person conversations, and other audio), and speech speed adjustment (which allows users to adjust the speed of voice playback based on their listening habits or the difficulty of the voice content to better understand the content) are displayed, while the rhythm adjustment and music style selection options are hidden.

[0112] For audio in the ambient sound category, the options of ambient noise suppression (for some ambient sounds containing noise, such as street noise, construction site noise, etc., it can help users suppress background noise and highlight the main ambient sound characteristics), scene classification, and sound effects addition (you can add some special effects to ambient sound, such as echo, reverberation, etc. to create a different atmosphere or enhance the listening experience) are displayed, and the speech-to-text and rhythm adjustment options are hidden.

[0113] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A method for dynamically adjusting subtitle and translation modes according to audio input type, characterized in that: The following steps are involved: Step S1: collecting audio data from multiple sources and preprocessing the collected audio data; Step S2: Building and training an audio classification model based on an intelligent audio signal classification algorithm to classify the audio data; Step S3: Input the current audio into the audio classification model to determine the category of the current audio; Step S4: Constructing a mapping relationship between audio categories, subtitles, and translation modes, matching the corresponding subtitles and translation modes according to the current audio category, and displaying the subtitles and translation content in real time; Step S5: Based on the subtitles and translation mode matched to the current audio, the relevant option settings are automatically hidden.

2. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 1, characterized in that: In step S1, a unique identification code is assigned to each audio acquisition source, and the unique identification code is associated with the relevant information of the audio acquisition device or data source and stored.

3. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 1, characterized in that: In step S2, a hidden Markov model is used to classify audio from different sources. The specific process includes: Define the model structure: determine the state set based on the audio category, and determine the observation set based on the audio feature extraction method; Initialize model parameters: set the probability of the audio data being in each state at the initial moment, define the probability of transition between states, and determine the probability of generating different observations in each state; Label the preprocessed audio data, determine the category of each audio sample, and divide it into training set and validation set; Perform feature extraction on the audio data in the training set to obtain an observation sequence; The Baum-Welch algorithm is used to iteratively estimate the model parameters to maximize the log-likelihood function of the training data. Use the validation set to evaluate the trained model.

4. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 3, characterized in that: In step S2, Mel-frequency cepstral coefficients are selected for feature extraction, the extracted features are normalized, and the feature vectors of each frame of audio after feature extraction and normalization are arranged in sequence to form an observation sequence.

5. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 3, characterized in that: In step S2, the model parameters are iteratively estimated using the Baum-Welch algorithm, and the specific process includes: Initialize model parameters; For each observation sequence in the training set, the forward-backward algorithm is used to assist in calculating its probability under the current model parameters, and the expected number of occurrences of each state and state transition is calculated; Update model parameters based on expected number of occurrences; Repeat the above steps and continuously update the model parameters until the change in the log-likelihood function is less than a preset threshold or the preset maximum number of iterations is reached.

6. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 1, characterized in that: In step S3, the category of the current audio is determined, and the specific process includes: Extract features from the current audio data to obtain an observation sequence; Use the Viterbi algorithm to infer the most likely hidden state sequence based on the obtained observation sequence and the trained audio classification model; Determine the category to which the audio belongs based on the inferred hidden state sequence.

7. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 1, characterized in that: The step S4 further includes detecting whether the category of the current audio has changed at a preset time interval. If so, dynamically adjusting the subtitle and translation modes according to the audio category determined in real time, and displaying the subtitle and translation content after the adjustment mode.

8. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 1, characterized in that: In step S5, the front-end framework is used to implement the option setting hiding logic.

9. The method of dynamically adjusting subtitle and translation modes according to audio input type according to claim 8, characterized in that: The step S5 further includes dynamically adjusting the subtitles and translation mode according to the audio category determined in real time when a change in the audio category is detected, and updating the display status of related option settings according to the re-matched subtitles and translation mode.

Citation Information

Patent Citations

  • Methods and systems for performing synchronization of audio with corresponding textual transcriptions and determining confidence values of the synchronization

    CN103003875A

  • Audio file display method, terminal and computer storage medium

    CN107729315A

  • Methods and systems of recommending media assets to users based on content of other media assets

    CN109074391A

  • Audio environment display method and device

    CN110099332A

  • Acoustic scene classification model training method and device, intelligent terminal and storage medium

    CN114627895A