Video subtitle information generation method and device, equipment, storage medium and program product
By capturing real-time audio data from the audio output stream and performing language recognition and translation, the problem that subtitle generation in the prior art cannot meet real-time performance is solved, and real-time generation and convenient display of video subtitles are realized.
Patent Information
- Application Number
- CN202510575629.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
AI Technical Summary
The existing subtitle generation method relies on pre-made translated text and cannot meet real-time needs such as live video streams and conference video streams.
Real-time audio data is captured from the audio output stream, audio language and subtitle information are obtained through audio recognition, and video subtitles are generated based on the language, and the pre-trained end-side AI model is used for real-time translation and display.
Real-time generation of video subtitles is realized, which improves the real-time and convenience of subtitles generation, reduces translation costs, and reduces user viewing delays.
Smart Images

Figure CN120455767A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, apparatus, device, storage medium, and program product for generating video subtitle information. Background Art
[0002] With the development of globalization, the trend of multilingual communication and multicultural integration is becoming increasingly prominent, and the demand for cross-language communication is also increasing. As an important form of media communication, video often requires subtitles in multiple languages to meet the needs of different audiences.
[0003] Traditional subtitle translation relies on pre-made translation texts and cannot meet the subtitle needs of video streams with high real-time requirements, such as live video streams and conference video streams. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device, equipment, storage medium and program product for generating video subtitle information, aiming to solve the technical problem that the existing subtitle generation method relies on pre-made translation text and cannot meet the real-time requirements of the video.
[0005] To achieve the above objectives, the present application proposes a method for generating video subtitle information, the method comprising:
[0006] Capture real-time audio data from the audio output stream;
[0007] Performing audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data;
[0008] Video subtitle information is determined based on the audio language and the first subtitle information.
[0009] In addition, to achieve the above-mentioned purpose, the present application also proposes a video subtitle information generation device, the video subtitle information generation device comprising:
[0010] Audio acquisition module, used to capture real-time audio data from the audio output stream;
[0011] an audio recognition module, configured to perform audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data;
[0012] The subtitle generation module is used to determine video subtitle information based on the audio language and the first subtitle information.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a video subtitle information generation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the video subtitle information generation method as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the video subtitle information generation method described above are implemented.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the video subtitle information generation method as described above are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart of the first embodiment of the method for generating video subtitle information of the present application is provided;
[0019] Figure 2 A flowchart of the second embodiment of the method for generating video subtitle information of this application is provided;
[0020] Figure 3 This is a schematic diagram of the first acoustic model structure in one implementation of the present application;
[0021] Figure 4 A flowchart of the third embodiment of the method for generating video subtitle information of this application is provided;
[0022] Figure 5 This is a schematic diagram of a scenario flow in one implementation of the method for generating video subtitle information of this application;
[0023] Figure 6 This is a schematic diagram of the module structure of the video subtitle information generating device according to an embodiment of the present application;
[0024] Figure 7Schematic diagram of the device structure of the hardware operating environment involved in the video subtitle information generation method in the embodiment of the present application.
[0025] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0026] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0027] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0028] The main solution of the embodiment of the present application is: capturing real-time audio data from the audio output stream; performing audio recognition on the real-time audio data to obtain the audio language and first subtitle information of the real-time audio data; and determining the video subtitle information based on the audio language and the first subtitle information.
[0029] The present application provides a solution that directly captures real-time audio data from the audio output stream of a video playback device, and obtains video subtitle information corresponding to a preset language by identifying and translating the real-time audio data. This greatly improves the real-time performance of subtitle generation, eliminates the need for pre-production of subtitles for the video, improves the real-time performance and practicality of the subtitle generation method, and reduces the labor cost required for translation.
[0030] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a video playback device, or a server or computer connected to the video playback device, or an electronic device or virtual device capable of implementing the above functions. The following describes this embodiment and the following embodiments using a video subtitle information generation device (hereinafter referred to as the generation device) as an example.
[0031] It can be understood that the generating device in the embodiment of the present application can be a video playback device for video playback, or a video processing unit or video processing device placed in the video playback device, or a computer or server connected to the video playback device, etc. The embodiment of the present application is not limited to this.
[0032] It should be understood that the above-mentioned video playback device is also a device that can perform functions such as video playback and video projection, such as a television, a mobile phone, a computer, etc., and the embodiments of the present application are not limited to this.
[0033] Based on this, the embodiment of the present application provides a method for generating video subtitle information, referring to Figure 1 , Figure 1A flowchart of the first embodiment of the method for generating video subtitle information of this application is provided.
[0034] In this embodiment, the method for generating video subtitle information includes steps S10 to S30:
[0035] Step S10: capturing real-time audio data from the audio output stream.
[0036] It should be noted that when a video playback device plays a video, it can synchronously decode audio data from a video file or other audio source to obtain an audio output stream, and output the audio output stream to an audio playback device such as a speaker or headphones through an audio output interface, thereby realizing sound playback.
[0037] It can be understood that the above-mentioned real-time audio data is also the audio data used for playback. In actual applications, the generating device can capture the real-time audio data from the audio output stream through the audio capture interface in the video playback device, and then realize the generation of video subtitle information based on the real-time audio data.
[0038] In some implementations of the present application, the generating device may first obtain the audio management permission of the video playback device. Upon successfully obtaining the audio management permission, the generating device may capture real-time audio data from the audio output stream of the playback device. Specifically, the step of capturing real-time audio data from the audio output stream includes: obtaining the audio management permission of the video playback device; and capturing real-time audio data from the audio output stream of the video playback device based on the audio management permission.
[0039] It should be noted that the aforementioned audio management permissions may be required for recording audio on a video playback device. On some operating systems (such as Android and iOS), since recording audio involves user privacy and can be a sensitive operation, audio management permissions (such as recording permissions) may be obtained from the video playback device before capturing real-time audio data from the audio output stream.
[0040] Exemplarily, the video playback device of the embodiment of the present application can be an Android device installed with an Android operating system. The generating device of the embodiment of the present application can utilize the recording permission of the video playback device and obtain real-time audio data in the audio output stream through a preset interface.
[0041] It should be noted that the preset interface can be the AudioRecord interface of an Android device. The AudioRecord interface is a class provided by the Android platform for audio recording. Based on the AudioRecord interface, the generating device can obtain audio data through the underlying audio hardware (such as a microphone) of the video playback device.
[0042] In a specific implementation, the generating device can utilize the recording permission of the video playback device and capture real-time audio data from the audio output stream of the video playback device through the AudioRecord interface.
[0043] Step S20: Perform audio recognition on the real-time audio data to obtain the audio language and first subtitle information of the real-time audio data.
[0044] It should be noted that the audio language may be a language such as Chinese, English, Japanese, French, etc. obtained through speech content recognition from real-time audio data, and this embodiment of the present application does not limit this. The first subtitle information may be text information converted based on the speech content in the real-time audio data.
[0045] In some implementations of the embodiments of the present application, a pre-trained end-side AI model can be used to convert real-time audio data into text. The end-side AI model can be a speech recognition model based on deep learning, such as RNN, Transformer, etc., which is not limited by the embodiments of the present application.
[0046] In a specific implementation, the generating device can perform audio recognition on the captured real-time audio data to determine the voice content in the real-time audio data, and then analyze the voice content to determine the audio language and the first subtitle information.
[0047] Step S30: Determine video subtitle information based on the audio language and the first subtitle information.
[0048] It should be noted that the above-mentioned preset language may be a language type set according to the current language environment.
[0049] For example, when the current language environment is Chinese, the preset language can be set to Chinese; when the current language environment is English, the preset language can be set to English; when the current language environment is French, the preset language can be set to French.
[0050] In some implementations of the embodiments of the present application, the setting of the above-mentioned preset language can be set by the manufacturer of the video playback device or the generation device, or can be set by the user according to usage requirements, and the embodiments of the present application are not limited to this.
[0051] It is understandable that when the audio language of the real-time audio data is a preset language, the first subtitle information may not be translated, and the video subtitle information corresponding to the real-time audio data may be obtained by directly performing standardization and other processing based on the first subtitle information.
[0052] In some implementations of the embodiments of the present application, the preset language may include multiple languages, such as a first preset language and a second preset language. When the audio language is the first preset language, the first subtitle information corresponding to the first preset language may be translated into the second initial video subtitle information corresponding to the second preset language; when the audio language is the second preset language, the second initial video subtitle information corresponding to the second language may be translated into the first initial video subtitle information corresponding to the first preset language. When the first initial video subtitle information and the second initial video subtitle information are obtained, these initial video subtitle information may be standardized to obtain the video subtitle information.
[0053] It is understandable that when the video subtitle information corresponding to the real-time audio is obtained, the video subtitle information can be displayed.
[0054] In some implementations of the present application, when translating the first subtitle information, a device-side translation model can be used to translate the identified first subtitle information into video subtitle information corresponding to a preset language, and the video subtitle information is overlaid and displayed on the video. The device-side translation model can be a neural network-based machine translation model, such as Seq2Seq or Transformer.
[0055] The embodiments of the present application capture real-time audio data from an audio output stream; perform audio recognition on the real-time audio data to obtain the audio language and first subtitle information of the real-time audio data; and determine video subtitle information based on the audio language and first subtitle information. By integrating language recognition and translation technology, subtitles are generated quickly, reducing viewing delays for users. Furthermore, by eliminating the need for pre-production of subtitles, the convenience and adaptability of video streaming playback are improved.
[0056] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , Figure 2 A flowchart of the second embodiment of the method for generating video subtitle information of this application is provided.
[0057] like Figure 2 As shown, in the embodiment of the present application, the step of performing audio recognition on the real-time audio data to obtain the audio language and first subtitle information of the real-time audio data includes:
[0058] Step S21: storing the real-time audio data into a ring buffer structure.
[0059] It should be noted that the ring buffer structure can be a data structure used for data buffering. It treats the buffer as a ring structure. When data is written to the end of the buffer, it automatically wraps around to the beginning of the buffer to continue writing, forming a ring loop. At the same time, data can also be read from the ring buffer in a circular manner. This data structure allows data to be stored and read in a circular manner in the ring buffer, thereby achieving efficient use of cache space. By storing real-time audio data in a ring buffer, continuous audio stream processing can be achieved while avoiding memory overflow, ensuring the continuity of the audio stream while improving audio reading and processing efficiency.
[0060] Step S22: performing standardization processing on the buffered audio data in the ring buffer to obtain standard audio data.
[0061] It should be noted that, in order to improve the processing accuracy and processing compatibility of audio and simplify the audio processing operation, the embodiment of the present application can perform standardization processing on the buffered audio data in the ring buffer, thereby converting the digital range of the buffered audio data from one representation to another more universal representation. Specifically, the standardization processing at least includes: floating-point conversion; the step of performing standardization processing on the buffered audio data in the ring buffer to obtain standard audio data includes: reading the buffered audio data in the ring buffer to obtain audio data to be processed; performing floating-point conversion on the audio data to be processed to obtain standard audio data.
[0062] It should be noted that the buffered audio data in the ring buffer can be read sequentially to obtain the audio data to be processed. This audio data is then converted to floating-point and normalized to the range of [-1.0, 1.0]. This normalized data can reduce numerical instability during model training and inference, thereby improving model stability. Furthermore, the standardized audio data obtained through normalization can accelerate model convergence and make the model perform more consistently when processing different audio inputs, thereby improving inference efficiency.
[0063] In some implementations of the embodiments of the present application, since audio data is typically stored in the form of 16-bit integers, its range is -32768 to 32767. Therefore, the embodiments of the present application can divide the audio data to be processed by a preset value of 32768.0f to achieve floating-point conversion of the audio data and normalize the audio data to be processed to the range of [-1.0, 1.0].
[0064] Step S23: Perform audio recognition on the standard audio data to obtain the audio language and first subtitle information.
[0065] It is understandable that after obtaining the standardized audio data, audio recognition can be performed on the standard audio data to obtain the audio language and the first subtitle information.
[0066] In some implementations of the embodiments of the present application, an acoustic model (speech recognition model) based on deep learning can be used to perform audio recognition on standard audio data. Specifically, it can be a speech recognition model based on a long short-term memory network, a convolutional neural network or other networks. The embodiments of the present application are not limited to this.
[0067] In some implementations of the embodiments of the present application, before the step of performing audio recognition on the standard audio data to obtain the audio language and the first subtitle information, the step further includes:
[0068] Acquire pre-trained voice data and pre-trained text data, wherein the pre-trained voice data and the pre-trained text data correspond to each other;
[0069] Preprocessing the pre-trained voice data and the pre-trained text data respectively to obtain pre-processed voice data and pre-processed text data;
[0070] The preprocessed speech data and the preprocessed text data are input into a first acoustic model for training to obtain a speech recognition model.
[0071] In an embodiment of the present application, a large amount of multi-language pre-trained voice data and corresponding pre-trained text data can be collected, and these data can cover different language types, accents and contexts.
[0072] Furthermore, the pre-trained speech data and pre-trained text data can be pre-processed. Specific pre-processing methods may include noise reduction, segmentation, feature extraction (such as MFCC, Mel spectrogram), etc. The pre-processed speech data and pre-processed text data obtained after pre-processing can be input into the first acoustic model for training, and loss evaluation is performed using a corresponding loss function to obtain a speech recognition model.
[0073] Specifically, the first acoustic model used in the embodiment of the present application can be constructed based on a network such as a recurrent neural network, a long short-term memory network, or a transformer, and the embodiment of the present application is not limited to this. The loss function used in the embodiment of the present application can be a CTC (Connectionist Temporal Classification) loss function or a cross-entropy loss function to optimize the speech-to-text conversion capability of the speech recognition model.
[0074] In the embodiment of the present application, audio recognition can be performed on standard audio data based on a pre-trained speech recognition model to obtain the audio language and corresponding first subtitle information. Specifically, the step of performing audio recognition on the standard audio data to obtain the audio language and first subtitle information includes: performing audio recognition on the standard audio data based on the speech recognition model to obtain the audio language and first subtitle information.
[0075] It should be noted that the speech recognition model of the embodiment of the present application can be used to perform audio recognition on standard audio data or real-time audio data, and then obtain the first subtitle information corresponding to the audio data and the audio language corresponding to the first subtitle information.
[0076] In an embodiment of the present application, the recognition process of standard audio data can be divided into a text recognition process and a language recognition process, that is, the recognition process of standard audio data can include: performing text recognition on the standard audio data to obtain first subtitle information; performing language recognition on the first subtitle information to obtain the audio language.
[0077] It should be noted that the speech recognition models used in the above-mentioned text recognition process and language recognition process can be the same model or different models, and the embodiments of the present application do not limit this.
[0078] In the embodiments of the present application, the training process of the speech recognition model for language recognition is taken as an example to illustrate the training process of each model in the embodiments of the present application. The training of other models can refer to this process, and the embodiments of the present application are not limited to this.
[0079] Specifically, if Figure 3 As shown, Figure 3 This is a schematic diagram of the first acoustic model structure in one implementation of the present application. The first acoustic model of the embodiment of the present application may include: an input layer, an embedding layer, an encoder, a pooling layer, a classification layer, an activation function layer and an output layer.
[0080] The step of inputting the preprocessed speech data and the preprocessed text data into a first acoustic model for training to obtain a speech recognition model includes:
[0081] The preprocessed text data is input into the word segmenter of the input layer for word segmentation processing to obtain text segmentation; the text segmentation is converted into a dense vector representation through the embedding layer; the context features of the dense vector are extracted through the encoder to obtain sequence features; the sequence features are feature pooled through the pooling layer to obtain pooling features of a preset length; the pooling features are mapped to language categories through the classification layer to obtain classification layer output; the classification layer output is converted through the activation function layer to obtain a language probability distribution; the language probability distribution is predicted through the output layer to obtain the target language corresponding to the preprocessed speech data; the cross entropy loss is determined according to the target language and the actual language of the preprocessed speech data; the model parameters of the first acoustic model are iterated according to the cross entropy loss; when the iteration termination condition is met, the first acoustic model that meets the iteration termination condition is used as the speech recognition model.
[0082] It should be noted that the pre-processed text data may be text data corresponding to the pre-processed voice data and obtained by translating the pre-processed voice data. By inputting the pre-processed text data into the input layer for word segmentation processing, the segmented text data can be obtained.
[0083] It should be explained that the input layer of the embodiment of the present application may include a pre-trained word segmenter, through which special tokens (such as [CLS] and [SEP]), padding, truncation, etc. can be added to the input text, thereby converting the input text into token IDs (text segmentation).
[0084] It should be noted that the embedding layer can convert text segmentation into dense vector representations. The specific output shape of the embedding layer can be expressed as: (batch_size, sequence_length, hidden_size). Among them, batch_size is the batch size, which is used to indicate the number of samples processed in one training or inference; sequence_length is the sequence length, which is used to indicate the number of elements contained in the input sequence in each sample; hidden_size is the embedding dimension, which is used to indicate the dimension size of the dense vector mapped to each input element.
[0085] It can be understood that dense vector representation is a data representation method used to map text, words or other objects to a continuous numerical vector space of fixed dimension. Unlike sparse vectors, dense vectors can efficiently encode the semantic information of text segmentation in a compact numerical form.
[0086] It should be noted that the encoder used in the embodiments of the present application may be a Transformer encoder, which may include a multi-layer attention mechanism and a feedforward neural network. The Transformer encoder can extract contextual features of dense vectors to obtain a feature sequence, whose specific output can be expressed as: (batch_size, sequence_length, hidden_size).
[0087] It should be explained that the above-mentioned pooling layer can be used to aggregate sequence features into vectors of preset length (pooling features). The preset length can be selected according to the needs of actual applications, and the embodiments of the present application are not limited to this.
[0088] Specifically, the vector of [CLS] token can usually be used as the representation of the entire sequence, and the output shape can be: (batch_size, hidden_size).
[0089] It should be noted that the classification layer can map the pooled features to language categories. Specifically, a fully connected layer can be used in the classification layer to map the hidden_size dimension features to the num_classes dimension (number of languages) to obtain the language probability distribution. The output shape of the classification layer can be expressed as: (batch_size, num_classes).
[0090] It should be noted that the activation function layer in the embodiments of the present application can use a Softmax function to convert the output of the classification layer into a probability distribution. The Predicted Language Code node in the output layer can use the argmax function to find the category index with the highest probability from the language probability distribution output by the classification layer. Based on this index, the corresponding target language and the language code corresponding to the target language are obtained from a predefined list of language codes, for example, "zh" for Chinese and "en" for English.
[0091] It should be noted that the target language is also the predicted language obtained by the model. Based on the predicted language and the actual language corresponding to the preprocessed text data, the model loss assessment can be performed.
[0092] Specifically, the loss function used in this application for loss evaluation can be cross-entropy loss. Cross-entropy loss can be used to measure the difference between the probability distribution predicted by the model and the true label distribution.
[0093] Specifically, the formula for cross entropy loss can be shown as follows:
[0094]
[0095] Among them, N represents the number of samples, C represents the number of categories (number of languages), and y i,j represents the true label of the (i, j)th sample, p i,j It is used to represent the probability that each sample of the i-th class predicted by the model belongs to the j-th class.
[0096] It is understood that the model parameters of the first acoustic model can be iterated based on the calculated cross-entropy loss. When the iteration termination condition is met, the first acoustic model that meets the iteration termination condition can be used as the speech recognition model. If the iteration termination condition is not met, the process returns to the step of inputting the preprocessed text data into the word segmenter of the input layer for word segmentation processing, and further iteration is performed to obtain text word segmentation.
[0097] In some implementations of the embodiments of the present application, the above-mentioned iterative termination conditions may be: the number of iterations is greater than the preset maximum number of iterations, the performance on the validation set no longer improves, early stopping (3. If the validation set performance does not improve within several consecutive rounds (such as patience = 5), then the training is terminated early), and the training loss and the validation loss are both regionally stable (that is, the amplitude of change is less than the preset loss threshold). When any one or more of the above-mentioned iterative termination conditions are met, the model can be considered to have converged, that is, the training is completed.
[0098] This embodiment of the present application stores real-time audio data in a circular buffer to obtain buffered audio data; standardizes the buffered audio data in the circular buffer to obtain standard audio data; and performs audio recognition on the standard audio data to obtain the audio language and first subtitle information. Because the audio data is stored and processed in a circular buffer, the continuity of the audio stream is ensured while improving audio processing efficiency. Standardization also improves the stability, inference efficiency, and generalization capabilities of the model.
[0099] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first embodiment and / or the second embodiment can be referred to the above introduction and will not be described in detail later. Figure 4 , Figure 4 This is a flowchart of Example 3 of the method for generating video subtitle information of this application.
[0100] like Figure 4 As shown, in the embodiment of the present application, the step of determining the video subtitle information based on the audio language and the first subtitle information includes:
[0101] Step S31, when the audio language is not a preset language, translating the first subtitle information through a text translation model to obtain subtitle information in a preset language;
[0102] Step S32, marking repeated text information in the preset language subtitle information, performing redundancy processing on the repeated text information, and obtaining video subtitle information;
[0103] Step S33: When the audio language is a preset language, repeated text information in the first subtitle information is marked, and redundancy removal is performed on the repeated text information to obtain video subtitle information.
[0104] It should be noted that when the audio language is not the preset language, the first subtitle information can be inferred using a pre-trained text translation model to obtain the translated text. Due to grammatical and cultural differences, the translated text may contain duplicate text information. The duplicate text information can be marked and de-redundant processed to obtain the video subtitle information. For the first subtitle information in the preset language, de-redundant processing can be performed directly on the marked duplicate text information.
[0105] Specifically, the text translation model used in the embodiments of this application can use a framework such as TensorFlow Lite or PyTorch Mobile. By embedding the pre-trained text translation model into the video playback device, on-device reasoning can be achieved. By performing reasoning on the device side, network dependence is reduced, the delay in displaying video subtitles is reduced, and user privacy is protected because data does not need to be uploaded externally.
[0106] Furthermore, the importance of each weight during training can be analyzed, allowing for pruning operations to remove weights that have little impact on the text translation model's output, thereby reducing model complexity and computational requirements. Simultaneously, model weights can be compressed from 32-bit floating-point numbers to 8-bit integers, reducing the model's memory footprint and improving computational efficiency. Furthermore, knowledge distillation techniques can be used to leverage existing large models to inform text translation model training, thereby improving performance.
[0107] In some implementations of the embodiments of the present application, the generating device of the embodiments of the present application may be a video playback device, and the trained text translation model and speech recognition model may be converted into a lightweight format through a framework such as TensorFlow Lite or PyTorch Mobile to adapt to the end-side device. By embedding the lightweight text translation model and speech recognition model into the video playback device, end-side reasoning can be achieved. By performing reasoning on the device end, the dependence on the network is reduced, the delay in displaying video subtitles is reduced, and since there is no need to upload data to the outside, the privacy of the user is protected.
[0108] In some implementations of the embodiments of the present application, the video subtitle information generation process of the embodiments of the present application can be as follows: Figure 5 As shown, Figure 5 This is a schematic diagram of a scenario flow in one implementation of the method for generating video subtitle information of this application.
[0109] Reference Figure 5 In an embodiment of the present application, the user can first open the subtitle setting interface to select a preset language, so that the generation device can generate video subtitles corresponding to the preset language. When video subtitle generation is required, the generation device can obtain audio management authority from the video playback device. When the audio management authority is successfully obtained, the generation device can capture real-time audio data from the audio output stream of the video playback device, and then perform audio language recognition based on the real-time audio data. The audio language recognition method can be based on a pre-trained speech recognition model or other methods, and the embodiment of the present application does not limit this.
[0110] Furthermore, the generating device can check whether the automatic speech recognition technology (ASR) and the translation side AI model of the corresponding language in the real-time audio data exist. The translation side AI model may include the above-mentioned speech recognition model and text translation model. If it does not exist, the corresponding model file can be downloaded from the cloud for deployment; if it exists, it can be checked whether the model version is the latest. If the model version is not the latest, it can be downloaded from the cloud to update the model version.
[0111] Furthermore, the generating device can load the corresponding ASR and translation end-side AI models, and obtain standardized audio data by preprocessing the real-time audio data. The text to be translated is obtained by performing speech recognition on the standardized audio data. The text to be translated is converted into an input format acceptable to the text translation model (such as word embedding or subword embedding) by performing preprocessing operations such as word segmentation and removing redundant symbols, and then the text to be translated after the preprocessing operation is input into the text translation model. The text translation model can perform inference through the encoder-decoder architecture and use the attention mechanism to weight the important parts of the input text to be translated to improve the translation quality.
[0112] It is understood that the encoder can encode the input text to be translated into a context vector to capture semantic information. The decoder can generate a text sequence in a preset language based on the context vector, that is, obtain the translated text.
[0113] It should be noted that the translation results output by the text translation model can be post-processed such as removing redundant information and formatting the text, and the subtitle display format can be adjusted based on the subtitle style set by the user (such as language, font, color, etc.), and the adjusted video subtitles can be displayed.
[0114] In this embodiment, when the audio language is not the preset language, the text translation model is used to translate the first subtitle information to obtain subtitle information in the preset language. The subtitle information in the preset language is then subjected to text normalization to obtain video subtitle information. Because the first subtitle information is translated using a pre-trained on-device text translation model, there is no need to upload the text to be translated, which improves translation speed and protects user privacy.
[0115] This application also provides a video subtitle information generation device, please refer to Figure 6 , Figure 6 This is a schematic diagram of the module structure of the video subtitle information generation device according to an embodiment of the present application, wherein the video subtitle information generation device includes:
[0116] An audio acquisition module 10 is used to capture real-time audio data from an audio output stream;
[0117] An audio recognition module 20 is configured to perform audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data;
[0118] The subtitle generating module 30 is configured to translate the first subtitle information into video subtitle information corresponding to the preset language when the audio language is not the preset language.
[0119] The video subtitle information generation device provided in this application utilizes the video subtitle information generation method described in the aforementioned embodiments, resolving the technical issue that existing subtitle generation methods rely on pre-created translation text and fail to meet the real-time requirements of videos. Compared to the prior art, the video subtitle information generation device provided in this application achieves the same beneficial effects as the video subtitle information generation method described in the aforementioned embodiments. Other technical features of the video subtitle information generation device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0120] The present application provides a video subtitle information generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the video subtitle information generation method of the above-mentioned embodiment 1.
[0121] Reference below Figure 7 , which shows a schematic structural diagram of a video subtitle information generating device suitable for implementing the embodiments of the present application. The video subtitle information generating device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The video subtitle information generating device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0122] like Figure 7As shown, the video subtitle information generating device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the video subtitle information generating device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, such as a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003, such as a magnetic tape or hard disk; and communication device 1009. The communication device 1009 can allow the video subtitle information generating device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a video subtitle information generating device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or provided instead.
[0123] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0124] The video subtitle information generation device provided in this application utilizes the video subtitle information generation method described in the aforementioned embodiment, resolving the technical issue that existing subtitle generation methods rely on pre-created translation text and fail to meet the real-time requirements of videos. Compared to the prior art, the video subtitle information generation device provided in this application achieves the same beneficial effects as the video subtitle information generation method described in the aforementioned embodiment. Other technical features of this video subtitle information generation device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0125] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0126] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0127] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the video subtitle information generation method in the above embodiment.
[0128] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0129] The computer-readable storage medium may be included in the video subtitle information generating device; or may exist independently without being assembled into the video subtitle information generating device.
[0130] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the video subtitle information generating device, the video subtitle information generating device:
[0131] Capture real-time audio data from the audio output stream;
[0132] Performing audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data;
[0133] Video subtitle information is determined based on the audio language and the first subtitle information.
[0134] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0135] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0136] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0137] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for generating video subtitles. This method addresses the technical issue that existing subtitle generation methods rely on pre-created translation text and fail to meet the real-time requirements of videos. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the method for generating video subtitles provided in the aforementioned embodiments and are not further elaborated here.
[0138] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned method for generating video subtitle information when executed by a processor.
[0139] The computer program product provided in this application can address the technical problem that existing subtitle generation methods rely on pre-created translation text and cannot meet the real-time requirements of videos. Compared with the existing technology, the beneficial effects of the computer program product provided in this application are the same as those of the video subtitle information generation method provided in the above embodiment, and will not be elaborated here.
[0140] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for generating video subtitle information, characterized in that: The method comprises: Capture real-time audio data from the audio output stream; Performing audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data; When the audio language is not a preset language, the first subtitle information is translated into video subtitle information corresponding to the preset language.
2. The method for generating video subtitle information according to claim 1, wherein: The step of performing audio recognition on the real-time audio data to obtain the audio language and first subtitle information of the real-time audio data includes: storing the real-time audio data into a ring buffer structure; Performing floating-point conversion on the real-time audio data in the ring buffer structure to obtain standard audio data; Audio recognition is performed on the standard audio data to obtain an audio language and first subtitle information.
3. The method for generating video subtitle information according to claim 2, wherein: Before the step of performing audio recognition on the standard audio data to obtain the audio language and the first subtitle information, the method further includes: Acquire pre-trained voice data and pre-trained text data, wherein the pre-trained voice data and the pre-trained text data correspond to each other; Preprocessing the pre-trained voice data and the pre-trained text data respectively to obtain pre-processed voice data and pre-processed text data; Inputting the preprocessed speech data and the preprocessed text data into a first acoustic model for training to obtain a speech recognition model; Accordingly, the step of performing audio recognition on the standard audio data to obtain the audio language and the first subtitle information includes: Audio recognition is performed on the standard audio data based on a speech recognition model to obtain an audio language and first subtitle information.
4. The method for generating video subtitle information according to claim 3, wherein: The first acoustic model includes: an input layer, an embedding layer, an encoder, a pooling layer, a classification layer, an activation function layer, and an output layer; The step of inputting the preprocessed speech data and the preprocessed text data into a first acoustic model for training to obtain a speech recognition model includes: Inputting the preprocessed text data into the word segmenter of the input layer for word segmentation processing to obtain text segmentation; Converting the text segmentation into a dense vector representation through the embedding layer; Extracting context features represented by the dense vector through the encoder to obtain sequence features; Performing feature pooling on the sequence features through the pooling layer to obtain pooled features of a preset length; Performing language category mapping on the pooled features through the classification layer to obtain a classification layer output; The output of the classification layer is converted through the activation function layer to obtain a language probability distribution; Performing language prediction on the language probability distribution through the output layer to obtain the target language corresponding to the preprocessed speech data; Determining a cross entropy loss based on the target language and the true language of the preprocessed speech data; Iterating model parameters of the first acoustic model according to the cross entropy loss; When the iteration termination condition is met, the first acoustic model that meets the iteration termination condition is used as the speech recognition model.
5. The method for generating video subtitle information according to claim 1, wherein: The step of determining video subtitle information based on the audio language and the first subtitle information includes: When the audio language is not a preset language, translating the first subtitle information through a text translation model to obtain subtitle information in a preset language; Repeated text information in the preset language subtitle information is marked, and redundancy removal processing is performed on the repeated text information to obtain video subtitle information.
6. The method for generating video subtitle information according to claim 5, wherein: The step of determining the video subtitle information based on the audio language and the first subtitle information further includes: When the audio language is a preset language, repeated text information in the first subtitle information is marked, and redundancy removal is performed on the repeated text information to obtain video subtitle information.
7. A video subtitle information generating device, characterized in that: The video subtitle information generating device includes: Audio acquisition module, used to capture real-time audio data from the audio output stream; an audio recognition module, configured to perform audio recognition on the real-time audio data to obtain an audio language and first subtitle information of the real-time audio data; The subtitle generation module is used to determine video subtitle information based on the audio language and the first subtitle information.
8. A video subtitle information generating device, characterized in that: The device includes: a memory, a processor, and a video subtitle information generation program stored in the memory and executable on the processor, wherein the video subtitle information generation program is configured to implement the steps of the video subtitle information generation method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium stores a video subtitle information generation program, which, when executed by a processor, implements the steps of the video subtitle information generation method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method for generating video subtitle information according to any one of claims 1 to 6 are implemented.