Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

360 results about "Subtitle" patented technology

Subtitles are text derived from either a transcript or screenplay of the dialog or commentary in films, television programs, video games, and the like, usually displayed at the bottom of the screen, but can also be at the top of the screen if there is already text at the bottom of the screen. They can either be a form of written translation of a dialog in a foreign language, or a written rendering of the dialog in the same language, with or without added information to help viewers who are deaf or hard of hearing to follow the dialog, or people who cannot understand the spoken dialogue or who have accent recognition problems.

Multi-language cross-culture communication auxiliary method and system based on large model

The invention provides a multi-language cross-culture communication assisting method and system based on a large model. The method comprises the following steps: receiving a source language audio stream during a call, calling a multi-language sound frequency harmonic modulation feature library to extract fundamental frequency harmonic intensity distribution and tone turning features, and generating a cultural acoustic fingerprint vector; based on the vector, controlling a microphone array phase difference, directionally enhancing a fundamental frequency harmonic component of a speaker and suppressing noise, and outputting a high signal-to-noise ratio spectrogram; analyzing the pronunciation rhythm and tone turning characteristics of the spectrogram, capturing the pitch jump and duration of the syllable boundary, and generating an acoustic culture label; associating the spectrogram with a target semantic library, matching harmonic distribution and a cultural context rule based on a large model, and outputting a cultural interpretation prompt containing an ambiguity resolution suggestion; and generating a calibration result according to the acoustic tag and the semantic prompt, and overlapping the dynamic floating subtitles to the face area of the speaker in the video conference picture. According to the invention, cultural tone ambiguity in multi-language communication is eliminated.
Owner:LUSTER LIGHTWAVE CO LTD

Multi-modal time sequence alignment AI video translation method and system

The invention relates to the technical field of subtitle translation, in particular to a multi-modal time sequence alignment AI video translation method and system, and the method comprises the steps: 1, carrying out the multi-modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint feature data and visual time sequence data; 2, performing cross-language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint feature data and the target language text; generating a mouth shape animation matched with the target language voice based on the lip key point data and the limb action time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the mouth shape animation and the limb action sequence through a cross-modal time sequence encoder, and dynamically adjusting the layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.
Owner:HANGZHOU BAOMIHUA TECH CO LTD

Video and voice automatic translation method based on pre-training model

The invention belongs to the technical field of speech translation, and particularly relates to a video speech automatic translation method based on a pre-training model, and the method comprises the following steps: 1, preprocessing video and audio data; step 2, voice recognition and language detection; step 3, machine translation and text post-processing; step 4, speech synthesis and audio mixing; step 5, synchronizing video processing and subtitles; step 6, quality control and multi-dimensional evaluation; 7, carrying out model iteration and data closed loop; and step 8, system deployment and engineering implementation. Through deep fusion of efficient transfer learning of the pre-training model and the multi-modal technology, a high-precision, low-cost and easy-to-expand video speech translation solution is constructed, the time and labor cost of globalized content production is greatly reduced, the cross-language communication efficiency is improved, immersive multi-language experience is provided, and the method is suitable for popularization and application. And a data-driven continuous optimization mechanism is established, so that the system performance is improved along with the increase of the use scale.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Captioning videos with multiple cross-modality teachers

Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.
Owner:SNAP INC

Speech synthesis method and device, electronic equipment, storage medium and program product

The invention relates to a speech synthesis method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: acquiring a source language dubbing audio corresponding to a source language subtitle text of a source video and a target language subtitle text translated by the source language subtitle text; extracting audio emotion features from the source language dubbing audio by using an emotion extractor, wherein the audio emotion features represent emotions expressed by the source language dubbing audio; converting the target language subtitle text into a phoneme sequence, and encoding the phoneme sequence by using a text encoder to obtain text encoding features; fusing the audio emotion features with the text coding features to obtain emotion text features; and generating a target language audio based on the emotion text features by using a decoder, the target language audio being used as a dubbing audio of the source video in the target language. Therefore, the high-quality target language audio with the source dubbing audio emotion can be automatically and efficiently generated, the cost is low, and the efficiency is high.
Owner:YOUKU CULTURE TECH (BEIJING) CO LTD

Video dubbing method and device, electronic equipment and storage medium

The invention provides a video dubbing method and device, electronic equipment and a storage medium, and belongs to the technical field of video processing, and the method comprises the steps: separating a first audio in a to-be-dubbed video, and carrying out the sentence-by-sentence text conversion, and obtaining a first subtitle text with a time code; translating the first subtitle text into a second subtitle text with the same time code; normalizing the second subtitle text to obtain a third subtitle text according to the difference between the dubbing duration of the second subtitle text and the subtitle display duration; generating a third audio corresponding to the third subtitle text; and finally, synthesizing all the third audios, the background audio of the video to be dubbed and the silent video to obtain a final target video. According to the invention, under the condition that the audio duration and the subtitle display duration are different in the video dubbing process, the subtitle text is normalized, so that the damage to the original video file is avoided, and the high quality of video dubbing is ensured.
Owner:ANHUI IFLYREC TECH CO LTD

Character loss compensation processing device and method for video subtitle extraction

The invention discloses a character loss compensation processing device and method for video subtitle extraction, which effectively solve the problem of character loss caused by video image quality, complex background interference, OCR (Optical Character Recognition) limitation, dynamic subtitle change and the like. According to the method, key technologies such as self-adaptive subtitle denoising, semantic compensation and multi-frame fusion are adopted, and the completeness and accuracy of subtitle extraction are remarkably improved. The scheme is suitable for various types of movies, television dramas, short videos, conference videos and the like. The core innovation of the method is that lost or wrong characters can be effectively detected, compensated and corrected by intelligently analyzing the context and integrating multi-frame information, so that the subtitle output quality is greatly improved, and the effect is remarkable particularly in a scene that a single character is easy to lose.
Owner:SUZHOU XIAOTONG TECH CO LTD

Short play video editing method, system and device based on multiple modes and medium

The invention discloses a multi-mode-based short episode video editing method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing of an original short episode video, and obtaining a second short episode video, subtitles and a subtitle timestamp; analyzing the subtitles, and generating a short episode abstract according to an analysis result of the subtitles; according to the short episode abstract and the highlight editing prompt, performing plot analysis and editing on the second short episode video, and performing editing to form a first video set; scoring videos in the first video set according to a multi-dimensional scoring rule, and screening out M highlight video clips before scoring; and adjusting the timestamps of the M highlight video clips according to the subtitle timestamps, and outputting the adjusted M highlight video clips. According to the method and the device, the full link from short video input to short play mixed video output is constructed, a user only needs to input the to-be-edited short play video, the mixed video can be automatically generated, and the video editing efficiency is improved while the video editing threshold is reduced.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Video editing method and device, electronic equipment and storage medium

The invention discloses a video editing method and device, electronic equipment and a storage medium. The method comprises the following steps: segmenting a to-be-edited video to obtain a plurality of video clips; understanding text generation and subtitle generation are carried out on each video clip, and a clip understanding text and a clip subtitle of each video clip are obtained; determining a target description text based on fragment understanding texts and / or fragment subtitles corresponding to a plurality of target video fragments selected from the plurality of video fragments; and performing video synthesis based on the plurality of target video clips and the target description text to obtain a target video corresponding to the to-be-edited video. According to the method provided by the invention, the video editing efficiency is relatively high, and the video editing cost is relatively low.
Owner:SHENZHEN SIYUAN ELECTRONICS TECH CO LTD

Short drama translation and explanation method based on multi-mode neural network

The invention belongs to the technical field of multimedia, and particularly relates to a short drama translation and explanation method based on a multi-mode neural network, which realizes full-process automatic processing from subtitle extraction to culture adaptation translation by integrating advanced technical means such as an OCR (optical character recognition) model, optical flow field analysis, an LSTM (long short term memory) network and a graph neural network. Particularly, the subtitles in a complex background can be accurately recognized and processed, subtitle traces are effectively removed, plot conflicts are understood through a deep learning algorithm, and then translation content conforming to target culture is generated. The method not only improves the accuracy of subtitle extraction, but also enhances the cultural adaptability of the translated content, so that the multimedia content can be more naturally and smoothly understood and accepted under different cultural backgrounds, and the watching experience of global users is greatly improved.
Owner:CHENGDU ROAD YOUYOU TECHNOLOGY CO LTD

Video content textualization method and system based on multi-modal fusion

The invention provides a video content textualization method based on multi-modal fusion, which comprises the following steps of: 1, dynamically identifying effective modal information existing in a video, including subtitle information detection, audio information detection and key frame information sampling; step 2, carrying out subtitle extraction on the subtitle information by adopting an OCR (Optical Character Region) enhancement method based on regional clustering; generating a voice text for the audio information by adopting a multi-engine collaborative transcription and weight fusion strategy; generating a descriptive text for the key frame information; step 3, performing space-time alignment and semantic fusion on the subtitles, the voice text, the descriptive text and the video time axis; and feeding back and adjusting a strategy of key frame information sampling according to a fusion result, and forming self-adaptive feedback. According to the textualization method, multi-modal information can be adaptively fused, and video total elements are covered, so that the robustness and comprehensiveness of content extraction are improved.
Owner:WUHAN UNIV

Video processing method and device, equipment and storage medium

The invention provides a video processing method and device, equipment and a storage medium, and the method comprises the steps: segmenting a video material into a plurality of semantically coherent video clips, automatically extracting related video clips from the video material according to the text content provided by a user, and generating corresponding subtitles and voice broadcast ports matched with the related video clips, and finally synthesizing a section of finished video with clear logic and coherent content. According to the scheme, full-process automation from video material generation to final video file generation is realized, a large number of video materials can be quickly processed, the tedious manual editing process in traditional video production is reduced, the dependence on manual intervention is greatly reduced, the operation cost is reduced, the video production process is simplified, the video production efficiency is improved, and the user experience is improved. And the quality and the efficiency of finished products are also improved.
Owner:CTRIP TRAVEL NETWORK TECH SHANGHAI0

Multi-modal content review method and device, medium and computer program

The invention provides a multi-mode content review method and device, a medium and a computer program, and relates to the technical field of content review. The multi-modal content review method comprises the following steps: calling a special recognition model to recognize subtitles and voice in target review content to obtain a character recognition result and a voice recognition result; the special identification model is a model obtained based on learning of a content sample and a subtitle label and a voice label of the content sample; and calling an artificial intelligence large model, performing fusion analysis and reasoning on the character recognition result, the voice recognition result and the video picture of the target review content according to the violation review rule, and outputting a violation behavior recognition result. The universality of the content review method is improved based on the good cognitive ability of the artificial intelligence large model; the artificial intelligence large model and the special small model are organically combined, the understanding cognitive ability of the artificial intelligence large model and the perception ability of the special small model are fully exerted, and accurate and effective examination of the content is realized.
Owner:BEIJING ZHONGKE PATTEK TECH CO LTD

Description video generation method and device based on large model, equipment and medium

The invention provides an explanation video generation method, device and equipment based on a large model and a medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of multimodality, natural language processing, computer vision, deep learning and the like. The method comprises the following steps: acquiring a plurality of subtitle texts and corresponding first timestamps in a to-be-processed video; based on the first timestamps of the plurality of subtitle texts, determining at least one subtitle-free fragment in the to-be-processed video and a corresponding second timestamp; performing visual content understanding on the at least one subtitle-free fragment by using a first multi-mode large model to obtain at least one subtitle complemented text corresponding to the at least one subtitle-free fragment; generating commentaries for the to-be-processed video by using a large language model based on the plurality of subtitle texts and the corresponding first timestamps and the at least one subtitle complemented text and the corresponding second timestamps; and generating a commentary video based on the commentary.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Voice, picture and subtitle alignment method and system for real-time voice translation synthesis

The invention provides a voice, picture and subtitle alignment method and system for real-time voice translation synthesis, and relates to the technical field of voice processing. According to the method, millisecond alignment of translated voice and video is realized through fragment-level parallel processing, compared with a traditional scheme, the delay is greatly reduced, and the problem that translated voice pictures are asynchronous during live broadcast is solved; the method comprises the following steps: automatically generating real-time voice translation according to live broadcast content, and respectively processing to obtain a corresponding srt file and a translated m3u8 fragment; therefore, there is no need to manually translate and generate subtitles in advance, and there is no need to replace the audio of the original video in advance, thereby saving the manpower cost, and providing a more universal live broadcast stream subtitle scheme. Besides, when the languages are switched, the player can automatically select the starting time point, so that the method supports multiple playing modes such as online live broadcast and live broadcast playback.
Owner:CHENGDOU HUAQIYUN TECH CO LTD

Video processing method and device

The embodiment of the invention provides a video processing method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of multimedia. The video processing method comprises the following steps: playing a target video; wherein the target video comprises a current video frame embedded with a target subtitle; under the condition that a first preset condition is triggered, displaying a modifiable identifier for the target subtitle on the current video frame; and popping up an editing entry for modifying the target subtitle in response to the triggering of the region where the modifiable identifier is located, receiving current editing content through the editing entry; and uploading the current edited content to a server, so that the server replaces the target subtitle with the current edited content. According to the technical scheme provided by the embodiment of the invention, the published video can be quickly corrected. And the man-machine interaction efficiency is improved through a convenient mode of modifying the identifier and editing the entry.
Owner:SHANGHAI BILIBILI TECH CO LTD

System and method for dubbing by automatically aligning time axis

The invention provides a system and a method for automatically aligning time axis dubbing, and the system comprises a silent video extraction module which is used for removing the audio of an original video; the OCR subtitle extraction module is used for generating a subtitle file in an SRT format; the AI translation module translates the subtitles into a target language; the subtitle erasing module is used for positioning and erasing original subtitles in the video; the TTS speech synthesis module is used for generating a dubbing file; the time axis alignment module is used for dynamically adjusting the speech speed, the video speed and the subtitle time; and the video synthesis module is used for integrating the silent video, dubbing and subtitles. According to the method, OCR subtitle extraction, AI subtitle translation and generative dubbing technologies are combined, and the voice speed of the voice is dynamically adjusted to keep the voice aligned with the time axis of the video and the time axis of the subtitle, so that the quality of the translation play can be improved, the cost can be greatly reduced, the product marketing is accelerated, and great convenience and competitive advantages are brought to content creators and issuers.
Owner:SHENZHEN MAPLE LEAF INTERACTIVE TECHNOLOGY CO LTD

Short drama subtitle translation system based on artificial intelligence

The invention provides a short play subtitle translation system based on artificial intelligence, and relates to the technical field of artificial intelligence, and the system comprises a subtitle recognition module, an erasing module, a translation module and an output module, and can automatically recognize time information, position coordinates and visual style parameters of subtitles in a short play video. And generating a picture sequence without original subtitles in combination with the erasing processing, and translating the recognized original subtitle text into target language subtitles. The system generates target rendering parameters based on explicit style parameters and implicit style embedding vectors, realizes high restoration of fonts, strokes, shadows, gradient, transparency, textures and dynamic special effects, and performs adaptive adjustment according to target subtitle text features. Therefore, the visual consistency and culture adaptability of translated subtitles are improved in a short drama scene with complicated subtitle styles, frequent dynamic changes and obvious cross-culture differences, the audience impression is improved, and the manual post-processing workload is reduced.
Owner:XIAN LINGXIANG BIRD CULTURE COMM CO LTD

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Subtitle processing method and device, equipment, medium and program product

The embodiment of the invention discloses a subtitle processing method and device, equipment, a medium and a program product. The method comprises the following steps: acquiring a first subtitle file; extracting the time axis in each piece of subtitle information to obtain subtitle content in each piece of subtitle information; obtaining translation prompt information of the first subtitle file, and translating each subtitle content under the prompt of the translation prompt information to obtain the translated content of each subtitle content; and performing reconstruction processing on the translation content of each subtitle content and the time axis corresponding to each subtitle content to generate a second subtitle file. By adopting the embodiment of the invention, the translation accuracy and translation efficiency of subtitle translation can be improved.
Owner:TENCENT TECH (BEIJING) CO LTD

Virtual digital human multimedia teaching interaction method and system and storage medium

The invention discloses a virtual digital human multimedia teaching interaction method and system and a storage medium, and the method comprises the steps: calculating an audio mouth shape time difference, a mouth shape subtitle time difference, an audio subtitle time difference and a video mouth shape time difference based on a same-window data group on the basis of a unified time aperture and a window, comparing the time differences with a time difference threshold value, and outputting a dislocation alarm result, unified constraint on four types of key alignment relationships is realized; when an alarm is triggered, selecting a degradation strategy according to a dislocation alarm result and generating an execution record, and summarizing the dislocation alarm result and the execution record to form an evidence chain data packet; when the cumulative number of times in the preset window number exceeds a diffusion blocking threshold value, recording output is switched into a video stream which retains audio and subtitles and does not overlap mouth shapes; and generating a correction release stream based on the evidence chain data packet, and replacing the recorded broadcast corresponding time slice and retaining the Hash verification value when the time sequence consistency index meets the release threshold, thereby realizing drivable treatment action, suppressible diffusion and replaceable correction.
Owner:ULEARNING

Simultaneous interpretation method and system based on large model and electronic equipment

The invention discloses a simultaneous interpretation method and system based on a large model and electronic equipment, and the method comprises the steps: extracting bilingual parallel corpora related to terms from professional resources based on a standardized professional dictionary, obtaining qualified corpora through data enhancement processing and manual screening, and constructing a multi-level corpus according to the levels of words, sentences and paragraphs; the method comprises the following steps: receiving an input audio stream in real time, extracting acoustic features through preprocessing, inputting a pre-established large-scale speech recognition model, and carrying out incremental decoding on the acoustic features in a sliding window mode; and calling a sentence boundary prediction network to judge a pause point, and outputting a text stream with a timestamp. Performing fine tuning on the large-scale speech recognition model by using a multi-level corpus, translating a text stream based on the fine-tuned large-scale speech recognition model, and constraining term translation according to a standardized professional dictionary; and synchronously displaying the audio output in the translation result and the subtitles. According to the scheme, the terminology recognition and translation accuracy is improved, and simultaneous interpretation delay is reduced.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Description video synthesis method and device, equipment and medium

The invention discloses a commentary video synthesis method and device, equipment and a medium, and the method comprises the steps: obtaining a commentary content text of an original video, and generating a subtitle data set corresponding to the commentary content text and an audio data set of the subtitle data set, the subtitle data set comprises a plurality of subtitle display time periods arranged according to a time sequence and subtitle texts corresponding to the subtitle display time periods; determining a corresponding image display time period according to each subtitle display time period in the subtitle data set; corresponding to each image display time period, determining a matched image frame sub-sequence in a preset image material library according to a caption text in the corresponding caption display time period in the caption data set so as to construct a video image frame sequence; and synthesizing a corresponding target explanation video according to the subtitle data set, the audio data set and the video image frame sequence. According to the method, the subtitles, the audios and the corresponding image frame sequences accurately matched with the explanation content text are automatically generated, so that the production efficiency and quality of the explanation video are remarkably improved.
Owner:GUANGZHOU OVERSEAS KANGBAZI NETWORK TECHNOLOGY CO LTD

Video brief introduction automatic generation method, system and device and storage medium

The invention discloses a video brief introduction automatic generation method, system and device and a storage medium, and the method comprises the steps: extracting a plurality of first keywords in a first subtitle text, and extracting a plurality of second keywords in a second subtitle text; constructing a popularity keyword data set based on the plurality of first keywords, and screening the plurality of second keywords based on the popularity keyword data set to obtain a plurality of first target keywords; screening the plurality of second images to obtain a screened second image; inputting the screened second images into a trained image text generation model to obtain a plurality of second image description texts; extracting a plurality of second target keywords in each second image description text; and generating a video brief introduction according to the plurality of first target keywords and the plurality of second target keywords. The accuracy of the generated video brief introduction can be improved.
Owner:UNICOM WOYUEDU TECH CULTURE CO LTD +1

Video multi-track synthesis method and system for intelligent sub-mirror calibration

The invention relates to the field of video processing, and discloses a video multi-track synthesis method for intelligent sub-lens calibration, which comprises the following steps: S1, importing subtitles and identifying; s2, carrying out split mirror calibration; s3, picture analysis and reasoning; s4, material capturing and adjusting: capturing and preprocessing external materials based on a pre-selection result, dynamically adjusting material positions through an interactive interface, binding role attributes, and supplementing scene and article elements; s5, operating an editor; and S6, exporting and synchronizing. In the subtitle processing link, by means of a pysrt library of Python and a multi-thread technology, the subtitle extraction efficiency is greatly improved compared with traditional manual operation; according to the intelligent mirror splitting calibration based on the Transform architecture, automatic mirror splitting of the long caption text can be completed within seconds. From caption import identification, sub-lens calibration, material pre-selection and track arrangement, the system realizes full-process intelligentization, the image analysis and reasoning module automatically screens materials and performs pre-selection and sorting, and the editor module automatically arranges the materials, so that the manual operation amount is greatly reduced, and the synthesis efficiency is remarkably improved.
Owner:JIANGXI HIP HOP TECHNOLOGY CO LTD

Video subtitle erasing method and device, equipment and storage medium

The invention discloses a video subtitle erasing method, device and equipment and a storage medium, and relates to the field of digital image processing, and the method comprises the steps: detecting subtitles of a subtitle video to be erased, merging timestamps of the same subtitles in the subtitle video to be erased, and determining a subtitle fragment set and a subtitle-free fragment set; splitting the subtitle segment into independent shot segments by using a preset lens splitting algorithm, and analyzing video frames of the independent shot segments to obtain a first frame and a tail frame; determining a target reference frame based on the first frame and the tail frame, generating an expanded independent shot segment according to the target reference frame and the independent shot segment, and segmenting a target character mask; and generating an erased independent lens segment according to the expanded independent lens segment and the target character mask through a preset erasure algorithm, performing a preset post-processing optimization operation on the erased independent lens segment to obtain a target independent lens segment, and integrating the target independent lens segment and the subtitle-free segment set to generate a target video. According to the method and the device, the video subtitles can be accurately erased.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Subtitle display method and device, electronic equipment and computer readable storage medium

The embodiment of the invention discloses a subtitle display method and device, electronic equipment and a computer readable storage medium. The method comprises the steps of determining a target font of a target character corresponding to a subtitle to be displayed; the target font corresponding to the target character is inquired in an index table, and the index table comprises a plurality of characters, a plurality of fonts corresponding to the characters and storage paths of the characters of all the fonts; when the target font corresponding to the target character is queried in the index table, obtaining a target storage path of the target character of the target font; obtaining a target character of the target font according to the target storage path; and displaying the subtitles according to the target characters of the target font. Therefore, according to the scheme, the target character of the target font can be quickly obtained based on the index table, so that the subtitle display speed is increased.
Owner:SHENZHEN TCL DIGITAL TECH CO LTD

Captioning videos with multiple cross-modality teachers

Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.
Owner:SNAP INC

Automatic subtitle translation method and device based on large model and storage medium

The invention relates to the technical field of video data processing, in particular to a caption automatic translation method and device based on a large model and a storage medium. The method comprises the following steps: acquiring and preprocessing subtitle data and video data to obtain a time axis index table; constructing a semantic model based on the time axis index table to generate a semantic vector; performing feature extraction on the video data to obtain a voice feature vector and a video feature vector, and performing feature fusion on the semantic vector, the voice feature vector and the video feature vector to construct a fusion feature vector; constructing a terminology library and generating a preliminary translation result according to the fusion feature vector; and constructing a subtitle evaluation model according to the preliminary translation result to obtain adaptive subtitles. According to the invention, accurate translation of the subtitle data in the video is realized.
Owner:BEIJING SIMAILI MEDIA TECHNOLOGY CO LTD