Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

233 results about "Subtitle" patented technology

Subtitles are text derived from either a transcript or screenplay of the dialog or commentary in films, television programs, video games, and the like, usually displayed at the bottom of the screen, but can also be at the top of the screen if there is already text at the bottom of the screen. They can either be a form of written translation of a dialog in a foreign language, or a written rendering of the dialog in the same language, with or without added information to help viewers who are deaf or hard of hearing to follow the dialog, or people who cannot understand the spoken dialogue or who have accent recognition problems.

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Video dubbing method and device, electronic equipment and storage medium

The invention provides a video dubbing method and device, electronic equipment and a storage medium, and belongs to the technical field of video processing, and the method comprises the steps: separating a first audio in a to-be-dubbed video, and carrying out the sentence-by-sentence text conversion, and obtaining a first subtitle text with a time code; translating the first subtitle text into a second subtitle text with the same time code; normalizing the second subtitle text to obtain a third subtitle text according to the difference between the dubbing duration of the second subtitle text and the subtitle display duration; generating a third audio corresponding to the third subtitle text; and finally, synthesizing all the third audios, the background audio of the video to be dubbed and the silent video to obtain a final target video. According to the invention, under the condition that the audio duration and the subtitle display duration are different in the video dubbing process, the subtitle text is normalized, so that the damage to the original video file is avoided, and the high quality of video dubbing is ensured.
Owner:ANHUI IFLYREC TECH CO LTD

Video editing method and device, electronic equipment and storage medium

The invention discloses a video editing method and device, electronic equipment and a storage medium. The method comprises the following steps: segmenting a to-be-edited video to obtain a plurality of video clips; understanding text generation and subtitle generation are carried out on each video clip, and a clip understanding text and a clip subtitle of each video clip are obtained; determining a target description text based on fragment understanding texts and / or fragment subtitles corresponding to a plurality of target video fragments selected from the plurality of video fragments; and performing video synthesis based on the plurality of target video clips and the target description text to obtain a target video corresponding to the to-be-edited video. According to the method provided by the invention, the video editing efficiency is relatively high, and the video editing cost is relatively low.
Owner:SHENZHEN SIYUAN ELECTRONICS TECH CO LTD

Video content textualization method and system based on multi-modal fusion

The invention provides a video content textualization method based on multi-modal fusion, which comprises the following steps of: 1, dynamically identifying effective modal information existing in a video, including subtitle information detection, audio information detection and key frame information sampling; step 2, carrying out subtitle extraction on the subtitle information by adopting an OCR (Optical Character Region) enhancement method based on regional clustering; generating a voice text for the audio information by adopting a multi-engine collaborative transcription and weight fusion strategy; generating a descriptive text for the key frame information; step 3, performing space-time alignment and semantic fusion on the subtitles, the voice text, the descriptive text and the video time axis; and feeding back and adjusting a strategy of key frame information sampling according to a fusion result, and forming self-adaptive feedback. According to the textualization method, multi-modal information can be adaptively fused, and video total elements are covered, so that the robustness and comprehensiveness of content extraction are improved.
Owner:WUHAN UNIV

Description video generation method and device based on large model, equipment and medium

The invention provides an explanation video generation method, device and equipment based on a large model and a medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of multimodality, natural language processing, computer vision, deep learning and the like. The method comprises the following steps: acquiring a plurality of subtitle texts and corresponding first timestamps in a to-be-processed video; based on the first timestamps of the plurality of subtitle texts, determining at least one subtitle-free fragment in the to-be-processed video and a corresponding second timestamp; performing visual content understanding on the at least one subtitle-free fragment by using a first multi-mode large model to obtain at least one subtitle complemented text corresponding to the at least one subtitle-free fragment; generating commentaries for the to-be-processed video by using a large language model based on the plurality of subtitle texts and the corresponding first timestamps and the at least one subtitle complemented text and the corresponding second timestamps; and generating a commentary video based on the commentary.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Short drama subtitle translation system based on artificial intelligence

The invention provides a short play subtitle translation system based on artificial intelligence, and relates to the technical field of artificial intelligence, and the system comprises a subtitle recognition module, an erasing module, a translation module and an output module, and can automatically recognize time information, position coordinates and visual style parameters of subtitles in a short play video. And generating a picture sequence without original subtitles in combination with the erasing processing, and translating the recognized original subtitle text into target language subtitles. The system generates target rendering parameters based on explicit style parameters and implicit style embedding vectors, realizes high restoration of fonts, strokes, shadows, gradient, transparency, textures and dynamic special effects, and performs adaptive adjustment according to target subtitle text features. Therefore, the visual consistency and culture adaptability of translated subtitles are improved in a short drama scene with complicated subtitle styles, frequent dynamic changes and obvious cross-culture differences, the audience impression is improved, and the manual post-processing workload is reduced.
Owner:XIAN LINGXIANG BIRD CULTURE COMM CO LTD

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Virtual digital human multimedia teaching interaction method and system and storage medium

The invention discloses a virtual digital human multimedia teaching interaction method and system and a storage medium, and the method comprises the steps: calculating an audio mouth shape time difference, a mouth shape subtitle time difference, an audio subtitle time difference and a video mouth shape time difference based on a same-window data group on the basis of a unified time aperture and a window, comparing the time differences with a time difference threshold value, and outputting a dislocation alarm result, unified constraint on four types of key alignment relationships is realized; when an alarm is triggered, selecting a degradation strategy according to a dislocation alarm result and generating an execution record, and summarizing the dislocation alarm result and the execution record to form an evidence chain data packet; when the cumulative number of times in the preset window number exceeds a diffusion blocking threshold value, recording output is switched into a video stream which retains audio and subtitles and does not overlap mouth shapes; and generating a correction release stream based on the evidence chain data packet, and replacing the recorded broadcast corresponding time slice and retaining the Hash verification value when the time sequence consistency index meets the release threshold, thereby realizing drivable treatment action, suppressible diffusion and replaceable correction.
Owner:ULEARNING

Simultaneous interpretation method and system based on large model and electronic equipment

The invention discloses a simultaneous interpretation method and system based on a large model and electronic equipment, and the method comprises the steps: extracting bilingual parallel corpora related to terms from professional resources based on a standardized professional dictionary, obtaining qualified corpora through data enhancement processing and manual screening, and constructing a multi-level corpus according to the levels of words, sentences and paragraphs; the method comprises the following steps: receiving an input audio stream in real time, extracting acoustic features through preprocessing, inputting a pre-established large-scale speech recognition model, and carrying out incremental decoding on the acoustic features in a sliding window mode; and calling a sentence boundary prediction network to judge a pause point, and outputting a text stream with a timestamp. Performing fine tuning on the large-scale speech recognition model by using a multi-level corpus, translating a text stream based on the fine-tuned large-scale speech recognition model, and constraining term translation according to a standardized professional dictionary; and synchronously displaying the audio output in the translation result and the subtitles. According to the scheme, the terminology recognition and translation accuracy is improved, and simultaneous interpretation delay is reduced.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Description video synthesis method and device, equipment and medium

The invention discloses a commentary video synthesis method and device, equipment and a medium, and the method comprises the steps: obtaining a commentary content text of an original video, and generating a subtitle data set corresponding to the commentary content text and an audio data set of the subtitle data set, the subtitle data set comprises a plurality of subtitle display time periods arranged according to a time sequence and subtitle texts corresponding to the subtitle display time periods; determining a corresponding image display time period according to each subtitle display time period in the subtitle data set; corresponding to each image display time period, determining a matched image frame sub-sequence in a preset image material library according to a caption text in the corresponding caption display time period in the caption data set so as to construct a video image frame sequence; and synthesizing a corresponding target explanation video according to the subtitle data set, the audio data set and the video image frame sequence. According to the method, the subtitles, the audios and the corresponding image frame sequences accurately matched with the explanation content text are automatically generated, so that the production efficiency and quality of the explanation video are remarkably improved.
Owner:GUANGZHOU OVERSEAS KANGBAZI NETWORK TECHNOLOGY CO LTD

Video multi-track synthesis method and system for intelligent sub-mirror calibration

The invention relates to the field of video processing, and discloses a video multi-track synthesis method for intelligent sub-lens calibration, which comprises the following steps: S1, importing subtitles and identifying; s2, carrying out split mirror calibration; s3, picture analysis and reasoning; s4, material capturing and adjusting: capturing and preprocessing external materials based on a pre-selection result, dynamically adjusting material positions through an interactive interface, binding role attributes, and supplementing scene and article elements; s5, operating an editor; and S6, exporting and synchronizing. In the subtitle processing link, by means of a pysrt library of Python and a multi-thread technology, the subtitle extraction efficiency is greatly improved compared with traditional manual operation; according to the intelligent mirror splitting calibration based on the Transform architecture, automatic mirror splitting of the long caption text can be completed within seconds. From caption import identification, sub-lens calibration, material pre-selection and track arrangement, the system realizes full-process intelligentization, the image analysis and reasoning module automatically screens materials and performs pre-selection and sorting, and the editor module automatically arranges the materials, so that the manual operation amount is greatly reduced, and the synthesis efficiency is remarkably improved.
Owner:JIANGXI HIP HOP TECHNOLOGY CO LTD

Video subtitle erasing method and device, equipment and storage medium

The invention discloses a video subtitle erasing method, device and equipment and a storage medium, and relates to the field of digital image processing, and the method comprises the steps: detecting subtitles of a subtitle video to be erased, merging timestamps of the same subtitles in the subtitle video to be erased, and determining a subtitle fragment set and a subtitle-free fragment set; splitting the subtitle segment into independent shot segments by using a preset lens splitting algorithm, and analyzing video frames of the independent shot segments to obtain a first frame and a tail frame; determining a target reference frame based on the first frame and the tail frame, generating an expanded independent shot segment according to the target reference frame and the independent shot segment, and segmenting a target character mask; and generating an erased independent lens segment according to the expanded independent lens segment and the target character mask through a preset erasure algorithm, performing a preset post-processing optimization operation on the erased independent lens segment to obtain a target independent lens segment, and integrating the target independent lens segment and the subtitle-free segment set to generate a target video. According to the method and the device, the video subtitles can be accurately erased.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Captioning videos with multiple cross-modality teachers

Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.
Owner:SNAP INC

Automatic subtitle translation method and device based on large model and storage medium

The invention relates to the technical field of video data processing, in particular to a caption automatic translation method and device based on a large model and a storage medium. The method comprises the following steps: acquiring and preprocessing subtitle data and video data to obtain a time axis index table; constructing a semantic model based on the time axis index table to generate a semantic vector; performing feature extraction on the video data to obtain a voice feature vector and a video feature vector, and performing feature fusion on the semantic vector, the voice feature vector and the video feature vector to construct a fusion feature vector; constructing a terminology library and generating a preliminary translation result according to the fusion feature vector; and constructing a subtitle evaluation model according to the preliminary translation result to obtain adaptive subtitles. According to the invention, accurate translation of the subtitle data in the video is realized.
Owner:BEIJING SIMAILI MEDIA TECHNOLOGY CO LTD

Foreign language teaching video generation method and apparatus

The present invention provides a foreign language teaching video generation method and apparatus. The foreign language teaching video generation method comprises: acquiring video subtitle text corresponding to a video to be processed; on the basis of the video subtitle text and a preset explanation generation rule, using a large language model to generate foreign language teaching explanation text of the video subtitle text; generating a corresponding explanation audio on the basis of the foreign language teaching explanation text; determining a time information relationship between the foreign language teaching explanation text and the explanation audio; on the basis of the time information relationship between the foreign language teaching explanation text and the explanation audio, generating a display style corresponding to the explanation audio; and on the basis of the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio, generating an explanation video corresponding to the video to be processed. In the present invention, a target teaching video corresponding to a video to be processed can be automatically generated, and the generated teaching video is highly targeted, so that the personalized requirements of users can be met.
Owner:JIANG QIUSHI

Multi-mode video subtitle identification method and system, electronic equipment and storage medium

The invention provides a multi-mode video subtitle recognition method and system, electronic equipment and a storage medium, and relates to the technical field of video processing, and the method comprises the steps: carrying out the audio and video track separation of a to-be-recognized video, and obtaining an audio file and a video file; carrying out human voice track and background sound track separation on the audio file to obtain a human voice track audio; performing subtitle recognition on the human voice track audio by adopting an automatic voice recognition method with a timestamp to obtain a first subtitle text; performing subtitle area detection on the video file according to the visual language model to obtain a subtitle area external frame; performing frame-by-frame subtitle recognition on the video file by adopting an optical character recognition method according to the external frame of the subtitle area to obtain a second subtitle text; and performing subtitle fusion on the first subtitle text and the second subtitle text according to the time axis to obtain a subtitle recognition result. According to the invention, the completeness and accuracy of subtitle recognition are improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Video processing method and device, equipment, storage medium and program product

The invention provides a video processing method and device, equipment, a storage medium and a program product. Comprising the following steps: based on eye movement data of a first object for a first video frame sequence in a target video and a second video frame sequence in the target video, determining gaze probability distribution of the first object on each region of the second video frame sequence; determining a subtitle priority of a subtitle text corresponding to the second video frame sequence; the subtitle priority is used for representing the importance degree of the subtitle text in the target video; performing video frame identification on each video frame in the second video frame sequence in sequence to obtain a contour area of each second object in the video frames; determining subtitle parameters of the subtitle text in the second video frame sequence based on the gaze probability distribution, the subtitle priority and the contour region; and performing subtitle rendering on the second video frame sequence based on the subtitle parameters. According to the invention, accurate layout and dynamic adaptation of the subtitle text in the target video can be realized.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and Method for AI-Powered Generation and Delivery of Video Clips

A system, a method and a process are for AI-powered generation and delivery of video clips. The processor is configured to: classify a first video content item including a first video file comprising video frames associated with timestamps and a first subtitle file comprising subtitle text associated with the timestamps, wherein the first video content item is classified with narrative classifiers by: executing a natural language processing (NLP) model with the first subtitle file as input, the NLP model including a dialogue analysis for identifying first narrative elements from dialogue included in the first subtitle file and associating the first narrative elements with first timestamps, executing an image recognition model with the first video file as input, the image recognition model including an object identification analysis for identifying second narrative elements from objects or persons portrayed in the video frames of the first video file and associating the second narrative elements with second timestamps, combining a first output of the NLP model with a second output of the image recognition model, and generating a first set of timestamps associated with the narrative classifiers; define one or more segments within the first video content item, each segment comprising a starting timestamp and an ending timestamp defining a duration and having one or more of the narrative classifiers associated therewith; and generate a video clip including one or more of the segments based on prioritization rules in which some narrative classifiers are associated with a priority for inclusion in the video clip, the one or more segments selected for inclusion in the video clip so that a combined duration of the one or more segments is less than a set time value, the set time value being less than a full duration of the first video content item.
Owner:PARAMOUNT GLOBAL INC

Video subtitle identification method, device, equipment, medium and program product

The embodiment of the invention provides a video subtitle recognition method and device, equipment, a medium and a program product. The video subtitle recognition method comprises the steps that text content and text observation features of all video frames in a target video are acquired; subtitle recognition prompt information is constructed based on the text content and the text observation features, and the subtitle recognition prompt information is used for indicating a subtitle recognition model to determine subtitles in the text content according to the text observation features; based on the subtitle recognition prompt information, the subtitle recognition model is called to generate a subtitle recognition result, and the subtitle recognition result comprises a subtitle spatial-temporal feature and a subtitle text. The subtitle recognition prompt information is constructed based on the text content and the text observation characteristics of each video frame, so that the subtitle recognition model can perform context correlation analysis on the text content in combination with the text observation characteristics, and the text segments with correlation are recognized as subtitles in the continuous frames, so that the subtitle recognition accuracy is improved.
Owner:XINGIN INFORMATION TECH (SHANGHAI) CO LTD

Video embedded subtitle detection method and device, electronic equipment and storage medium

The invention relates to the technical field of computer vision, and provides a video embedded subtitle detection method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an audio stream associated with a video stream, carrying out the voice recognition of the audio stream, and generating first time sequence data comprising a transliteration text and corresponding time information; extracting a video frame sequence from the video stream, and performing character recognition on each video frame in the video frame sequence to generate second time sequence data comprising a recognition text and corresponding position information; and performing cross-modal time sequence alignment and text content comparison based on the first time sequence data and the second time sequence data, and determining target subtitle information embedded in the video stream according to a comparison result. According to the method, the cross-modal analysis of the audio and the video is introduced, so that the scene characters in the embedded subtitles and the video images can be effectively distinguished, the false detection rate is greatly reduced, and the subtitle detection accuracy is remarkably improved.
Owner:ANHUI FEISHU INFORMATION TECHNOLOGY CO LTD

Graphical User Interface for Audio and Video Chat on Electronic Devices

ActiveCN309763814SGraphical user interfaceVideo chat
1. Name of the product in this design: Audio and video chat graphical user interface for electronic devices. 2. Purpose of this design: An electronic device. 3. The key design feature of this product is its graphical user interface. 4. The image or photo that best illustrates the design points: Design 1 Interface Change State Diagram 1. 5. Design 1 is designated as the basic design. 6. Purpose of the graphical user interface: for audio and video chat. 7. Description of the changing states of the graphical user interface: Clicking "Voice Call" in the main view of Design 1 will result in the interface changing state diagram 1 of Design 1. In the interface changing state diagram 1 of Design 1, the real-time collected voice information will be converted into subtitles for display, resulting in the interface changing state diagram 2 of Design 1. Clicking "Video Call" in the main view of Design 2 will bring up the Design 2 interface change state diagram 1. In the Design 2 interface change state diagram 1, the real-time collected voice information will be converted into subtitles for display, resulting in the Design 2 interface change state diagram 2. 8. Other situations requiring explanation: In each view, "XX" represents text or characters, and each view uses color blocks to represent variable content screens.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

English movie and television reading difficulty grading method and system based on natural language processing

The invention provides an English movie and television reading difficulty grading method and system based on natural language processing, and relates to the technical field of English movie and television grading, and the method comprises the steps: firstly collecting multi-mode subtitle data of a target English movie and television resource, the data comprising language layer information and physical layer information; on the basis of language layer information, complexity features of a source language and comparison features of Chinese and English translation are extracted in parallel through a natural language processing technology, and meanwhile playing feature information of physical layer information is converted to obtain adaptive feature vectors; and then fusion feature representation corresponding to each fragment of the target English film and television resource is generated through a multi-modal cooperation mechanism, and after sequential dynamic modeling is carried out by a recurrent neural network, a modeling result is subjected to grading processing to obtain a difficulty grading result. The English film and television reading difficulty can be accurately and efficiently graded.
Owner:BEIJING INFORMATION TECH COLLEGE

Automatic in-game subtitles and closed captions

ActiveUS12427413B2Video gamesSpeech recognitionClosed captioningSubtitle
An approach is provided for a gaming overlay application to provide automatic in-game subtitles and / or closed captions for video game applications. The overlay application accesses an audio stream and a video stream generated by an executing game application. The overlay application processes the audio stream through a text conversion engine to generate at least one subtitle. The overlay application determines a display position to associate with the at least one subtitle. The overlay application generates a subtitle overlay comprising the at least one subtitle located at the associated display position. The overlay application causes a portion of the video stream to be displayed with the subtitle overlay.
Owner:ATI TECHNOLOGIES ULC +1

A method, apparatus, device and medium for erasing video subtitles

This invention discloses a video subtitle erasure method. It involves acquiring a fidelity stream and a computational stream of the video to be processed, where the computational stream consists of multiple computational frames. Each computational frame includes a target detection region. For each computational frame, the method detects whether a target subtitle to be erased exists within each target detection region. If so, it acquires an initial repair mask for the target subtitle. It then performs structural texture restoration on the initial repair masks to acquire corresponding target repair masks. Based on the fidelity stream, the initial repair masks, and the corresponding target repair masks, it obtains the subtitle-erased video. This method effectively improves the computational efficiency of acquiring the subtitle-erased video and enhances its quality.
Owner:BEIJING YUNSHANG TECH CO LTD

Subtitle processing methods and devices

This disclosure relates to a subtitle processing method and apparatus. The method includes: during the editing of a multimedia material segment, obtaining subtitle text corresponding to the audio and timestamp information of audio segments corresponding to each text element in the subtitle text through speech recognition; determining the material segment matching the text element in the multimedia material segment based on the timestamp information of the audio segment corresponding to each text element; and then synthesizing each text element with the matching material segment within the specified time to obtain a target multimedia material with a subtitle text appearing word by word in an animation effect. The solution of this disclosure can achieve a subtitle animation effect where the corresponding text subtitle appears when a certain word is spoken; furthermore, user input commands can automatically generate dynamic subtitles, simplifying user operation and improving user experience.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Display device and subtitle language identification method

The application provides a display device and a subtitle language recognition method. The method can acquire a subtitle corresponding to a media video in response to a play instruction of the media video, extract visual texture features of the subtitle in a case where the subtitle is an image subtitle, perform language recognition according to the visual texture features to obtain a language recognition result, extract syntax topology features of the subtitle in a case where the subtitle is a text subtitle, perform language recognition according to the syntax topology features to obtain a language recognition result, write the language recognition result into metadata corresponding to the media video, control a display to play the media video and the subtitle, and display a language identifier corresponding to the subtitle according to the language recognition result in the metadata corresponding to the media video. The method recognizes a language by analyzing features of different languages in physical forms and syntax structures, so that the language identifier of the subtitle can be correctly displayed when the media video is played.
Owner:HISENSE ELECTRONICS TECH SHENZHEN CO LTD

Real-time sign language-subtitle-voice three-dimensional synchronous generation system for barrier-free drama

The invention relates to the technical field of computer vision and natural language processing, in particular to a barrier-free drama-oriented real-time sign language-subtitle-voice three-dimensional synchronous generation system. Comprising a multi-mode drama content collection module, a drama semantic and emotion deep analysis module, a three-dimensional emotional content generation module, a millisecond-level synchronous calibration module, a personalized demand adaptation module, a multi-terminal output module and a feedback iteration module which are linked in sequence. The multi-mode drama content acquisition module supports offline theaters and online live broadcast / recorded broadcast scenes and can acquire line audios, actor performance data and scene auxiliary data, the deep binding of role personalization, plot emotion, scene atmosphere and barrier-free content is realized for the first time, the pain points of'action stiffness and emotion missing 'of a general system are solved, and the system has the advantages of being high in practicability and high in practicability. The sign language / voice / subtitle drama adaptation degree is improved to 95% or above, the overall time delay is controlled within 100 ms through a synchronous calibration algorithm and is far lower than the standard of 200 ms in the industry, and the drama continuity is guaranteed.
Owner:方锦瑶

Multi-language subtitle display method, system and device

The embodiment of the invention provides a multi-language subtitle display method, system and device, and the multi-language subtitle display method is applied to terminal equipment, and comprises the steps: receiving a first language selected by a user for playing a video; acquiring subtitle configuration information of the played video; the subtitle configuration information comprises resource information of subtitles of multiple languages of the played video; loading a subtitle file of the subtitles of the first language based on resource information of the subtitles of the first language in the subtitle configuration information; decoding a subtitle file of the subtitles of the first language to obtain a first decoding result; and rendering and displaying a first subtitle of the first language according to the first decoding result and the first playing progress of the played video.
Owner:SWEET POTATO TECHNOLOGY (SHANGHAI) CO LTD

Automated Media Packaging, Validation, and Delivery System

The present disclosure relates to a cloud-based system designed to automate the packaging, validation, processing, and delivery of media content and its supporting items. The system processes video, audio, subtitles, artwork, and metadata according to predefined specifications, ensuring compliance with technical and qualitative requirements for various endpoints. The system performs validation, error correction, transcoding, file conversion, and packaging based on saved profiles or templates. By leveraging cloud-based workflows and optional human oversight, the system ensures efficient and accurate delivery of media content to any destination.
Owner:PANTOJA PAULETTE

Cross-modal structure alignment method and system for video subtitle generation

The invention provides a cross-modal structure alignment method and system for video subtitle generation, and belongs to the technical field of video subtitle generation. The method aims at solving the problems that in an existing attention mechanism, structural compatibility in a multi-modal or language generation scene is not considered, and noise is generated in the modal fusion process. Structural compatibility in a multi-modal or language generation scene is considered, feature loss after text features and visual features are fused can be reduced, noise generated by modal fusion is reduced, natural mismatch of cross-modal fusion in a semantic mapping space is relieved, and the method is suitable for being popularized and applied. Therefore, the modeling capability of the model on the high-order semantic relationship is improved, and the negative influence generated after multi-modal fusion is reduced.
Owner:HARBIN ENG UNIV