Video subtitle identification method, device, equipment, medium and program product
By acquiring the text content and features of video frames, constructing subtitle recognition prompts, and calling the model to generate results, the problems of completeness and accuracy in dynamic subtitle recognition are solved, and efficient recognition and differentiation of dynamic subtitles are achieved.
Patent Information
- Application Number
- CN202512054824.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to accurately identify the complete spatiotemporal trajectory of dynamic subtitles in videos and cannot distinguish between dynamic subtitles and static scene text, leading to incomplete recognition and misclassification, which affects the accuracy and reliability of subsequent processing.
By acquiring the text content and text observation features of video frames, subtitle recognition prompts are constructed, and a subtitle recognition model is called to generate subtitle recognition results, including subtitle spatiotemporal features and subtitle text, thereby realizing cross-frame association and dynamic subtitle recognition.
It improves the completeness and accuracy of dynamic subtitle recognition, ensures the accurate output of the spatiotemporal characteristics and text content of the subtitles, and supports the efficiency and reliability of subsequent processing.
Smart Images

Figure CN121937984A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to a video subtitle recognition method, apparatus, device, medium, and program product. Background Technology
[0002] With the widespread application of video content in social media, education, and other fields, textual information, especially dynamically changing subtitles, has become a crucial carrier for conveying key information. For example, real-time narration subtitles in short videos, scrolling announcements in live streams, news marquees at the bottom of news videos, and annotations appearing alongside individuals in vlogs are all typical examples of dynamic subtitles, broadly carrying core semantic information. Intelligent video processing technology urgently needs to accurately identify and understand this dynamic text to support advanced applications such as accessibility, copyright management, and intelligent editing.
[0003] However, existing methods are mostly based on single-frame recognition and simple tracking, which are difficult to deal with rapid movement, deformation or occlusion of subtitles, and are prone to trajectory breakage. At the same time, they lack semantic understanding of text functions and cannot effectively distinguish between dynamic subtitles and static scene text, resulting in incomplete recognition and misclassification, which affects the accuracy and reliability of subsequent processing.
[0004] Therefore, there is an urgent need for a method that can accurately identify dynamic subtitles in videos. Summary of the Invention
[0005] In view of this, embodiments of this specification provide a video subtitle recognition method. One or more embodiments of this specification also relate to a video subtitle recognition device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, a video subtitle recognition method is provided, comprising:
[0007] Obtain the text content and text observation features of each video frame in the target video;
[0008] Subtitle recognition prompt information is constructed based on the text content and the text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features;
[0009] Based on the subtitle recognition prompt information, the subtitle recognition model is invoked to generate subtitle recognition results, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text.
[0010] According to a second aspect of the embodiments of this specification, a video subtitle recognition device is provided, comprising:
[0011] The acquisition module is configured to acquire the text content and text observation features of each video frame in the target video;
[0012] The construction module is configured to construct subtitle recognition prompt information based on the text content and the text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features;
[0013] The calling module is configured to call the subtitle recognition model to generate subtitle recognition results based on the subtitle recognition prompt information, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text.
[0014] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0015] Memory and processor;
[0016] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the video subtitle recognition method described above.
[0017] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the video subtitle recognition method described above.
[0018] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the video subtitle recognition method described above.
[0019] One embodiment of this specification implements the acquisition of text content and text observation features of each video frame in a target video; constructing subtitle recognition prompt information based on the text content and text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine subtitles in the text content according to the text observation features; and generating subtitle recognition results based on the subtitle recognition prompt information, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text. By constructing subtitle recognition prompt information based on the text content and text observation features of each video frame, the subtitle recognition model can combine the text observation features to perform contextual correlation analysis on the text content, identifying related text segments as subtitles in consecutive frames, thereby improving the completeness and accuracy of dynamic subtitle recognition. Attached Figure Description
[0020] Figure 1This is a flowchart illustrating a video subtitle recognition method provided in one embodiment of this specification;
[0021] Figure 2a This is a schematic diagram of a task material selection interface provided in one embodiment of this specification;
[0022] Figure 2b This is a schematic diagram of a video loading interface provided in one embodiment of this specification;
[0023] Figure 2c This is a schematic diagram illustrating the updating of a video loading interface according to one embodiment of this specification;
[0024] Figure 2d This is a schematic diagram of an updated video loading interface provided in one embodiment of this specification;
[0025] Figure 2e This is a schematic diagram of a video browsing interface provided in one embodiment of this specification;
[0026] Figure 2f This is a flowchart illustrating the processing procedure of a video subtitle recognition method provided in one embodiment of this specification;
[0027] Figure 3 This is a schematic diagram of the structure of a video subtitle recognition device provided in one embodiment of this specification;
[0028] Figure 4 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0029] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0030] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0031] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0032] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0033] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0034] LLM: Large Language Model, refers to a neural network model trained on large-scale text data that possesses deep semantic understanding and generation capabilities.
[0035] OCR: Optical Character Recognition, is a technology that detects and recognizes text information from images or video frames.
[0036] ROI: Region of Interest, in an image or video frame, is the area that is of particular interest because it contains specific information (such as text).
[0037] Spatiotemporal localization refers to the precise determination of the appearance time, duration, and spatial coordinates of an object (such as a subtitle) in both time and space dimensions.
[0038] Selective processing refers to the technique of applying differentiated processing (such as blurring, preservation, and enhancement) to text based on its semantic function and spatiotemporal attributes, rather than applying a globally uniform processing.
[0039] Video has become a core medium for information dissemination, and dynamic text, such as scrolling news, tracking annotations, and vlog subtitles, widely carries key semantic information, playing a vital role in content creation, copyright protection, and accessibility. Intelligent recognition and processing of dynamic text in videos is a crucial technological step in achieving structuring of video content and enhancing user experience.
[0040] Existing technologies face two major challenges when processing dynamic text in videos: First, it is difficult to accurately track the complete spatiotemporal trajectory of dynamic subtitles. When the text moves quickly, deforms, or is briefly obscured, the trajectory is prone to breakage, resulting in large errors in the start and end time positioning. Second, it is difficult to distinguish text with different functions. For example, dynamic subtitles and static scene text may look similar, but they have different uses and require differentiated processing.
[0041] Currently, text tracking typically employs a combination of independent single-frame recognition and optical flow or bounding box matching, relying on preset rules or independent classification models to determine text functionality. This approach treats spatial detection, temporal tracking, and semantic understanding in isolation, lacking a unified modeling capability.
[0042] Because existing methods do not fully integrate temporal context and semantic information, they have poor generalization ability in complex scenarios, weak adaptability to new styles, and require frequent rule updates or model retraining, resulting in incomplete dynamic caption recognition, high category misclassification rate, and affecting the accuracy and reliability of subsequent processing.
[0043] To address the aforementioned problems, this specification provides a video subtitle recognition method. This specification also relates to a video subtitle recognition device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0044] See Figure 1 , Figure 1 A flowchart of a video subtitle recognition method according to an embodiment of this specification is shown, specifically including the following steps 102-106.
[0045] Step 102: Obtain the text content and text observation features of each video frame in the target video.
[0046] Text content refers to the text strings identified from video frames. It is the basic information for understanding the semantics of the image. For example, readable text such as "Today's weather is sunny" and "Click to learn more" are extracted by optical character recognition technology.
[0047] Text observation features refer to the set of spatial and temporal attributes corresponding to the text identified in a video frame. They include two dimensions: text coordinates and timestamps. They are used to describe the location of the text in the frame and the time when it appears, and are the basic input unit for building a time series analysis model.
[0048] In practical applications, the system first samples the target video frame by frame. This can be done by extracting video frames at fixed time intervals or by selecting keyframes based on motion intensity or scene transition detection results. For each extracted video frame, the system calls the local OCR engine to perform text recognition, extracting one or more text blocks. For each recognized text block, the system records its corresponding text content and extracts its two-dimensional coordinates within the current frame, forming text coordinates to represent its specific area in the image. Simultaneously, the system determines its timestamp based on the frame's playback time in the video, serving as a time identifier for the text's appearance. The content of all text blocks and their corresponding text observation features are associated and stored one by one. Subsequently, the system organizes the recognition results of each frame in chronological order, forming a data sequence containing multiple frames of text content and corresponding text observation features. For this step, one option is to run a lightweight OCR model directly on the device to complete the recognition; another option is to send the extracted image data to an edge computing node to perform OCR processing and return the results; yet another option is to use a multi-language supported OCR service to adapt to different language scenarios.
[0049] For example, in a video editing assistance task, the system receives a 60-second target video containing scrolling real-time information captions at the bottom and a fixed program name watermark at the top. The system extracts frames from the video at a rate of 10 frames per second, obtaining a total of 600 video frames. For each video frame, the system calls a locally deployed OCR component to recognize the text content and generates bounding box coordinates for each recognized text block, such as "stock market up 2%" located in the [50, 900, 750, 940] region in a certain frame. Simultaneously, the system labels the frame with a timestamp based on its time position; for example, frame 101 corresponds to 10.1 seconds. For identical or similar text appearing consecutively in multiple frames, the system retains a complete record of each occurrence, including content, coordinates, and timestamp. The recognition results of all frames are aggregated into an ordered dataset, serving as the input source for the next step of constructing prompt information.
[0050] Furthermore, the text observation features include text spatial coordinates and text timestamps; obtaining the text content and text observation features of each video frame in the target video includes: sampling multiple video frames from the target video at fixed time intervals; performing text block recognition on the target video frames to extract at least one text block, wherein the target video frame is any one of the multiple video frames; determining the text content of the target text block, and determining the text spatial coordinates and text timestamps corresponding to the text content, wherein the target text block is any one of at least one text block, the text spatial coordinates represent the position region of the target text block in the target video frame, and the text timestamp represents the time position of the target video frame in the target video.
[0051] In practical applications, the system first samples the target video frame by frame. This can be done by extracting video frames at fixed time intervals, such as one frame every 100 milliseconds, or by automatically adapting the sampling frequency to the video frame rate. For each extracted video frame, the system calls the local OCR engine to perform text block recognition, detecting one or more text regions in the image. For each recognized text block, the system maps its image region to standardized text content and records the bounding box information of the text block within the current frame, forming text spatial coordinates. Simultaneously, the system calculates the absolute time position of the frame based on its playback time in the video, generating a corresponding text timestamp. All recognized text blocks are bound to their content, spatial coordinates, and timestamp, and stored as the recognition result of a single frame. Subsequently, the system organizes the recognition results of multiple video frames in chronological order, forming an ordered data sequence. For this step, one option is to use a lightweight OCR model to directly complete the recognition on the terminal device to reduce the dependence on external services; another option is to upload the extracted image data to the edge server to perform OCR processing and send the results back; yet another option is to combine a motion detection mechanism to dynamically adjust the sampling density and increase the sampling frequency during periods of drastic changes in the image to capture more details.
[0052] In the embodiments of this specification, by sampling and extracting the content, spatial coordinates and timestamps of text blocks in each frame at fixed time intervals, a raw observation dataset with spatiotemporal attributes is constructed, providing a complete and structured input foundation for subsequent cross-frame association and dynamic pattern recognition.
[0053] For example, in a short video processing task, the system receives a 45-second target video containing scrolling captions at the bottom and a fixed title at the top. The system sets the sampling interval to 200 milliseconds and extracts a total of 225 video frames. For each video frame, the system runs the locally integrated OCR component to recognize the text content. In frame 30 (corresponding to 6.0 seconds), the system detects two text blocks: one located in a rectangular area of [80, 920, 720, 960], with the content "Today we will introduce"; the other located in an area of [10, 10, 200, 50], with the content "Life tips". The system records the text content, text spatial coordinates, and timestamp 6000 for each of these two text blocks. As subsequent frames are processed, the same scrolling caption gradually moves to the left, and the system retains the position and content record for each appearance. The recognition results of all frames are summarized into a time-sorted list, with each entry containing a triplet of text content, coordinates, and timestamp, serving as the raw input for the next step of constructing prompt information.
[0054] Further, the text content of the target text block is determined, and the text spatial coordinates and text timestamps corresponding to the text content are determined, including: content matching of text blocks in multiple video frames based on the text content of the target text block; aggregating text blocks with the same text content into the same text instance, and recording multiple text timestamps and text spatial coordinates corresponding to the text instance; and determining the integration time interval and integration spatial region of the text instance based on the multiple text timestamps and text spatial coordinates corresponding to the text instance.
[0055] A text instance is a logical object formed by the system by aggregating the same or highly similar text content in multiple frames. It represents a text entity that appears continuously in the video, and its lifecycle is defined by the start and end times and the range of changes in its position.
[0056] The integrated time interval refers to the time range spanned by a text instance from its first appearance to its last disappearance. It is determined by the minimum and maximum values among multiple text timestamps and is used to describe the complete duration of the text's presence in the video.
[0057] The integrated spatial region refers to the spatial location range covered by a text instance during its existence. It can be the union of all text spatial coordinates or the envelope region of its movement path, used to reflect the overall activity area of the text in the picture.
[0058] In practical applications, the system first acquires the text content, text spatial coordinates, and text timestamp of each text block in multiple video frames. Then, the system compares these text blocks according to their text content, using string similarity calculation to determine if they are different occurrences of the same text. For text blocks with identical content or meeting a preset similarity threshold, the system groups them into the same text instance. Each text instance maintains a list recording its corresponding text timestamp and text spatial coordinates in each frame. After matching and aggregating all frames, the system statistically analyzes the timestamp set of the text instance, taking the earliest timestamp as the start time and the latest timestamp as the end time to generate the integrated time interval for the text instance. Simultaneously, the system merges the spatial coordinate sets of the text instance, calculating its bounding rectangle or motion trajectory envelope to form an integrated spatial region. For this step, one optional approach is to use precise string matching in the content matching stage to improve efficiency; another optional approach is to introduce a fuzzy matching algorithm to handle minor differences in content caused by OCR recognition errors; yet another optional approach is to combine visual features such as text font and color to assist in determining whether they belong to the same instance, improving matching accuracy.
[0059] In the embodiments of this specification, by performing cross-frame matching and aggregation based on text content, text instances with complete time span and spatial coverage are generated, enabling the system to restore the real existence cycle and movement pattern of dynamic text, providing a reliable basis for subsequent identification of whether it is a continuously moving subtitle.
[0060] For example, in a news broadcast video processing task, the system has extracted the OCR results of the bottom scroll bar from 150 consecutive frames. The text "International oil prices rose 3% today" appears repeatedly between frames 20 and 45, but due to scrolling displacement, its position gradually shifts to the left in each frame. The system binds the identified text content in each frame to its coordinates and timestamp and initiates a content matching process. Through string comparison, the system confirms that these 26 occurrences all belong to the same semantic content, and thus aggregates them into a single text instance. The timestamp set of this instance covers from 2.0 seconds to 4.5 seconds, and the system determines its integrated time interval as [2000ms, 4500ms]. Its spatial coordinates gradually move from [800, 900, 1500, 940] to [100, 900, 800, 940] in each frame. The system calculates the minimum bounding rectangle of all coordinates, obtaining the integrated spatial region as [100, 900, 1500, 940]. The complete spatiotemporal information of the text instance is preserved for subsequent determination of whether it constitutes a complete dynamic subtitle segment.
[0061] Step 104: Construct subtitle recognition prompt information based on text content and text observation features. The subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features.
[0062] Subtitle recognition prompts are structured natural language instructions generated by the system based on text content and text observation features. They include task objectives, input data format descriptions, and analysis requirements, and are used to guide the subtitle recognition model to make coherence judgments and functional classifications of the text by combining spatiotemporal information.
[0063] In practical applications, the system first integrates the text content from each video frame acquired in the preceding steps with the corresponding text observation features, forming a data sequence arranged in chronological order. Each item in this sequence corresponds to a text record within one or more frames, including its content, coordinates, and timestamp. The system converts this data sequence into a machine-readable structured format, such as key-value pairs or nested objects, and adds necessary context labels to clarify the meaning of the fields. Subsequently, the system loads a preset prompt template, which includes character setting statements, task description statements, and output format constraints. The system embeds the converted structured data into a specified position in the template, generating complete subtitle recognition prompt information. For this step, one optional approach is to encapsulate the input data in JSON format and insert it as a code block into the natural language prompt; another optional approach is to simplify the time-series data into text paragraph descriptions, such as "Text A appears at time t1, position p1; Text B appears at time t2, position p2," and attach analysis instructions. The generated prompts can include explicit instructions on subtitle aggregation rules, start and end time annotation requirements, and conditions for distinguishing between dynamic and static text, ensuring that the subtitle recognition model can reason based on the complete context.
[0064] For example, in a caption extraction task, the system has obtained 30 consecutive frames of OCR results from a news video. Each frame contains a portion of the text in the bottom scrollbar, along with its coordinates and timestamp. The system organizes this data into an ordered list, with each entry containing three fields: "timestamp," "text content," and "bounding box coordinates." The system organizes this data using standard JSON format and embeds it into a pre-designed prompt template. The template begins by stating "You are a professional video caption analysis engine," followed by the task description: "Based on the provided time-series text data, identify all complete dynamic caption segments, label their start and end times, and determine whether they are fixed identifiers in the scene." The system inserts the JSON data after "Input data is as follows:" and appends the output format requirements: "Please return the results in list form, with each entry containing caption text, start time, end time, and type label." Finally, a complete prompt message is generated, ready to be submitted to a remote large language model service for processing.
[0065] Furthermore, the method constructs subtitle recognition prompts based on text content and text observation features, including: generating a structured data sequence based on the text content and text observation features of each video frame, wherein the structured data sequence includes multiple frame entries, each frame entry corresponds to a video frame, and corresponds to at least one text content and at least one text observation feature of the text content identified in the video frame; obtaining a subtitle recognition prompt template, wherein the subtitle recognition prompt template includes role setting instructions, task description instructions, and output format requirement information, wherein the role setting instructions are used to define the role of the subtitle recognition model, the task description instructions are used to instruct the subtitle recognition model to recognize dynamic subtitles and static scene text in the structured data sequence, and the output format requirement information is used to specify the structured return format of the subtitle recognition results; and embedding the structured data sequence into the subtitle recognition prompt template to form subtitle recognition prompts.
[0066] Structured data sequences are collections of multi-frame text information generated by the system and organized in chronological order. Each entry corresponds to a video frame and contains one or more text contents identified within that frame and their corresponding text observation features, providing a complete and ordered data foundation for subsequent analysis.
[0067] The subtitle recognition prompt template is a pre-defined natural language framework that includes role setting instructions, task description instructions, and output format requirements. Role setting instructions clarify the functional role of the subtitle recognition model in this task; task description instructions explain the specific analysis objectives to be performed; and output format requirements constrain the structure of the returned results to ensure that downstream systems can directly parse them.
[0068] The subtitle recognition prompt information is the final input formed by embedding the structured data sequence into the subtitle recognition prompt template. It is used to guide the subtitle recognition model to complete the task of recognizing and classifying dynamic subtitles.
[0069] In practical applications, the system first integrates the text content of each video frame obtained in the preceding steps with the corresponding text observation features to form a data sequence sorted by playback time. This sequence consists of multiple frame entries, each corresponding to a sampled video frame and containing information on all identified text blocks within that frame. For each text block, the system records its text content, text spatial coordinates, and text timestamp, organized as key-value pairs or object arrays. Subsequently, the system loads a pre-designed subtitle recognition prompt template, which contains three core parts: first, a role setting instruction, such as "You are a professional video subtitle analysis engine"; second, a task description instruction, such as "Based on the provided time-series text data, identify all complete dynamic subtitle segments and distinguish which are fixed scene texts"; and third, output format requirements, such as "Please return the results in JSON format, including the fields: start_time, end_time, text, bbox, type". The system converts the aforementioned structured data sequence into a suitable data representation, such as a JSON string or code block, and inserts it into the specified position in the template, for example, "Input data is as follows: json...". The final step generates a complete, machine- and human-readable natural language prompt, i.e., caption recognition prompt information. For this step, one possible approach is to encapsulate the structured data sequence in standard JSON format and embed it as an independent field in the template. Another possible approach is to convert each frame of data into a natural language description sentence, such as "At 1.2 seconds, the text 'Welcome to watch' appeared at the bottom of the screen, at position [50,900,750,940]", and then concatenate all sentences before embedding them in the template. Yet another possible approach is to adjust the prompt style according to the target model's language preferences, such as using a more concise command-style language or adding example samples (few-shot prompting) to improve understanding accuracy.
[0070] In the embodiments of this specification, by organizing temporal text data into a structured sequence and embedding prompt templates with clear roles, tasks and format constraints, the subtitle recognition model can understand the analysis intent within a unified framework, thereby supporting cross-frame semantic and spatiotemporal joint reasoning and improving the standardization and usability of the output results.
[0071] For example, in a video editing assistance task, the system has completed OCR processing of a documentary, obtaining recognition results for 60 video frames. Each frame contains an average of two texts: one is the scrolling narration at the bottom, and the other is the fixed column name at the top of the screen. The system organizes the data of each frame into an entry, each entry containing two fields: "timestamp" and "texts". The "texts" field is an array storing attributes such as "content" and "coordinates" for multiple text blocks. The system organizes this structured data sequence in JSON format and loads a preset prompt template. The template begins by stating: "You are a video content understanding assistant. Please assist in identifying dynamic subtitles and static markers in the following data." It then describes the task: "Please determine which texts are dynamic subtitles that move or update over time, and which are scene texts with fixed positions and unchanged content." The output format is then specified: "Return a list, each element containing four fields: text, start, end, bbox, and category." The system inserts the JSON data after "Input data:" to form the final prompt. For example, in a certain frame, "Evolutionary history of the Earth" appears in [100, 890, 700, 930], and its position remains unchanged for multiple consecutive frames; while "Approximately 4.5 billion years ago, the Earth's crust began to form" shifts to the left frame by frame, and the system retains all its records. After the entire prompt message is generated, it is ready to be submitted to the remote language model service for processing.
[0072] Furthermore, obtaining the subtitle recognition prompt template includes: obtaining subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information. The subtitle timing aggregation information instructs the subtitle recognition model to aggregate consecutively occurring text content into subtitle segments. The subtitle attribute annotation information instructs the subtitle recognition model to annotate each subtitle segment with a start timestamp, end timestamp, and text type. The category discrimination information instructs the subtitle recognition model to distinguish between subtitles and static scene text. Based on the subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information, a task description instruction is constructed. Based on preset role setting instructions, output format requirements, and task description instructions, a subtitle recognition prompt template is constructed.
[0073] Subtitle timing aggregation information refers to the instruction content used to guide the subtitle recognition model to merge text segments that appear continuously across frames and have related content or position into complete semantic units. Its function is to enable the model to recognize the entire process of the same subtitle from appearance to disappearance, forming a coherent subtitle segment.
[0074] Subtitle attribute annotation information refers to the instructions used to instruct the subtitle recognition model to add metadata such as start timestamp, end timestamp, and text type to each recognized subtitle segment, supporting precise management of the subtitle lifecycle and functional categories.
[0075] Category discrimination information refers to the judgment rules used to guide the subtitle recognition model to distinguish between dynamic subtitles and static scene text based on the text's positional change patterns, frequency of occurrence, and contextual features. This ensures that the model can accurately identify the target object to be processed and eliminate interference items.
[0076] The task description instruction is a natural language or structured statement generated by the system based on the combination of subtitle timing aggregation information, subtitle attribute annotation information and category discrimination information. It is used to clearly inform the subtitle recognition model of the specific analysis task to be performed and is the core logical part of the prompt template.
[0077] The role setting instruction is a preset text used to define the identity of the subtitle recognition model in this task, such as "You are a professional video content analysis engine", which is used to enhance the model's understanding of the seriousness and professionalism of the task and improve the consistency of output.
[0078] The output format requirements are preset technical constraints used to specify the structure of the subtitle recognition results. For example, they may require specific fields to be included in JSON format to ensure that the response data can be directly parsed and used by downstream modules, avoiding the uncertainty brought by free text.
[0079] In practical applications, the system first loads a predefined set of task elements, from which it retrieves subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information. This information is stored in parameterized form, which can be key-value pairs in a configuration file or records in a database. The system converts these three types of information into natural language expressions. For example, subtitle timing aggregation information is converted to "Please aggregate text blocks that move continuously or have progressive content into a complete subtitle segment"; subtitle attribute annotation information is converted to "Affix the start time, end time, and type label (such as 'narration' or 'bullet screen') to each subtitle segment"; and category discrimination information is converted to "Determine whether the text moves or updates over time; if not, classify it as static scene text." Subsequently, the system concatenates and integrates these three descriptions to form a complete task description instruction. Simultaneously, the system reads preset role setting instructions and output format requirements, such as "You are a senior video semantic understanding expert," and "Please return the results in list format, where each element must contain five fields: text, start_time, end_time, bbox, and category." The system organizes task description instructions, role setting instructions, and output format requirements in a fixed order, typically starting with role settings, centering the task description, and ending with format requirements to construct the final subtitle recognition prompt template. For this step, one option is to use a template variable replacement mechanism to dynamically inject different task parameters at runtime, enabling reuse across multiple scenarios. Another option is to introduce a version control mechanism to manage different versions of the prompt template, facilitating rollback and comparative testing. Yet another option is to support a visual editing interface, allowing technicians to configure task description instructions by dragging and dropping components and automatically generating corresponding prompt templates.
[0080] In the embodiments of this specification, by structuring task elements such as subtitle timing aggregation, attribute annotation, and category discrimination into prompt templates, refined guidance of language model behavior is achieved, improving the completeness, standardization, and task adaptability of model output.
[0081] For example, during the initialization of a general subtitle recognition service, the system needs to build subtitle recognition prompt templates applicable to various video types. The system loads standard task elements from the configuration center: subtitle timing aggregation information is defined as "treating identical or similar text that appears in two or more consecutive frames and whose horizontal displacement is less than a threshold as the same subtitle segment"; subtitle attribute annotation information is defined as "each subtitle segment needs to be annotated with start time, end time, spatial coordinates, and type (dynamic / static)"; category discrimination information is defined as "if the text remains unchanged in position across multiple frames and there are no content updates, it is determined to be static scene text". The system translates these rules into natural language and combines them into task description instructions: "Based on the provided time-series text data, identify all complete dynamic subtitle segments, aggregate consecutively appearing text, label their start time, end time, location area, and type, and distinguish which are fixed background texts." Next, the system adds a role setting instruction: "You are a professional video content structuring engine," and output format requirements: "Return a JSON array, where each object contains the fields text, start_time, end_time, bbox, and category." This ultimately forms a complete, clear, and structurally consistent prompt template, which is saved to the local resource directory for subsequent tasks.
[0082] Step 106: Based on the subtitle recognition prompt information, call the subtitle recognition model to generate subtitle recognition results, whereby the subtitle recognition results include the subtitle spatiotemporal features and the subtitle text.
[0083] The subtitle recognition result is the output data returned by the subtitle recognition model, which contains information about one or more subtitle segments identified. Each segment includes fields such as subtitle text and subtitle spatiotemporal features, which describe the content of the subtitle and its time interval and spatial location.
[0084] In practical applications, the system sends the completed subtitle recognition prompts to a remote subtitle recognition model service via a network interface. This interface supports standard communication protocols; the system submits the prompts in the form of a request message and waits for the server's response. Upon receiving the returned response data, the system parses it and extracts the structured content that conforms to a preset format. If the response is in JSON format, the system reads top-level fields such as "subtitle_list" or "results," iterates through each entry, and obtains information such as the subtitle text, start time, end time, type label, and spatial coordinate range. All valid entries are summarized into the final subtitle recognition result. For this step, one option is to use a synchronous call mechanism, blocking and waiting for the model to return the result in the current task flow; another option is to use an asynchronous task queue mode, submitting the request first and recording the task ID, and then obtaining the processed result through polling or callback notification. The system can also interface with multiple different subtitle recognition model services, selecting the appropriate model to execute the call based on resource availability or response quality.
[0085] For example, in a video content processing task, the system has completed the pre-processing OCR and prompt information construction for a variety show, and is ready to submit it to the cloud-based language model for subtitle recognition. The system sends a POST request to the specified API endpoint via HTTPS, sending the prompt information containing the timing data of scrolling subtitles within 30 seconds as the request body. After receiving the request, the server schedules computing resources in the background to perform inference and returns a JSON response approximately two seconds later. Upon receiving the response, the system first verifies its integrity and signature validity, and then begins parsing. The system extracts the array under the "detected_subtitles" field from the response, and reads the "text", "start_time", "end_time", "bbox", and "category" values for each subtitle item by item. For example, it identifies the text "Welcome to this episode" appearing at the bottom of the screen from 5.2 seconds to 7.8 seconds, categorized as "Dynamic Subtitles". All successfully parsed entries are organized into a unified data structure for use in subsequent image processing or storage tasks.
[0086] Furthermore, based on the subtitle recognition prompt information, the subtitle recognition model is invoked to generate subtitle recognition results, including: sending the subtitle recognition prompt information to the subtitle recognition model service in the cloud via the application programming interface for spatiotemporal semantic joint analysis; receiving the response data returned from the cloud, parsing the response data, and obtaining structured subtitle recognition results, wherein the subtitle recognition results include a dynamic subtitle list and a static scene text list, and the dynamic subtitle list includes at least one subtitle text and the text type and spatiotemporal features corresponding to the subtitle text.
[0087] Cloud-based caption recognition model services refer to large-scale language models or dedicated inference services deployed on remote servers. These services have the ability to process natural language prompts, perform sequence modeling and logical inference, and can perform dynamic judgment, category classification and spatiotemporal range recognition of text based on input data.
[0088] The response data is the result information returned by the subtitle recognition model service after completing inference. It is usually in JSON or other structured formats and contains the recognized subtitle fragments and their attribute fields. It needs to be parsed before it can be used by downstream tasks.
[0089] The subtitle recognition result is the final output extracted from the response data, including a dynamic subtitle list and a static scene text list. The dynamic subtitle list contains one or more text segments with start and end times and spatial location variation patterns; each entry includes the subtitle text, text type, and corresponding spatiotemporal features of the subtitle, used to describe its periodicity and motion characteristics.
[0090] In practical applications, the system encapsulates the completed subtitle recognition prompts into a standard request body and sends it to the designated cloud-based subtitle recognition model service interface via network protocol. This request includes necessary authentication information, content type declarations, and timeout settings to ensure communication security and stability. Upon receiving the request, the server schedules computing resources in the background to execute inference tasks. It performs spatiotemporal semantic joint analysis, combining the temporal data and instructions in the prompts to determine which text blocks constitute continuously moving dynamic subtitles and which belong to fixed-position, unchanging scene text. It also labels the start time, end time, spatial trajectory, and functional category of each dynamic subtitle. After analysis, the server organizes the results into a pre-formatted data object and returns a response. Upon receiving the response data, the system first verifies its completeness and validity, then initiates the parsing process, reading top-level fields such as "results" or "subtitle_output," traversing the array entries, and extracting the content, text type label, start time, end time, and bounding box information for each subtitle text. All valid entries that meet the format requirements are categorized as dynamic subtitles or static scene text and summarized into structured subtitle recognition results. For this step, one option is to use a synchronous call mode, which blocks and waits for the service response in the current task flow, suitable for scenarios with high real-time requirements; another option is to return the task ID immediately after submitting the task, and then query the status by polling and obtain the result after completion, suitable for long videos or high-load environments; yet another option is to configure multiple backup model service endpoints and automatically switch when the main service is unavailable, improving the system's fault tolerance.
[0091] In the embodiments of this specification, by calling the cloud model service to perform spatiotemporal semantic joint analysis, a technical path is realized that complex reasoning can be completed without local training, enabling the system to accurately identify complete segments of dynamic subtitles and distinguish their functional differences from static text.
[0092] For example, in a variety show processing task, the system has completed the construction of prompt information and is ready to submit it to the cloud-based language model service. The video includes scrolling bullet comments at the bottom (dynamic subtitles) and fixed brand logos in the stage background (static scene text). The system packages the prompt information containing 60 frames of OCR data into a JSON request body and sends it to the API address "https: / / ai.example.com / v1 / subtitle-analyze" via HTTPS protocol. After receiving the request, the server starts inference and returns a structured response. After receiving the response, the system parses its content and determines that the "dynamic_subtitles" array contains a record: "text:'This is too funny', start:12500, end:13800, bbox:[50,500,750,540], type:comment"; while in the "static_elements" list, it identifies "text:'Happy Camp', position:[300,10,500,60], category:logo". The system extracts and organizes these data one by one into a unified data structure for subsequent image processing modules to call.
[0093] Furthermore, the spatiotemporal features of the subtitles include the subtitle time segment and the subtitle spatial coordinates. After generating the subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information, the process also includes: responding to the subtitle processing task request, determining the video timestamp of the target video frame, wherein the target video frame is any one of multiple video frames, and the subtitle processing task includes the subtitle to be processed; querying the target subtitle time segment and target subtitle spatial coordinates corresponding to the subtitle to be processed in the subtitle recognition result; determining whether the video timestamp matches the target subtitle time segment; if so, performing image processing operations in the target video frame based on the target subtitle spatial coordinates.
[0094] Subtitle spatiotemporal characteristics refer to the composite attributes that describe the periodicity and location range of subtitles. They are composed of subtitle time segments and subtitle spatial coordinates, and are used to fully depict the coverage area of the subtitle in both time and space dimensions.
[0095] The subtitle time segment refers to the time interval spanned by a subtitle from its first appearance to its final disappearance, including the start timestamp and the end timestamp, used to determine whether the subtitle is visible at a certain moment.
[0096] Subtitle spatial coordinates refer to the position area of a subtitle in the picture. They are usually represented by a rectangular bounding box and are used to identify the specific distribution range of the subtitle within the video frame. They are the basis for performing pixel-level operations.
[0097] The target video frame refers to the specific video frame for which image operations need to be performed during the captioning task, and its time position is uniquely determined by the video timestamp.
[0098] Subtitles to be processed refer to subtitle objects that require specific image operations as specified by the user or system. They can be specified based on subtitle text, type labels, or location information and are key inputs that trigger the processing flow.
[0099] In practical applications, the system enters a waiting state after completing subtitle recognition. When a subtitle processing task request is received, the subsequent process begins. This request contains the identification information of the subtitle to be processed, such as the specific text content or category label. The system first parses the request, determines the target processing object, and searches for a matching entry in the subtitle recognition results, extracting the target subtitle time segment and target subtitle spatial coordinates. Subsequently, the system acquires the current target video frame to be processed and reads its playback time in the video, i.e., the video timestamp. The system compares this video timestamp with the target subtitle time segment to determine whether it falls between the start and end times (inclusive). If the result is yes, it means that the subtitle is in a display state in this frame, and the system continues to perform the specified image processing operation within the screen area of the target video frame based on the target subtitle spatial coordinates; if the result is no, the frame is skipped or marked as not requiring processing. For this step, one option is to use a closed interval method for time matching to ensure that the first and last frames are not missed; another option is to introduce a buffer mechanism to extend the time threshold before the start time and after the end time to cope with model inference delay or frame sampling error; yet another option is to support batch processing mode, load multiple subtitles to be processed and their spatiotemporal parameters at one time, scan and perform operations frame by frame according to a unified rule, and improve the overall processing efficiency.
[0100] In the embodiments of this specification, by combining the dual matching mechanism of subtitle time segment and spatial coordinates, the existence status of subtitles is accurately determined, so that image processing operations only take effect within the appropriate time and location range, thereby improving the accuracy and security of processing.
[0101] For example, in a video watermark removal task, the system has identified the scrolling sponsor subtitle at the bottom of a variety show: "This program is sponsored by XX brand." The subtitle recognition result indicates a start time of 30.2 seconds and an end time of 59.8 seconds, with spatial coordinates located in the [50, 910, 750, 950] region at the bottom of the screen. When the user initiates a request to "blur the subtitle," the system responds to this task and begins traversing each keyframe of the original video. For frame 302 (corresponding to time 30.2 seconds), the system reads its video timestamp 30200, finds the start and end times of the subtitle to be [30200, 59800], and determines that 30200 falls within this range, resulting in a match. The system then extracts the subtitle's spatial coordinates [50, 910, 750, 950] and performs a Gaussian blur operation on the corresponding area of that frame. For frame 600 (time 60.0 seconds), the system reads timestamp 60000, determines it exceeds the end time of 59800, classifies it as a mismatch, and skips processing. Throughout the process, the system only performs operations on frames within the display time frame, avoiding accidental modification of irrelevant images while ensuring that all relevant frames are processed correctly.
[0102] Furthermore, image processing operations are performed in the target video frame based on the target subtitle spatial coordinates, including at least one of the following: performing special effects adjustments in the processing area determined based on the target subtitle spatial coordinates; performing blurring processing in the processing area determined based on the target subtitle spatial coordinates; obtaining replacement text based on the target video and subtitle text, performing blurring processing in the processing area determined based on the target subtitle spatial coordinates, and adding replacement text.
[0103] The target subtitle spatial coordinates refer to the location area of the subtitle to be processed in the picture. It is usually represented by a rectangular bounding box and is used to define the specific distribution range of the subtitle within the video frame. It is the basis for performing pixel-level operations.
[0104] The processing area refers to the sub-region of the screen defined by the target subtitle space coordinates. The system applies specified image transformation operations within this area to ensure that the processing behavior only acts on the target object and avoids affecting surrounding unrelated pixels.
[0105] Special effects adjustment refers to applying visual enhancement or modification operations to the image content within the processing area, such as brightness adjustment, contrast optimization, edge sharpening, or stylized rendering, to improve the display effect of subtitles or adapt to specific creative needs.
[0106] Blur processing refers to smoothing the image content within a processing area. Methods such as Gaussian blur, mean blur, or motion blur are used to reduce the sharpness of the area. It is often used for tasks such as privacy protection, copyright protection, or weakening of interfering information.
[0107] Replacement text refers to new text content generated to replace the original subtitles. It can be generated based on the target video's metadata, user input, or an external language model, and is used to achieve subtitle updates, translation replacements, or content corrections.
[0108] In practical applications, after confirming that the current target video frame requires processing of a specific subtitle, the system first delineates a clear processing area within the frame based on the target subtitle spatial coordinates obtained in previous steps. Then, the system selects the corresponding image manipulation strategy based on the received subtitle processing task type. If the task requires special effects adjustments, the system performs color enhancement, contrast optimization, or local sharpening transformations on the pixels within the processing area to improve readability or visual integration. If the task requires blurring, the system applies a preset blurring algorithm within the area to make the original subtitle content unrecognizable, achieving the purpose of masking. If the task requires content replacement, the system first obtains the replacement text, which can come from user configuration, official subtitle files, or an automatically generated service. Then, based on the blurring process, the replacement text is redrawn within the same processing area with an appropriate font, size, and style to maintain layout consistency. For this step, one option is to use dynamic blur intensity control, automatically adjusting the blur kernel size according to the speed of the subtitle movement to make the processing result more natural; another option is to automatically detect the background color and adjust the text outline or shadow when adding replacement text to improve contrast and readability; yet another option is to support multilingual replacement, automatically loading the corresponding language translation text for overlay based on the target audience region.
[0109] In the embodiments described in this specification, by performing differentiated image operations based on precise spatial coordinates, fine-grained control of the subtitle area is achieved, enabling the system to flexibly select processing methods according to task requirements, and maintain the overall quality of the image while completing the target intervention.
[0110] For example, in an accessibility viewing assistance task, the system has identified a low-quality, automatically generated hard-coded subtitle in an instructional video: "This is a key knowledge point," with the subtitle's time range from 120.5 seconds to 126.3 seconds and spatial coordinates [100, 890, 700, 940]. When the user requests to "optimize the subtitle," the system responds by checking if the timestamp in each frame falls within the aforementioned range. For frame 1205 (time 120.5 seconds), the system confirms a match, extracts the spatial coordinates, and delineates the processing area. The system first performs a slight blurring of this area to reduce the visibility of the original erroneous text. Then, it retrieves the manually proofread replacement text from local resources: "This is a key knowledge point," and redraws it in the same position using a standard font, white as the primary color, and a black outline, ensuring clarity and a consistent style. For subsequent frames, the system continues this process until the subtitle ends. Throughout the process, the system did not make any modifications to other areas of the screen, including the faces of the people, the content on the whiteboard, and the logo in the corner, all of which remained unchanged.
[0111] Furthermore, the spatiotemporal features of the subtitles include the subtitle time segment and the subtitle spatial coordinates; after generating the subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information, it also includes: verifying whether there is an overlap in the subtitle time segment; if so, then integrating the subtitle spatial coordinates of each overlapping subtitle.
[0112] Overlapping subtitles refer to subtitles whose time segments overlap on the timeline. This means that multiple subtitles are determined to be displayed within the same time period, which may cause issues such as overlapping processing areas or resource contention.
[0113] Integration processing refers to performing unified and coordinated operations on multiple captions that have temporal overlap, including adjusting their time segments, merging spatial regions, or reallocating processing priorities, to ensure that subsequent image operations can be performed in an orderly manner.
[0114] In practical applications, after generating the subtitle recognition results, the system initiates a verification process, comparing the subtitle time segments of all dynamic subtitle entries in the results pairwise. The system sequentially reads the time segments of each pair of subtitles, determining whether they overlap—that is, whether the end time of the preceding subtitle is greater than the start time of the following subtitle, and whether the start time of the preceding subtitle is less than the end time of the following subtitle. If the determination result indicates overlap, the pair of subtitles is marked as an overlapping subtitle group and enters the integration processing stage. The system analyzes the relative positional relationship of each overlapping subtitle in the screen based on its subtitle spatial coordinates. If the spatial regions do not intersect and are far apart, their original time segments are retained; if the spatial regions are close or there is visual interference, the time segments are fine-tuned, such as shortening the display duration of one subtitle or introducing a fade-in / fade-out transition mechanism. Another option is to merge the spatial coordinates of multiple overlapping subtitles into a unified processing area, and then perform overall blurring or masking operations on the combined area. Yet another option is to set priority rules based on subtitle type tags, and when overlap occurs, prioritize the preservation of the integrity of high-priority subtitles, while performing time clipping or position avoidance processing on low-priority subtitles.
[0115] In the embodiments of this specification, by verifying whether there is overlap in the subtitle time segments and performing corresponding integration processing, orderly management of multiple subtitle co-occurrence scenarios is achieved, avoiding processing anomalies caused by time conflicts and improving the stability and rationality of system behavior.
[0116] For example, in a news broadcast video, the system identified two continuously scrolling bottom captions: the first read "Domestic News Flash: 12 New Cases Today," with a time interval of [15.2s, 18.7s]; the second read "International News: European and American Stock Markets Generally Rise," with a time interval of [18.5s, 22.0s]. The system detected a 0.2-second overlap between the two (from 18.5s to 18.7s). At this point, the system read the spatial coordinates of the two captions and found that the first caption was located at [50, 900, 750, 940], and the second at [60, 905, 760, 945]. Their positions were close and partially overlapped, causing visual confusion. The system initiated an integration process, employing a time fine-tuning strategy to advance the end time of the first caption from 18.7 seconds to 18.6 seconds, and delay the start time of the second caption from 18.5 seconds to 18.6 seconds, thus eliminating the overlapping area. Simultaneously, the system updates the corresponding fields in the subtitle recognition results and records this adjustment in the log. For subsequent blurring or replacement tasks performed based on these results, the corrected time segment will be used to ensure that the processing actions do not simultaneously affect two adjacent subtitle areas, thus avoiding image anomalies.
[0117] Furthermore, after generating the subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information, the process also includes: matching the subtitle text with a preset content understanding feature library; determining the subtitle text that conforms to the content processing rules and its corresponding content understanding category identifier and subtitle spatiotemporal features based on the matching result, and generating the content understanding classification result; generating content understanding classification prompt information based on the content understanding classification result, and sending the content understanding classification prompt information to the front end for display.
[0118] Content understanding feature library refers to a pre-set dataset for text pattern recognition, which contains specific keywords, fixed expressions or structured language templates. It can be organized according to semantic categories and associated with corresponding content understanding category identifiers, serving as the basis for text matching and classification.
[0119] Content understanding category identifiers refer to pre-defined labels for texts of different semantic categories, used to distinguish subtitle texts from differences in language patterns or functional attributes, such as classifying them into categories A, B, C, etc., to support subsequent differentiated processing logic.
[0120] Content understanding classification results refer to the structured output generated by the system based on the matching process between subtitle text and content understanding feature library. It includes all successfully matched subtitle texts and their corresponding content understanding category identifiers, subtitle spatiotemporal features and other fields, which are used to support the condition judgment and process control of downstream tasks.
[0121] Content understanding and categorization prompts refer to converting the content understanding and categorization results into an information format that can be parsed by the front end. This includes the text content of the matching subtitles, the time range of appearance, the screen location area, and the associated category identifier, which are used for visual display or interactive guidance in the user interface.
[0122] In practical applications, after completing subtitle recognition, the system initiates a content understanding comparison process. First, the system extracts the subtitle text of all dynamic subtitle entries from the subtitle recognition results and compares them one by one with a locally or remotely stored content understanding feature library. The comparison process employs exact matching or multimodal matching mechanisms, supporting wildcards, synonym substitution, and variant avoidance detection to ensure effective recognition of variant expressions. When a subtitle text is found to match an entry in the feature library, the system records the matching relationship and determines its corresponding content understanding category identifier according to the preset mapping rules in the feature library. Simultaneously, the system retains the spatiotemporal features of the corresponding subtitle, including start time, end time, and spatial coordinates, forming a complete matching record. All successfully matched entries are summarized as the content understanding classification result. Subsequently, the system converts this result into a format suitable for front-end display, such as a JSON structure or message object, which contains the text, time segment, location area, and content understanding category identifier for each matched subtitle. This content understanding classification prompt is sent to the front-end management platform via a communication channel and displayed in the user interface as a timeline marker, pop-up reminder, or list highlight. For this step, one possible approach is to use an incremental matching mechanism to trigger comparison in real time when new subtitle recognition results are generated, thereby achieving streaming processing; another possible approach is to introduce a weighted scoring model to comprehensively calculate the matching confidence based on the number of matches, context, and content understanding category identifiers, which is used for sorting and priority display; yet another possible approach is to support a feedback loop mechanism, allowing users to confirm or correct the matching results and send the corrected samples back to optimize the feature library.
[0123] In the embodiments of this specification, by automatically matching the subtitle text with the content understanding feature library and generating a structured output with spatiotemporal positioning, efficient recognition and accurate positioning of specific language patterns are achieved, thereby improving the response efficiency and processing accuracy of downstream tasks for target subtitles.
[0124] For example, in a user-uploaded short video processing task, the system has completed subtitle recognition, identifying the scrolling subtitle segment at the bottom: "Click the link to receive high rebates." The system initiates the content understanding comparison process, loading a preset content understanding feature library containing keywords such as "high rebates," "earn money by brushing orders," and "internal channels," and associating them with category identifiers "Category A" and "Category B" respectively. The system scans the subtitle text word by word, finding that "high rebates" perfectly matches an entry in the feature library, thus marking it as a successful match and associating it with the content understanding category identifier "Category A." Simultaneously, the system extracts the spatiotemporal features corresponding to this subtitle: start time 8.3 seconds, end time 10.7 seconds, spatial coordinates located at the bottom of the screen [60, 910, 740, 950]. This record is added to the content understanding classification results. The system further generates content comprehension and categorization prompts, including the fields: {"text":"Click the link to receive a high rebate","start_time":8300,"end_time":10700,"bbox":[60,910,740,950],"category":"Category A"}. This information is pushed to the front-end workbench in real time via a WebSocket connection, displayed as a highlighted bar on the video timeline, along with accompanying explanations. Users can directly click to jump to the corresponding time segment to view the relevant content without manually searching for and matching subtitles.
[0125] One embodiment of this specification implements the acquisition of text content and text observation features of each video frame in a target video; constructing subtitle recognition prompts based on the text content and text observation features, wherein the subtitle recognition prompts are used to instruct the subtitle recognition model to determine subtitles in the text content according to the text observation features; and generating subtitle recognition results based on the subtitle recognition prompts, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text. By constructing subtitle recognition prompts based on the text content and text observation features of each video frame, the subtitle recognition model can combine the text observation features to perform contextual correlation analysis on the text content, identifying related text segments as subtitles in consecutive frames, thereby improving the completeness and accuracy of dynamic subtitle recognition.
[0126] The following describes the front-end interaction process for video subtitle recognition provided in the embodiments of this specification. Please refer to [link to documentation]. Figure 2a , Figure 2a This is a schematic diagram of a task material selection interface provided in one embodiment of this specification, such as... Figure 2a As shown, the task material selection interface can be displayed, which includes multiple media materials, as well as the video generation control "One-Click Video Generation".
[0127] In response to the media material selection operation in the task material selection interface and the triggering of "one-click video creation," multiple task materials for the content generation task in this embodiment can be obtained. Based on these multiple task materials, the video subtitle recognition and analysis method provided in this embodiment can be executed to perform spatiotemporal positioning and functional classification of the text content in the video, providing subtitle structured information for subsequent video editing or copywriting generation. Figure 2b For example, Figure 2b This is a schematic diagram of a video loading interface provided in one embodiment of this specification. The client can respond to the "One-Click Video Generation" trigger operation of the video generation control in the task material selection interface to obtain the selected media material. The server can generate a video inference process based on the selected media material. The client can determine and display the video inference process 302a on the video loading interface. The inference process 302a includes at least one of the following inference information: a highlight segment 302b in at least one media material, a material content description 302g of at least one media material, a video theme 302c, a video content description 302d, and a video content summary 302e. The inference information is determined based on the selected at least one media material. The highlight segment 302b, material content description 302g, etc., in the at least one media material displayed in the above inference process 302a can be generated by other modules, while the dynamic subtitles identified in the video and their spatiotemporal position, functional category, etc., are determined based on the video subtitle recognition and analysis method provided in the embodiments of this specification.
[0128] by Figure 2b For example, the reasoning process includes a summary of the video content, 302e, which is "An Unforgettable Trip, with everyday narrative text, accompanied by relaxing music, and packaged in a simple, everyday style." Here, the video title is "An Unforgettable Trip"; the video text is "everyday narrative text," which can be understood as a video text type; the background music is "accompanied by relaxing music," which can be understood as a video background music type; the video style is "simple, everyday style packaging," which can be understood as a video style type. The video voiceover can be understood as a video voiceover type, such as "funny voice."
[0129] In addition, the video loading interface includes a command input field, through which update commands for at least one type of inference information are received. In response to the update command, the inference process updated based on the update command is displayed or dynamically displayed on the video loading interface. If the user wishes to adjust the recognition results or processing method of subtitles in the video (e.g., correcting the subtitle time range, changing the subtitle type determination, etc.), they can input an update command through the command input field. The system will then re-execute subtitle recognition and analysis based on the update command and display the updated subtitle information on the video loading interface. Continuing... Figure 2bFor example, the video loading interface includes a command input control 316a. Clicking the command input control 316a allows the client to respond to a trigger operation on the command input control 316a by pulling up the keyboard 318a in the video loading interface. Figure 2c As shown, Figure 2c This is a schematic diagram illustrating the update of a video loading interface according to one embodiment of this specification. When the command input control 316a is clicked, the keyboard changes from a hidden state to a raised state. An update command can be entered in the command input area 318b using the keyboard 318a. The entered update command can be displayed in the command input area 318b. The associated position of the command input area 318b may also include an input confirmation control 318c. The client can respond to the trigger operation of the input confirmation control 318c to confirm the entered update command and display the updated reasoning process in the video loading interface, such as... Figure 2d As shown, Figure 2d This is a schematic diagram of an updated video loading interface provided in one embodiment of this specification.
[0130] Continue with Figure 2b For example, the video loading interface also includes a video viewing control 302f. The client can respond to trigger operations on the video viewing control 302f and display, as shown below. Figure 2e The video browsing interface shown is as follows. Figure 2e This is a schematic diagram of a video browsing interface provided in one embodiment of this specification. The video browsing interface includes a video title, "An Unforgettable Trip," and a visual style, namely, a "simple style." For example... Figure 2e As shown, the video browsing interface also includes a video update control called "Regenerate". If you are not satisfied with the recognition results or processing effect of the subtitles in the current video, you can click the "Regenerate" control. The system will re-execute the subtitle recognition and analysis process and display the updated subtitle processing results (such as the corrected subtitle timeline or category labels) in the video browsing interface.
[0131] The following is in conjunction with the appendix Figure 2f Taking the video subtitle recognition method provided in this manual as an example in a video subtitle removal scenario, this paper further explains the video subtitle recognition method. Specifically, Figure 2f The present specification illustrates a flowchart of a video subtitle recognition method according to an embodiment, which includes the following steps.
[0132] Step 202: Extract and format local text metadata, perform OCR processing, and construct structured data.
[0133] Specifically, this step is performed on the user's local computer, and the core task is to extract text information from the video and construct it into a format suitable for LLM to understand.
[0134] First, a standard OCR tool (e.g., the open-source PaddleOCR library) is used to process each frame of the video, or frames sampled at fixed intervals. From each frame, all text blocks are extracted:
[0135] Spatial coordinates: [x_min, y_min, x_max, y_max].
[0136] Text content (text_content).
[0137] Timestamp: The time in milliseconds that this frame is in the video.
[0138] After OCR recognition, the OCR results from multiple consecutive frames are organized into a structured, time-series text sequence. For example, a JSON array where each object represents one frame of data, as shown below:
[0139]
[0140]
[0141] Step 204: Build the Prompt and call the Large Language Model API.
[0142] Specifically, by designing a Prompt, a general-purpose LLM can be guided to complete professional spatiotemporal analysis tasks.
[0143] First, the structured data generated in step 202 is embedded into a Prompt containing explicit instructions. This Prompt consists of three parts: role settings, task description and output format requirements, and input data. A specific example is shown below:
[0144] System Prompt: "You are a professional video content analyst. Your task is to identify continuous dynamic captions from a series of text data extracted from video frames, each with timestamps and coordinates, and to distinguish them from static scene text."
[0145] Task Description and Output Format Requirements: "Please analyze the following 'frame_sequence' data. Identify all text segments belonging to the same dynamic subtitle and aggregate them into a complete subtitle segment. For each identified subtitle segment, please provide the start timestamp, end timestamp, the complete text content after merging, and determine its type as 'dynamic_subtitle'. For independent scene text with fixed positions, please determine it as 'static_scene_text'. Please return the final result in JSON format, as follows: {"dynamic_subtitles":[...],"static_scene_texts":[...]}".
[0146] Input data: Attach the JSON data generated in step 202 as text.
[0147] The completed Prompt is used as input and sent to the Large Language Model Service API via an HTTP request.
[0148] Step 206: Parse the API return results.
[0149] Specifically, the local program receives the JSON result returned by LLMAPI, which is the precise processing instruction.
[0150] Parse the JSON string returned by the API to obtain a structured list of subtitle segments and scene text. For example:
[0151]
[0152] Step 208: Perform selective processing.
[0153] Specifically, the process iterates through each frame of the video and renders it according to the parsed instructions: obtains the timestamp of the current frame; checks if the time range of dynamic_subtitles (representing dynamic subtitles) covers the current frame; if it covers it, applies processing, such as Gaussian blur, to the representative_roi area of the subtitle; for areas marked as static_scene_text (representing static scene text), no operation is performed to ensure that they are clearly visible; finally, the processed video frame is generated and output.
[0154] Through steps 202-208 above, a zero-shot video dynamic text processing solution was constructed by designing a Prompt. The temporal text features extracted by local OCR are encapsulated as structured input, and guided by natural language prompts, a cloud-based large language model performs spatiotemporal trajectory analysis and semantic discrimination. This achieves accurate recognition and functional classification of dynamic subtitles, and outputs standardized results for downstream execution. This "perception-decision" separation architecture balances lightweight front-end with strong back-end inference capabilities. It achieves efficient cloud collaboration by transmitting only lightweight text data, reducing local computing power requirements while achieving high generalization, strong adaptability, and stable integrability.
[0155] Corresponding to the above method embodiments, this specification also provides embodiments of a video subtitle recognition device. Figure 3 A schematic diagram of a video subtitle recognition device according to one embodiment of this specification is shown. Figure 3 As shown, the device includes:
[0156] The acquisition module 302 is configured to acquire the text content and text observation features of each video frame in the target video;
[0157] Module 304 is configured to construct subtitle recognition prompt information based on text content and text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features;
[0158] Module 306 is configured to call the subtitle recognition model to generate subtitle recognition results based on subtitle recognition prompt information. The subtitle recognition results include subtitle spatiotemporal features and subtitle text.
[0159] Optionally, the text observation features include text spatial coordinates and text timestamps; correspondingly, the acquisition module 302 is further configured to sample multiple video frames from the target video at fixed time intervals; perform text block recognition on the target video frames to extract at least one text block, wherein the target video frame is any one of the multiple video frames; determine the text content of the target text block, and determine the text spatial coordinates and text timestamps corresponding to the text content, wherein the target text block is any one of at least one text block, the text spatial coordinates represent the position region of the target text block in the target video frame, and the text timestamp represents the time position of the target video frame in the target video.
[0160] Optionally, the acquisition module 302 is further configured to perform content matching on text blocks in multiple video frames based on the text content of the target text block; aggregate text blocks with the same text content into the same text instance, and record multiple text timestamps and text spatial coordinates corresponding to the text instance; and determine the integrated time interval and integrated spatial region of the text instance based on the multiple text timestamps and text spatial coordinates corresponding to the text instance.
[0161] Optionally, the construction module 304 is further configured to generate a structured data sequence based on the text content and text observation features of each video frame, wherein the structured data sequence includes multiple frame entries, each frame entry corresponds to a video frame, and corresponds to at least one text content and at least one text observation feature of the text content identified in the video frame; obtain a subtitle recognition prompt template, wherein the subtitle recognition prompt template includes a role setting instruction, a task description instruction, and output format requirement information, the role setting instruction is used to define the role of the subtitle recognition model, the task description instruction is used to instruct the subtitle recognition model to recognize dynamic subtitles and static scene text in the structured data sequence, and the output format requirement information is used to specify the structured return format of the subtitle recognition result; embed the structured data sequence into the subtitle recognition prompt template to form subtitle recognition prompt information.
[0162] Optionally, the construction module 304 is further configured to acquire subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information. The subtitle timing aggregation information is used to instruct the subtitle recognition model to aggregate continuously occurring text content into subtitle segments. The subtitle attribute annotation information is used to instruct the subtitle recognition model to annotate the start timestamp, end timestamp, and text type of each subtitle segment. The category discrimination information is used to instruct the subtitle recognition model to distinguish between subtitles and static scene text. Based on the subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information, a task description instruction is constructed. Based on the preset role setting instruction, output format requirement information, and task description instruction, a subtitle recognition prompt template is constructed.
[0163] Optionally, the calling module 306 is further configured to send the subtitle recognition prompt information to the subtitle recognition model service in the cloud via the application programming interface for spatiotemporal semantic joint analysis; receive the response data returned by the cloud, parse the response data, and obtain a structured subtitle recognition result, wherein the subtitle recognition result includes a dynamic subtitle list and a static scene text list, and the dynamic subtitle list includes at least one subtitle text and the text type and spatiotemporal features corresponding to the subtitle text.
[0164] Optionally, the spatiotemporal features of the subtitles include a subtitle time segment and subtitle spatial coordinates. Correspondingly, the video subtitle recognition device also includes a task processing module, configured to respond to a subtitle processing task request by: determining the video timestamp of a target video frame, wherein the target video frame is any one of multiple video frames, and the subtitle processing task includes subtitles to be processed; querying the target subtitle time segment and target subtitle spatial coordinates corresponding to the subtitles to be processed in the subtitle recognition results; determining whether the video timestamp matches the target subtitle time segment; and if so, performing image processing operations in the target video frame based on the target subtitle spatial coordinates.
[0165] Optionally, the task processing module is further configured to perform special effects adjustments in the processing area determined based on the target subtitle spatial coordinates; perform blurring in the processing area determined based on the target subtitle spatial coordinates; obtain replacement text based on the target video and subtitle text, perform blurring in the processing area determined based on the target subtitle spatial coordinates, and add replacement text.
[0166] Optionally, the spatiotemporal features of the subtitles include subtitle time segments and subtitle spatial coordinates; correspondingly, the video subtitle recognition device also includes an integration module, configured to verify whether there is overlap in the subtitle time segments; if so, integration processing is performed based on the subtitle spatial coordinates of each overlapping subtitle.
[0167] Optionally, the video subtitle recognition device also includes a classification module, configured to match subtitle text with a preset content understanding feature library; determine subtitle text that conforms to content processing rules and its corresponding content understanding category identifier and subtitle spatiotemporal features based on the matching results, and generate content understanding classification results; generate content understanding classification prompt information based on the content understanding classification results, and send the content understanding classification prompt information to the front end for display.
[0168] This video subtitle recognition device uses an acquisition module 302 to obtain the text content and text observation features of each video frame in the target video, providing basic data for subsequent analysis. A construction module 304 constructs subtitle recognition prompts based on the text content and text observation features, structuring the position, time, and other observation features of the text across multiple frames with their corresponding content, and embedding analysis instructions to guide the model to focus on the spatiotemporal changes in the text. A calling module 306 inputs this prompt information into the subtitle recognition model, enabling it to perform contextual analysis of the text content in consecutive frames during inference, combining text observation features to identify text segments with positional continuity and content evolution relationships as dynamic subtitles. Thus, through the collaborative work of these modules, the completeness and accuracy of dynamic subtitle recognition are improved.
[0169] The above is an illustrative scheme of a video subtitle recognition device according to this embodiment. It should be noted that the technical solution of this video subtitle recognition device and the technical solution of the video subtitle recognition method described above belong to the same concept. For details not described in detail in the technical solution of the video subtitle recognition device, please refer to the description of the technical solution of the video subtitle recognition method described above.
[0170] Figure 4 A structural block diagram of a computing device 400 according to one embodiment of this specification is shown. The components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 via a bus 430, and a database 450 is used to store data.
[0171] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0172] In one embodiment of this specification, the above-described components of the computing device 400 and Figure 4 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 4 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0173] Computing device 400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). Computing device 400 can also be a mobile or stationary server.
[0174] The processor 420 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the video subtitle recognition method described above.
[0175] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computing device embodiments are basically similar to the video subtitle recognition method embodiments, so the description is relatively simple; relevant parts can be referred to in the description of the video subtitle recognition method embodiments.
[0176] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the video subtitle recognition method described above.
[0177] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the video subtitle recognition method embodiments, so the description is relatively simple; relevant parts can be referred to in the description of the video subtitle recognition method embodiments.
[0178] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the video subtitle recognition method described above.
[0179] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the video subtitle recognition method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the video subtitle recognition method described above.
[0180] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0181] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0182] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0183] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0184] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A video subtitle recognition method, characterized in that, include: Obtain the text content and text observation features of each video frame in the target video; Subtitle recognition prompt information is constructed based on the text content and the text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features; Based on the subtitle recognition prompt information, the subtitle recognition model is invoked to generate subtitle recognition results, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text.
2. The method according to claim 1, characterized in that, The text observation features include text spatial coordinates and text timestamps; obtaining the text content and text observation features of each video frame in the target video includes: Sample multiple video frames from the target video at fixed time intervals; Text block recognition is performed on the target video frame to extract at least one text block, wherein the target video frame is any one of the plurality of video frames; The text content of the target text block is determined, and the text spatial coordinates and text timestamp corresponding to the text content are determined. The target text block is any one of the at least one text block. The text spatial coordinates represent the position area of the target text block in the target video frame, and the text timestamp represents the time position of the target video frame in the target video.
3. The method according to claim 2, characterized in that, The process of determining the text content of the target text block and determining the text spatial coordinates and text timestamp corresponding to the text content includes: Content matching is performed on text blocks in multiple video frames based on the text content of the target text block; Text blocks with the same text content are aggregated into the same text instance, and multiple text timestamps and text spatial coordinates corresponding to the text instance are recorded. Based on the multiple text timestamps and text spatial coordinates corresponding to the text instance, the integrated time interval and integrated spatial region of the text instance are determined.
4. The method according to claim 1, characterized in that, The construction of subtitle recognition prompt information based on the text content and the text observation features includes: A structured data sequence is generated based on the text content and text observation features of each video frame. The structured data sequence includes multiple frame entries, each frame entry corresponds to a video frame, and corresponds to at least one text content identified in the video frame and the text observation features of the at least one text content. Obtain a subtitle recognition prompt template, wherein the subtitle recognition prompt template includes a role setting instruction, a task description instruction, and output format requirement information. The role setting instruction is used to define the role of the subtitle recognition model. The task description instruction is used to instruct the subtitle recognition model to recognize dynamic subtitles and static scene text in the structured data sequence. The output format requirement information is used to specify the structured return format of the subtitle recognition result. The structured data sequence is embedded into the subtitle recognition prompt template to form subtitle recognition prompt information.
5. The method according to claim 4, characterized in that, The process of obtaining the subtitle recognition prompt template includes: The system acquires subtitle timing aggregation information, subtitle attribute annotation information, and category discrimination information. The subtitle timing aggregation information is used to instruct the subtitle recognition model to aggregate continuously occurring text content into subtitle segments. The subtitle attribute annotation information is used to instruct the subtitle recognition model to annotate each subtitle segment with a start timestamp, end timestamp, and text type. The category discrimination information is used to instruct the subtitle recognition model to distinguish between subtitles and static scene text. Based on the subtitle timing aggregation information, the subtitle attribute annotation information, and the category discrimination information, a task description instruction is constructed. Based on the preset character setting instructions and output format requirements, as well as the task description instructions, a subtitle recognition prompt template is constructed.
6. The method according to claim 1, characterized in that, The step of generating a subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information includes: The subtitle recognition prompt information is sent to the subtitle recognition model service in the cloud via the application programming interface for spatiotemporal semantic joint analysis; The system receives response data returned from the cloud, parses the response data, and obtains structured subtitle recognition results. The subtitle recognition results include a dynamic subtitle list and a static scene text list. The dynamic subtitle list includes at least one subtitle text and the text type and spatiotemporal features corresponding to the subtitle text.
7. The method according to claim 1, characterized in that, The spatiotemporal features of the subtitles include subtitle time segments and subtitle spatial coordinates; after generating subtitle recognition results by calling the subtitle recognition model based on the subtitle recognition prompt information, the method further includes: In response to a subtitle processing task request, the video timestamp of the target video frame is determined, wherein the target video frame is any one of the plurality of video frames, and the subtitle processing task includes subtitles to be processed; In the subtitle recognition results, query the target subtitle time segment and target subtitle spatial coordinates corresponding to the subtitle to be processed; Determine whether the video timestamp matches the target subtitle time segment; If so, then image processing operations are performed in the target video frame based on the target subtitle space coordinates.
8. The method according to claim 7, characterized in that, The image processing operation performed in the target video frame based on the target subtitle spatial coordinates includes at least one of the following: Special effects adjustments are made in the processing area determined based on the target subtitle spatial coordinates; Blurring is performed in the processing area determined based on the target subtitle spatial coordinates; Replacement text is obtained based on the target video and the subtitle text. The replacement text is then blurred and added to the processing area determined based on the spatial coordinates of the target subtitle.
9. The method according to claim 1, characterized in that, The spatiotemporal characteristics of the subtitles include subtitle time segments and subtitle spatial coordinates; After generating the subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information, the method further includes: Verify whether the subtitle time segments overlap; If so, then the subtitle space coordinates of each overlapping subtitle will be used for integration processing.
10. The method according to claim 1, characterized in that, After generating the subtitle recognition result by calling the subtitle recognition model based on the subtitle recognition prompt information, the method further includes: The subtitle text is matched with a preset content understanding feature library; Based on the matching results, determine the subtitle text that conforms to the content processing rules, its corresponding content understanding category identifier and subtitle spatiotemporal features, and generate content understanding classification results; Based on the content understanding and classification results, a content understanding and classification prompt message is generated and sent to the front end for display.
11. A video subtitle recognition device, characterized in that, include: The acquisition module is configured to acquire the text content and text observation features of each video frame in the target video; The construction module is configured to construct subtitle recognition prompt information based on the text content and the text observation features, wherein the subtitle recognition prompt information is used to instruct the subtitle recognition model to determine the subtitle in the text content according to the text observation features; The calling module is configured to call the subtitle recognition model to generate subtitle recognition results based on the subtitle recognition prompt information, wherein the subtitle recognition results include subtitle spatiotemporal features and subtitle text.
12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the video subtitle recognition method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, It stores a computer program / instruction that, when executed by a processor, implements the steps of the video subtitle recognition method according to any one of claims 1-10.
14. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the steps of the video subtitle recognition method according to any one of claims 1-10.
Citation Information
Cited By
In-picture text translation replacement method and device
CN122160555A