Video processing method and device

By displaying the modifiable logo and pop-up editing portal during video playback, users can directly modify subtitles when watching videos, solving the problem of difficult to quickly correct video subtitles in the prior art and improving human-computer interaction efficiency.

CN120091183APending Publication Date: 2025-06-03SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510246393.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to quickly correct published video subtitle errors, and human-computer interaction is inefficient.

Method used

The icon can be modified by triggering a preset condition (such as video pause) display during video playback, and the editing portal pops up to receive the modified content entered by the user and upload it to the server to replace the target subtitles.

Benefits of technology

It enables users to directly modify subtitles when watching videos, without relying on video producers, quickly correcting published video subtitles errors, and improving human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091183A_ABST
    Figure CN120091183A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of multimedia. The video processing method comprises the following steps: playing a target video; wherein the target video comprises a current video frame embedded with a target subtitle; under the condition that a first preset condition is triggered, displaying a modifiable identifier for the target subtitle on the current video frame; and popping up an editing entry for modifying the target subtitle in response to the triggering of the region where the modifiable identifier is located, receiving current editing content through the editing entry; and uploading the current edited content to a server, so that the server replaces the target subtitle with the current edited content. According to the technical scheme provided by the embodiment of the invention, the published video can be quickly corrected. And the man-machine interaction efficiency is improved through a convenient mode of modifying the identifier and editing the entry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of multimedia technology, and in particular, to a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the rapid development of Internet technology and the explosive growth of video content, video has become an important medium for people to obtain information, entertainment, and learning. To enhance the viewing experience of videos, subtitles, as an important part of video content, can not only help viewers better understand the video content but also meet the needs of multilingual users. However, errors often occur in video subtitles, which not only affect the user's viewing experience but may also mislead viewers.

[0003] Currently, the traditional way to modify video subtitles usually depends on video producers to complete during the production stage. This method cannot quickly correct already released videos, and the human-computer interaction efficiency is low.

[0004] It should be noted that the above content is not necessarily prior art and does not limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above technical problems.

[0006] One aspect of embodiments of the present application provides a video processing method for a client, and the method includes: Playing a target video; wherein the target video includes a current video frame embedded with target subtitles; When a first preset condition is triggered, displaying a modifiable identifier for the target subtitles on the current video frame; and In response to the area where the modifiable identifier is located being triggered, popping up an editing entry for modifying the target subtitles; Receiving current editing content through the editing entry; and Uploading the current editing content to a server so that the server replaces the target subtitles with the current editing content.

[0007] Optionally, the first preset condition includes video pause; when the first preset condition is triggered, displaying a modifiable identifier for the target subtitles on the current video frame includes: In response to a pause operation on the target video, identifying the position of the target subtitles in the current video frame; Generate a modifiable identifier for the target subtitle according to the position; Wherein, the modifiable identifier includes a prompt box enclosing the target subtitle.

[0008] Optionally, the server replaces the target subtitle with the current edited content, including: Obtain the display area and time period occupied by the target subtitle in the target video; According to the display area and the time period, update the current edited content to the display area to cover the target subtitle; Wherein, the area occupied by the current edited content is the same as the area occupied by the target subtitle.

[0009] Optionally, the time period includes a start time and an end time; obtaining the display area and time period occupied by the target subtitle in the target video includes: Obtain the current timestamp corresponding to the current video frame; Identify frame by frame forward based on the current timestamp to obtain the start time when the target subtitle appears; and Identify frame by frame backward based on the current timestamp to obtain the end time when the target subtitle disappears; Determine the time period according to the start time and the end time.

[0010] Optionally, according to the display area and the time period, updating the current edited content to the display area to cover the target subtitle includes: Generate a text layer for covering the display area according to the display area and the time period; Update the current edited content to the text layer.

[0011] Another aspect of the embodiments of the present application provides a video processing method for a server, and the method includes: Obtain the display area, time period and current edited content uploaded by the client; wherein, the target subtitle is located within the display area; Generate a text layer for covering the display area according to the display area and the time period; Update the current edited content to the text layer.

[0012] Optionally, updating the current edited content to the text layer includes: When the text layer is generated, load the current edited content to the text layer; When the text layer is removed, end the display of the current edited content.

[0013] Another aspect of the embodiments of the present application provides a video processing device for a client, and the device includes: A playback module for playing a target video; wherein, the target video includes a current video frame embedded with a target subtitle; A display module for displaying a modifiable identifier for the target subtitle on the current video frame when a first preset condition is triggered; and A pop-up module for popping up an editing entry for modifying the target subtitle in response to the area where the modifiable identifier is located being triggered; A receiving module for receiving current editing content through the editing entry; and An uploading module for uploading the current editing content to a server so that the server replaces the target subtitle with the current editing content.

[0014] Another aspect of the embodiments of the present application provides a video processing device for a server, and the device includes: An acquisition module for acquiring a display area, a time period, and current editing content uploaded by a client; wherein, the target subtitle is located within the display area; A generation module for generating a text layer for covering the display area according to the display area and the time period; An update module for updating the current editing content to the text layer.

[0015] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0016] Another aspect of the embodiments of the present application provides a computer-readable storage medium, and computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0017] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0018] The embodiments of the present application adopting the above technical solutions may include the following advantages: During the process of playing a video frame embedded with a target subtitle, if a preset condition is triggered, a modifiable identifier is displayed. By triggering the area where the modifiable identifier is located, an editing entry can be popped up to receive the current editing content input by a user (such as a video viewer). The server replaces the target subtitle with the current editing content. It can be seen that in the embodiments of the present application, the user can directly modify the target subtitle (such as an incorrect subtitle) while watching the video, without relying on the video producer, and can quickly correct the published video. And through this shortcut of the modifiable identifier and the editing entry, the user can quickly locate and modify the target subtitle, improving the human-computer interaction efficiency. Description of the Drawings

[0019] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0020] Figure 1 Schematically shows the operating environment diagram of the video processing method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the video processing method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 The sub-step flowchart of step S202 in; Figure 4 Schematically shows Figure 2 The sub-step flowchart of step S208 in; Figure 5 Schematically shows Figure 4 The sub-step flowchart of step S400 in; Figure 6 Schematically shows Figure 4 The sub-step flowchart of step S402 in; Figure 7 Schematically shows the flowchart of the video processing method according to Embodiment 2 of the present application; Figure 8 Schematically shows Figure 7 The sub-step flowchart of step S704 in; Figure 9 Schematically shows the exemplary application flowchart of the embodiments of the present application; Figure 10 Schematically shows the video player screen diagram; Figure 11A display diagram schematically showing a start time and an end time; Figure 12 An interaction diagram schematically showing corrected subtitles; Figure 13 An effect diagram schematically showing the correction of these subtitles; Figure 14 A block diagram schematically showing a video processing apparatus according to Embodiment 3 of the present application; and Figure 15 A block diagram schematically showing a video processing apparatus according to Embodiment 4 of the present application; and Figure 16 A schematic diagram of the hardware architecture of a computer device according to Embodiment 5 of the present application. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0024] First, the following provides explanations of the terms involved in the present application: OCR (Optical Character Recognition) technology: A technology that converts the text in an image into editable text. It can scan documents, pictures, etc., recognize the text content therein, and convert it into a text format that can be processed by a computer.

[0025] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of this application by those skilled in the art, the related technologies are described below: The applicant has learned that subtitle errors during video playback are a common problem. These errors may be caused by various reasons, such as subtitle file encoding issues, data corruption during transmission, or human input errors, etc. Incorrect subtitles not only affect the viewing experience of the audience but may also lead to misunderstandings or inaccurate information transmission. Currently, the common subtitle correction method is that video creators correct the subtitles in video editing software, then re-encode and export the video, and finally upload it to a video playback platform or software, and only then can users see the corrected subtitle effect. This processing method is inefficient, especially when the number of subtitles is large or the video content is complex, it is difficult for video creators to discover and correct all incorrect subtitles in a timely manner.

[0026] For this reason, the embodiments of this application provide a video processing technical solution. In this technical solution, when a video viewer discovers a subtitle error, they can manually pause the video, use OCR technology to recognize the subtitles in the current frame and box out each subtitle. The video viewer selects the incorrect subtitle and enters the new subtitle in the pop-up keyboard input box. The system records the incorrect subtitle and its position, and traverses the video frames forward and backward to determine the subtitle correction time period (start time and end time). Then, the correction information (including the correction time period, position, and new subtitle) is uploaded to the server. The video player will render the corrected new subtitle during the correction time period and at the position, improving the subtitle quality. See the following for details.

[0027] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0028] As Figure 1 shown, the operating environment diagram includes: a server 2, a client 4, and a video production end 8.

[0029] The server 2 can connect to the client 4 and the video production end 8 through a network.

[0030] The server 2 can be a single server, a server cluster, or a cloud computing service center.

[0031] The server 2 can provide services such as updated subtitles to the client.

[0032] The server 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The server 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and their combinations, etc., or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0033] The client 4 and the video production end 8 can be separately configured to access the content and services of the service platform 2. The client 4 and the video production end 8 can each separately include an electronic device with a built-in or external display panel, such as a mobile device, a tablet device, a laptop computer, a workstation, a virtual reality device, a gaming device, a digital streaming device, a vehicle terminal, a smart TV, a set-top box, etc., or can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load the virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.

[0034] The client 4 can be associated with one or more viewers. A single viewer can also use one or more of the clients 4 to access the service platform 2. The client 4 can travel to various locations and use different networks to access the service platform 2.

[0035] The video production end 8 can be associated with one or more video creators. A single video creator can also use the video production end 8 to access the service platform 2. The video production end 8 can travel to various locations and use different networks to access the service platform 2. It should be noted that the client 4 and the video production end 8 can be the same or different devices.

[0036] The client / video production end 8 can include an interface. The interface can include a touchpad, a touch screen, a mouse, a keyboard, or other sensing elements. For example, the input element can be configured to receive viewer instructions, and the viewer instructions can cause the client to perform various operations, such as pausing a video, playing a video, generating (producing) a video, etc.

[0037] Note that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different requirements.

[0038] The technical solutions of the present application will be introduced below through multiple embodiments with the client and the server as the execution subjects respectively. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0039] Embodiment 1 This method embodiment can be executed in the client 4. The client will be used as the execution subject of this process below.

[0040] Figure 2 A flowchart of a video processing method according to Embodiment 1 of the present application is schematically shown.

[0041] As shown Figure 2 in the figure, the video processing method may include steps S200 to S208, where: Step S200, play a target video; where the target video includes a current video frame embedded with a target subtitle; Step S202, when a first preset condition is triggered, display a modifiable identifier for the target subtitle on the current video frame; and Step S204, in response to the area where the modifiable identifier is located being triggered, pop up an editing entry for modifying the target subtitle; Step S206, receive current editing content through the editing entry; and Step S208, upload the current editing content to a server so that the server replaces the target subtitle with the current editing content.

[0042] In the video processing method provided in this embodiment, during the process of playing a video frame embedded with a target subtitle, if a preset condition is triggered, a modifiable identifier is displayed. By triggering the area where the modifiable identifier is located, an editing entry can be popped up to receive the current editing content input by a user (such as a video viewer). The server replaces the target subtitle with the current editing content. It can be seen that in the embodiments of the present application, the user can directly modify the target subtitle (such as an incorrect subtitle) when watching the video without relying on the video producer, and can quickly correct the published video. And through this shortcut of the modifiable identifier and the editing entry, the user can quickly locate and modify the target subtitle, improving the human-computer interaction efficiency.

[0043] The following will combine Figure 2 to elaborate in detail on each step in steps S200 to S208 and optional other steps.

[0044] Step S200 , play a target video; where the target video includes a current video frame embedded with a target subtitle.

[0045] On the client 4, start a video player to load and play the target video. The target video is video content produced and published by video creators (such as UP owners, bloggers, film and television production companies, etc.). The target video can be various types such as on-demand videos, live videos, user-generated content (UGC), advertising videos, etc. The target video is produced and published by video creators through professional tools or platforms, and multiple subtitles are embedded therein.

[0046] The multiple subtitles include target subtitles. The target subtitles in the embodiments of this application refer to the subtitles embedded in the target video that may be incorrect, inaccurate, or require further optimization. The target subtitles may include the following types: (1) Original subtitles: Subtitles manually added by the video creator or automatically generated through speech recognition technology.

[0047] (2) Translated subtitles: Multilingual subtitles generated through machine translation or manual translation.

[0048] (3) Bullet subtitles: The bullet content sent by users in real time while watching the video, which is displayed in a scrolling form on the video screen.

[0049] (4) Annotation subtitles: Auxiliary subtitles used to explain the video content or provide additional information, such as term explanations, background descriptions, etc.

[0050] Step S202 , when the first preset condition is triggered, a modifiable identifier for the target subtitle is displayed on the current video frame.

[0051] The modifiable identifier is used to indicate the position of the target subtitle and prompt the user (it should be noted that the user mentioned later is the video viewer) that it can be modified. The modifiable identifier may include, but is not limited to, the following various visual cues: (1) Highlight box: For example, a semi-transparent highlight box is displayed around the target subtitle.

[0052] (2) Icon: For example, an editable icon (such as a pencil icon) is displayed near the target subtitle, and the user can trigger the editing function by clicking the icon.

[0053] (3) Underline: For example, a dynamic underline is displayed below the target subtitle.

[0054] (4) Background color block: For example, a semi-transparent color block is added to the background area of the target subtitle to highlight the subtitle content.

[0055] (5) Dynamic effect: For example, a blinking, zooming, or fading animation is added to the target subtitle to attract the user's attention.

[0056] The first preset condition is the condition for triggering the display of the modifiable identifier. The first preset condition includes: (1) Automatic display: When it is detected that the target subtitle may be incorrect (such as through semantic analysis, user feedback, or subtitle confidence score), the modifiable identifier is automatically displayed.

[0057] (2) Manual trigger: When the user performs a specific operation (such as double-clicking on the subtitle area, long-pressing the subtitle, or using a shortcut key), the modifiable identifier is displayed.

[0058] (3) Condition trigger: When the video playback state changes (such as the video pauses, the playback speed decreases, or the user switches the subtitle language), a modifiable identifier is displayed.

[0059] In an optional embodiment, the first preset condition includes video pause; as Figure 3 shown, step S202 may include: Step S300, in response to a pause operation on the target video, identify the position of the target subtitle in the current video frame; Step S302, generate a modifiable identifier for the target subtitle according to the position; Wherein, the modifiable identifier includes a hint box enclosing the target subtitle.

[0060] Specifically, text information can be extracted from the image of the video frame through image recognition technology (such as OCR), and the position of the target subtitle can be determined by analyzing the layout and boundaries of the text. In some embodiments, image segmentation technology can also be used to analyze features such as pixel distribution and color contrast of the image, extract the contour of the subtitle area, and locate the position of the subtitle.

[0061] The hint box can be drawn using border lines (such as solid lines, dashed lines, or dotted lines). In some embodiments, a semi-transparent background color block is added inside the hint box to highlight the subtitle content. In some embodiments, a flashing, zooming, or fading animation can be added to the hint box to attract the user's attention. In some embodiments, if there are multiple subtitles in the current video frame, independent hint boxes can be generated for each subtitle, and the user is supported to modify them separately.

[0062] In this embodiment, by accurately positioning the position of the target subtitle and generating a modifiable identifier, the user can quickly discover the modifiable subtitle (i.e., the target subtitle) when the video pauses, thereby improving the convenience of subtitle modification and the user experience.

[0063] Step S204 , in response to the area where the modifiable identifier is located being triggered, a modification entry for the target subtitle is popped up.

[0064] When the user triggers the area where the modifiable identifier is located (such as clicking or touching), a modification entry for the target subtitle is popped up. The modification entry may include the following functions: (1) Display the original content of the target subtitle; (2) Provide an input box for the user to input the modified content; (3) Support adjusting the display style of the subtitle (such as font, color, size, etc.); (4) Provide a preview function so that the user can view the modified subtitle effect in real time.

[0065] Step S206 , receive the current editing content through the editing entry.

[0066] Combined with Figure 12 , the editing entry can be a pop-up window, a floating toolbar or an independent editing interface. In some embodiments, the editing entry may include the following functional modules to support users in editing the target subtitle: (1) Text editing module: including an input box, a keyboard, a translation assistance tool, etc. Users can modify the text content of the target subtitle through the text editing module.

[0067] (2) Style adjustment module: Users can customize the display style of the subtitle through the style adjustment module, including: Font selection: Provide a variety of font options, and users can select different font styles according to their needs.

[0068] Font size adjustment: Adjust the size of the subtitle to adapt to different video content and display requirements.

[0069] Color selection: Provide a rich color option, and users can customize the font color and background color of the subtitle.

[0070] Alignment method: Support multiple alignment methods such as left alignment, center alignment, and right alignment.

[0071] Special effect addition: Provide special effect options such as shadows, strokes, gradients, etc. to enhance the visual effect of the subtitle.

[0072] (3) Timeline adjustment module: Users can precisely control the display time of the subtitle through the timeline adjustment module, including: Timeline slider: Provide a visual slider, and users can adjust the start time and end time of the subtitle by dragging the slider.

[0073] Time input box: Allow users to directly enter specific time values to adjust the display time of the subtitle.

[0074] Synchronized playback function: When adjusting the timeline, the video can automatically play the corresponding time period to help users more intuitively adjust the subtitle time.

[0075] Automatic alignment function: Provide an automatic alignment function to automatically adjust the subtitle timeline according to speech recognition or audio waveform to ensure subtitle-audio synchronization.

[0076] The editing entry is used to receive the current editing content input by the user. The current editing content may include, but is not limited to: (1)Modification of the target subtitle text; (2)Adjustment of the subtitle display style; (3)Adjustment of the subtitle timeline (such as start time and end time).

[0077] Step S208 , upload the current edited content to the server so that the server replaces the target subtitle with the current edited content.

[0078] When the user confirms that the modification is completed, upload the current edited content to the server. After the server receives the current edited content, it replaces the original target subtitle with the current edited content and updates the subtitle file of the target video. Synchronize the updated subtitle file to all clients so that other users can see the corrected subtitles (the content of the corrected subtitles is the current edited content) when watching the same video. In some embodiments, different modification permissions can be set according to different user roles (such as ordinary users, administrators) to improve the security and controllability of subtitle modification.

[0079] In an alternative embodiment, as Figure 4 shown, step S208 may include: Step S400, obtain the display area and time period occupied by the target subtitle in the target video; Step S402, update the current edited content to the display area according to the display area and the time period to cover the target subtitle; wherein, the area occupied by the current edited content is the same as the area occupied by the target subtitle.

[0080] Determine the display area occupied by the target subtitle in the current video frame through image recognition technology or subtitle file parsing. The display area includes the coordinate range of the target subtitle (such as the upper left corner coordinates and the lower right corner coordinates) and the width and height of the subtitle. In some embodiments, if the target subtitle is a multi-line subtitle, the display area of each line of the subtitle can be obtained separately.

[0081] The time period during which the target subtitle is displayed can be inferred through the timestamp information of the video and the display logic of the subtitle. In some embodiments, the time period during which the target subtitle is displayed can also be parsed from the subtitle file. In some embodiments, if the target subtitle is a segmented subtitle (such as multiple sentences of dialogue), the time period of each segment of the subtitle can be obtained separately.

[0082] During the process of updating the current edited content to the display area, if the text length of the current edited content is inconsistent with the target subtitle, the font size, line spacing, and character spacing of the current edited content can be adjusted to make the area it occupies the same as the area occupied by the target subtitle. In some embodiments, it is also possible to make the area it occupies the same as the area occupied by the target subtitle by automatically wrapping lines or adjusting the layout.

[0083] In some embodiments, the display style of the current edited content can be adjusted according to the display style of the target subtitle (such as font, color, background) to make it coordinated with the video picture.

[0084] Combined with Figure 13 , the current edited content can be embedded in the display area through image processing technology to cover the target subtitle.

[0085] In this embodiment, by covering the display area and the corresponding time period occupied by the target subtitle with the current edited content of the user, an effective integration of the current edited content and the video picture is achieved, thereby significantly improving the accuracy of subtitle modification and the user experience.

[0086] In an alternative embodiment, the time period includes a start time and an end time; as Figure 5 shown, step S400 may include: Step S500, obtaining the current timestamp corresponding to the current video frame; Step S502, performing frame-by-frame recognition forward based on the current timestamp to obtain the start time when the target subtitle appears; and Step S504, performing frame-by-frame recognition backward based on the current timestamp to obtain the end time when the target subtitle disappears; Step S506, determining the time period according to the start time and the end time.

[0087] Specifically, the time period can be expressed as an interval between the start time and the end time (such as 00:01:23 to 00:01:28). The time period includes a start time and an end time. It can be combined with Figure 11 , and by using image recognition technology (such as OCR) or subtitle file parsing, each frame when the target subtitle appears is detected to determine the start time and the end time.

[0088] For determining the start time, perform frame-by-frame recognition of the subtitle content in the video frame forward based on the current video frame. When the target subtitle is first detected, record the timestamp corresponding to this frame as the start time. In some embodiments, if the target subtitle is gradually displayed (such as a fade-in effect), the first frame when the subtitle is fully displayed is used as the start time. Based on the current timestamp, perform frame-by-frame recognition of the subtitle content in the video frame backward.

[0089] For determining the end time, the subtitle content in the video frame is identified frame by frame backward based on the current video frame. When it is detected that the target subtitle completely disappears, the time stamp corresponding to this frame is recorded as the end time. In some embodiments, if the target subtitle fades out gradually (such as a fade - out effect), the last frame when the subtitle completely disappears is used as the end time.

[0090] In some embodiments, a machine - learning model can be used to predict the start time of the subtitle appearance and the end time of the subtitle disappearance, reducing the computational complexity of frame - by - frame recognition and improving the efficiency.

[0091] In this embodiment, by identifying frame by frame forward and backward based on the current time stamp corresponding to the target subtitle, the start time and the end time of the target subtitle are respectively determined, thereby accurately delimiting the time period during which the target subtitle is displayed, and thus improving the accuracy of subsequent subtitle correction. At the same time, this automated processing method reduces the operation complexity of the user and improves the efficiency of subtitle correction.

[0092] In an alternative embodiment, as Figure 6 shown, step S402 may include: Step S600, generating a text layer for covering the display area according to the display area and the time period; Step S602, updating the current editing content to the text layer.

[0093] The display area refers to the spatial range occupied by the target subtitle in the target video, usually represented by a rectangular box. This rectangular box contains the position, size, and boundary information of the target subtitle.

[0094] The time period refers to the time range during which the target subtitle appears in the target video, including the start time and the end time of the subtitle. The time period is used to determine the display duration and appearance timing of the subtitle in the video.

[0095] The text layer is an independent layer in the video player or editing software for rendering and displaying subtitles. This text layer can be a transparent and editable graphic layer. The size and position of the text layer can match the display area of the target subtitle, thereby improving the accuracy of covering the target subtitle. In some embodiments, the transparency, background color, and border style of the text layer can be set to coordinate with the video picture of the target video.

[0096] In some embodiments, during the process of updating the current editing content to the text layer, the display style of the current editing content can be adjusted according to the display style (such as font, color, size) of the target subtitle.

[0097] In this embodiment, by generating a text layer that matches the display area of the target subtitle, the current edited content can accurately cover the original subtitle.

[0098] Embodiment 2 This method embodiment can be executed in the server 2. Hereinafter, the server is taken as the execution subject of this process. It should be noted that the technical details and technical effects in this embodiment can be referred to, introduced, or combined with Embodiment 1.

[0099] Figure 7 Schematically shows a flowchart of a video processing method according to Embodiment 2 of the present application.

[0100] As Figure 7 shown, the video processing method may include steps S700 to S704, where: Step S700, obtaining the display area, time period, and current edited content uploaded by the client; wherein, the target subtitle is located within the display area; Step S702, generating a text layer for covering the display area according to the display area and the time period; Step S704, updating the current edited content to the text layer.

[0101] In some embodiments, the current edited content can be detected for sensitive words through a pre-trained NLP (Natural Language Processing) model, and if there is non-compliant content, the update is prohibited. In some embodiments, the video background complexity can be detected through OCR, and the text position can be automatically adjusted to reduce the occlusion of key pictures. In some embodiments, by analyzing the video background color, the text layer can be dynamically adjusted (such as dark color, light color, or complementary color), and stroke or local blur processing can be added to reduce the interference to the user's viewing.

[0102] In this embodiment, the server generates a text layer covering the display area according to the configuration data (i.e., the display area, time period, and current edited content) uploaded by the client, and finally updates the edited content to the text layer, realizing the effective replacement of the target subtitle (error subtitle) and improving the subtitle editing efficiency.

[0103] In an optional embodiment, as Figure 8 shown, step S704 may include: Step S800, when the text layer is generated, loading the current edited content to the text layer; Step S802, when the text layer is removed, ending the display of the current edited content.

[0104] Specifically, if the target video is played to the start time corresponding to the recorded target subtitle, a text layer is generated according to the recorded size and position (display area). The current edited content is loaded into the text layer. For example, the text layer can be set to have a solid black filled background color, and the current edited content can be set to white. Subsequently, the generated text layer and the edited content are overlaid on the recorded display area and continuously displayed until the end time corresponding to the target subtitle, and then the text layer and the current edited content are removed. In some embodiments, the user can be allowed to remove the text layer, for example, to quickly clear the current subtitle when a correction error is found.

[0105] In this embodiment, by dynamically generating a text layer and loading the current edited content, and synchronously ending the display when removing the layer, the fluency of subtitle correction and the user experience are improved.

[0106] To make the present application easier to understand, the following provides an exemplary application in combination with Figures 9 to 13 an exemplary application is provided.

[0107] Step S11, a video creator publishes a target video embedded with multiple subtitles. The multiple subtitles include incorrect subtitles.

[0108] Step S12, the server receives the target video and forwards it to the client.

[0109] Step S13, a video viewer plays the target video.

[0110] Step S14, in response to pausing the target video, the subtitles and positions in the current video frame are recognized in real time through OCR technology, and the recognized subtitles are marked with a rectangular box. The rectangular box is used to prompt the subtitles for which the video viewer can perform modification operations.

[0111] Step S15A, in response to the video viewer selecting an incorrect subtitle, a keyboard and an input box are popped up to receive the edited content of the video viewer.

[0112] Step S15B, in response to the video viewer selecting an incorrect subtitle, the display area and time period of the incorrect subtitle in the target video are obtained. The time period includes a start time and an end time.

[0113] (1) Regarding obtaining the start time, it can be recognized frame by frame forward from the current video frame corresponding to the incorrect bullet screen. When the incorrect subtitle first appears is detected, the time stamp corresponding to this frame can be recorded as the start time.

[0114] (2) Regarding obtaining the end time, it can be recognized frame by frame backward from the current video frame corresponding to the incorrect bullet screen. When the incorrect subtitle disappears is detected, the time stamp corresponding to this frame can be recorded as the start time.

[0115] Step S16: When the video viewer performs an upload operation, upload the current edited content, display area, and time period to the server.

[0116] Step S17: If the target video plays to the recorded start time, the server generates a text box based on the pre-edited content, display area, and time period, and sends the text box and the current edited content to each client.

[0117] Step S18: Each client covers the display area of the incorrect subtitle with the text box; and updates the current edited content onto the text box.

[0118] Step S19: If the target video plays to the recorded end time, each client removes the text box and ends the display of the current edited content.

[0119] Based on the position, start time, and end time of the incorrect subtitle, as well as the corrected information such as the new subtitle edited by the video viewer, the instant update and sharing of the corrected information are realized. Through this mechanism, the corrected new subtitle can be quickly synchronized to other users, thus significantly improving the real-time performance of the system and the interactivity between users.

[0120] Embodiment III Figure 14 The block diagram of the video processing device according to Embodiment III of the present application is schematically shown. This device is used for the client and can be divided into one or more program modules. One or more program modules are stored in the storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 14 shown, the device 1000 may include: a playback module 1100, a display module 1200, a pop-up module 1300, a receiving module 1400, and an upload module 1500, where: The playback module 1100 is used to play the target video; where the target video includes the current video frame embedded with the target subtitle; The display module 1200 is used to display a modifiable identifier for the target subtitle on the current video frame when a first preset condition is triggered; and The pop-up module 1300 is used to pop up an editing entry for modifying the target subtitle in response to the area where the modifiable identifier is located being triggered; The receiving module 1400 is used to receive the current edited content through the editing entry; and The upload module 1500 is used to upload the current edited content to the server so that the server replaces the target subtitle with the current edited content.

[0121] In an alternative embodiment, the first preset condition includes video pause; the display module 1200 is further configured to: In response to a pause operation on the target video, identify the position of the target subtitle in the current video frame; Generate a modifiable identifier for the target subtitle according to the position; Wherein, the modifiable identifier includes a hint box enclosing the target subtitle.

[0122] In an alternative embodiment, the upload module 1500 is further configured to: Obtain the display area and time period occupied by the target subtitle in the target video; Update the current edited content to the display area according to the display area and the time period to cover the target subtitle; Wherein, the area occupied by the current edited content is the same as the area occupied by the target subtitle.

[0123] In an alternative embodiment, the time period includes a start time and an end time; obtaining the display area and time period occupied by the target subtitle in the target video includes: Obtain the current timestamp corresponding to the current video frame; Identify frame by frame forward based on the current timestamp to obtain the start time when the target subtitle appears; and Identify frame by frame backward based on the current timestamp to obtain the end time when the target subtitle disappears; Determine the time period according to the start time and the end time.

[0124] In an alternative embodiment, updating the current edited content to the display area according to the display area and the time period to cover the target subtitle includes: Generate a text layer for covering the display area according to the display area and the time period; Update the current edited content to the text layer.

[0125] Embodiment IV Figure 15 Schematically shows a block diagram of a video processing apparatus according to Embodiment IV of the present application. The apparatus is for a server and can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. AsFigure 15 As shown, the device 2000 may include: an acquisition module 2100, a generation module 2200, and an update module 2300, where: The acquisition module 2100 is configured to acquire a display area, a time period, and current edited content uploaded by a client; wherein, the target caption is located within the display area; The generation module 2200 is configured to generate a text layer for covering the display area according to the display area and the time period; The update module 2300 is configured to update the current edited content to the text layer.

[0126] In an alternative embodiment, the generation module 2200 is further configured to: When generating the text layer, load the current edited content to the text layer; When removing the text layer, end the display of the current edited content.

[0127] Embodiment 5 Figure 16 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the video processing method according to Embodiment 5 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 16 shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Wherein: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the video processing method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.

[0128] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0129] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0130] It should be noted that Figure 16 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively.

[0131] In this embodiment, the video processing method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0132] Embodiment Six The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video processing method in the embodiments are implemented.

[0133] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the video processing method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0134] Embodiment VII The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.

[0135] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps of them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0136] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A video processing method, characterized in that: For a client, the method comprises: Playing a target video; wherein the target video includes a current video frame embedded with a target subtitle; In the case where the first preset condition is triggered, displaying a modifiable mark for the target subtitle on the current video frame; and In response to the area where the modifiable mark is located being triggered, an editing entry for modifying the target subtitle is popped up; Receiving the current editing content through the editing entry; and The current edited content is uploaded to a server, so that the server replaces the target subtitles with the current edited content.

2. The method according to claim 1, characterized in that The first preset condition includes pausing the video; when the first preset condition is triggered, displaying a modifiable mark for the target subtitle on the current video frame includes: In response to a pause operation on the target video, identifying a position of the target subtitle in the current video frame; generating a modifiable identifier for the target subtitle according to the position; The modifiable mark includes a prompt box containing the target subtitle.

3. The method according to claim 1, characterized in that The server replaces the target subtitle with the current edited content, comprising: Obtaining a display area and time period occupied by the target subtitle in the target video; According to the display area and the time period, updating the current editing content to the display area to cover the target subtitles; The area occupied by the current editing content is the same as the area occupied by the target subtitle.

4. The method according to claim 3, characterized in that: The time period includes a start time and an end time; obtaining the display area and time period occupied by the target subtitle in the target video includes: Get the current timestamp corresponding to the current video frame; Based on the current timestamp, forward frame-by-frame identification is performed to obtain the start time when the target subtitle appears; and Based on the current timestamp, identify backward frame by frame to obtain the end time of disappearance of the target subtitle; The time period is determined according to the start time and the end time.

5. The method according to claim 3, characterized in that: According to the display area and the time period, updating the current editing content to the display area to cover the target subtitles includes: Generate a text layer for covering the display area according to the display area and the time period; Update the current editing content to the text layer.

6. A video processing method, characterized in that: For the server, the method includes: Obtaining the display area, time period and current editing content uploaded by the client; wherein the target subtitle is located in the display area; Generate a text layer for covering the display area according to the display area and the time period; Update the current editing content to the text layer.

7. The method according to claim 6, characterized in that Updating the current editing content to the text layer includes: When the text layer is generated, the current editing content is loaded into the text layer; When the text layer is removed, the display of the current editing content ends.

8. A video processing device, characterized in that: For a client, the device comprises: A playback module, used for playing a target video; wherein the target video includes a current video frame embedded with a target subtitle; A display module, configured to display a modifiable mark for the target subtitle on the current video frame when a first preset condition is triggered; and A pop-up module, configured to pop up an editing entry for modifying the target subtitle in response to the area where the modifiable mark is located being triggered; A receiving module, used for receiving the current editing content through the editing entry; and The uploading module is used to upload the current editing content to the server, so that the server replaces the target subtitles with the current editing content.

9. A video processing device, characterized in that: For a server, the device comprises: An acquisition module, used to acquire the display area, time period and current editing content uploaded by the client; wherein the target subtitle is located in the display area; A generating module, used for generating a text layer for covering the display area according to the display area and the time period; An updating module is used to update the current editing content to the text layer.

10. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-line dynamic subtitle adding method and device, medium and program product

    CN120897094A

  • Video subtitle synchronization method and device, electronic equipment and storage medium

    CN121665051A