Subtitle processing method and device, electronic equipment and readable storage medium

By splitting and special effects processing the video subtitles, the problem of OpenGL rendering complexity is solved and the efficient highlighting of subtitles is achieved.

CN120730006APending Publication Date: 2025-09-30BEIJING MEISHE NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510868476.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing method of using OpenGL to render text in real time is relatively complex and technically difficult.

Method used

The subtitles to be processed in the video are split to obtain multiple words and their time range in the video timeline. The target words are determined according to the playback time and special effects are performed during display.

Benefits of technology

It realizes the subtitle highlighting of the word spoken in the speech without the need for complex interaction or data conversion, which reduces the difficulty and improves the efficiency and effect of subtitle processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120730006A_ABST
    Figure CN120730006A_ABST
Patent Text Reader

Abstract

The invention provides a subtitle processing method and device, electronic equipment and a readable storage medium, and the method comprises the steps: splitting a to-be-processed subtitle in a video, obtaining a plurality of words, obtaining a time range corresponding to each word in a time axis of the video, obtaining a current playing moment of the video, and storing the current playing moment in the video; and if the playing moment is within the time range corresponding to the word, determining the word as a target word, and when the target word is displayed, performing display special effect adding processing on the target word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to a subtitle processing method, device, electronic device and readable storage medium. Background Art

[0002] In some video post-processing software, voice is often recorded in real time during video playback, and the voice is directly converted into text and added to the preview video. When the voice speaks a word, the word is highlighted.

[0003] In the related art, OpenGL is used to render each character in real time according to the characters and the corresponding time information. When the current character is rendered, a special color processing is performed on the current character and the character is displayed.

[0004] However, the method of using OpenGL to render text in real time is relatively complicated and technically difficult. Summary of the Invention

[0005] The present application aims to provide a subtitle processing method, device, electronic device and readable storage medium, which at least solves the problem that the method of using OpenGL to render text in real time in the prior art is relatively complex and technically difficult.

[0006] In a first aspect, an embodiment of the present application discloses a subtitle processing method, the method comprising:

[0007] Splitting the subtitles to be processed in the video to obtain a plurality of words and a time range corresponding to each word in the time axis of the video; the endpoints of the time range include the start time and the end time of the playback of the word;

[0008] Obtaining the current playback time of the video, and if the playback time is within the time range corresponding to the word, determining the word as a target word;

[0009] When displaying the target word, a display special effect processing is added to the target word.

[0010] In a second aspect, an embodiment of the present application discloses a subtitle processing device, the device comprising:

[0011] A splitting module is used to split the subtitles to be processed in the video to obtain multiple words and the time range corresponding to each word in the time axis of the video; the endpoints of the time range include the start time and the end time of the playback of the word;

[0012] a determination module, configured to obtain a current playback time of the video, and if the playback time is within a time range corresponding to the word, determine the word as a target word;

[0013] The processing module is used to add display special effects to the target words when displaying the target words.

[0014] In a third aspect, an embodiment of the present application further discloses an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] In summary, in an embodiment of the present application, the subtitles to be processed in the video are split to obtain multiple words and the time range corresponding to each word in the timeline of the video. If the current playback moment of the video is in the time range corresponding to the word, the word is determined as the target word. When displaying the target word, the target word is added with display special effects. The entire process does not require complex interactions or data conversion, and the various parts work closely together, which greatly reduces the difficulty of highlighting the word by adding display special effects when the voice speaks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In the attached figure:

[0018] Figure 1 This is a flowchart of the steps of a subtitle processing method provided by an embodiment of the present application;

[0019] Figure 2 This is a flowchart of another subtitle processing method provided by an embodiment of the present application;

[0020] Figure 3 This is a block diagram for determining a background area provided by an embodiment of the present application;

[0021] Figure 4 This is a flowchart of subtitle effect settings for video editing software provided by an embodiment of the present application;

[0022] Figure 5 This is a block diagram of a subtitle processing device provided by an embodiment of the present application;

[0023] Figure 6 is a block diagram of an electronic device according to an embodiment of the present application;

[0024] Figure 7 This is a block diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0027] Figure 1 This is a flowchart of a subtitle processing method provided by an embodiment of the present application, see Figure 1 , the method may include the following steps:

[0028] Step 101: Split the subtitles to be processed in the video to obtain multiple words and the time range corresponding to each word in the timeline of the video; the endpoints of the time range include the start and end playback times of the words.

[0029] For example, the subtitles to be processed are presented as text in a video, helping viewers understand the video content. These texts may be character dialogues, narration, or explanatory text. Subtitle splitting is the process of breaking the complete subtitle text into individual words according to specific rules. This facilitates more detailed analysis and processing of the subtitles.

[0030] For example, the time range corresponding to a word in the video timeline can be understood as the time period during which each word is highlighted during video playback. The start time is the time point when a word begins to be highlighted in the video. The end time is the time point when a word stops being highlighted in the video.

[0031] For example, by splitting subtitles and determining the time range of each word, the text content in the video can be precisely managed. This provides a foundation for subsequent operations such as adding special effects and interactive design for specific words, greatly enriching the expressiveness of the video.

[0032] Step 102: Obtain the current playing time of the video. If the playing time is within the time range corresponding to the word, determine the word as the target word.

[0033] For example, the current playback moment of a video refers to the playback progress of the video player at a certain moment, usually expressed in the form of a timestamp. The time range corresponding to a word is the time interval in which each word is highlighted in the subtitles. For example, a word appears at the 10th second of the video and disappears at the 15th second, and its corresponding time range is from 10 seconds to 15 seconds. The target word is usually the content that needs to be highlighted and paid attention to in the current picture or plot. Specifically, when the current playback moment of the video is within the time range corresponding to a certain word, the word is identified as the target word.

[0034] For example, by monitoring the playback progress of the CNvStreamingContext object context, the current playback time of the video is obtained, and the target word is determined according to the playback time and the time range corresponding to the word.

[0035] For example, after identifying the target word, the system can capture the subtitle words that match the current time during video playback in real time, thereby accurately locating specific content. This positioning provides a basis for subsequent operations such as highlighting the target word and adding special effects, helping viewers quickly grasp important information in the video.

[0036] Step 103: When displaying the target word, add display special effects to the target word.

[0037] For example, display special effects processing uses graphics rendering, style adjustment and other technologies to change the visual features of the target word, making it different from other ordinary subtitles and thus highlighting it.

[0038] For example, adding display special effects processing to the target words means that when the video is played to a specific period, the target words in the subtitles are displayed by changing their visual presentation to enhance their prominence in the picture in order to attract the audience's attention.

[0039] For example, display effects may include: color changes, font style adjustments, etc. For example, by changing the color of the target word, it can be made more conspicuous among other subtitles, and through bright and contrasting colors, the audience can be guided to focus on the target content. In educational videos, when explaining key knowledge points, the corresponding target word color is set to red, while other subtitles remain white, so that the audience can quickly focus on the key content in red. Font style adjustment changes the appearance of the text and highlights the target word by bolding, italicizing or underlining the font of the target word. In movie subtitles, in order to emphasize the important lines of the character, the relevant words are bolded, and the audience can perceive the importance of the lines without paying special attention.

[0040] In an embodiment of the present application, the subtitles to be processed in the video are split to obtain multiple words and the time range corresponding to each word in the timeline of the video. If the current playback moment of the video is in the time range corresponding to the word, the word is determined as the target word. When the target word is displayed, the target word is added with display effects. The entire process does not require complex interactions or data conversions, and the various parts work closely together, which greatly reduces the difficulty of highlighting a word by adding display effects when the voice speaks.

[0041] Figure 2 This is a flowchart of another subtitle processing method provided by this application, see Figure 2 , the method may include the following steps:

[0042] Step 201: Split the subtitles to be processed in the video to obtain multiple words and the time range corresponding to each word in the timeline of the video; the endpoints of the time range include the start and end times of playing the words.

[0043] This step may be specifically referred to the above step 101 and will not be described in detail here.

[0044] Optionally, every two adjacent words in the plurality of words are separated by a preset symbol, and step 201 may specifically include:

[0045] Sub-step 2011: constructing a regular expression according to the preset symbol, and performing regular matching on the subtitles to be processed using the regular expression to obtain a plurality of words;

[0046] Sub-step 2012: obtaining the total display duration of the subtitle to be processed in the timeline of the video, and determining the display duration of each word based on the total display duration;

[0047] Sub-step 2013: Determine the time range corresponding to the word in the time axis of the video based on the start time of the subtitle to be processed in the time axis of the video and the display duration of the word.

[0048] For substeps 2011-2013, the preset symbols are pre-set, specific symbols used to separate adjacent words in the subtitles. Common preset symbols include punctuation marks such as spaces, commas, and periods, which can be used to split subtitle text. Regular expressions are a powerful text matching tool that, by defining specific patterns, enables operations such as searching, replacing, and splitting text. During subtitle processing, regular expressions are used to match and split subtitle text based on the preset symbols. The total display duration is the duration from the start to the end of the subtitle's display on the video timeline. The word display duration is the length of time each word is highlighted in the video, which is allocated according to a specific rule from the total display duration. The start time is the time point at which the subtitle begins to display on the video timeline. The time range is the period between the start and end of each word's display on the video timeline, determined by the start time and display duration.

[0049] Take the following paragraph as an example. It is divided into multiple sentences according to punctuation marks:

[0050] How do you do!

[0051] My name is jack.

[0052] What's your name?

[0053] Next, using the example of "How do you do!", the software development kit (SDK) can be used to split "How do you do!" into four words, "How," "do," "you," and "do!", using spaces. If "How do you do!" is added to the timeline as a single subtitle, the entire sentence appears together. If you want each word in this sentence to appear with a different background color than the other words, you need to split these four words into four subtitle objects and add them to the timeline according to the time range corresponding to each word in the video's timeline. The addition process is as follows:

[0054] NvsTimelineCaptioncaption1=timeline.addCaption(0,500000,"How"),

[0055] NvsTimelineCaptioncaption2=timeline.addCaption(500000,1000000,“do”),NvsTimelineCaptioncaption3=timeline.addCaption(1000000,1500000,“you”),

[0056] NvsTimelineCaptioncaption4=timeline.addCaption(1500000,2000000,“do!”),

[0057] For example, let's take the words "How do you do!". Each pair of adjacent words is separated by a space. We can construct the regular expression \s. In regular expressions, \s represents a whitespace character, including spaces, tabs, and newlines. Using this regular expression, we can perform a regular match on the subtitle text "How do you do!"

[0058] For example, let's take the subtitle "How do you do!", which has a total display duration of 8 seconds in the video timeline. To determine the display duration of each word, we can use an even distribution strategy. Since the subtitle is divided into four words, the display duration of each word is 8 ÷ 4 = 2 seconds. Of course, in practice, the display duration of each word can be flexibly adjusted based on factors such as word importance and speech rate. For example, if the subtitle begins playing at the 10th second of the video, based on a 2-second display duration for each word, we can determine the time range for each word: "How": 10th to 12th seconds, "Do": 12th to 14th seconds, "You": 14th to 16th seconds, and "Do!": 16th to 18th seconds. Using this method, when the video reaches the corresponding time, special effects such as color change and amplification can be applied to specific words. For example, changing the color of "You" at the 14th second can attract the viewer's attention and enhance the presentation effect.

[0059] Step 202: If the playback time is greater than or equal to the start playback time of the word and less than or equal to the end playback time of the word, then the word is used as the target word.

[0060] This step may be specifically referred to the above step 102 and will not be described in detail here.

[0061] Step 203: When displaying the target word, add display special effects to the target word.

[0062] This step may be specifically referred to the above step 103 and will not be described in detail here.

[0063] Optionally, step 203 may specifically include:

[0064] Sub-step 2031: determining a first background region where the target word is located, and second background regions corresponding to words other than the target word in a plurality of words;

[0065] Sub-step 2032: Fill the first background area with a target color, and fill the second background area with a color other than the target color.

[0066] With respect to sub-steps 2031 and 2032, the background area of ​​the target word refers to a rectangular or other shaped area used to carry the word in visual presentation. For simplicity, it is assumed that a rectangular area is created with each word as the center, which just accommodates the word as the background area. In the screen coordinate system, this area is determined by the coordinates of its upper left corner and lower right corner. For each word in the multiple words other than the target word, a rectangular area is also created around it to just accommodate the word as the background area.

[0067] For example, a target color, such as red, is selected to fill the first background area where the target word "do" is located. This makes the target word more visually prominent. A color other than the target color, such as white, is selected to fill the second background area corresponding to the non-target words "How," "you," and "do!"

[0068] The black box represents the background area, see Figure 3 , Figure 3 The word "do" in the video has a black frame, indicating its background area. To fill the background area with the desired color, call the setBackgroundColor method in the NvsTimelineCaption class. For example, call caption2.setBackgroundColor(inputColor) to set the background color for "do." All other words are given a different, uniform color. Repeat this process for each sentence. This way, throughout the video, only the currently mentioned word will be highlighted.

[0069] Optionally, step 203 may specifically include:

[0070] Sub-step 2033: determining the outline of the target word;

[0071] Sub-step 2034: Stroke the outline of the target word.

[0072] For sub-steps 2033-2034, determining the outline of the target word refers to defining the spatial boundaries occupied by the target word in the display area. On the screen or other display media, each word has a corresponding coordinate position. Usually, a two-dimensional coordinate system is established with the upper left corner as the origin (0,0), and the position of the word is determined by the horizontal and vertical coordinates. Different fonts have different properties such as character height and width. It is necessary to obtain the measurement information of the font used for the target word, such as the ascending height, descending height, spacing, etc., in order to accurately calculate the space occupied by the word. Based on the coordinates of the word and the font measurement information, calculate the boundary range that can completely contain the word. Usually, the minimum rectangle that surrounds the target word (it may also be a more complex polygon or other shape) is used to represent its outline.

[0073] First, select a stroke color, such as black. Also, determine the stroke width, which determines the thickness of the outline. In practice, choose an appropriate color and width based on visual effects and design requirements. Stroke the outline of the target word. This will highlight the target word and make it easier for users to identify it.

[0074] For example, the setOutlineColor method in the NvsTimelineCaption class can be called to stroke the target word, that is, the caption2.setOutlineColor(Color inputColor) method can be used to stroke the target word.

[0075] Optionally, step 203 may specifically include:

[0076] Sub-step 2035: determining the first position of the target word and the second positions of the words other than the target word in the plurality of words;

[0077] Sub-step 2036: setting the font color of the word at the first position to a target color, and setting the font color of the word at the second position to a color other than the target color.

[0078] For sub-steps 2035 and 2036, in the two-dimensional coordinate system of the screen, each word has its specific drawing position, usually represented by the coordinates of the upper left corner of the word. When drawing subtitles, we will record the starting drawing position of each word. Taking the drawing order from left to right and from top to bottom as an example, when drawing the first word "How", set its upper left corner coordinates to (x1, y1). The position of each subsequent word will be determined based on the width of the previous word and the pre-set spacing. For example, "do" follows "How", and its upper left corner coordinates are (x1+textlength("How",font), y1). And so on, the position of each word, including the target word and other words, can be determined.

[0079] To highlight target words, select a target color, such as red, and select other colors, such as black, for non-target words. You can set the font color for target words using the caption2.setOutlineColor(Color inputColor) method.

[0080] Optionally, step 203 may specifically include:

[0081] Sub-step 2037: determining the third position of the target word and the fourth positions of the words other than the target word in the plurality of words;

[0082] Sub-step 2038: setting the font of the word at the third position to the target font, and setting the font of the word at the fourth position to a font other than the target font.

[0083] For substeps 2037 and 2038, select a target font, such as Kaiti, to highlight the target words. Also select other fonts, such as Songti, for non-target words. Set the font based on the word's location. You can set the font for the target word using the caption2.setFontFamily(String family) method.

[0084] Optionally, the method further includes:

[0085] Step 204: Add a video file to the video track in the timeline area, and add an audio file corresponding to the video file to the audio track in the timeline area;

[0086] Step 205: Convert the audio in the audio file into the subtitles to be processed, and add the subtitles to be processed to the timeline area;

[0087] Step 206: Call the connection interface to bind the timeline area to the preview control, so as to preview the video content in the video file and the subtitles to be processed in the display area in response to the triggering of the preview control.

[0088] For steps 204-206, the SDK is first initialized and the CNvStreamingContext object context is obtained, i.e., the streaming media context object. Create a timeline object timeline by calling createTimeline through context. The timeline object is a data structure that describes a timeline. This timeline refers to a timeline, which includes a video track and an audio track. The video track is used to add local video files, and the audio track is used to add local audio files. The timeline can be understood as the content in the player. Establishing the timeline is to build up the content to be played by the player, so as to facilitate subsequent operations such as adding subtitles. If you want to preview the multimedia material selected by the user in real time on the client program, you need a control for display. The video editing tool provides a Livewindow control that can be directly used for preview. By using the connectTimelineWithLivewindow interface of the streaming media object context, timeline and livewindow are passed in as parameters, and the timeline and Livewindow controls can be connected. At this time, you can preview the video content on the client program.

[0089] For example, SDK tool users can first initialize the tool, that is, call the Init interface to obtain the streaming media context object context, create a timeline, and call context.createTimeline to obtain the timeline; create a preview control, call CNvLiveWindowWidget to instantiate the livewindow, preview the connection, call context.connectTimelineWithLivewindow(timeline, livewindow), create a video track, call timeline.appendVideoTrack to obtain videoTrack, add a video or image, and call videoTrack.appendVideoClip(video or image path).

[0090] For example, SDK tool users can initialize the interface by calling the initContext(Context context) interface and passing in the current instance object. Then, initialize the preview interface by calling the initTimeline(List videoList, List audioList, int width, int height) interface, passing in the video and audio files to be previewed and the width and height of the created timeline, for example, 1920*1080.

[0091] For example, call addCaptions(NvsTimeline timeline, Json jsonContent) to pass in the created timeline and text content, add the text interface, call setHightCaptionColor(int timestamp, Color color), pass in the timestamp and highlighted text color, set the highlighted text background color interface, call setCommonCaptionColor(int timestamp, Color color), pass in the timestamp and common text color, set the common text background color interface, call compileTimeline(NvsTimeline timeline, String path), pass in the timeline object to be exported and the export video path, and export the timeline interface. The overall preview effect of the current timeline is exported as a final video, and the editing function is now complete.

[0092] In an embodiment of the present application, the subtitles to be processed in the video are split to obtain multiple words and the time range corresponding to each word in the timeline of the video. If the current playback moment of the video is in the time range corresponding to the word, the word is determined as the target word. When the target word is displayed, the target word is added with display effects. The entire process does not require complex interactions or data conversions, and the various parts work closely together, which greatly reduces the difficulty of highlighting a word by adding display effects when the voice speaks.

[0093] See also Figure 4 , which shows a flowchart of subtitle effect settings for a video editing software provided by an embodiment of the present application, showing the main steps from creating a video project to exporting a video:

[0094] Initialize the video editing software: Use the SDK to create a timeline, then add video or image materials, and then add audio to build the basic framework of the video.

[0095] Speech-to-text: Use the speech-to-text function to obtain the text information corresponding to the speech in the video and related time information to prepare for adding subtitles.

[0096] Monitor playback progress: Get the playback time of the current timeline in real time to accurately locate the timing for subtitle display.

[0097] Set subtitle effect:

[0098] Determine whether to highlight: distinguish between currently displayed subtitles and non-currently displayed subtitles, and highlight the current subtitles.

[0099] More property settings: Further set the subtitle's background color, text color, stroke, font and other properties to enrich the subtitle's visual effects.

[0100] Export video: After completing the above operations, use the SDK to export the video with subtitle effects.

[0101] See also Figure 5 , which shows a subtitle processing device 30 provided in an embodiment of the present application, the subtitle processing device 30 includes:

[0102] The splitting module 301 is used to split the subtitles to be processed in the video to obtain multiple words and the time range corresponding to each word in the time axis of the video; the endpoints of the time range include the start time and the end time of the playback of the word;

[0103] A determination module 302 is configured to obtain a current playback time of the video, and if the playback time is within a time range corresponding to the word, determine the word as a target word;

[0104] The processing module 303 is configured to add display effects to the target words when displaying the target words.

[0105] Optionally, the determining module includes:

[0106] The first determining submodule is configured to use the word as the target word if the playback time is greater than or equal to the start playback time of the word and less than or equal to the end playback time of the word.

[0107] Optionally, the processing module includes:

[0108] A second determining submodule is configured to determine a first background region where the target word is located, and second background regions corresponding to words other than the target word in a plurality of words;

[0109] The filling submodule is configured to fill the first background area with a target color and fill the second background area with a color other than the target color.

[0110] Optionally, the processing module includes:

[0111] A third determination submodule is used to determine the outline of the target word;

[0112] The stroke submodule is used to stroke the outline of the target word.

[0113] Optionally, the processing module includes:

[0114] a fourth determining submodule, configured to determine a first position where the target word is located, and second positions where each of the words other than the target word is located in a plurality of words;

[0115] The first setting submodule is configured to set the font color of the word at the first position to a target color, and set the font color of the word at the second position to a color other than the target color.

[0116] Optionally, the processing module includes:

[0117] a fifth determining submodule, configured to determine a third position where the target word is located, and fourth positions where each of the words other than the target word is located among the plurality of words;

[0118] The second setting submodule is configured to set the font of the word at the third position to a target font, and set the font of the word at the fourth position to a font other than the target font.

[0119] Optionally, every two adjacent words in the plurality of words are separated by a preset symbol; and the splitting module includes:

[0120] A matching submodule, configured to construct a regular expression according to the preset symbol, and perform regular matching on the subtitles to be processed using the regular expression to obtain a plurality of words;

[0121] a sixth determining submodule, configured to obtain a total display duration of the subtitles to be processed in the timeline of the video, and determine a display duration of each word based on the total display duration;

[0122] The seventh determining submodule is configured to determine a time range corresponding to a word in the time axis of the video according to a start time of playing the subtitle to be processed and a display duration of the word in the time axis of the video.

[0123] Optionally, the device further includes:

[0124] A first adding module is used to add a video file to the video track in the timeline area, and to add an audio file corresponding to the video file to the audio track in the timeline area;

[0125] A second adding module is used to convert the audio in the audio file into the subtitles to be processed, and add the subtitles to be processed to the timeline area;

[0126] The preview module is used to call the connection interface to bind the timeline area with the preview control, so as to preview the video content in the video file and the subtitles to be processed in the display area in response to the triggering of the preview control.

[0127] In an embodiment of the present application, the subtitles to be processed in the video are split to obtain multiple words and the time range corresponding to each word in the timeline of the video. If the current playback moment of the video is in the time range corresponding to the word, the word is determined as the target word. When the target word is displayed, the target word is added with display effects. The entire process does not require complex interactions or data conversions, and the various parts work closely together, which greatly reduces the difficulty of highlighting a word by adding display effects when the voice speaks.

[0128] See also Figure 6 , electronic device 400 may include one or more of the following components: a processing component 402 , a memory 404 , a power component 406 , a multimedia component 408 , an audio component 410 , an input / output (I / O) interface 412 , a sensor component 414 , and a communication component 416 .

[0129] The processing component 402 generally controls the overall operation of the electronic device 400, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 402 may include one or more modules to facilitate interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate interaction between the multimedia component 408 and the processing component 402.

[0130] The memory 404 is used to store various types of data to support operations on the electronic device 400. Examples of such data include instructions for any application or method operating on the electronic device 400, contact data, phone book data, messages, pictures, multimedia, etc. The memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0131] The power supply assembly 406 provides power to the various components of the electronic device 400. The power supply assembly 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 400.

[0132] The multimedia component 408 includes an interface that provides an output interface between the electronic device 400 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of touch or slide actions, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0133] The audio component 410 is used to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio component 410 also includes a speaker for outputting audio signals.

[0134] The input / output I / O interface 412 provides an interface between the processing component 402 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0135] The sensor assembly 414 includes one or more sensors for providing various aspects of status assessment for the electronic device 400. For example, the sensor assembly 414 can detect the open / closed state of the electronic device 400, the relative positioning of components, such as the display and keypad of the electronic device 400. The sensor assembly 414 can also detect changes in the position of the electronic device 400 or a component of the electronic device 400, the presence or absence of user contact with the electronic device 400, the orientation or acceleration / deceleration of the electronic device 400, and temperature changes of the electronic device 400. The sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 414 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0136] The communication component 416 is used to facilitate wired or wireless communication between the electronic device 400 and other devices. The electronic device 400 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0137] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement a subtitle processing method provided in an embodiment of the present application.

[0138] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of an electronic device 400 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0139] Figure 7FIG is a block diagram of an electronic device 500 according to another embodiment of the present invention. For example, the electronic device 500 may be provided as a server. Figure 7 Electronic device 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by memory 532 for storing instructions executable by processing component 522, such as applications. The applications stored in memory 532 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 522 is configured to execute instructions to perform a subtitle processing method provided in an embodiment of the present application.

[0140] The electronic device 500 may further include a power supply component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 may operate based on an operating system stored in the memory 532, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0141] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0142] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A subtitle processing method, characterized in that: The method comprises: Splitting the subtitles to be processed in the video to obtain a plurality of words and a time range corresponding to each word in the time axis of the video; the endpoints of the time range include the start time and the end time of the playback of the word; Obtaining the current playback time of the video, and if the playback time is within the time range corresponding to the word, determining the word as a target word; When displaying the target word, a display special effect processing is added to the target word.

2. The method according to claim 1, characterized in that If the playback moment is within the time range corresponding to the word, determining the word as a target word includes: If the playback time is greater than or equal to the start playback time of the word and less than or equal to the end playback time of the word, the word is used as the target word.

3. The method according to claim 1, characterized in that The adding display special effects processing to the target word includes: Determining a first background area where the target word is located, and second background areas corresponding to words other than the target word in a plurality of words; The first background area is filled with a target color, and the second background area is filled with a color other than the target color.

4. The method according to claim 1, wherein The adding display special effects processing to the target word includes: determining an outline of the target word; Performing a stroke process on the outline of the target word.

5. The method according to claim 1, characterized in that The adding display special effects processing to the target word includes: Determining a first position where the target word is located, and second positions where each of the words other than the target word is located; The font color of the word at the first position is set to a target color, and the font color of the word at the second position is set to a color other than the target color.

6. The method according to claim 1, wherein The adding display special effects processing to the target word includes: determining a third position where the target word is located, and fourth positions where each of the words other than the target word is located; The font of the word at the third position is set as a target font, and the font of the word at the fourth position is set to a font other than the target font.

7. The method according to claim 1, characterized in that Every two adjacent words in the plurality of words are separated by a preset symbol; The subtitles to be processed in the video are split to obtain multiple words and the time range corresponding to each word in the time axis of the video, including: Constructing a regular expression according to the preset symbol, and performing regular matching on the subtitles to be processed using the regular expression to obtain a plurality of words; Obtaining a total display duration of the subtitle to be processed in the timeline of the video, and determining a display duration of each word according to the total display duration; The time range corresponding to the word in the time axis of the video is determined according to the start playing time of the subtitle to be processed in the time axis of the video and the display duration of the word.

8. The method according to claim 1, characterized in that The method further comprises: Adding a video file to the video track in the timeline area, and adding an audio file corresponding to the video file to the audio track in the timeline area; Converting the audio in the audio file into the subtitles to be processed, and adding the subtitles to be processed to the timeline area; The connection interface is called to bind the timeline area with the preview control, so as to preview the video content in the video file and the subtitles to be processed in the display area in response to triggering of the preview control.

9. A subtitle processing device, characterized in that: The device comprises: A splitting module is used to split the subtitles to be processed in the video to obtain multiple words and the time range corresponding to each word in the time axis of the video; the endpoints of the time range include the start time and the end time of the playback of the word; a determination module, configured to obtain a current playback time of the video, and if the playback time is within a time range corresponding to the word, determine the word as a target word; The processing module is used to add display special effects to the target words when displaying the target words.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Subtitle generation method, server, terminal device and system

    CN111835988A

  • Subtitle editing method and device, electronic equipment and storage medium

    CN114501159A

  • Subtitle processing method and device for video conference, electronic equipment and storage medium

    CN116156098A

  • Subtitle processing method and device

    CN117749965A

  • Video subtitle file generation method and device, video generation method and device and electronic equipment

    CN118803381A