Subtitle processing methods and devices

By performing speech recognition and timestamp matching on the audio of multimedia materials, subtitles with animated effects that jump out one character at a time are synthesized, solving the problem of low subtitle editing efficiency and realizing the automatic generation of dynamic subtitles and a user-friendly operating experience.

CN117749965BActive Publication Date: 2026-05-26BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2022-09-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies are inefficient for editing subtitles in batch text scenarios, requiring users to repeatedly adjust and preview to achieve the desired subtitle effect.

Method used

By performing speech recognition on the audio of multimedia materials, the subtitle text and its timestamp information are obtained, and the material segments are matched according to the timestamp information to synthesize subtitles with a word-by-word animated effect.

Benefits of technology

It enables automatic generation of dynamic subtitles, improves user experience, simplifies operation, and is suitable for various devices, especially mobile devices with smaller screens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117749965B_ABST
    Figure CN117749965B_ABST
Patent Text Reader

Abstract

This disclosure relates to a subtitle processing method and apparatus. The method includes: during the editing of a multimedia material segment, obtaining subtitle text corresponding to the audio and timestamp information of audio segments corresponding to each text element in the subtitle text through speech recognition; determining the material segment matching the text element in the multimedia material segment based on the timestamp information of the audio segment corresponding to each text element; and then synthesizing each text element with the matching material segment within the specified time to obtain a target multimedia material with a subtitle text appearing word by word in an animation effect. The solution of this disclosure can achieve a subtitle animation effect where the corresponding text subtitle appears when a certain word is spoken; furthermore, user input commands can automatically generate dynamic subtitles, simplifying user operation and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for subtitle processing. Background Technology

[0002] Subtitles in videos can help understand the video content, so they are often added when editing videos.

[0003] Currently, subtitle text is typically obtained by manually inputting it or by using subtitle recognition tools to identify the corresponding audio. Then, the subtitle text is repeatedly listened to and adjusted to segment it into numerous text fragments. These fragments are then combined with the video to add subtitles. For batch text scenarios like subtitles, if users want to achieve a specific subtitle effect, they need to repeatedly adjust, combine, and preview the segmented subtitle text. This method of subtitle editing is very inefficient. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a subtitle processing method and apparatus.

[0005] In a first aspect, embodiments of this disclosure provide a subtitle processing method, including:

[0006] During the editing process of multimedia materials, speech recognition is performed on the audio corresponding to the multimedia materials to obtain the subtitle text corresponding to the audio and the timestamp information of each text element of the subtitle text corresponding to the audio segment.

[0007] Based on the timestamp information of the audio segments corresponding to each text element, the multimedia material is matched with each material unit to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment matching the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline;

[0008] Each of the text elements is combined with a material fragment within a matching time range to obtain a target multimedia material with a subtitle text jumping out word by word animation effect.

[0009] Secondly, embodiments of this disclosure provide a subtitle processing apparatus, including:

[0010] The speech recognition module is used to perform speech recognition on the audio corresponding to the multimedia material during the editing process to obtain the subtitle text corresponding to the audio and the timestamp information of each text element of the subtitle text corresponding to the audio segment.

[0011] The matching module is used to match the timestamp information of the audio segments corresponding to each text element with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment that matches the text element on the editing timeline is the same as the time of the audio segment corresponding to the text element on the editing timeline;

[0012] The subtitle synthesis module is used to synthesize each of the text elements with a material segment within a matching time range to obtain a target multimedia material with a subtitle text jumping out word by word animation effect.

[0013] Thirdly, embodiments of this disclosure provide an electronic device, including: a memory and a processor; the memory is configured to store computer program instructions; the processor is configured to execute the computer program instructions, causing the electronic device to implement the subtitle processing method described in the first aspect.

[0014] Fourthly, embodiments of this disclosure provide a readable storage medium, including: computer program instructions, wherein at least one processor of an electronic device executes the computer program instructions, causing the electronic device to implement the subtitle processing method as described in the first aspect.

[0015] Fifthly, embodiments of this disclosure provide a computer program product, which is executed by an electronic device to enable the electronic device to implement the subtitle processing method described in the first aspect.

[0016] This disclosure provides a subtitle processing method and apparatus. The method includes: during the editing of a multimedia material segment, obtaining subtitle text corresponding to the audio and timestamp information of audio segments corresponding to each text element in the subtitle text through speech recognition; determining the material segment matching the text element in the multimedia material segment based on the timestamp information of the audio segments corresponding to each text element; and then synthesizing each text element with the matching material segment within the specified time to obtain a target multimedia material with a subtitle text appearing word by word in an animation effect. In this disclosure, the time of the material segment matching the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline, enabling the subtitle animation effect where the corresponding text subtitle appears when a certain word is spoken. Furthermore, user input commands can automatically generate dynamic subtitles, simplifying user operation and improving user experience. The method of this disclosure is applicable to various types of devices and has a wide range of applications. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a subtitle processing method provided in an embodiment of this disclosure;

[0020] Figure 2 A flowchart illustrating a subtitle processing method provided in another embodiment of this disclosure;

[0021] Figure 3 A flowchart illustrating a subtitle processing method provided in another embodiment of this disclosure;

[0022] Figures 4A to 4I This is a schematic diagram of the human-computer interaction interface provided in this disclosure;

[0023] Figure 5 This is a schematic diagram of the structure of a subtitle processing device provided in an embodiment of the present disclosure. Detailed Implementation

[0024] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0025] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0026] Subtitles help users understand video content, and different subtitle effects can represent more dimensions of content. For example, the corresponding text may only appear when the audio says a certain word. This type of subtitle effect is often used to represent voice-over in dramatic narratives, and to express the speaker's confident and enthusiastic emotions in talking videos. Achieving these specific subtitle effects often involves the user manually inputting subtitles, breaking the text down into individual characters, and then repeatedly listening to the audio for adjustments. Alternatively, a complete sentence can be input by the user, and keyframe masks can be used to achieve the effect of text appearing one character at a time. However, this method is not only inefficient for subtitle editing but also extremely inconvenient on mobile devices.

[0027] Based on this, embodiments of this disclosure provide a subtitle processing method and apparatus. The method includes: during the editing of multimedia materials, obtaining subtitle text corresponding to the audio and timestamp information of audio segments corresponding to each text element in the subtitle text through speech recognition; determining the material segment matching the text element in the multimedia material segment based on the timestamp information of the audio segments corresponding to each text element; and then synthesizing each text element with the matching material segment within the time frame to obtain a target multimedia material with a subtitle text appearing word by word in an animation effect. In this disclosure, the start time of the time range of the video frame image matching the text element is consistent with the start time of the audio segment corresponding to the text element, enabling the subtitle animation effect of the corresponding text subtitle appearing when a certain word is spoken. Furthermore, user input commands can automatically generate dynamic subtitles, simplifying user operation and improving user experience. The method of this disclosure is applicable to various types of devices and has a wide range of applications.

[0028] The method provided in this disclosure can be executed by an electronic device, which may be, but is not limited to, a tablet computer, a mobile phone (such as a foldable screen phone, a large screen phone, etc.), a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. This disclosure does not impose any restrictions on the specific type of electronic device.

[0029] Figure 1 This is a flowchart illustrating a subtitle processing method provided in one embodiment of this disclosure. Taking an electronic device executing the method of this embodiment as an example, the electronic device has an editing application installed, allowing the user to edit multimedia materials. Please refer to... Figure 1 As shown, the method in this embodiment includes:

[0030] S101. During the editing process of multimedia materials, speech recognition is performed on the audio corresponding to the multimedia materials to obtain the subtitle text corresponding to the audio and the timestamp information of each text element in the subtitle text corresponding to the audio segment.

[0031] Multimedia materials can be video footage recorded by the user in real time, previously edited video footage, or video footage stored on an electronic device. This disclosure does not limit these categories. They can also be audio footage, image footage, etc. This disclosure does not limit the type of multimedia materials. Furthermore, this disclosure does not limit the number of multimedia materials. If there are multiple multimedia materials, they can be arranged according to the import order and can be considered as a single unit.

[0032] Editing multimedia materials can be understood as pre-recording or importing multimedia materials or audio materials with audio, or adding background music to multimedia materials (such as video materials or image materials). Of course, the editing methods are not limited to these.

[0033] The subtitle text is obtained through text recognition of the audio corresponding to the currently edited multimedia material. The audio can be the original audio of the multimedia material or background music added by the user. Background music can be an audio file within the application, such as a complete song, a segment of a song, or a cut audio clip, etc. This disclosure does not limit its scope. When the multimedia material is audio, speech recognition can be performed on the multimedia material itself.

[0034] In some embodiments, the application can send audio to the middleware service via an electronic device. The middleware service calls a subtitle recognition tool to perform text recognition on the audio, and obtains the corresponding subtitle text and the timestamp information of the audio segment corresponding to each text element in the subtitle text. The timestamp information may include the start time and end time of the audio segment.

[0035] For example, if the total duration of the audio corresponding to a multimedia clip is 7 seconds, and the subtitle text obtained by speech recognition of the audio is: "I'm very happy today," there are a total of 7 text elements. Each text element corresponds to a 1-second duration of the audio clip. Therefore, the correspondence between each text element and the timestamp information of the corresponding audio clip is shown in Table 1 below:

[0036] Table 1

[0037] Text elements timestamp information of audio clips I 00:00—00:01 now 00:01—00:02 sky 00:02—00:03 very 00:03—00:04 open 00:04—00:05 Heart 00:05—00:06 ah 00:06—00:07

[0038] The above example uses Chinese as the language of the audio, and the corresponding text elements are in units of characters. If the audio uses other languages, the text elements are in units of words. For example, if the audio uses English, the text elements are in units of English words.

[0039] In some embodiments, the application may perform speech recognition in response to user input commands. This disclosure does not limit the implementation method of the commands that trigger speech recognition. In some embodiments, the commands for speech recognition may include, but are not limited to, operations such as clicking, double-clicking, long-pressing, and swiping. For example, when an area / control for adding recognition subtitles to multimedia materials is set on a page of the application, the command for speech recognition may be an operation received on that area / control.

[0040] S102. Match the timestamp information of the audio segments corresponding to each text element with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment that matches the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline.

[0041] If the multimedia material is image / video material, the material segment that matches the text element can be understood as an image / video segment, which includes image / video frames synthesized with the text element. If the multimedia material is audio material, the material segment that matches the text can be understood as an audio segment, which includes one or more speech units synthesized with the text element.

[0042] Since the subtitle processing method provided in this disclosure aims to achieve the effect of corresponding text subtitles appearing when a certain word is spoken, when determining the matching material segment based on the timestamp information of the audio segment corresponding to each text element, the time of the audio segment corresponding to the text element on the editing timeline is consistent with the time of the material segment on the editing timeline.

[0043] Here, "consistent on the editing timeline" can be understood as the starting time of the material segment matching the text element on the editing timeline being consistent with the starting time of the audio segment corresponding to the text element on the editing timeline.

[0044] The timing of the disappearance of text elements in the subtitles can be flexible. They can disappear at the end of their corresponding audio segment, or when the sentence (or text segment of a specified length) they belong to reaches the end position, or after a preset duration after the end of their corresponding audio segment. This disclosure does not impose any limitations.

[0045] Therefore, on the editing timeline, the end time of the material segment that matches the text element can be equal to the end time of the audio segment corresponding to that text element. Using this method, the subtitle text has a word-by-word jump animation effect, and the text elements that appear earlier will disappear as their corresponding audio segments reach their end time.

[0046] The end time of the time range of the video frame image that matches the text element can be later than the end time of the audio segment corresponding to the text element. In this way, the text elements will appear one by one, and the text elements that appear earlier will remain for a period of time after their corresponding audio segments end before disappearing.

[0047] The speed at which text elements switch depends on the speaking speed of the audio object.

[0048] S103. Combine each text element with the material fragments within the matching time range to obtain the target multimedia material with the subtitle text jumping out word by word animation effect.

[0049] When combining text elements with matching material clips, a pre-defined first subtitle animation style can be applied to the text elements. Subtitles automatically added to multimedia materials will automatically carry the subtitle effect corresponding to the first subtitle animation style during generation, meeting user needs for subtitle effects and reducing post-processing. The first subtitle animation style can include one or more of the text element's entrance style, exit style, and loop style.

[0050] Steps S102 and S103 can be automatically implemented by calling a dynamic subtitle resource package (also known as a subtitle animation resource package). The subtitle text and the timestamp information of each text element included in the subtitle text are passed into the dynamic subtitle resource package. The dynamic subtitle resource package applies a pre-set subtitle animation style to each text element in batches and overlays the text elements with the preset subtitle animation style onto the matching material segment, thereby adding subtitles with a word-by-word animation effect using the first subtitle animation style to the multimedia material.

[0051] The method of this embodiment can achieve the subtitle animation effect when a certain word is spoken; in addition, the user input command can realize the automatic generation of dynamic subtitles, which is simple for users and helps to improve the user experience. Moreover, the method of this embodiment can be applied to various types of devices and has a wide range of applications. In the case of batch text, it can also quickly add subtitles with specified effects to multimedia material clips on mobile devices with small screens.

[0052] pass Figure 1 The method of the illustrated embodiment allows users to further edit the content of the subtitle text after adding subtitles to multimedia material clips. This editing can include, but is not limited to, deleting text elements, inserting new text elements, and replacing text elements. Figure 2 A flowchart illustrating a subtitle processing method provided in another embodiment of this disclosure. Please refer to [link / reference]. Figure 2 As shown, the method in this embodiment is Figure 1 Based on the illustrated embodiment, it also includes:

[0053] S104. In response to the text deletion instruction, delete the corresponding text element from the subtitle text to obtain the updated subtitle text.

[0054] Among them, to delete the text element in the subtitle text, just keep the remaining text elements and the timestamp information of the remaining text elements, so as to obtain the updated subtitle text and the timestamp information of each text element in the updated subtitle text.

[0055] Exemplarily: Assume that the subtitle text before deletion is: Today (00:00 - 00:01) is (00:01 - 00:02) really (00:02 - 00:04) happy (00:04 - 00:05) ah (00:06 - 00:07), and the content in the brackets represents the timestamp information of the audio segment corresponding to the text element.

[0056] After deleting the last text element "ah", the updated subtitle text is: Today (00:00 - 00:01) is (00:01 - 00:02) really (00:02 - 00:04) happy (00:04 - 00:05), and the content in the brackets represents the timestamp information of the audio segment corresponding to the text element.

[0057] If deleting text elements at other positions, just process them in a similar way.

[0058] S105. In response to the text insertion instruction, insert a new text element into the subtitle text to obtain the updated subtitle text.

[0059] The text insertion performed in this step is to insert a new text element without deleting the existing text elements in the subtitle text. In some embodiments, different processing methods can be configured according to the insertion position of the new text element. In some embodiments, if the insertion position of the new text element is in the middle or at the end of the subtitle text, merge the new text element with the adjacent previous text element and share the timestamp of the audio segment corresponding to the adjacent previous text element; if the insertion position of the new text element is at the very front of the subtitle text, merge the new text element with the first text element of the subtitle text and share the timestamp of the audio segment corresponding to the first text element.

[0060] Exemplarily, assume that before inserting text, the subtitle text is: Today (00:00 - 00:01) is (00:01 - 00:02) really (00:02 - 00:04) happy (00:04 - 00:05) ah (00:06 - 00:07), and the content in the brackets represents the timestamp information of the audio segment corresponding to the text element.

[0061] Case 1: After inserting the text element "的" after the text element "真", the updated subtitle text is: 今(00:00-00:01)天(00:01-00:02)真的(00:02-00:04)开(00:04-00:05)心(00:05-00:06)啊(00:06-00:07), where the content in the parentheses represents the timestamp information of the audio segment corresponding to the text element.

[0062] By comparison, it can be seen that after inserting the new text element, "真的" shares the timestamp information (00:02-00:04) of the audio segment originally corresponding to "真".

[0063] Case 2: After inserting the new text element "哈哈" before the text element "今", the updated subtitle text is: 哈哈今(00:00-00:01)天(00:01-00:02)真(00:02-00:04)开(00:04-00:05)心(00:05-00:06)啊(00:06-00:07), where the content in the parentheses represents the timestamp information of the audio segment corresponding to the text element.

[0064] By comparison, it can be seen that after inserting the new text element, "哈哈我" shares the timestamp (00:00-00:01) of the audio segment originally corresponding to "我".

[0065] S106. In response to the text replacement instruction, use the replacement text to replace one or more text elements in the subtitle text to obtain the updated subtitle text.

[0066] When replacing, the timestamp information corresponding to the replacement text is the same as the timestamp information of the audio segment corresponding to the replaced text element. In one replacement, the replacement text can include one or more text elements. The replacement text can be understood as a whole, and the number of replaced text elements can also be one or more consecutively located text elements.

[0067] Exemplarily, assume that before inserting the text, the subtitle text is: 今(00:00-00:01)天(00:01-00:02)真(00:02-00:04)开(00:04-00:05)心(00:05-00:06)啊(00:06-00:07), where the content in the parentheses represents the timestamp information of the audio segment corresponding to the text element.

[0068] Suppose we replace "happy" with "sad" and "ah" with "ah". The updated subtitle text after the replacement is: Today (00:00-00:01) day (00:01-00:02) is really (00:02-00:04) sad (00:04-00:06) ah (00:06-00:07). The text in parentheses represents the timestamp information of the audio segment corresponding to the text element.

[0069] The comparison shows that after the replacement, "sad" uses the sum of the timestamps of the audio segments corresponding to the original "happy" (00:04-00:06); "ah" uses the timestamps of the audio segments corresponding to the original "ah" (00:06-00:07).

[0070] You can choose one or more of the above editing methods to edit the subtitle text as needed.

[0071] S107. Based on the timestamp information of the audio segments corresponding to each text element in the updated subtitle text, determine the material segments in the multimedia material that match each of the text elements respectively.

[0072] S108. Combine each of the text elements with the material segments within the corresponding time period to re-add subtitles to the multimedia material.

[0073] Steps S107 to S108 are respectively related to the aforementioned Figure 1 In the illustrated embodiment, steps S102 and S103 are implemented in a similar manner, and can be referred to the foregoing. Figure 1 Detailed description of the illustrated embodiment.

[0074] If the update is automatically achieved by calling the dynamic subtitle resource package, the updated subtitle text and the timestamp information of each text element included in the updated subtitle text will be re-entered into the dynamic subtitle resource package. The dynamic subtitle resource package will then re-apply the pre-set subtitle animation style to each text element included in the updated subtitle text in batches, and overlay the text elements with the preset subtitle animation style onto the matching material clips, thereby adding subtitles to the multimedia material again.

[0075] The method in this embodiment can meet the user's need to adjust the subtitle content when adding subtitles to multimedia materials. For the updated subtitle text, it can automatically generate subtitles with specified subtitle effects, which is convenient for users and helps to improve the user experience.

[0076] pass Figure 1 The method shown in the embodiment allows users to adjust the current subtitle animation style after adding subtitles to multimedia materials, so as to obtain subtitle effects that meet the user's expectations. Figure 3This is a schematic flowchart illustrating a subtitle processing method according to another embodiment of this disclosure. Please refer to [link / reference]. Figure 3 As shown, the method in this embodiment includes:

[0077] S301. During the editing process of multimedia materials, speech recognition is performed on the audio corresponding to the multimedia materials to obtain the subtitle text corresponding to the audio and the timestamp information of each text element in the subtitle text corresponding to the audio segment.

[0078] S302. Match the timestamp information of the audio segments corresponding to each text element with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment that matches the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline.

[0079] S303. Apply the specified first subtitle animation style to each text element in batches, and combine the text elements with the first subtitle animation style with the material fragments within the matching time to obtain the target multimedia material with subtitle text using the first subtitle animation style to jump out word by word animation effect.

[0080] In this embodiment, steps S301 to S302 are respectively connected with... Figure 1 Steps S101 to S103 in the illustrated embodiment are similar and can be referred to accordingly. Figure 1 The detailed description of the illustrated embodiment will not be repeated here. It should be noted that the first subtitle animation style can be understood as the application's default subtitle animation style.

[0081] S304. Responding to the subtitle animation style switching instruction, apply the second subtitle animation style to each text element in batches, and synthesize the text elements with the second subtitle animation style with the matching material fragments within the time frame to obtain the target multimedia material with subtitle text using the second subtitle animation style to jump out word by word with animation effect.

[0082] The application can provide users with a page for editing subtitle animation styles via electronic devices. The page can display one or more areas or controls corresponding to subtitle animation styles that users can select. Users can input subtitle animation style switching commands by operating the areas or controls corresponding to the subtitle animation styles.

[0083] If implemented automatically by calling a dynamic subtitle resource package, the updated subtitle text and the timestamp information of each text element included in the updated subtitle text are re-entered into the dynamic subtitle resource package. The dynamic subtitle resource package then applies the user-specified second subtitle animation style to each text element included in the updated subtitle text in batches, and overlays the text elements with the second subtitle animation style onto the matching material clips, thereby adding subtitles to the multimedia material again.

[0084] The method in this embodiment can meet the user's needs for adjusting the subtitle effect in the later stage, and supports batch editing of subtitle animation styles, with high subtitle processing efficiency.

[0085] Based on the foregoing description, this disclosure will use electronic devices as examples, along with accompanying drawings and application scenarios, to illustrate the subtitle processing method provided by this disclosure. For ease of explanation, Figures 4A-4I In this example, a mobile phone is used as the electronic device, and a video editing application (referred to as Application 1) is installed on the phone. Furthermore, the multimedia materials imported by the user into Application 1 are video materials.

[0086] Please see Figures 4A-4I , Figures 4A-4I This is a schematic diagram of a human-computer interaction interface provided in an embodiment of this disclosure.

[0087] Application 1 can be displayed on a mobile phone as shown in the example below. Figure 4A The user interface 11 shown is used to display a multimedia material editing page (hereinafter referred to as the editing page). Application 1 performs a set of functions on the editing page, such as previewing and playing the editing effects of multimedia materials, adding background music to multimedia materials, adding filters, stickers, text, etc. to multimedia materials.

[0088] Reference Figure 4A As shown, the user interface 11 includes area a1, which is a preview area for the editing effects of multimedia materials; the user interface 11 also includes area a2, in which multimedia materials and other editing materials added during the editing process can be displayed according to the timeline. The user interface 11 also includes area a3, in which multiple editing function entry points can be provided to the user. For example, area a3 includes control 101, which is used to access the text function collection page of application 1. This text function collection page includes multiple controls, each corresponding to a different text function.

[0089] For example, when application 1 receives a user's message... Figure 4A After performing an operation such as clicking control 101 in the user interface 11 shown, application 1 can display, exemplarily, on the mobile phone, as shown in the example. Figure 4BThe user interface 12 shown in the image displays a collection page of text functions provided by application 1. The collection page of text functions can provide users with entry points for various text functions. Users can enter the corresponding text function operation page through the entry point to add text content to multimedia materials.

[0090] The user interface 12 includes an area a4, which contains entry points for the functions of creating new text, text templates, subtitle recognition, lyrics recognition, stickers, and drawing pen. The control 102 shown in the user interface 12 is the entry point for the subtitle recognition function.

[0091] Application 1 received the user's... Figure 4B After performing an operation such as clicking control 102 in the user interface 12 shown, application 1 can display, exemplarily, on the mobile phone, as shown... Figure 4C The user interface 13 shown is used to display the subtitle recognition panel provided by application 1. The subtitle recognition panel can provide users with options for recognition type, language selection entry, switch to mark invalid segments, switch to dynamic subtitles, and switch to clear existing subtitles at the same time.

[0092] Dynamic subtitles refer to the function of adding subtitles to multimedia materials with an animated effect where the subtitle text appears one character at a time. Specifically, when the dynamic subtitle switch is off, the added subtitles appear as single subtitles in the form of sentence segments; when the dynamic subtitle switch is on, the added subtitles appear as a one-character jump effect, that is, the text elements in the subtitle text appear one by one and the text elements are displayed at the beginning of the corresponding audio segment.

[0093] In some embodiments, the user's selection can be remembered, and when the subtitle recognition panel is opened, the on / off state of the dynamic subtitles when the user last exited subtitle recognition can be displayed, which is more in line with the user's usage habits. When the dynamic subtitle function is updated for the first time in application 1, the dynamic subtitle switch can be in the off state, as shown in user interface 13.

[0094] Application 1 receives user information such as Figure 4C After performing an operation such as clicking the on / off button for dynamic subtitles in the user interface 13 shown, the following will be displayed: Figure 4D The user interface 14 shown is in the on state of the dynamic subtitles.

[0095] The user interface 14 also includes a control 103, which is used to instruct the user to start speech recognition and add subtitles with a word-by-word animation effect. In response to a user performing an action such as clicking control 103 in the user interface 14, application 1 displays, for example, on the mobile phone... Figure 4EThe user interface 15 shown has its caption recognition panel closed, and displays prompts, such as animations and text, in area a4 to inform the user that a dynamic caption animation is currently being created. To minimize the obstruction of the animation and text from the preview screen displayed in area a1, area a4 can be positioned above area a1. It should be understood that area a4 can also be positioned in other locations, and this disclosure does not limit this.

[0096] In conjunction with the foregoing, the user's switching on and off of the dynamic subtitles and the operation of control 103 trigger application 1 to perform speech recognition on the audio corresponding to the multimedia material and automatically add dynamic subtitles with a word-by-word pop-up animation effect.

[0097] Once the dynamic caption animation is created, application 1 can display it on the phone as an example, such as... Figure 4F The user interface 16 shown can display prompts in area a4, such as the text "Recognition successful, subtitles have been automatically generated".

[0098] Afterwards, users can click the preview play button to preview the subtitle effect in area a1. If it meets the user's expectations, the edited multimedia material can be exported as the target video for publication or saving.

[0099] Combination Figures 4A to 4F The interactive process shown in this disclosure provides users with a dynamic subtitle switch in the pre-recognition stage, making it easier for users to use. Furthermore, it remembers the on / off state of the dynamic subtitle switch when the user last exited the subtitle recognition panel, so the user does not need to perform any operation when using it again, thus further improving efficiency and reducing the need for excessive user intervention.

[0100] To better meet user needs, Application 1 also provides users with the ability to add dynamic subtitles or modify the animation style of existing subtitles in subsequent stages.

[0101] For example, in Figure 4F Based on the user interface 16 shown, area a2 displays icons corresponding to multimedia materials and subtitle text according to a timeline. By operating on the icons of the subtitle text displayed in area a2 (such as clicking), the subtitle can be edited again. When application 1 receives a user's click operation on any text segment contained in a subtitle in area a2 of user interface 16, application 1 can display, for example, on the mobile phone... Figure 4G The user interface shown is 17.

[0102] In user interface 17, area a1 displays a text box 104 corresponding to the subtitle text. The text box 104 contains the text content corresponding to the current preview position, which can be one or more sentences (i.e., text fragments). Area a1 can also display controls for manipulating the text box, such as rotation and copying. Users can also zoom in or out of the text box using two fingers, and the size of the text elements in the text box will change with the size of the text box. User interface 17 also includes area a5, which displays the subtitle editing function collection page. This page provides entry points for various editing functions for editing the currently added subtitles, such as: batch editing subtitles, subtitle splitting, copying subtitles, editing subtitles, deleting subtitles, adding on-screen text, and subtitle animation styles. User interface 17 also includes control 105, which is used to access the subtitle animation panel to add subtitle effects (including dynamic subtitle effects) to the current subtitle or modify the subtitle animation style used by the current subtitle.

[0103] Application 1 receives an operation from the user in the user interface 17, such as clicking control 105, and then displays the following: Figure 4H The user interface 18 shown includes region a6.

[0104] In this example, user A6 in region A displays a subtitle animation panel. This panel includes a tag 106 for setting animation styles, as well as font tags, style tags, decorative text tags, and text template tags. In some embodiments, it can be as follows: Figure 4H As shown, entering the subtitle animation style panel will default to label 106 and display the relevant content of label 106. In other embodiments, it can default to other labels, and application 1 will display the relevant content of label 106 after receiving a user's click operation on label 106.

[0105] Reference Figure 4H As shown, area a6 also includes a dynamic subtitle switch 107, which can add a subtitle effect with text elements displayed one by one to the current subtitle by operating the dynamic subtitle switch 107.

[0106] In some embodiments, if the user has added dynamic subtitles during the pre-processing stage, this area can be displayed as "on"; if the user has not used dynamic subtitles during the pre-processing stage, this area can be displayed as "off". The user can toggle the on / off state of the dynamic subtitle switch 107 displayed in the user interface 18 to "on". Figure 4H In the embodiment shown, the dynamic subtitle switch 107 is in the off state.

[0107] In addition, area a6 also includes: label 108 for setting the subtitle entrance style, label 109 for setting the subtitle exit style, label 110 for setting the subtitle loop style, label 111 for setting the dynamic subtitle animation style, and area a7. Area a7 is used to display the content of the corresponding label based on the currently positioned label. In some cases, when the dynamic subtitle switch 107 is off, content related to any label can be displayed by default, for example... Figure 4H The user interface 17 shown defaults to displaying the content corresponding to label 108.

[0108] Specifically, when application 1 receives an operation (such as a click) performed by the user on the dynamic subtitle switch 107 in the user interface 18, and the dynamic subtitle switch 107 switches from the off state to the on state, application 1 can display, for example, on the mobile phone, the following: Figure 4I The user interface 19 shown is referenced. Figure 4I As shown, in user interface 19, the dynamic subtitle switch 107 is on, and label 111 is selected. Area a7 is used to display one or more dynamic subtitle animation styles that are available for user selection. The display icons corresponding to various dynamic subtitle animation styles can be arranged sequentially from left to right, and users can view them back and forth by swiping the screen left and right. Among them, the default dynamic subtitle animation style in application 1 can be displayed in the first position from left to right, so that users can clearly see which dynamic subtitle animation style application 1 uses by default.

[0109] The area a7 may also include a disable button 112. The disable button can be located on the far left of area a7, or it can be located in other positions; this disclosure does not limit this. When the user clicks the disable button 112, the dynamic subtitle effect is turned off, and the dynamic subtitle switch 107 will switch to the off state.

[0110] Suppose a user clicks on the second dynamic subtitle animation style from left to right in area a7. This is equivalent to inputting a subtitle animation style switching command into application 1. Application 1 responds to the subtitle animation style switching command by applying the second dynamic subtitle style to each text element included in the subtitle text. The user can switch the dynamic subtitle animation style multiple times until they obtain the subtitle effect that meets their expectations.

[0111] exist Figure 4H The user interface 18 shown and Figure 4IBased on the user interface 19 shown, area a5 also includes area a8, which displays a text editing box. Users can use this text editing box to delete text elements in the subtitle text, insert new text, or replace existing text elements. User operations on the text editing box are equivalent to inputting delete, insert, and replace commands to application 1. When editing the text content in the text editing box in area a8, the edited text content is simultaneously displayed in the text box 104 shown in area a1, allowing users to preview the edited subtitle content and its display effect in the video frame image of the multimedia material clip.

[0112] As shown above Figures 4F to 4I The embodiment shown allows users to add dynamic subtitles and adjust the dynamic subtitle style by setting a dynamic subtitle switch and a dynamic subtitle animation style label in the subtitle animation style panel during the post-processing stage.

[0113] It should be noted that the above Figures 4A to 4I The interactive interface diagram shown is not intended to limit the subtitle processing method provided in this disclosure. It should be understood that the styles, triggering methods, etc. of some controls, panels, and labels can be flexibly adjusted according to requirements.

[0114] Figure 5 This is a schematic diagram of a subtitle processing apparatus provided according to an embodiment of this disclosure. Please refer to [link / reference]. Figure 5 As shown, the device 500 provided in this embodiment includes:

[0115] The speech recognition module 501 is used to perform speech recognition on the audio corresponding to the multimedia material during the editing process to obtain the subtitle text corresponding to the audio and the timestamp information of each text element of the subtitle text corresponding to the audio segment.

[0116] The matching module 502 is used to match the timestamp information of the audio segments corresponding to each text element with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment that matches the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline.

[0117] The subtitle synthesis module 503 is used to synthesize each of the text elements with material fragments within a matching time range to obtain target multimedia material with subtitle text jumping out word by word animation effect.

[0118] In some embodiments, on the editing timeline, the start time of the time segment to which the material segment matching the text element belongs is the same as the start time of the audio segment corresponding to the text element; and on the editing timeline, the end time of the time segment to which the material segment matching the text element belongs is the same as the end time of the audio segment corresponding to the text element, or the end time of the time segment to which the material segment matching the text element belongs is later than the end time of the audio segment corresponding to the text element.

[0119] In some embodiments, the subtitle synthesis module 503 is specifically used to apply a specified first subtitle animation style to each of the text elements in batches, and to synthesize the text elements with the first subtitle animation style with the material fragments within the matching time period to obtain the target multimedia material with subtitle text using the first subtitle animation style to jump out word by word with an animation effect.

[0120] Optionally, the device 500 also includes a subtitle text update module 504.

[0121] In some embodiments, the subtitle text update module 504 is used to respond to a text deletion instruction by deleting the corresponding text element from the subtitle text to obtain the updated subtitle text.

[0122] Accordingly, the matching module 502 is also used to determine the material segments in the multimedia material that match each of the text elements according to the timestamp information of the audio segments corresponding to each text element in the updated subtitle text.

[0123] The subtitle synthesis module 503 is also used to synthesize each of the text elements included in the updated subtitle text with the material fragments within the matching time period, so as to re-add subtitles with the subtitle text jumping out animation effect for the multimedia material.

[0124] In some embodiments, the subtitle text update module 504 is further configured to respond to a text insertion instruction by inserting new text elements into the subtitle text to obtain updated subtitle text.

[0125] Accordingly, the matching module 502 is also used to match the timestamp information of the audio segments corresponding to each text element in the updated subtitle text with each material unit in the multimedia material to determine the material segments in the multimedia material that match each of the text elements respectively; wherein, the newly added text element and the adjacent text element share the timestamp information of the audio segments corresponding to the adjacent text element.

[0126] The subtitle synthesis module 503 is also used to combine each of the text elements included in the updated subtitle text with the material segments within the matching time range, so as to re-add subtitles with the subtitle text jumping out animation effect for the multimedia material.

[0127] In some embodiments, if the newly added text element is inserted at the very beginning of the subtitle text, the newly added text element is merged with the first text element in the subtitle text, sharing the timestamp of the audio segment corresponding to the first text element; if the newly added text element is inserted in the middle or at the very end of the subtitle text, the newly added text element is merged with the adjacent preceding text element, sharing the timestamp of the audio segment corresponding to the preceding text element.

[0128] In some embodiments, the subtitle text update module 504 is further configured to respond to a text replacement instruction by replacing one or more text elements in the subtitle text with replacement text to obtain an updated subtitle text.

[0129] Accordingly, the matching module 502 is further configured to determine the material segments in the multimedia material that match each of the text elements according to the timestamp information of the audio segments corresponding to each text element in the updated subtitle text; wherein, the timestamp information of the audio segment corresponding to the replaced text element is used.

[0130] The subtitle synthesis module 503 is also used to synthesize each of the text elements included in the updated subtitle text with the material fragments within the matching time period, so as to re-add subtitles with the subtitle text jumping out animation effect for the multimedia material.

[0131] In some embodiments, the subtitle synthesis module 503 is further configured to respond to a subtitle animation style switching instruction, apply a second subtitle animation style to each of the text elements in batches, and synthesize the text elements with the second subtitle animation style with material fragments within a matching time range to obtain a target multimedia material with subtitle text using the second subtitle animation style to jump out word by word with an animation effect.

[0132] In some embodiments, the audio corresponding to the multimedia material is the original audio included in the multimedia material or background music added to the multimedia material.

[0133] The subtitle processing device provided in this embodiment can be used to execute the technical solution of any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and can be referred to the detailed description of the foregoing method embodiments. For the sake of brevity, it will not be repeated here.

[0134] For example, this disclosure provides an electronic device including: one or more processors; a memory; and one or more computer programs; wherein the one or more computer programs are stored in the memory; and when the one or more processors execute the one or more computer programs, the electronic device causes the subtitle processing method of the foregoing embodiments to be implemented.

[0135] For example, this disclosure provides a chip system applied to an electronic device including a memory and a sensor; the chip system includes: a processor; when the processor executes the caption processing method of the preceding embodiments.

[0136] For example, this disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program being processed by a processor to cause an electronic device to implement the subtitle processing method of the preceding embodiments.

[0137] For example, this disclosure provides a computer program product that, when run on a computer, causes the computer to execute the subtitle processing method described in the preceding embodiments.

[0138] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0139] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A subtitle processing method, characterized in that, include: During the editing process of multimedia materials, speech recognition is performed on the audio corresponding to the multimedia materials to obtain the subtitle text corresponding to the audio and the timestamp information of each text element of the subtitle text corresponding to the audio segment. Based on the timestamp information of the audio segments corresponding to each text element, the multimedia material is matched with each material unit to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment matching the text element on the editing timeline is consistent with the time of the audio segment corresponding to the text element on the editing timeline; Each of the text elements is combined with a matching time segment to obtain a target multimedia material with a subtitle text jumping out word by word animation effect; In response to a text insertion command, a new text element is inserted into the subtitle text to obtain the updated subtitle text; Based on the timestamp information of the audio segments corresponding to each text element in the updated subtitle text, the multimedia material is matched with each material unit to determine the material segments in the multimedia material that match each text element; wherein, the newly added text element is merged with the adjacent text element and shares the timestamp information of the audio segment corresponding to the adjacent text element; The updated subtitle text includes each of the text elements and is combined with the material fragments within the matching time range to re-add subtitles with a word-by-word animation effect to the multimedia material; If the newly added text element is inserted at the very beginning of the subtitle text, then the newly added text element is merged with the first text element in the subtitle text, and they share the timestamp of the audio segment corresponding to the first text element. If the newly added text element is inserted in the middle or at the end of the subtitle text, the newly added text element is merged with the adjacent preceding text element and shares the timestamp of the audio segment corresponding to the preceding text element.

2. The method according to claim 1, characterized in that, On the editing timeline, the start time of the time segment matching the text element is the same as the start time of the audio segment corresponding to the text element; and on the editing timeline, the end time of the time segment matching the text element is the same as the end time of the audio segment corresponding to the text element, or the end time of the time segment matching the text element is later than the end time of the audio segment corresponding to the text element.

3. The method according to claim 1, characterized in that, The process of combining each text element with a matching time-limited material segment to obtain a target multimedia material with a text-to-text animation effect includes: A specified first subtitle animation style is applied to each of the text elements in batches. The text elements with the first subtitle animation style are combined with the material segments within the matching time period to obtain the target multimedia material with subtitle text using the first subtitle animation style to jump out word by word animation effect.

4. The method according to claim 1, characterized in that, Also includes: In response to a text deletion command, the corresponding text elements are deleted from the subtitle text to obtain the updated subtitle text; Based on the timestamp information of the audio segments corresponding to each text element in the updated subtitle text, determine the material segments in the multimedia material that match each of the text elements respectively; The updated subtitle text includes each of the text elements, which are then combined with the corresponding material segments within the same time frame to re-add subtitles with a word-by-word animation effect to the multimedia material.

5. The method according to claim 1, characterized in that, Also includes: In response to a text replacement instruction, one or more text elements in the subtitle text are replaced with replacement text to obtain updated subtitle text; Based on the timestamp information of the audio segments corresponding to each text element in the updated subtitle text, the material segments in the multimedia material that match each of the text elements are determined; wherein, the replacement text corresponds to the timestamp information of the audio segment corresponding to the replaced text element; The updated subtitle text includes each of the text elements, which are then combined with the corresponding material segments within the same time frame to re-add subtitles with a word-by-word animation effect to the multimedia material.

6. The method according to claim 3, characterized in that, Also includes: In response to the subtitle animation style switching command, the second subtitle animation style is applied to each of the text elements in batches. The text elements with the second subtitle animation style are combined with the material segments within the matching time range to obtain the target multimedia material with subtitle text using the second subtitle animation style to jump out word by word with animation effect.

7. The method according to any one of claims 1 to 6, characterized in that, The audio corresponding to the multimedia material is either the original audio included in the multimedia material or background music added to the multimedia material.

8. A subtitle processing device, characterized in that, include: The speech recognition module is used to perform speech recognition on the audio corresponding to the multimedia material during the editing process to obtain the subtitle text corresponding to the audio and the timestamp information of each text element of the subtitle text corresponding to the audio segment. The matching module is used to match the timestamp information of the audio segments corresponding to each text element with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the time of the material segment that matches the text element on the editing timeline is the same as the time of the audio segment corresponding to the text element on the editing timeline; The subtitle synthesis module is used to synthesize each of the text elements with a material segment within a matching time range to obtain a target multimedia material with a subtitle text jumping out word by word animation effect; The device further includes: a subtitle text update module, used to respond to a text insertion command and insert new text elements into the subtitle text to obtain updated subtitle text; The matching module is further configured to match the timestamp information of the audio segments corresponding to each text element in the updated subtitle text with each material unit in the multimedia material to determine the material segments in the multimedia material that match each text element; wherein, the newly added text element and the adjacent text element share the timestamp information of the audio segments corresponding to the adjacent text element. The subtitle synthesis module is also used to combine each of the text elements included in the updated subtitle text with the material fragments within the matching time range, so as to re-add subtitles with the subtitle text jumping out animation effect for the multimedia material; If the newly added text element is inserted at the very beginning of the subtitle text, then the newly added text element is merged with the first text element in the subtitle text, and they share the timestamp of the audio segment corresponding to the first text element. If the newly added text element is inserted in the middle or at the end of the subtitle text, the newly added text element is merged with the adjacent preceding text element and shares the timestamp of the audio segment corresponding to the preceding text element.

9. An electronic device, characterized in that, include: Memory and processor; The memory is configured to store computer program instructions; The processor is configured to execute the computer program instructions, causing the electronic device to implement the subtitle processing method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, include: Computer program instructions; At least one processor of the electronic device executes the computer program instructions, causing the electronic device to implement the subtitle processing method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The electronic device executes the computer program product, causing the electronic device to implement the subtitle processing method as described in any one of claims 1 to 7.