Video processing method, device, equipment, medium and program product

By displaying candidate subtitle templates in the video processing method and applying special effect materials to generate target videos, the problem of single subtitle form is solved, and rich subtitle effects and better user experience is achieved.

CN120301997APending Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346292.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video subtitles are single, affecting the display effect and user experience of video content.

Method used

By displaying candidate subtitle templates in the video processing method, the target subtitle template is determined based on the subtitle text and label information, and the target video is generated using special effects materials.

Benefits of technology

It enriches the video subtitle effect, improves the efficiency of subtitle editing, enhances the vividness and display effect of the video content, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301997A_ABST
    Figure CN120301997A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method and device, equipment, a medium and a program product. The method comprises the following steps: displaying a to-be-edited video and a first control in a first page; in response to a trigger operation of the first control, determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the labeling information of the subtitle text, displaying a subtitle style page, and displaying the at least one candidate subtitle template on the subtitle style page; and in response to a trigger operation of the at least one candidate subtitle template, determining a target subtitle template in the at least one candidate subtitle template, generating a target video according to the at least one special effect material in the target subtitle template and the to-be-edited video, and displaying the target video in the first page. According to the embodiment of the invention, at least one special effect material can be added to the to-be-edited video by applying the candidate subtitle template, the subtitle effect editing efficiency can be improved, the subtitle effect of the to-be-edited video is enriched, and the problem of single subtitle form in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to computer technology, and in particular, to a video processing method, apparatus, device, medium, and program product. Background Art

[0002] With the development of computer technology, videos have become an important information transmission medium due to the intuitiveness and richness of their information content. Subtitles of videos can help viewers better watch videos. In related technologies, it is usually necessary to manually add video subtitles. Currently, video subtitles only include text information, and the subtitle form is relatively single, which affects the display effect of video content and the user experience is not good. Summary of the Invention

[0003] Embodiments of the present disclosure provide a video processing method, apparatus, device, medium, and program product, which can solve the problem of single subtitle form and enrich the display effect of video content.

[0004] In a first aspect, an embodiment of the present disclosure provides a video processing method, including:

[0005] Display a video to be edited and a first control on a first page;

[0006] In response to a trigger operation of the first control, determine at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, display a subtitle style page, and display the at least one candidate subtitle template on the subtitle style page, where the candidate subtitle template includes at least one special effect material;

[0007] In response to a trigger operation of the at least one candidate subtitle template, determine a target subtitle template in the at least one candidate subtitle template, generate a target video according to the at least one special effect material in the target subtitle template and the video to be edited, and display the target video on the first page.

[0008] In a second aspect, an embodiment of the present disclosure further provides a video processing apparatus, including:

[0009] A first display module, configured to display a video to be edited and a first control on a first page;

[0010] A second display module, configured to, in response to a trigger operation of the first control, determine at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, display a subtitle style page, and display the at least one candidate subtitle template on the subtitle style page, where the candidate subtitle template includes at least one special effect material;

[0011] A video generation module, configured to determine a target subtitle template from the at least one candidate subtitle template in response to a trigger operation for the at least one candidate subtitle template, generate a target video based on at least one special effect material in the target subtitle template and the video to be edited, and display the target video on the first page.

[0012] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:

[0013] One or more processors;

[0014] A storage device for storing one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the video processing method according to any embodiment of the present disclosure.

[0016] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the video processing method according to any embodiment of the present disclosure when executed by a computer processor.

[0017] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program, where the computer program implements the video processing method according to any embodiment of the present disclosure when executed by a processor.

[0018] An embodiment of the present disclosure provides a video processing method. By displaying a video to be edited and a first control on a first page, in response to a trigger operation of the first control, determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, displaying the at least one candidate subtitle template on a subtitle style page, in response to a trigger operation of the at least one candidate subtitle template, determining a target subtitle template from the at least one candidate subtitle template, generating a target video based on at least one special effect material in the target subtitle template and the video to be edited, and displaying the target video. Since the embodiment of the present disclosure determines a candidate subtitle template associated with the video subtitle based on the subtitle text and annotation information of the video to be edited, thus obtaining a candidate subtitle template with a higher degree of fit with the video subtitle. By applying the candidate subtitle template, at least one special effect material can be added to the video to be edited, which can improve the editing efficiency of subtitle effects, enrich the subtitle effects of the video to be edited, solve the problem of single subtitle form in the related art, and make the video content more vivid through rich subtitle effects, thereby enriching the video display effect and enhancing the user experience. Description of the Drawings

[0019] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original and elements are not necessarily drawn to scale.

[0020] Figure 1 Schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0021] Figure 2 Schematic diagram of an interaction provided by an embodiment of the present disclosure;

[0022] Figure 3 Schematic diagram of a first page provided by an embodiment of the present disclosure;

[0023] Figure 4 Another schematic diagram of an interaction provided by an embodiment of the present disclosure;

[0024] Figure 5 Another schematic diagram of an interaction provided by an embodiment of the present disclosure;

[0025] Figure 6 Another schematic diagram of an interaction provided by an embodiment of the present disclosure;

[0026] Figure 7 Schematic flowchart of another video processing method provided by an embodiment of the present disclosure;

[0027] Figure 8 Schematic flowchart of another video processing method provided by an embodiment of the present disclosure;

[0028] Figure 9 Architecture diagram of a subtitle style prediction model provided by an embodiment of the present disclosure;

[0029] Figure 10 Schematic structural diagram of a video processing device provided by an embodiment of the present disclosure;

[0030] Figure 11 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specific Embodiments

[0031] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not used to limit the protection scope of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0033] As used herein, the term "including" and its variants are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0034] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0038] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0039] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0040] It should be understood that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0041] It should be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0042] Figure 1 The figure is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of adding subtitles to a video. This method can be executed by a video processing device, and the device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, and the electronic device can be a mobile terminal, a PC or a server, etc.

[0043] As Figure 1 shown, the method includes:

[0044] S110. Display a video to be edited and a first control on a first page.

[0045] Among them, the video to be edited represents a video to which subtitle special effects are to be added. The video to be edited may include subtitles, or the video to be edited may not include subtitles. In subsequent processing, the subtitles are determined by identifying the audio information of the target object in the video to be edited. For example, the video to be edited may include a video of a voiceover, a video clip of a film and television work, or an explanatory video of a film and television work, etc.

[0046] The first page includes a preview page or a clip page. Among them, the preview page includes at least one preview area. The clip page includes a preview area and a clip track area, and a clip track is displayed in the clip track area.

[0047] The first control is used to trigger a subtitle style adjustment event. For example, the first control may be a control in the function area at the bottom of the first page. The first page is divided into a video preview area, a subtitle preview area and a function area according to a preset layout style. The video to be edited is displayed in the video preview area. The ordinary subtitles of the video to be edited are displayed in the subtitle preview area. Among them, the ordinary subtitles refer to the subtitle text to which special effects are added.

[0048] Exemplarily, the first control, the video to be edited, and the subtitles of the video to be edited are displayed on the first page, wherein the subtitles of the video to be edited are associated with the subtitle text. Through the triggering operation of the first control, the subtitle style page can be invoked to intuitively display at least one candidate subtitle template. In addition, the video subtitles displayed in the video preview area are the same as the selected subtitle text in the subtitle preview area, and the video voice content corresponding to the selected subtitle text is played synchronously, so that the video content can be displayed from both visual and auditory aspects.

[0049] Optionally, if the subtitle text is non-Mandarin content, a fourth control is displayed on the first page, and the fourth control is used to trigger a conversion event for converting non-Mandarin content to Mandarin content.

[0050] Optionally, a batch editing control can also be displayed at the bottom of the first page. In response to the triggering operation of the batch editing control, a subtitle editing page is displayed. The subtitle editing page can be in full-screen form. The subtitle text of the video to be edited is displayed on the subtitle editing page, so as to implement the editing operation of the video subtitles corresponding to any video frame.

[0051] Optionally, in response to the triggering operation of a single-sentence subtitle in the subtitle preview area on the first page, an input box and a virtual keyboard are displayed, and the single-sentence subtitle to be edited is displayed in the input box.

[0052] Optionally, an all-delete control can also be displayed at the bottom of the first page. In response to the triggering operation of the all-delete control, a preset prompt pop-up window is displayed to prompt the user to confirm whether to perform the operation of deleting the video subtitles of the video to be edited.

[0053] Optionally, displaying the first control, the video to be edited, and the subtitles of the video to be edited on the first page includes: displaying the video to be edited and a second control on a second page, wherein the second control is used to trigger an event for adding subtitles to the video to be edited. In response to the triggering operation of the second control, the subtitle text is determined according to the audio characteristics of the video to be edited, and the annotation information is obtained by annotating the subtitle text. The first page is displayed, and the first control, the video to be edited, and the subtitles of the video to be edited are displayed on the first page.

[0054] Among them, the second page can be an editing page. The video to be edited and a second control are displayed on the second page. In response to the triggering operation of the second control, the first page is displayed. The first page includes a video preview area, a subtitle preview area, and a function area. In response to the triggering operation of the second control, the page to be edited is also input into a pre-trained video understanding model to obtain the subtitle text of the video to be edited and the annotation information of the subtitle text. The first page is displayed, a first control is displayed in the function area of the first page, the video to be edited is displayed in the video preview area of the first page, and the subtitles of the video to be edited are displayed in the subtitle preview area of the first page.

[0055] Among them, the video understanding model can be a multimodal model for outputting subtitle text and annotation information for a video. The annotation information refers to the annotation results of specific content in the subtitle text. The specific content can be the content preset according to the business scenario. Fine-grained marking of the subtitle text is achieved through text annotation. For example, marking of the subtitle text can be performed based on at least one of the text content, prosodic features in the audio information of the target object, and linguistic features. Among them, the prosodic features include fundamental frequency, speech intensity, and speech rate, etc. The linguistic features include intonation and phoneme, etc. Through fine-grained annotation of the subtitle text, similar special effect materials can be matched from the special effect material library based on the annotation information, increasing the degree of association between the special effect materials and the video content, avoiding the situation where the special effect materials are separated from the video content, and improving the video quality.

[0056] S120. In response to the triggering operation of the first control, at least one candidate subtitle template is determined according to the subtitle text of the video to be edited and the annotation information of the subtitle text, the subtitle style page is displayed, and the at least one candidate subtitle template is displayed on the subtitle style page.

[0057] Among them, the candidate subtitle template includes at least one special effect material. The subtitle template can be a collection of special effect materials. One subtitle template corresponds to one special effect material package. The special effect material package includes at least one special effect material file. Optionally, the at least one candidate subtitle template includes subtitle templates of different preset styles. The subtitle style page includes at least one type of candidate subtitle template of a preset style. The style type of the subtitle template can be preset according to the business scenario. Placeholder identifiers corresponding to the number of candidate subtitle templates of each style are set on the subtitle style page. Optionally, the subtitle style page also includes multiple subtitle style editing items, for example, editing items such as color, style, or font. Exemplarily, the template effect option is also included in the editing items such as color, style, and font. After editing the color, style, or font of the subtitle, in response to the triggering operation of the template effect option, the subtitle can present the subtitle effect of the target subtitle template.

[0058] Among them, the special effect material represents a material with a specific effect. For example, the special effect material includes stickers and / or text templates. The stickers include text or images, etc., and the text templates include text with specific effects.

[0059] Exemplarily, in response to a trigger operation of the first control, a subtitle style page is displayed, and at least one candidate subtitle template is displayed on the subtitle style page. It should be noted that the screen is divided into upper and lower parts according to a preset split-screen layout strategy. The video preview area in the first page is displayed in the upper half of the screen. The subtitle style page is displayed in the lower half of the screen. Thus, during the video preview process, based on the trigger operation of at least one candidate subtitle template in the subtitle style page, the video subtitle effect is edited.

[0060] Figure 2 An interaction schematic diagram provided by an embodiment of the present disclosure. As Figure 2 shown, in the first page 210, a video to be edited 220 and a first control 230 are displayed. In response to the trigger operation of the first control 230, a subtitle style page 240 is displayed to present the effect of simultaneously displaying the video preview area in the first page 210 and the subtitle style page 240. At least one candidate subtitle template 250 is displayed in the subtitle style page 240.

[0061] S130. In response to the trigger operation of the at least one candidate subtitle template, determine a target subtitle template among the at least one candidate subtitle templates, generate a target video according to at least one special effect material in the target subtitle template and the video to be edited, and display the target video on the first page.

[0062] Among them, the target subtitle template represents the subtitle template in the selected state among the at least one candidate subtitle templates. The target video represents a video carrying at least one special effect material in the target subtitle template.

[0063] Exemplarily, according to the timestamp corresponding to the subtitle text, determine the addition position of at least one special effect material in the target subtitle template in the video to be edited. Add the corresponding special effect material to the video to be edited based on the addition position to obtain a target video, so that the video frames of the target video carry the corresponding subtitle special effects.

[0064] Optionally, displaying the target video includes: displaying the target video in the video preview area of the first page, and displaying the subtitle special effects corresponding to the target subtitle template in the target video.

[0065] Figure 3 A schematic diagram of a first page provided by an embodiment of the present disclosure. As Figure 3As shown, the video to be edited 320 is displayed in the video preview area of the first page 310, and at least one candidate subtitle template 340 is displayed in the subtitle style page 330. The at least one candidate subtitle template 340 includes subtitle template A, subtitle template C, subtitle template R, subtitle template B, and subtitle template O. In response to a trigger operation on the at least one candidate subtitle template 340, subtitle template C is determined as the target subtitle template. The target video 350 is displayed in the video preview area of the first page 310, and the subtitle special effect 360 corresponding to subtitle template C is displayed in the target video 350. For example, the subtitle special effect 360 includes Figure 3 the clouds and stars in

[0066] Optionally, the first page includes a preview page or a clip page.

[0067] In some embodiments, the displaying of the target video on the first page includes: displaying the target video in the preview page. In response to an editing operation on the special effect material in the target video, the special effect material is edited, where the editing operation includes at least one of a shifting operation, a deleting operation, and a flipping operation.

[0068] Figure 4 Another interaction schematic diagram provided by the embodiments of the present disclosure. As Figure 4 shown, a subtitle style page 400 and a first page 410 are displayed on the screen. The target video 420 and a save control 430 are displayed on the first page 410. In response to a trigger operation on the save control 430, a second page 440 is displayed. The target video 420 is displayed in the second page 440, and the special effect material 450 is displayed in the video frame of the target video 420. In response to a trigger operation on the special effect material 450, a preset control is displayed. The preset control includes a subtitle editing control 460 and a subtitle deleting control. In response to a trigger operation on the subtitle editing control 460, the first page 410 is displayed, the target video 420, subtitle text, and a first control 470 are displayed on the first page 410, and the special effect material 450 is displayed in the video frame of the target video 420. In response to an editing operation on the special effect material 450, at least one of a shifting process, a deleting process, and an inversion process is performed on the special effect material 450 displayed on the first page 410.

[0069] Optionally, a third control is displayed on the preview page, where the third control is used to trigger an editing event for the target video. In response to the triggering operation of the third control, the editing page is displayed, and an editing track is displayed on the editing page, where the editing track includes a text track and / or a sticker track. In response to an editing operation on the text track and / or the sticker track, the caption text in the text track is edited, and / or the sticker effect in the sticker track is edited. Among them, the third control may be a editing control, etc.

[0070] Figure 5 Another interaction schematic diagram provided by an embodiment of the present disclosure. As Figure 5 shown, an editing control 520 is displayed on the preview page 510. In response to the triggering operation of the editing control 520, the editing page 530 is displayed. An editing track is displayed in the editing track area 540 on the editing page 530. The editing track includes a text track and / or a sticker track.

[0071] In some other embodiments, displaying the target video on the first page includes:

[0072] The target video and an editing track are displayed on the editing page, where the editing track includes a special effect track, the special effect track includes a text track and / or a sticker track, the text track includes caption text, and the sticker track includes a sticker effect.

[0073] Figure 6 Another interaction schematic diagram provided by an embodiment of the present disclosure. As Figure 6 shown, the video 620 to be edited, an editing track 630, and a second control 640 are displayed on the editing page 610. The editing track 630 includes a video track. In response to the triggering operation of the second control 640, the preview page 650 is displayed. The video 620 to be edited, caption text, and a first control 660 are displayed on the preview page 650. In response to the triggering operation of the first control 660, the caption style page 670 is displayed. The caption style page 670 includes at least one candidate caption template 680. In response to the triggering operation of at least one candidate caption template 680, the target video 680 and a confirmation control 690 are displayed on the preview page 650. In response to the triggering operation of the confirmation control 690, the editing page 610 is displayed. The target video, the video track, the text track, and the sticker track are displayed on the editing page 610.

[0074] The technical solution of the embodiment of the present disclosure displays an unedited video and a first control on a first page, responds to the triggering operation of the first control, determines at least one candidate subtitle template according to the subtitle text of the unedited video and the annotation information of the subtitle text, displays at least one candidate subtitle template on the subtitle style page, responds to the triggering operation of at least one candidate subtitle template, determines a target subtitle template among at least one candidate subtitle template, generates a target video according to at least one special effect material in the target subtitle template and the unedited video, and displays the target video. Since the embodiment of the present disclosure determines the candidate subtitle template associated with the video subtitle based on the subtitle text and annotation information of the unedited video, thus, a candidate subtitle template with a higher degree of fit with the video subtitle is obtained. By applying the candidate subtitle template, at least one special effect material can be added to the unedited video, which can improve the editing efficiency of the subtitle effect, enrich the subtitle effect of the unedited video, solve the problem of single subtitle form in the related art, and make the video content more vivid through the rich subtitle effect, thereby enriching the video display effect and enhancing the user experience.

[0075] Figure 7 FIG. 4 is a schematic flowchart of another video processing method provided by an embodiment of the present disclosure. On the basis of the above embodiments, the embodiment of the present disclosure further defines the determination method of the candidate subtitle template and the application method of the target subtitle template.

[0076] As Figure 7 shown, the method includes:

[0077] S710. Display an unedited video and a first control on a first page.

[0078] S720. In response to the triggering operation of the first control, determine the target subtitle content in the subtitle text according to the annotation information of the subtitle text, and determine the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library.

[0079] Among them, the target subtitle content represents the text content corresponding to the annotation information in the subtitle text. For example, the annotation information includes at least one of style, price, location, item, style representation, etc. For each subtitle text, extract the text content corresponding to the annotation information in the subtitle text as the target subtitle content of the current subtitle text. The preset style represents the special effect material style of the subtitle candidate template. The preset styles include different style types such as variety show sense, parent-child style, daily, professional, etc.

[0080] For each subtitle text, calculate the similarity between the target subtitle content of the current subtitle text and the special effect materials of a preset style in the special effect material library. Obtain at least one special effect material similar to the target subtitle content from the special effect materials of the preset style as the target special effect material of the current subtitle text. Determine at least one candidate subtitle template of the preset style according to the target special effect materials of each subtitle in the subtitle text. Optionally, each candidate subtitle template of the preset style includes: at least one target special effect material of each subtitle text of the video to be edited. For example, the candidate subtitle template with a variety show sense includes at least one target special effect material with a variety show style of each subtitle text of the video to be edited.

[0081] Exemplarily, determining the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library includes: determining a first feature vector according to the target subtitle content. Determining a second feature vector according to the special effect name and the special effect preview image of the special effect material of the preset style. Determining the target special effect material according to the similarity between the first feature vector and the second feature vector.

[0082] Through the embodiments of the present disclosure, the similarity matching can be performed with the target subtitle content from two aspects of the name and the preview image of the special effect material, so as to obtain a text template and a sticker special effect with a high similarity to the target subtitle content, thereby improving the richness of the special effect materials and avoiding missing the special effect materials with a high similarity to the target subtitle content, and improving the recall rate of the special effect materials.

[0083] For example, perform vectorization processing on the target subtitle content to obtain a first feature vector. Perform vectorization processing on the special effect name of the special effect material of the preset style to obtain a name vector. Determine the description text corresponding to the special effect preview image of the special effect material of the preset style, and perform vectorization processing on the description text to obtain a preview image vector. Concatenate the name vector and the preview image vector to obtain a second feature vector. Determine the target special effect material in the special effect materials of the preset style according to the cosine similarity between the first feature vector and the second feature vector.

[0084] In some embodiments, determine the special effect materials of the preset style whose cosine similarity associated with each subtitle text exceeds a preset similarity threshold, sort the determined special effect materials according to the cosine similarity, and determine the target special effect materials in the special effect materials of the preset style according to the sorting result. For example, in the case of descending order, select n special effect materials from the head of the sorting result of the special effect materials of the preset style as the target special effect materials. In the case of ascending order, select n special effect materials from the tail of the sorting result of the special effect materials of the preset style as the target special effect materials. Wherein, n is a positive integer associated with the number of placeholder identifiers of the candidate subtitle template of the same style. Or, n is a preset positive integer.

[0085] S730. Determine at least one candidate subtitle template according to the target special effect materials of the preset style.

[0086] Exemplarily, divide the n target special effect materials of the preset style into multiple groups according to the number of placeholder identifiers of the candidate subtitle templates of the preset style, and obtain at least one candidate subtitle template of the preset style. For example, according to the number of placeholder identifiers of the candidate subtitle templates of the variety show style, allocate at least one target special effect material of the variety show style corresponding to each subtitle text to each candidate subtitle template of the variety show style, and obtain the candidate subtitle templates corresponding to the number of placeholder identifiers. Each candidate subtitle template includes the target special effect materials of the variety show style corresponding to all subtitle texts of the video to be edited.

[0087] S740. Display a subtitle style page, and display the at least one candidate subtitle template on the subtitle style page, where the candidate subtitle template includes at least one special effect material.

[0088] S750. In response to the trigger operation of the at least one candidate subtitle template, determine the target subtitle template of the preset style.

[0089] Exemplarily, in response to the selection operation of at least one candidate subtitle template, determine the selected candidate subtitle template as the target subtitle template. Convert the target subtitle template to the selected state, and display a preset prompt text at the display position of the video to be edited on the first page to prompt that the target subtitle template is currently being applied to the video to be edited.

[0090] S760. Determine at least one target special effect material corresponding to the subtitle text according to the similarity between the first feature vector corresponding to the subtitle text and the second feature vector corresponding to the target subtitle template.

[0091] Exemplarily, for each subtitle text, determine at least one target special effect material corresponding to the current subtitle text according to the cosine similarity between the first feature vector corresponding to the current subtitle text and the second feature vector of the target special effect material in the target subtitle template.

[0092] For example, for the target special effect materials such as stickers and text templates in the target subtitle template, respectively use the type of target special effect material with the highest cosine similarity as at least one target special effect material corresponding to the current subtitle text. If there are the second feature vectors of the sticker and the text template whose cosine similarities with the first feature vector of the current subtitle text exceed the preset similarity threshold, respectively determine at least one target special effect material of the current subtitle text according to the sticker and the text template with the highest cosine similarity.

[0093] S770. Add at least one target special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

[0094] S780. Display the target video on the first page.

[0095] Based on the annotation information of each subtitle text with annotation information, the technical solution of the embodiment of the present disclosure determines the target subtitle content in the subtitle text, determines the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library, and then determines at least one candidate subtitle template according to the target special effect material of the preset style, and displays at least one candidate subtitle template on the subtitle style page, so as to realize matching the corresponding candidate subtitle template in the special effect material library based on the video content. By determining at least one target special effect material corresponding to the subtitle text according to the similarity between the first feature vector corresponding to the subtitle text and the second feature vector corresponding to the target subtitle template, the adaptability of the special effect material to the video content can be improved, and the situation that the special effect material is separated from the video content can be avoided. Based on the corresponding relationship between the subtitle text and the special effect material, the special effect material is added to the video to be edited to obtain the target video, so as to apply the special effect material of the target subtitle template to the entire video to be edited, and the subtitle editing efficiency is improved.

[0096] Figure 8 It is a schematic flowchart of another video processing method provided by the embodiment of the present disclosure. On the basis of the above embodiments, the embodiment of the present disclosure further limits the determination method of the candidate subtitle template and the application method of the target subtitle template.

[0097] As Figure 8 shown, the method includes:

[0098] S810. Display the video to be edited and the first control on the first page.

[0099] S820. In response to the trigger operation of the first control, determine the subtitle style prediction result according to the video to be edited, the subtitle text and the annotation information.

[0100] Wherein, the subtitle style prediction result includes the subtitle text and the recommended special effect material, and the recommended special effect material includes the recommended special effect name and the recommended special effect preview image. In some embodiments, the subtitle style prediction result includes the subtitle text, the recommended special effect material, and the corresponding relationship between the subtitle text and the recommended special effect material.

[0101] Exemplarily, in response to a triggering operation of the first control, the video to be edited, the subtitle text of the video to be edited, and the annotation information of the subtitle text are input into a pre-trained subtitle style prediction model to obtain the subtitle style prediction result output by the subtitle style prediction model. Among them, the subtitle style prediction model may include a text preprocessing module, a visual encoding module, a vision-language adaptation module, and a large language model. The large language model can be a deep learning model that performs unsupervised training on a large amount of text data to learn natural language patterns, grammatical structures, and semantic relationships. The large language model captures long-range dependencies in the text through the self-attention mechanism. The text preprocessing module is used to convert the text into a sequence and perform vectorization processing. The visual encoding module is used to convert video frames into visual feature vectors. The vision-language adaptation module is used to map the visual feature vectors to the text feature space.

[0102] The training sample set can be determined by video samples whose subtitle special effects meet the expectations. The large language model is fine-tuned using the training sample set so that the large language model can learn the design elements and rules of the special effect materials in the video samples and generate subtitle style prediction results that meet the expectations based on the design elements and rules. For example, the timestamp, subtitle text, and preset prompt words of the video sample are input into the text preprocessing module, and the text preprocessing module performs vectorization processing on the timestamp, subtitle text, and preset prompt words of the video sample to obtain text feature vectors. Among them, the preset prompt words may include subtitle style description information, etc. The video sequence of the video sample is input into the visual encoding module, and the visual encoding module converts the video frames into visual feature vectors and outputs the visual feature vectors to the vision-language adaptation module. The vision-language adaptation module maps the visual feature vectors to the text feature space to obtain the visual feature mapping result. The text feature vectors are input into the large language model through the text preprocessing module, and the visual feature mapping result is input into the large language model through the vision-language adaptation module to obtain the subtitle style prediction result output by the large language model. The prediction loss is calculated according to the subtitle style prediction result and the subtitle special effect information of the video sample. The model parameters of the large language model and the vision-language adaptation module are adjusted according to the prediction loss to complete the fine-tuning training of the subtitle style prediction model.

[0103] Figure 9 This is an architecture diagram of a subtitle style prediction model provided by an embodiment of the present disclosure. As Figure 9As shown, the timestamp, subtitle text, subtitle style description information, and video sequence are determined based on the video sample 910. Among them, the subtitle style description information is associated with the subtitle special effect in the video sample 910. The subtitle style description information, the preset prompt words, the timestamp, and the subtitle text are input into the text preprocessing module 920. The video sequence is input into the visual encoding module 930. The text feature vector is output through the text preprocessing module 920. The visual feature vector is input into the vision-language adaptation module 940 through the visual encoding module 930. The visual feature vector is mapped to the text feature space through the vision-language adaptation module 940 to obtain the visual feature mapping result. The text feature vector and the visual feature mapping result are input into the large language model 950, and the large language model learns the design elements and rules of the special effect materials in the video sample to generate the subtitle style prediction result of the video sample.

[0104] S830. Determine the third feature vector corresponding to the recommended special effect material according to the recommended special effect name and the recommended special effect preview image.

[0105] S840. Determine the fourth feature vector corresponding to the special effect material according to the special effect name and the special effect preview image of the special effect materials in the special effect material library.

[0106] Exemplarily, the recommended special effect name, the recommended special effect preview image, the recommended name and the recommended preview image of the special effect materials in the special effect material library are input into the pre-trained special effect mapping model. Through the special effect mapping model, the recommended special effect name, the recommended special effect preview image, the special effect name and the special effect preview image of the special effect materials in the special effect material library are mapped to the same vector. The feature vectors of the recommended feature name and the recommended special effect preview image in the same vector are concatenated to obtain the third feature vector corresponding to the recommended special effect material. The feature vectors of the special effect name and the special effect preview image of the special effect material are concatenated to obtain the fourth feature vector corresponding to the special effect material.

[0107] In some embodiments, the special effect mapping model is a multi-modal model trained based on pairs of special effect materials with similar special effects. For example, pairs of special effect materials with similar special effects are input into the special effect mapping model to be trained, and the similarity prediction result of the pairs of special effect materials is output through the special effect mapping model. Then, the contrast loss is determined according to the true value and the prediction result of the similarity of the pairs of special effect materials, and the special effect mapping model is trained according to the contrast loss.

[0108] By mapping the recommended special effect material to the special effect material in the special effect material library that the subtitle style prediction model has not been trained on through the special effect mapping model, the subtitle style prediction ability of the subtitle style prediction model can be transferred to the special effect material library that has not participated in the training of this model, improving the generalization ability of the subtitle style prediction result.

[0109] S850. In response to the similarity between the third feature vector and the fourth feature vector meeting a preset similarity condition, update the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector to obtain an updated subtitle style prediction result.

[0110] Exemplarily, the updating the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector includes: replacing the special effect material corresponding to the third feature vector in the subtitle style prediction result with the special effect material corresponding to the fourth feature vector.

[0111] For example, according to the similarity between the third feature vector and the fourth feature vector, determine the fourth feature vector associated with the maximum similarity in a preset special effect material library. Replace the recommended special effect material corresponding to the third feature vector with the special effect material corresponding to the fourth feature vector associated with the maximum similarity.

[0112] S860. Determine the candidate subtitle template according to the updated subtitle style prediction result.

[0113] Exemplarily, determine the candidate subtitle template according to the special effect material in the updated subtitle style prediction result.

[0114] S870. In response to a trigger operation on the at least one candidate subtitle template, determine the candidate subtitle template associated with the updated subtitle style prediction result as the target subtitle template.

[0115] S880. Determine at least one special effect material corresponding to the subtitle text according to the updated subtitle style prediction result.

[0116] Since the subtitle style prediction result includes the correspondence between the subtitle text and the recommended special effect material, and the subtitle style prediction result is updated by updating the special effect material, the updated subtitle style prediction result also includes the correspondence between the subtitle text and the special effect material. Based on the correspondence between the subtitle text and the special effect material in the updated subtitle style prediction result, determine at least one special effect material corresponding to the subtitle text.

[0117] S890. Add at least one special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

[0118] S8100. Display the target video on the first page.

[0119] The technical solution of the embodiment of the present disclosure determines a subtitle style prediction result based on the video to be edited, the subtitle text, and the annotation information, maps the recommended special effect materials in the subtitle style prediction result to similar special effect materials in the material library, and obtains an updated subtitle style prediction result, which improves the generalization ability of the subtitle style prediction result. Determine at least one candidate subtitle template according to the updated subtitle style prediction result. In response to the triggering operation of at least one candidate subtitle template, determine the candidate subtitle template associated with the updated subtitle style prediction result as the target subtitle template, and based on the corresponding relationship between the subtitle text and the special effect material in the updated subtitle style prediction result, add the special effect material to the video to be edited to obtain the target video, so as to apply the special effect material of the target subtitle template to the entire video to be edited, improving the subtitle editing efficiency.

[0120] Figure 10 FIG. 4 is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a mobile terminal, a PC, or a server, etc.

[0121] As Figure 10 shown, the device includes: a first display module 1010, a second display module 1020, and a video generation module 1030.

[0122] The first display module 1010 is configured to display the video to be edited and a first control on a first page;

[0123] The second display module 1020 is configured to, in response to the triggering operation of the first control, determine at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, display a subtitle style page, and display the at least one candidate subtitle template on the subtitle style page, where the candidate subtitle template includes at least one special effect material;

[0124] The video generation module 1030 is configured to, in response to the triggering operation of the at least one candidate subtitle template, determine a target subtitle template in the at least one candidate subtitle template, generate a target video according to the at least one special effect material in the target subtitle template and the video to be edited, and display the target video on the first page.

[0125] Optionally, the determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text includes:

[0126] Determine the target subtitle content in the subtitle text according to the annotation information of the subtitle text, and determine the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library;

[0127] Determine at least one candidate subtitle template according to the target special effect material of the preset style.

[0128] Optionally, the determining the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library includes:

[0129] Determine a first feature vector according to the target subtitle content;

[0130] Determine a second feature vector according to the special effect name and special effect preview image of the special effect material of the preset style;

[0131] Determine the target special effect material according to the similarity between the first feature vector and the second feature vector.

[0132] Optionally, the determining the target subtitle template in the at least one candidate subtitle template in response to the triggering operation of the at least one candidate subtitle template, and generating a target video according to at least one special effect material in the target subtitle template and the video to be edited includes:

[0133] In response to the triggering operation of the at least one candidate subtitle template, determine the target subtitle template of the preset style;

[0134] Determine at least one target special effect material corresponding to the subtitle text according to the similarity between the first feature vector corresponding to the subtitle text and the second feature vector corresponding to the target subtitle template;

[0135] Add at least one target special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

[0136] Optionally, the determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text includes:

[0137] Determine a subtitle style prediction result according to the video to be edited, the subtitle text and the annotation information, where the subtitle style prediction result includes the subtitle text and recommended special effect materials, and the recommended special effect materials include a recommended special effect name and a recommended special effect preview image;

[0138] Determine a third feature vector corresponding to the recommended special effect material according to the recommended special effect name and the recommended special effect preview image;

[0139] Determine a fourth feature vector corresponding to the special effect material according to the special effect name and the special effect preview image of the special effect material in the special effect material library;

[0140] In response to the similarity between the third feature vector and the fourth feature vector meeting a preset similarity condition, update the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector to obtain an updated subtitle style prediction result;

[0141] Determine the candidate subtitle template according to the updated subtitle style prediction result.

[0142] Optionally, the updating the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector includes:

[0143] Replace the special effect material corresponding to the third feature vector in the subtitle style prediction result with the special effect material corresponding to the fourth feature vector.

[0144] Optionally, the determining the target subtitle template in the at least one candidate subtitle template and generating a target video according to at least one special effect material in the target subtitle template and the video to be edited in response to the triggering operation of the at least one candidate subtitle template includes:

[0145] In response to the triggering operation of the at least one candidate subtitle template, determine the candidate subtitle template associated with the updated subtitle style prediction result as the target subtitle template;

[0146] According to the updated subtitle style prediction result, determine at least one special effect material corresponding to the subtitle text;

[0147] Add at least one special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

[0148] Optionally, the displaying the video to be edited and the first control on the first page includes:

[0149] Display the first control, the video to be edited and the subtitle of the video to be edited on the first page, wherein the subtitle of the video to be edited is associated with the subtitle text.

[0150] Optionally, the displaying the first control, the video to be edited and the subtitle of the video to be edited on the first page includes:

[0151] Display the video to be edited and the second control on the second page, wherein the second control is used to trigger an event of adding subtitles to the video to be edited;

[0152] In response to the triggering operation of the second control, determine the subtitle text according to the audio feature of the video to be edited, and perform annotation on the subtitle text to obtain the annotation information;

[0153] Display the first page, and display the first control, the video to be edited, and the subtitles of the video to be edited on the first page.

[0154] Optionally, the first page includes a preview page or a clip page.

[0155] Optionally, displaying the target video on the first page includes:

[0156] Display the target video on the preview page;

[0157] In response to an editing operation on the special effect material in the target video, perform an editing process on the special effect material, where the editing operation includes at least one of a shifting operation, a deleting operation, and a flipping operation.

[0158] Optionally, it further includes:

[0159] Display a third control on the preview page, where the third control is used to trigger an editing event for the target video;

[0160] In response to a triggering operation of the third control, display the clip page, and display a clip track on the clip page, where the clip track includes a text track and / or a sticker track;

[0161] In response to an editing operation on the text track and / or the sticker track, perform an editing process on the subtitle text in the text track and / or perform an editing process on the sticker special effect in the sticker track.

[0162] Optionally, displaying the target video on the first page includes:

[0163] Display the target video and the clip track on the clip page, where the clip track includes a special effect track, the special effect track includes a text track and / or a sticker track, the text track includes subtitle text, and the sticker track includes a sticker special effect.

[0164] The video processing device provided by the embodiments of the present disclosure can execute the video processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0165] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0166] Figure 11A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The following refers to Figure 11 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server in Figure 11 ) 1100 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0167] As Figure 11 shown, the electronic device 1100 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1101, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage device 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An editing / output (I / O) interface 1105 is also connected to the bus 1104.

[0168] Generally, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 shows the electronic device 1100 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.

[0169] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0170] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0171] The electronic device provided in the embodiment of the present disclosure and the video processing method provided in the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0172] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the video processing method provided in the above embodiment is implemented.

[0173] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0174] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0175] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0176] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:

[0177] display the video to be edited and the first control on the first page;

[0178] In response to a triggering operation on the first control, at least one candidate subtitle template is determined according to the subtitle text of the video to be edited and the annotation information of the subtitle text, a subtitle style page is displayed, and the at least one candidate subtitle template is displayed on the subtitle style page, where the candidate subtitle template includes at least one special effect material;

[0179] In response to a triggering operation on the at least one candidate subtitle template, a target subtitle template in the at least one candidate subtitle template is determined, and a target video is generated according to the at least one special effect material in the target subtitle template and the video to be edited, and the target video is displayed on the first page.

[0180] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0182] The units involved in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of a unit does not constitute a limitation on the unit itself.

[0183] The functions described above herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0184] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] The above description is only a preferred embodiment of the present disclosure and an illustration of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.

[0186] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0187] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video processing method, characterized in that, Including: Displaying a video to be edited and a first control on a first page; In response to a trigger operation of the first control, determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, displaying a subtitle style page, and displaying the at least one candidate subtitle template on the subtitle style page, wherein the candidate subtitle template includes at least one special effect material; In response to a trigger operation of the at least one candidate subtitle template, determining a target subtitle template among the at least one candidate subtitle template, generating a target video according to the at least one special effect material in the target subtitle template and the video to be edited, and displaying the target video on the first page.

2. The method according to claim 1, wherein The step of determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text includes: Determining target subtitle content in the subtitle text according to the annotation information of the subtitle text, and determining a target special effect material of a preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library; Determining at least one candidate subtitle template according to the target special effect material of the preset style.

3. The method according to claim 2, wherein The step of determining the target special effect material of the preset style according to the similarity between the target subtitle content and the special effect materials of the preset style in the special effect material library includes: Determining a first feature vector according to the target subtitle content; Determining a second feature vector according to the special effect name and the special effect preview image of the special effect material of the preset style; Determining the target special effect material according to the similarity between the first feature vector and the second feature vector.

4. The method according to claim 3, characterized in that The step of, in response to a trigger operation of the at least one candidate subtitle template, determining a target subtitle template among the at least one candidate subtitle template, and generating a target video according to the at least one special effect material in the target subtitle template and the video to be edited includes: In response to a trigger operation of the at least one candidate subtitle template, determining a target subtitle template of the preset style; Determining at least one target special effect material corresponding to the subtitle text according to the similarity between the first feature vector corresponding to the subtitle text and the second feature vector corresponding to the target subtitle template; Adding at least one target special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

5. The method according to claim 1, wherein The step of determining at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text includes: Determining a subtitle style prediction result according to the video to be edited, the subtitle text and the annotation information, wherein the subtitle style prediction result includes the subtitle text and recommended special effect materials, and the recommended special effect materials include a recommended special effect name and a recommended special effect preview image; Determining a third feature vector corresponding to the recommended special effect material according to the recommended special effect name and the recommended special effect preview image; Determining a fourth feature vector corresponding to the special effect material in the special effect material library according to the special effect name and the special effect preview image of the special effect material; In response to the similarity between the third feature vector and the fourth feature vector meeting a preset similarity condition, update the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector to obtain an updated subtitle style prediction result; Determine the candidate subtitle template according to the updated subtitle style prediction result.

6. The method according to claim 5, wherein The updating the subtitle style prediction result based on the special effect material corresponding to the fourth feature vector includes: Replacing the special effect material corresponding to the third feature vector in the subtitle style prediction result with the special effect material corresponding to the fourth feature vector.

7. The method according to claim 5, wherein The determining the target subtitle template in the at least one candidate subtitle template and generating a target video according to at least one special effect material in the target subtitle template and the video to be edited in response to a triggering operation on the at least one candidate subtitle template includes: In response to a triggering operation on the at least one candidate subtitle template, determining the candidate subtitle template associated with the updated subtitle style prediction result as the target subtitle template; Determining at least one special effect material corresponding to the subtitle text according to the updated subtitle style prediction result; Adding at least one special effect material corresponding to the subtitle text to the video to be edited to obtain the target video.

8. The method according to claim 1, characterized in that, The displaying the video to be edited and a first control on the first page includes: Displaying the first control, the video to be edited, and the subtitle of the video to be edited on the first page, wherein the subtitle of the video to be edited is associated with the subtitle text.

9. The method according to claim 8, wherein The displaying the first control, the video to be edited, and the subtitle of the video to be edited on the first page includes: Displaying the video to be edited and a second control on a second page, wherein the second control is used to trigger an event of adding a subtitle to the video to be edited; In response to a triggering operation on the second control, determining the subtitle text according to the audio feature of the video to be edited, and obtaining annotation information by annotating the subtitle text; Displaying the first page, and displaying the first control, the video to be edited, and the subtitle of the video to be edited on the first page.

10. The method according to claim 1, characterized in that, The first page includes a preview page or a clip page.

11. The method according to claim 10, characterized in that, The displaying the target video on the first page includes: Displaying the target video on the preview page; In response to an editing operation on a special effect material in the target video, performing an editing process on the special effect material, wherein the editing operation includes at least one of a shifting operation, a deleting operation, and a flipping operation.

12. The method according to claim 10, wherein Further includes: Displaying a third control on the preview page, wherein the third control is used to trigger an editing event for the target video; In response to a triggering operation on the third control, displaying the clip page, and displaying a clip track on the clip page, wherein the clip track includes a text track and / or a sticker track; In response to an editing operation on the text track and / or the sticker track, performing an editing process on the subtitle text in the text track and / or performing an editing process on the sticker effect in the sticker track.

13. The method according to claim 10, characterized in that, The displaying the target video on the first page includes: The target video and a clip track are displayed on the clip page, wherein the clip track includes a special effect track, the special effect track includes a text track and / or a sticker track, the text track includes subtitle text, and the sticker track includes sticker special effects.

14. A video processing device, characterized in that, It includes: A first display module, configured to display a video to be edited and a first control on a first page; A second display module, configured to, in response to a trigger operation of the first control, determine at least one candidate subtitle template according to the subtitle text of the video to be edited and the annotation information of the subtitle text, display a subtitle style page, and display the at least one candidate subtitle template on the subtitle style page, wherein the candidate subtitle template includes at least one special effect material; A video generation module, configured to, in response to a trigger operation of the at least one candidate subtitle template, determine a target subtitle template among the at least one candidate subtitle templates, generate a target video according to the at least one special effect material in the target subtitle template and the video to be edited, and display the target video on the first page.

15. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of claims 1-13.

16. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the video processing method according to any one of claims 1-13 when executed by a computer processor.

17. A computer program product comprising a computer program, characterized in that, The computer program implements the video processing method according to any one of claims 1-13 when executed by a processor.