Video editing method and apparatus, device, and storage medium

By identifying target audio segments from video audio signals, matching and adding visual materials, the problem of low material collection efficiency in traditional video editing is solved, achieving intelligent material matching and efficient editing.

WO2026157444A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-11-12
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

In traditional video editing, users need to collect and add materials themselves, resulting in low editing efficiency.

Method used

The system identifies target audio segments from the audio signals of video content, matches visual materials with the text content of the target audio segments, adds corresponding visual materials in the editing interface, and generates or selects suitable materials from the material library using a generative model.

Benefits of technology

It enables intelligent matching and efficient addition of materials, improving the efficiency of media editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025134362_30072026_PF_FP_ABST
    Figure CN2025134362_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video editing method and apparatus, a device, and a computer-readable storage medium. The method comprises: presenting an editing interface for video content, the editing interface comprising an editing control; in response to selection of the editing control, determining a target audio segment from an audio signal of the video content; and in the editing interface, adding a visual material to at least one image frame of the video content, the at least one image frame being determined on the basis of the target audio segment, and the visual material being determined on the basis of text content corresponding to the target audio segment.
Need to check novelty before this filing date? Find Prior Art

Description

Video editing methods, apparatus, devices and storage media

[0001] This application claims priority to Chinese Patent Application No. 202510127621.4, filed on January 27, 2025, entitled "Video Editing Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to video editing methods, apparatus, devices, and computer-readable storage media. Background Technology

[0003] With the development of computer technology, people can share or access media content through internet platforms. Media editing tools have also gradually become important tools in people's lives; for example, people can use media editing tools to edit video content and share it through media platforms. Summary of the Invention

[0004] In a first aspect of this disclosure, a video editing method is provided. The method includes: presenting an editing interface of video content, the editing interface including editing controls; determining a target audio segment from an audio signal of the video content in response to selection of the editing controls; and adding visual material to at least one image frame of the video content in the editing interface, the at least one image frame being determined based on the target audio segment, and the visual material being determined based on text content corresponding to the target audio segment.

[0005] In a second aspect of this disclosure, an apparatus for video editing is provided. The apparatus includes: a presentation module configured to present an editing interface of video content, the editing interface including editing controls; a determination module configured to determine a target audio segment from an audio signal of the video content in response to a selection of the editing controls; and an adding module configured to add visual material to at least one image frame of the video content in the editing interface, the at least one image frame being determined based on the target audio segment, and the visual material being determined based on text content corresponding to the target audio segment.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figures 2A to 2C illustrate example interfaces according to some embodiments of the present disclosure;

[0013] Figure 3 illustrates a flowchart of an example process for video editing according to some embodiments of the present disclosure;

[0014] Figure 4 shows a schematic structural block diagram of an example device for video editing according to some embodiments of the present disclosure; and

[0015] Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0020] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0021] As mentioned above, media editing tools have gradually become important tools in people's lives. For example, people can use media editing tools to edit video content and share it through media platforms.

[0022] In the traditional editing process, people usually need to collect the necessary materials themselves and add them to the required places, which greatly affects the efficiency of media editing.

[0023] Embodiments of this disclosure propose a video editing scheme. The scheme includes: presenting an editing interface for video content, the editing interface including editing controls; determining a target audio segment from an audio signal of the video content in response to selection of the editing controls; and adding visual material to at least one image frame of the video content in the editing interface, the at least one image frame being determined based on the target audio segment, and the visual material being determined based on text content corresponding to the target audio segment.

[0024] In this way, embodiments of this disclosure can match corresponding materials (also known as B-roll materials) based on the narration of the video, and can add the corresponding elements to the video, thereby achieving intelligent material matching. Therefore, embodiments of this disclosure can improve the efficiency of media editing.

[0025] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0026] Example Environment

[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.

[0028] In this example environment 100, electronic device 110 may run a video editing application 120. Application 120 may be any suitable type of application for video editing, examples of which may include, but are not limited to, media editing applications or other suitable applications that provide media editing services. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.

[0029] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present a video editing interface 150.

[0030] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0031] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide background services for video editing applications 120 in electronic devices 110.

[0032] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.

[0033] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0034] The various example implementations of this disclosure will be described in detail below.

[0035] Example Interaction

[0036] Figures 2A to 2C illustrate example interfaces 200A to 200C according to some embodiments of the present disclosure. Interfaces 200A to 200C may be provided, for example, by the electronic device 110 shown in Figure 1.

[0037] As shown in Figure 2A, Figure 2A illustrates an example editing interface 200A according to some embodiments of the present disclosure. As an example, the editing interface 200A may also be referred to as a multi-track editing interface to support users in editing different tracks of a video draft, such as video tracks, audio tracks, effects tracks, etc.

[0038] As shown in the figure, the editing interface 200A can correspond to the video content 205. In some embodiments, the electronic device 110 can provide an editing panel 210 in the editing interface 200A, which can provide a set of controls for intelligently matching materials.

[0039] As shown in Figure 2A, the editing panel 210 can present a set of candidate styles 215, which can correspond to the packaging style of the matched material. For example, such packaging styles can correspond to different themes of video content, such as variety show style or technology style. Alternatively, such packaging styles can correspond to different visual effects, such as hand-drawn style.

[0040] Furthermore, users can access more configuration controls, as shown in Figure 2B, by clicking entry point 220. As shown in Figure 2B, the editing panel 210 can provide configuration controls 230 and 235.

[0041] As an example, the configuration control 230 can be used to configure the source of the acquired materials. For instance, the electronic device 110 can receive selections on the configuration control 230 and can present a set of preset material libraries. As an example, such a material library may include an image library, a video library, a chart library, an emoticon library, etc.

[0042] Additionally, the electronic device 110 can, for example, determine whether it supports generating new material using generative models through configuration control 230.

[0043] As another example, configuration control 235 can be used to configure whether to support obtaining materials from a commercial material library. In some scenarios, obtaining materials from a commercial material library requires consuming corresponding virtual resources.

[0044] Furthermore, the electronic device 110 can receive a user's selection of an editing control (e.g., button 225 or button 240) to trigger the acquisition of footage matching the video content. In some scenarios, the acquired footage may also be referred to as B-roll footage, or auxiliary footage, which can be inserted into A-roll footage (i.e., the video content) to provide background information or visual effects.

[0045] As shown in Figure 2C, the electronic device 110 can add visual material 245 to at least one image frame of the video content 205. This visual material 245 can be determined based on the text content corresponding to a target audio segment in the audio signal of the video content.

[0046] In some embodiments, the electronic device 110 can also add a material track corresponding to the visual material 245 in the editing interface 200C to support further editing operations on the visual material 245.

[0047] As an example, the electronic device 110 can provide a replacement entry corresponding to the visual material 245 to support users in further replacing the added visual material 245.

[0048] As another example, the electronic device 110 can also provide a rematch entry. Accordingly, the electronic device 110 can receive a selection of the rematch entry, thereby triggering the rematching and addition of new visual material based on the audio signal of the video content.

[0049] In some scenarios, the added visual material 245 can include various types. As an example, visual material 245 can include image material, which can be inserted into an appropriate position within an image frame. Furthermore, image material can be associated with a preset display hierarchy, for example, it can be overlaid above the main subject in the image frame or displayed below the main subject.

[0050] In other examples, visual material 245 may also include video material, which may be added to an image frame, for example, as picture-in-picture material.

[0051] In some other examples, visual material 245 may also include graphic material, such as icons that may correspond to spoken text and may be inserted in the appropriate position of an image frame to display the corresponding information in a visual style.

[0052] In some other examples, visual material 245 may also include emoticon material. Such emoticon material may include, for example, static or animated emoticons.

[0053] The following section will further describe the process of determining and encapsulating visual material 245.

[0054] In some embodiments, upon receiving a trigger on the editing control, a target audio segment for matching material can be determined from the audio signal of the video content. Specifically, the electronic device 110 or the server 130 can determine the reference text corresponding to the audio signal, also known as the spoken text.

[0055] Furthermore, the electronic device 110 or server 130 can provide reference text to the model to determine text fragments of the material to be matched from the reference text. For example, the electronic device 110 or server 130 can provide spoken text to the language model to support the determination of text fragments suitable for matching B-roll material from the spoken text.

[0056] As an example, a language model can determine the text fragments used for matching material from spoken text and can generate one or more search terms, also known as search terms, that correspond to the text fragments.

[0057] Furthermore, the target audio segment corresponding to the text segment can be determined based on the text segment of the material to be matched.

[0058] Additionally, server 130 may determine the corresponding visual material, for example, from the text fragment (i.e., the text content of the target audio fragment). In some embodiments, server 130 may determine one or more visual materials that match the text content from a material library.

[0059] In some embodiments, server 130 may retrieve one or more search terms corresponding to the text content. For example, such search terms may be automatically generated by a language model based on spoken text. Further, server 130 may search for visual materials from a material library based on the one or more search terms.

[0060] In some embodiments, the media library used for matching media may include a preset media library. Alternatively, such a media library may be determined based on user configuration information. For example, the media library may correspond to media sources configured by the user through configuration control 230.

[0061] In some embodiments, server 130 can also generate corresponding visual materials based on text content. Specifically, server 130 can, for example, provide a prompt (also called a hint) determined based on the text content to the material generation model to generate visual materials. As an example, such a material generation model can include an appropriate generative model, such as an image generation model, a video generation model, an expression generation model, a chart generation model, etc. Accordingly, the prompt can include, for example, prompt words provided to the generative model.

[0062] As an example, when the user specifies through the configuration control 230 that generative models can be used to generate materials, the server 130 can construct prompts for generating materials based on text content and preset prompt templates, and can obtain the generated visual materials.

[0063] In some embodiments, server 130 may also populate a material template with a set of information determined based on the text content to generate visual material. Taking chart material as an example, server 130 may determine a chart template (i.e., material template) suitable for displaying the text content, and may extract corresponding information from the text content to populate the chart template.

[0064] In some examples, the material template can be associated with a preset material style. Taking icon materials as an example, such a material style could include, for instance, a bar chart. Furthermore, the material template can include one or more parameters to be determined, and the values ​​of these parameters can be determined using information from the text content. For example, the data dimensions included in the bar chart and the corresponding values ​​for each data dimension can be determined from the text content.

[0065] As an example, if the text content includes "Sales volume in 2021 was X, and sales volume in 2022 was Y", server 130 can determine that the text content is suitable for display using chart materials. Furthermore, server 130 can obtain a chart template suitable for displaying a comparison of data from a chart template library, and can fill in "2021", "X", "2022", and "Y" into such a chart template to obtain the corresponding chart material.

[0066] In some embodiments, the type of visual material being matched may be determined based on user configuration information. For example, a user may specify through the editing panel 210 that the type of visual material they wish to acquire includes only image materials.

[0067] After obtaining the corresponding visual materials, the following will further introduce the specific process of creating packaging visual materials.

[0068] Specifically, the electronic device 110 or the server 130 can determine the effect (also known as the wrapping effect) corresponding to the visual material and can add the visual material with the applied effect to at least one video frame.

[0069] In some embodiments, the electronic device 110 or the server 130 can obtain the style parameters selected by the user and determine the effect applied to the visual material based on the style parameters.

[0070] Alternatively or additionally, the electronic device or server 130 may also determine the effect corresponding to the content of the visual material from a set of candidate effects.

[0071] In some embodiments, the effect applied to the visual material can indicate the display parameters of the visual material in at least one image frame. For example, such display parameters can indicate the placement, size, and angle of the visual material within the image frame.

[0072] In other embodiments, the effect applied to the visual material can indicate its hierarchical relationship with at least one image frame. For example, whether such a visual material is displayed above an image frame or as a lower layer than the main object in the image frame.

[0073] In other embodiments, the effects applied to the visual material can indicate a dynamic process of the visual material. For example, such a dynamic process can include various appropriate motion effects, such as zooming in, zooming out, blinking, lighting, entrance animations, exit animations, etc.

[0074] In some other embodiments, the effect corresponding to the visual material can indicate the image processing operation applied to the visual material. As an example, such image processing operation may include, but is not limited to: extracting the region corresponding to a preset object from the visual material, adding a border related to the visual material, applying a preset filter, etc.

[0075] As an example, when the text content includes a product name or brand name, and the visual material includes a product image or brand icon, server 130 can perform image cutout processing on the visual material to extract the area corresponding to the product or brand from the visual material. Furthermore, server 130 can add a preset border to the extracted area.

[0076] In some embodiments, the packaging effect corresponding to the visual material can also indicate the application of preset editing operations to one or more image frames of the video content 205. That is, such an effect can include the combined effect of A-roll and B-roll materials.

[0077] As an example, such editing operations could include shrinking the A-roll footage (i.e., the video frame content) and displaying the added B-roll footage in other areas, thereby shifting the focus of information from the A-roll footage to the B-roll footage.

[0078] Alternatively or additionally, the editing operations applied to A-roll content can also be related to the proportions of B-roll content. For example, if A-roll and B-roll content have the same proportions, the wrapping effects that can be applied may include, but are not limited to: covering A-roll content with B-roll content; displaying B-roll content in a picture-in-picture style; extracting the main object from A-roll content and overlaying it onto B-roll content, etc.

[0079] If the A-roll and B-roll assets have the same aspect ratio, the wrapping effects that can be applied can include displaying the A-roll and B-roll assets in a split-screen format, or displaying the B-roll assets in a picture-in-picture style.

[0080] In this way, the embodiments of this disclosure can provide a smoother material insertion effect and improve the efficiency of users obtaining information through B-roll materials.

[0081] In some embodiments, as shown in FIG2C, the electronic device 110 may also add subtitles 250 corresponding to the text content to the video content 205. This allows for a more prominent emphasis on the spoken text matching the B-roll material.

[0082] In some embodiments, if the video content 205 already includes subtitle material, the electronic device 110 can first remove the subtitle portion corresponding to the matched spoken text from the subtitle material, and then add subtitles corresponding to the spoken text in a preset subtitle style.

[0083] In this way, embodiments of this disclosure can match corresponding materials (also known as B-roll materials) based on the narration of the video, and can add the corresponding elements to the video, thereby achieving intelligent material matching. Therefore, embodiments of this disclosure can improve the efficiency of media editing.

[0084] In some embodiments, the addition of visual material 245 can also be performed based on cloud draft editing. Specifically, electronic device 110 can upload structured data (e.g., JSON data) corresponding to the local draft and the corresponding material resources to server 130.

[0085] Furthermore, server 130 can utilize cloud draft editing capabilities to add determined visual materials to the cloud draft, thereby obtaining an updated edited draft. After completing the draft editing, server 130 can send the updated draft to electronic device 110, thereby completing the matching and addition of visual materials. In this way, embodiments of this disclosure can reduce the local computing cost of electronic devices and support the same material addition logic on different platforms (e.g., mobile platforms or PC platforms).

[0086] Example process

[0087] Figure 3 shows a flowchart of an example process 300 for video editing according to some embodiments of the present disclosure. Process 300 can be implemented at electronic device 110. Process 300 will now be described with reference to Figure 1.

[0088] As shown in Figure 3, in box 310, the electronic device 110 presents an editing interface for video content, which includes editing controls.

[0089] In box 320, electronic device 110, in response to the selection of an editing control, determines a target audio segment from the audio signal of the video content.

[0090] In frame 330, electronic device 110 adds visual material to at least one image frame of video content in the editing interface. The at least one image frame is determined based on a target audio segment, and the visual material is determined based on the text content corresponding to the target audio segment.

[0091] In some embodiments, process 300 further includes: determining reference text corresponding to the audio signal; providing the reference text to the model to determine a text segment of the material to be matched from the reference text; and determining a target audio segment corresponding to the text segment.

[0092] In some embodiments, process 300 further includes: determining visual materials from a material library that match the text content based on the text content.

[0093] In some embodiments, process 300 further includes: determining at least one search term corresponding to the text content; and searching for visual materials from a material library based on the at least one search term.

[0094] In some embodiments, process 300 further includes: providing prompts to a material generation model to generate visual material, the prompts being determined based on text content; or filling a material template with a set of information determined based on text content to generate visual material, wherein the material template is determined based on text content.

[0095] In some embodiments, the visual material is also determined based on user configuration information, which indicates at least one of the following: the type of visual material to be acquired; and the source for acquiring the visual material.

[0096] In some embodiments, process 300 further includes: determining an effect corresponding to visual material; and adding visual material to at least one video frame to apply the effect.

[0097] In some embodiments, process 300 further includes: obtaining the style parameter selected by the user; and determining the effect corresponding to the style parameter.

[0098] In some embodiments, process 300 further includes: determining an effect corresponding to the visual content from a set of candidate effects based on the visual content.

[0099] In some embodiments, the effect indicates at least one of the following: the display parameters of the visual material in at least one image frame; the hierarchical relationship between the visual material and at least one image frame; the dynamic process of the visual material; and the image processing operation applied to the visual material.

[0100] In some embodiments, the effect also indicates that a preset editing operation is applied to one or more image frames of the video content.

[0101] In some embodiments, process 300 further includes adding subtitles to the video content that correspond to the text content.

[0102] In some embodiments, before adding a subtitle corresponding to the text content to at least one image frame, process 300 further includes: in response to the video content including subtitle material, removing the subtitle portion corresponding to the text content from the subtitle material.

[0103] In some embodiments, visual materials include at least one of the following: image materials; video materials; chart materials; emoticon materials.

[0104] Example devices and equipment

[0105] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an example apparatus 400 for video editing according to certain embodiments of this disclosure. Apparatus 400 may be implemented as or included in electronic device 110. The various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0106] As shown in Figure 4, the device 400 includes: a presentation module 410 configured to present an editing interface for video content, the editing interface including editing controls; a determination module 420 configured to determine a target audio segment from an audio signal of the video content in response to a selection of the editing controls; and an adding module 430 configured to add visual material to at least one image frame of the video content in the editing interface, the at least one image frame being determined based on the target audio segment, and the visual material being determined based on text content corresponding to the target audio segment.

[0107] In some embodiments, the apparatus 400 further includes a text determination module configured to: determine reference text corresponding to an audio signal; provide the reference text to a model to determine a text segment of material to be matched from the reference text; and determine a target audio segment corresponding to the text segment.

[0108] In some embodiments, the apparatus 400 further includes a material determination module configured to: determine visual materials from a material library that match the text content based on the text content.

[0109] In some embodiments, the apparatus 400 further includes a search term determination module configured to: determine at least one search term corresponding to the text content; and search for visual materials from a material library based on the at least one search term.

[0110] In some embodiments, the apparatus 400 further includes a providing module configured to: provide prompting information to a material generation model to generate visual material, the prompting information being determined based on text content; or fill a material template with a set of information determined based on text content to generate visual material, wherein the material template is determined based on text content.

[0111] In some embodiments, the visual material is also determined based on user configuration information, which indicates at least one of the following: the type of visual material to be acquired; and the source for acquiring the visual material.

[0112] In some embodiments, the apparatus 400 further includes an effect determination module configured to: determine an effect corresponding to visual material; and add visual material to at least one video frame to apply the effect.

[0113] In some embodiments, the device 400 further includes a parameter acquisition module, configured to: acquire style parameters selected by the user; and determine the effect corresponding to the style parameters.

[0114] In some embodiments, the apparatus 400 further includes an effect determination module configured to: determine an effect corresponding to the visual content from a set of candidate effects based on the visual content.

[0115] In some embodiments, the effect indicates at least one of the following: the display parameters of the visual material in at least one image frame; the hierarchical relationship between the visual material and at least one image frame; the dynamic process of the visual material; and the image processing operation applied to the visual material.

[0116] In some embodiments, the effect also indicates that a preset editing operation is applied to one or more image frames of the video content.

[0117] In some embodiments, the device 400 further includes a subtitle adding module configured to add subtitles corresponding to the text content to the video content.

[0118] In some embodiments, the apparatus 400 further includes a subtitle removal module configured to remove subtitle portions corresponding to text content from the subtitle material in response to the video content including subtitle material.

[0119] In some embodiments, visual materials include at least one of the following: image materials; video materials; chart materials; emoticon materials.

[0120] As shown in Figure 5, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors 510 or processing units, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0121] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0122] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0123] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0124] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0125] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0126] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0127] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0128] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0130] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A video editing method, comprising: An editing interface for presenting video content, the editing interface including editing controls; In response to a selection of the editing control, a target audio segment is determined from the audio signal of the video content; as well as In the editing interface, visual material is added to at least one image frame of the video content. The at least one image frame is determined based on the target audio segment, and the visual material is determined based on the text content corresponding to the target audio segment.

2. The method of claim 1, wherein determining the target audio segment from the audio signal of the video content comprises: Determine the reference text corresponding to the audio signal; Provide the reference text to the model to determine the text fragments of the material to be matched from the reference text; and Identify the target audio segment corresponding to the text segment.

3. The method according to claim 1 or 2, wherein the visual material is determined based on the following process: Based on the text content, visual materials that match the text content are determined from the material library.

4. The method according to claim 3, wherein determining the visual material matching the text content from the material library based on the text content comprises: Determine at least one search term corresponding to the text content; as well as Based on the at least one search term, the visual material is searched from the material library.

5. The method according to any one of claims 1 to 4, wherein the visual material is determined based on the following process: Provide prompts to the material generation model to generate the visual material, wherein the prompts are determined based on the text content; or The visual material is generated by filling a material template with a set of information determined based on the text content, wherein the material template is determined based on the text content.

6. The method according to any one of claims 1 to 5, wherein the visual material is further determined based on user configuration information, the configuration information indicating at least one of the following: The type of visual material to be acquired; Used to obtain the source of visual materials.

7. The method according to any one of claims 1 to 6, wherein adding visual material to at least one image frame of the video content comprises: Determine the effect corresponding to the visual material; as well as Add the visual material to the at least one video frame to which the effect is applied.

8. The method of claim 7, wherein determining the packaging effect corresponding to the visual material includes: Obtain the style parameters selected by the user; as well as Determine the effect corresponding to the style parameters.

9. The method according to claim 7 or 8, wherein determining the effect corresponding to the visual material includes: Based on the content of the visual material, the effect corresponding to the content of the material is determined from a set of candidate effects.

10. The method according to any one of claims 7 to 9, wherein the effect indicates at least one of the following: The display parameters of the visual material in the at least one image frame; The hierarchical relationship between the visual material and the at least one image frame; The dynamic process of the visual material; Image processing operations applied to the visual material.

11. The method of claim 10, wherein the effect further indicates the application of a preset editing operation to one or more image frames of the video content.

12. The method according to any one of claims 1 to 11, further comprising: Add subtitles to the video content that correspond to the text content.

13. The method of claim 12, wherein before adding a caption corresponding to the text content to the at least one image frame, the method comprises: In response to the video content including subtitle material, the subtitle portion corresponding to the text content is removed from the subtitle material.

14. The method according to any one of claims 1 to 13, wherein the visual material comprises at least one of the following: Image material; Video footage; Chart materials; Expression material.

15. An apparatus for video editing, comprising: The presentation module is configured to present an editing interface for video content, the editing interface including editing controls; The determination module is configured to determine a target audio segment from the audio signal of the video content in response to a selection of the editing control; as well as The addition module is configured to add visual material to at least one image frame of the video content in the editing interface, wherein the at least one image frame is determined based on the target audio segment, and the visual material is determined based on the text content corresponding to the target audio segment.

16. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processor.

17. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 14.

18. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 14.