Video editing method and apparatus, device, and storage medium
By matching video materials with audio rhythm information, highlight clips are automatically extracted and target videos are generated, which solves the problem of high threshold for video editing creation and improves video creation efficiency.
Patent Information
- Application Number
- PCT/CN2025/079099
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-25
- Publication Date
- 2025-09-04
AI Technical Summary
In the prior art, the creation threshold for video editing is relatively high, and it is difficult for non-professional personnel to generate creative video content, and the editing is high and the efficiency is low.
By dividing it into multiple segments based on the semantic information of the video material and matching it with the slots in the video template, the audio rhythm information is determined, the video content is adjusted to match the rhythm, and the target video is generated.
Automatically extract highlights of the material and complete the matching of audio and video, reducing the difficulty of video creation and improving work efficiency.
Smart Images

Figure CN2025079099_04092025_PF_FP_ABST
Abstract
Description
Video editing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Video Editing Method, Apparatus, Device and Storage Medium” filed on March 1, 2024, with application number 202410239113.0, the entire contents of which are incorporated herein by reference. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for video editing. Background Art
[0003] With the development of media content sharing platforms, more and more users are using them to publish videos depicting their daily lives. Video shooting and editing, among other creative activities, have become widely used by many groups. However, achieving high-quality video editing often presents a certain barrier to entry. Consequently, users desire to reduce the complexity of video editing and improve the efficiency of video creation. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for video editing is provided. The method includes: in response to obtaining a set of video materials input by a user, dividing the set of video materials into a plurality of video segments based on semantic information of the set of video materials; obtaining video content corresponding to the plurality of slots by dividing the plurality of video segments into a plurality of groups and associating the plurality of groups with a plurality of slots in a video template, wherein each group includes at least one video segment from the plurality of video segments; determining rhythm information of first audio content, wherein the rhythm information indicates a set of time points in the first audio content; adjusting the time length of at least one item of video content in the video content based on the rhythm information so that the adjusted at least one item of video content matches the rhythm information; and generating a target video based on the first audio content and the adjusted video content.
[0005] In a second aspect of the present disclosure, a device for video editing is provided. The device includes: a segmentation module configured to, in response to obtaining a set of video materials input by a user, segment the set of video materials into a plurality of video segments based on semantic information of the set of video materials; an extraction module configured to obtain video content corresponding to the plurality of slots by dividing the plurality of video segments into a plurality of groups and associating the plurality of groups with a plurality of slots in a video template, wherein each group includes at least one video segment from the plurality of video segments; a determination module configured to determine rhythm information of a first audio content, wherein the rhythm information indicates a set of time points in the first audio content; an adjustment module configured to adjust the time length of at least one video content in the video content based on the rhythm information so that the adjusted at least one video content matches the rhythm information; and a generation module configured to generate a target video based on the first audio content and the adjusted video content.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;
[0012] FIG2 illustrates a block diagram of an example process of video editing according to some embodiments of the present disclosure;
[0013] FIG3 is an example of adjusting a music card point according to some embodiments of the present disclosure;
[0014] FIG4 illustrates a flow chart of an example process for video editing according to some embodiments of the present disclosure;
[0015] FIG5 shows a schematic structural block diagram of an example apparatus for video editing according to some embodiments of the present disclosure; and
[0016] FIG6 illustrates a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0020] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0021] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0022] As mentioned above, with the development of media content sharing platforms, especially short video sharing platforms, there is a demand for editing and creating media content (e.g., video content). Generally speaking, editing and creating can involve combining background music, video materials, special effects, subtitles, transitions, and other content to form a short video that can be uploaded to the sharing platform. These videos have greatly enriched people's lives, not only broadening the ways of entertainment and leisure, but also becoming a way for people to obtain information.
[0023] However, the barriers to entry for video editing and creation are often high. For non-professionals, the resulting video often fails to deliver the desired, creative visuals. Therefore, users desire to reduce the complexity of video editing and improve the efficiency of video creation.
[0024] The embodiments of the present disclosure propose a video editing solution. According to this solution, a set of video materials input by a user can be obtained and, based on the semantic information of the set of video materials, video content corresponding to multiple slots in a video template can be extracted from them. Furthermore, rhythm information of the first audio content indicating a set of time points in the first audio content can be determined and, based on the rhythm information, the time length of at least one video content item in the video content can be adjusted so that the adjusted at least one video content item matches the rhythm information. Furthermore, a target video can be generated based on the first audio content and the adjusted video content.
[0025] In this way, the embodiments of the present disclosure can automatically extract the highlight segments of the material and complete the matching of audio and highlight segments, thereby reducing the difficulty of video creation and improving work efficiency.
[0026] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0027] Sample Environment
[0028] First, reference is made to FIG1 , which schematically illustrates a diagram of an example environment 100 in which example implementations according to the present disclosure may be implemented.
[0029] 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, a target application 120 is installed in an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110.
[0030] The target application 120 may be an application that can provide media data-related services to the user 102, including creation (e.g., shooting and / or editing), publishing, browsing, etc. of media data. In this document, "media data" may be content in various forms, including video, audio, image, image collection, text, etc.
[0031] For example, electronic device 110 can edit the received original media material through target application 120. In some embodiments, the original media material can be, for example, media material input by user 102. In some other embodiments, the original media material can also be media material collected by electronic device 110 using a collection device based on instructions from user 102. In some embodiments, the collection device can be configured to be connected to electronic device 110. In some other embodiments, the collection device can also be integrated within electronic device 110.
[0032] In the environment 100 of Figure 1, if the target application 120 is active, the electronic device 110 can present a page 140 of the target application 120 to the user 102. The page 140 can be any type of page that the target application 120 can provide, such as a media data presentation page, a content creation page, a content editing page, and the like.
[0033] In some embodiments, at least some of the functionality of the target application 120 may be implemented based on the model 131. The model 131 may be deployed, for example, in the server 130. For example, the electronic device 110 communicates with the server 130 to provide services for the target application 120. In other words, during the operation of the target application 120, the capabilities of one or more models (e.g., the model 131) may be invoked.
[0034] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 is various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0035] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0036] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0037] Video editing process
[0038] FIG2 shows a block diagram of an example process 200 for video editing according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 110. For example, the process 200 shown in FIG2 may be implemented by the target application 120 at the electronic device 110. The process 200 is described below with reference to FIG1.
[0039] As shown in FIG2 , the electronic device 110 may receive one or more video materials 201 input by the user. It should be understood that the one or more video materials 201 may be one or more videos shot by the user 102 using the electronic device 110, or may be one or more videos shot by other shooting devices of the user and transmitted to the electronic device 110.
[0040] After the electronic device 110 receives a set of video materials input by the user, in block 210 , the electronic device 110 divides the set of video materials into a plurality of video segments based on semantic information of the set of video materials.
[0041] In some embodiments, the electronic device 110 may, for example, process the set of video materials using the semantic slicing module 211 to obtain the plurality of video segments, each of which may be associated with the same semantic information. In other words, the electronic device 110 uses the semantic slicing module 211 to perform semantic slicing on all the materials input by the user. The semantics in each slice (i.e., video segment) are consistent.
[0042] For example, the semantic slicing module 211 can be implemented using a method based on a machine learning model, such as using Generic Event Boundary Detection (GEBD). Specifically, a convolutional neural network (CNN) can be used to extract features of each video frame. The structure context transformer (SC-Transformer) can calculate the group similarity of video frames to determine the differences between video frames. The fully convolutional network can then determine event boundaries based on group similarity, and then determine multiple video segments. It should be understood that the semantic slicing module 211 can also be implemented based on other machine learning models, and the solution disclosed herein is not limited to this.
[0043] Furthermore, at block 210, electronic device 110 may also generate multiple descriptive texts 212 for the multiple video segments using semantic slicing module 211. For example, after semantic slicing and descriptive text generation, the output content may be: from 0 to 25 seconds, two men are seen chatting; from 25 to 34 seconds, two men are seen having lunch.
[0044] In block 220 , based on the segmented multiple videos and the corresponding multiple description texts 202 , the electronic device 110 may extract video content corresponding to the multiple slots in the video template.
[0045] In some embodiments, the electronic device 110 may determine, from the plurality of video segments, a plurality of groups of video segments corresponding to the plurality of slots in the video template, wherein each slot corresponds to a group of video segments. The electronic device 110 may also obtain the video content corresponding to the plurality of slots in the video template by extracting, from each group of video segments, video content that matches the time length of the corresponding slot.
[0046] For example, the electronic device 110 may divide the multiple video segments into multiple groups based on the description texts of the multiple video segments, and each group may include at least one video segment from the multiple video segments.
[0047] For example, after obtaining the multiple segmented video segments and the corresponding multiple description texts 202, the electronic device 110 can divide the multiple segmented videos into several chapters (i.e., multiple groups) according to the multiple description texts. Each chapter can include one or more video segments. In addition, the electronic device 110 can also use the copy generation module 221 to generate corresponding chapter copy 222 for each chapter.
[0048] For example, the copy generation module 221 can generate chapter copy 222 based on a machine learning model. The input of the machine learning module can be one or more frames extracted from a video, and its output can be a description of the video. It should be understood that the copy generation module 221 can be implemented as any model / module capable of generating text from images, and the solutions of this disclosure are not limited in this respect.
[0049] During the chapter copy generation process, in some embodiments, the copy generation module 221 may also obtain prompts for generating the chapter copy 222. The prompts may, for example, relate to aspects such as shot grouping, slicing / chapter generation, copy generation, and copy rewriting.
[0050] Afterwards, the electronic device 110 further associates the plurality of groups with the plurality of slots in the video template to determine the plurality of video segments corresponding to the plurality of slots. For example, the electronic device 110 may associate the plurality of groups with the plurality of slots in the video template based on the generated chapter copy 222.
[0051] In some embodiments, the electronic device 110 may determine a target duration corresponding to a target slot among the plurality of slots. Based on the target duration, the electronic device 110 extracts at least one highlight segment from a set of video segments corresponding to the target slot, such that the cumulative duration of the at least one highlight segment is the target duration.
[0052] For example, the electronic device 110 can use the highlight extraction module 223 to extract highlight segments. The highlight extraction module 223 can obtain a video template 224. The video template 224 can be, for example, a video creation template preconfigured in the target application 120. It should be understood that a video creation template can include multiple slots. For example, one or more video effects can be pre-set in each slot. For example, multiple transition effects can be pre-set between slots.
[0053] After acquiring the video template 224, the highlight extraction module 223 can determine the duration of multiple slots. Based on the duration of each slot, the highlight extraction module 223 can extract the highlight segments corresponding to the slot duration from the set of video segments contained in the video group corresponding to the slot as the video content to be presented in that slot. Therefore, for each slot, the highlight extraction module 223 can extract the video content 225 to be presented in that slot from the corresponding groups.
[0054] For example, the highlight segment can be extracted based on the semantic relevance and picture evaluation of the video frames in a set of video segments. Specifically, the definition of a highlight segment can be a video segment that best conforms to specific semantics and has beautiful pictures. For example, a series of frames contained in a video can be scored. The scoring can be done from two dimensions: semantic relevance and aesthetic model. The highlight extraction module 223 can extract the video segment corresponding to the frame with the highest score as a highlight segment. In other words, the highlight extraction module 223 essentially extracts a portion of video frames with a target length that meet the semantic relevance and aesthetic model from a video clip with a certain length.
[0055] Continuing with FIG2 , at block 230, electronic device 110 matches audio content to video content corresponding to multiple slots in a video template. Electronic device 110 determines tempo information for a first audio content, where the tempo information indicates a set of time points in the first audio content, i.e., music check points 231. Based on the tempo information, electronic device 110 adjusts the duration of at least one item of video content in the video content so that the adjusted at least one item of video content matches the tempo information.
[0056] For example, the first audio content may include music content. The electronic device 110 may, for example, use a beat detection module to determine rhythm information of the music content. For example, the rhythm information may correspond to a set of beat points of the music content. The set of beat points may, for example, correspond to music card points of the video template 224.
[0057] In some embodiments, the electronic device 110 may increase or decrease the duration of at least one item of video content according to the rhythm information, so that a group of time points in the first audio content corresponds to the switching moments of different video contents.
[0058] For example, if the audio length between two beat points is 3 seconds and the video content is 4 seconds long, the video content can be compressed to 3 seconds.
[0059] In some embodiments, multiple video segments between two beat points can also be compressed or lengthened proportionally. For example, if the audio length between beat points is 3.5 seconds, the first video segment is 2 seconds (0-2 seconds), and the second video segment is 2 seconds (2-4 seconds), then the first and second video segments can be shortened by 0.25 seconds each.
[0060] As shown in Figure 3, the total duration of the first video content 311 and the second video content 312 is less than the duration corresponding to the first music card point 310. Therefore, the durations of the first video content 311 and the second video content 312 can be proportionally extended. The total duration of the third video content 321 and the fourth video content 322 is greater than the duration corresponding to the second music card point 320. Therefore, the durations of the third video content 321 and the fourth video content 322 can be proportionally compressed.
[0061] As shown in FIG. 2 , after matching the first audio content to the video content, the electronic device 110 generates a target video 202 based on the first audio content and the adjusted video content.
[0062] In some embodiments, the electronic device 110 may generate multiple text contents corresponding to the adjusted video content and generate the target video based on the first audio content, the text content, and the adjusted video content. For example, the target video may include multiple text contents and second audio content generated from the multiple text contents.
[0063] For example, the multiple text contents may be the chapter texts described above. When outputting the target video, since the duration of the video content has already been determined, the second audio content corresponding to the chapter text will also be adjusted to have a duration that matches the video content. For example, the audio portion of the second audio content corresponding to the corresponding text content has a first duration, and the video portion of the target video corresponding to the corresponding text content has a second duration, and the first duration should be the same as the second duration.
[0064] According to the content described above, the solution proposed in the present disclosure can automatically extract the highlight segments of the material and complete the matching of audio and highlight segments, thereby reducing the difficulty of video creation and improving work efficiency.
[0065] Example Process
[0066] FIG4 shows a flow chart of an example process 400 for video editing according to some embodiments of the present disclosure. The process 400 may be implemented at the electronic device 110. The process 400 is described below with reference to FIG1.
[0067] In block 410 , in response to obtaining a set of video materials input by a user, the electronic device 110 divides the set of video materials into a plurality of video segments based on semantic information of the set of video materials.
[0068] In block 420 , the electronic device 110 obtains video content corresponding to a slot by dividing the plurality of video segments into at least one group and associating the at least one group with a slot in a video template, wherein the group includes at least one video segment from the plurality of video segments.
[0069] In block 430 , the electronic device 110 determines tempo information of the first audio content, the tempo information indicating a set of time points in the first audio content.
[0070] In block 440 , the electronic device 110 adjusts a time length of at least one item of the video content based on the rhythm information, so that the adjusted at least one item of the video content matches the rhythm information.
[0071] In block 450 , the electronic device 110 generates a target video based on the first audio content and the adjusted video content.
[0072] In some embodiments, the first audio content includes music content, and determining the rhythm information of the first audio content includes: using a beat detection module to determine the rhythm information of the music content, the rhythm information corresponding to a set of beat points of the music content.
[0073] In some embodiments, adjusting the time length of at least one item of the video content based on the rhythm information includes: increasing or decreasing the time length of the at least one item of video content so that the set of time points in the first audio content corresponds to the switching moments of different video contents.
[0074] In some embodiments, obtaining the video content corresponding to the slot includes: determining a group of video segments corresponding to the slot in the video template from the multiple video segments; and obtaining the video content corresponding to the slot in the video template by extracting video content that matches the time length of the corresponding slot from the video segments.
[0075] In some embodiments, segmenting the set of video materials into a plurality of video segments based on semantic information of the set of video materials includes: processing the set of video materials using a semantic slicing module to obtain the plurality of video segments, wherein the video segments are associated with the same semantic information.
[0076] In some embodiments, the method further includes: generating, using the semantic slicing module, description texts about the plurality of video segments and dividing the plurality of video segments into at least one group based on the description texts.
[0077] In some embodiments, determining the video content corresponding to the slot in the video template by extracting video content that matches the time length of the corresponding slot from the video segment includes: determining a target time length corresponding to the target slot in the slot; and extracting the at least one highlight segment from a group of video segments corresponding to the target slot based on the semantic relevance and picture evaluation of the video frames in the group of video segments, the cumulative time length of the at least one highlight segment being the target time length.
[0078] In some embodiments, generating a target video based on the first audio content and the adjusted video content includes: generating text content corresponding to the adjusted video content; and generating the target video based on the first audio content, the text content, and the adjusted video content. The target video includes: the text content; or second audio content generated based on the text content.
[0079] In some embodiments, the audio portion of the second audio content corresponding to the corresponding text content has a first duration, the video portion of the target video corresponding to the corresponding text content has a second duration, and the first duration is the same as the second duration.
[0080] Example devices and equipment
[0081] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG5 shows a schematic structural block diagram of an example apparatus 500 for video editing according to certain embodiments of the present disclosure. Apparatus 500 may be implemented as or included in electronic device 110. Each module / component in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0082] As shown in Figure 5, the device 500 includes a segmentation module 510, which is configured to, in response to obtaining a group of video materials input by a user, divide the group of video materials into multiple video segments based on semantic information of the group of video materials; an acquisition module 520, which is configured to obtain video content corresponding to the slot by dividing the multiple video segments into at least one group and associating the at least one group to a slot in a video template, wherein the group includes at least one video segment among the multiple video segments; a determination module 530, which is configured to determine rhythm information of a first audio content, wherein the rhythm information indicates a set of time points in the first audio content; an adjustment module 540, which is configured to adjust the time length of at least one video content in the video content based on the rhythm information, so that the adjusted at least one video content matches the rhythm information; and a generation module 550, which is configured to generate a target video based on the first audio content and the adjusted video content.
[0083] In some embodiments, the first audio content includes music content, and the determination module 530 is further configured to: determine the rhythm information of the music content using a beat detection module, where the rhythm information corresponds to a set of beat points of the music content.
[0084] In some embodiments, the adjustment module 540 is further configured to increase or decrease the time length of the at least one video content so that the set of time points in the first audio content corresponds to the switching moments of different video contents.
[0085] In some embodiments, the acquisition module 520 is further configured to: determine a group of video segments corresponding to the slots in the video template from the multiple video segments; and obtain the video content corresponding to the slots in the video template by extracting video content that matches the time length of the corresponding slots from the video segments.
[0086] In some embodiments, the segment division module 510 is further configured to: process the set of video materials using a semantic slicing module to obtain the plurality of video segments, wherein the video segments are associated with the same semantic information.
[0087] In some embodiments, the apparatus 500 is further configured to: generate description texts about the plurality of video segments using the semantic slicing module, and divide the plurality of video segments into at least one group based on the description texts.
[0088] In some embodiments, the acquisition module 520 is further configured to: determine a target time length corresponding to a target slot in the slot; and extract the at least one highlight segment from a group of video segments corresponding to the target slot based on the semantic relevance and picture evaluation of the video frames in the group of video segments, the cumulative time length of the at least one highlight segment being the target time length.
[0089] In some embodiments, the generation module 550 is further configured to: generate text content corresponding to the adjusted video content; and generate the target video based on the first audio content, the text content, and the adjusted video content. The target video includes: the text content; or second audio content generated based on the text content.
[0090] In some embodiments, the audio portion of the second audio content corresponding to the corresponding text content has a first duration, the video portion of the target video corresponding to the corresponding text content has a second duration, and the first duration is the same as the second duration.
[0091] FIG6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 can be used to implement the electronic device 110 of FIG1 .
[0092] As shown in FIG6 , electronic device 600 is a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.
[0093] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.
[0094] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0095] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0096] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0097] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0098] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0100] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0101] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0102] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A video editing method, comprising: In response to obtaining a set of video materials input by a user, segmenting the set of video materials into a plurality of video segments based on semantic information of the set of video materials; Obtaining video content corresponding to a slot by dividing the plurality of video segments into at least one group and associating the at least one group with a slot in a video template, wherein the group includes at least one video segment from the plurality of video segments; determining tempo information of a first audio content, the tempo information indicating a set of time points in the first audio content; Based on the rhythm information, adjusting the time length of at least one item of the video content so that the adjusted at least one item of the video content matches the rhythm information; and A target video is generated based on the first audio content and the adjusted video content.
2. The method of claim 1 , wherein the first audio content comprises music content, and determining the tempo information of the first audio content comprises: The rhythm information of the music content is determined by using a beat detection module, where the rhythm information corresponds to a set of beat points of the music content.
3. The method according to claim 1, wherein adjusting the duration of at least one item of the video content based on the rhythm information comprises: The time length of the at least one video content is increased or decreased so that the set of time points in the first audio content corresponds to switching moments of different video contents.
4. The method according to claim 1, wherein obtaining the video content corresponding to the slot comprises: Determining, from the plurality of video segments, a set of video segments corresponding to the slots in the video template; as well as The video content corresponding to the slot in the video template is obtained by extracting the video content that matches the time length of the corresponding slot from the group of video segments.
5. The method according to claim 1, wherein dividing the set of video materials into a plurality of video segments based on semantic information of the set of video materials comprises: The set of video materials is processed using a semantic slicing module to obtain the plurality of video segments, wherein the video segments are associated with the same semantic information.
6. The method according to claim 4, further comprising: Generate descriptive text about the plurality of video segments using the semantic slicing module; as well as Based on the description text, the plurality of video segments are divided into the at least one group.
7. The method according to claim 4, wherein determining the video content corresponding to the slot in the video template by extracting video content matching the time length of the corresponding slot from the video segment comprises: determining a target time length corresponding to a target slot in the slots; as well as From a group of video segments corresponding to the target slot, based on the semantic relevance and picture evaluation of the video frames in the group of video segments, extract the at least one highlight segment, the cumulative time length of the at least one highlight segment being the target time length.
8. The method according to claim 1, wherein generating a target video based on the first audio content and the adjusted video content comprises: Generating text content corresponding to the adjusted video content; as well as Generate the target video based on the first audio content, the text content, and the adjusted video content, wherein the target video includes: the content of the text; or Second audio content is generated based on the text content.
9. The method according to claim 8, wherein the audio portion in the second audio content corresponding to the corresponding text content has a first duration, the video portion in the target video corresponding to the corresponding text content has a second duration, and the first duration is the same as the second duration.
10. A device for video editing, comprising: A segmentation module is configured to, in response to obtaining a set of video materials input by a user, segment the set of video materials into a plurality of video segments based on semantic information of the set of video materials; an extraction module configured to obtain video content corresponding to a plurality of slots by dividing the plurality of video segments into a plurality of groups and associating the plurality of groups with a plurality of slots in a video template, wherein each group includes at least one video segment from the plurality of video segments; a determination module configured to determine rhythm information of a first audio content, wherein the rhythm information indicates a set of time points in the first audio content; an adjustment module configured to adjust a time length of at least one item of video content in the video content based on the rhythm information, so that the adjusted at least one item of video content matches the rhythm information; as well as The generating module is configured to generate a target video based on the first audio content and the adjusted video content.
11. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.
12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by a processor.
Citation Information
Patent Citations
Method and system for automatically generating video
CN107360383A
Audio and video processing method and device
CN114390352A
Video editing method and electronic equipment
CN115134646A
Video content generation method and generation program
JP2020043454A
Systems and methods for audio track selection in video editing
US9838730B1