Video editing method and device, and storage medium
By acquiring user input information and generating recommendation information to automatically edit videos, the cumbersome video editing problem in existing technologies is solved, achieving a more efficient and better user experience.
Patent Information
- Application Number
- PCT/CN2025/105916
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
The existing video editing process is cumbersome, inefficient, and difficult, which negatively impacts the user experience.
By acquiring user input information, the system determines the target intent of recommended information to prompt the editing effect, and generates editing results based on this, thus simplifying user operations.
It reduces the difficulty of video editing and improves editing efficiency and experience.
Smart Images

Figure CN2025105916_08012026_PF_FP_ABST
Abstract
Description
Video editing method, device and storage medium
[0001] Cross-reference to Related Applications
[0002] The present application claims priority to the Chinese patent application No. 202410876718.0, filed on July 01, 2024, and entitled "Video editing method, device and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] Embodiments of the present disclosure relate to the technical field of computer and network communication, and particularly relate to a video editing method, device and storage medium. BACKGROUND
[0004] With the development of communication technology and the rise of mobile short videos, the demand for video creation is becoming more and more vigorous, and video creators are gradually spreading from professionals to the general public.
[0005] In a video editing scenario, a series of editing operations need to be performed on a current video draft, including but not limited to adding stickers, adding background music, material cutting and splicing, adding transition effects, adding subtitles, titles, or tags to the current video draft. SUMMARY
[0006] Embodiments of the present disclosure provide a video editing method, device and storage medium to simplify video editing operations, reduce video editing difficulty, improve video editing efficiency, and improve video editing experience.
[0007] In a first aspect, embodiments of the present disclosure provide a video editing method, comprising:
[0008] obtaining input information of a user on a current video editing task and / or a historical video editing task;
[0009] determining recommendation information for the current video editing task according to the input information, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task;
[0010] generating an editing result of the current video editing task according to the recommendation information and the input information.
[0011] In a second aspect, embodiments of the present disclosure provide a video editing device, comprising:
[0012] an input unit configured to obtain input information of a user on a current video editing task and / or a historical video editing task;
[0013] The recommendation unit is configured to determine recommendation information for the current video editing task according to the input information, wherein the recommendation information is used to prompt a target intention of an editing effect of the current video editing task.
[0014] The editing unit is configured to generate an editing result of the current video editing task according to the recommendation information and the input information.
[0015] In a third aspect, an electronic device is provided, which includes at least one processor and a memory.
[0016] The memory stores computer-executable instructions.
[0017] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video editing method as described in the first aspect and various possible designs of the first aspect.
[0018] In a fourth aspect, a computer-readable storage medium is provided, which stores computer-executable instructions, and when a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible designs of the first aspect is implemented.
[0019] In a fifth aspect, a computer program product is provided, which includes computer-executable instructions, and when a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible designs of the first aspect is implemented.
[0020] The video editing method, device and storage medium provided by the embodiments of the present disclosure can obtain input information of a user for a current video editing task and / or a historical video editing task, determine recommendation information for the current video draft according to the target intention, determine recommendation information for the current video editing task according to the input information, wherein the recommendation information is used to prompt a target intention of an editing effect of the current video editing task, and generate an editing result of the current video editing task according to the recommendation information and the input information. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] FIG. 1 is a scene example diagram of a video editing method provided by an embodiment of the present disclosure;
[0023] FIG. 2 is a flowchart of a video editing method according to an embodiment of the present disclosure;
[0024] FIG. 3 is a flowchart of a video editing method according to another embodiment of the present disclosure;
[0025] FIG. 4 is a structural block diagram of a video editing device according to an embodiment of the present disclosure;
[0026] FIG. 5 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0028] In a video editing scenario, a series of editing operations need to be performed on a video draft, including but not limited to adding a sticker, adding background music, material cutting and splicing, adding a transition effect, adding a subtitle, a title, or a label, etc. However, in the prior art, the user usually needs to manually complete the above operations, which leads to a tedious video editing process, low efficiency, high difficulty in video editing, and affects the video editing experience.
[0029] To solve the above technical problems, the video editing method of the present disclosure can obtain input information of a user on a current video editing task and / or a historical video editing task; determine recommendation information for the current video editing task according to the input information, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task; and generate an editing result of the current video editing task according to the recommendation information and the input information. In this embodiment, the recommendation information is recommended based on the input information of the user, and then the video editing is automatically performed based on the input information and the recommendation information, which simplifies the user operation, reduces the difficulty of video editing, improves the efficiency of video editing, and improves the video editing experience.
[0030] The video editing method provided by the embodiments of the present disclosure can be applied in an electronic device such as a terminal device or a server, as shown in FIG. 1. Taking the terminal device as an example, the terminal device can obtain input information of a user on a current video editing task and / or a historical video editing task. Further, the terminal device can determine a target intention of an editing effect of the current video editing task according to the input information, and determine recommendation information for a current video draft according to the target intention. The terminal device can generate a target editing instruction for the current video draft according to the recommendation information and the input information. Further, the terminal device can generate an editing result of the current video editing task based on the target editing instruction. For example, in a video editing scenario, a video editing model is provided, and the video editing model is called with the target editing instruction as input, so that the video editing model generates the editing result of the current video editing task according to the target editing instruction.
[0031] The video editing method of the present disclosure will be described in detail below in combination with specific embodiments.
[0032] Referring to FIG. 2, FIG. 2 is a flowchart of a video editing method provided by an embodiment of the present disclosure. The method of the present embodiment can be applied in an electronic device such as a terminal device or a server. The video editing method includes the following steps.
[0033] S201, obtaining input information of a user on a current video editing task and / or a historical video editing task.
[0034] In the present embodiment, in a video editing scenario, a method for automatically editing a video according to an editing instruction is provided. For example, a video editing model can be deployed in the video editing scenario, and a user can input an editing instruction to control the video editing model to automatically edit a video. The video editing model can be any known artificial intelligence model, such as an Artificial Intelligence Generated Content (AIGC) model or a Large Language Model (LLM). The editing instruction input by the user can be an instruction for directly generating a video or a video editing draft according to some materials to obtain a target video or a target video editing draft. Alternatively, the editing instruction input by the user can be an instruction for performing a specific editing function on a video draft, such as adding a sticker, adding background music, cutting and splicing materials, adding a transition effect, adding a caption, a title, or a label, etc. (i.e., the content of the editing instruction input by the user is the content of the caption, the title, or the label).
[0035] Optionally, the user can input the input information in any manner, and the input information can be a complete sentence or not (e.g., a sentence that the user has not completed inputting during inputting), one or more words, or a piece of video material (picture, video clip, music, etc.), or the like.
[0036] In this embodiment, during editing of the current video draft, input information input by the user for the current video draft can be received, or input information for a historical video draft input by the user can be acquired.
[0037] In this embodiment, the input information can be used to determine the recommendation information for the current video editing task, and the recommendation information can be used to prompt the target intention of the editing effect of the current video editing task, for example, can be used to supplement or perfect the input information, and generate a target editing instruction for the current video editing task.
[0038] In this embodiment, the input information can be used to determine the recommendation information for the current video editing task, and the recommendation information can be used to prompt the target intention of the editing effect of the current video editing task, for example, can be used to supplement or perfect the input information, and generate a target editing instruction for the current video editing task.
[0039] Optionally, the target intention specifically can include an editing scene and an editing function corresponding to the input information, wherein the editing scene specifically can define what content is edited, for example, a video clip is edited, a picture is edited, a text is edited, music is edited, or the like, and the editing function specifically can define how to edit, for example, cutting, splicing, adding a sticker, adding a subtitle, adding a title, adding a label, or the like, that is, in this embodiment, the editing scene and the editing function corresponding to the input information can be determined, and the editing scene and the editing function are determined as the target intention of the input information.
[0040] Optionally, the target intention can be determined according to the input information, and the recommendation information for the current video editing task can be determined according to the target intention.
[0041] In this embodiment, the input information can be used to determine the recommendation information for the current video editing task, and the recommendation information can be used to prompt the target intention of the editing effect of the current video editing task, for example, can be used to supplement or perfect the input information, and generate a target editing instruction for the current video editing task.
[0042] In this embodiment, the target intention can be determined only based on the input information of the user for the current video editing task, or can be determined in combination with the input information of the user for the current video editing task and the input information of the user for a historical video editing task, or can be determined only based on the input information of the user for the historical video editing task, for example, when the user has not input the input information for the current video editing task, the target intention of the user this time can be predicted according to the input information of the user for the historical video editing task.
[0043] The intention recognition of the input information can be implemented in any manner in this embodiment. Optionally, the target intention can be determined according to a preset rule based on the input information; or an intention analysis model is called with the input information as input, and the target intention is output.
[0044] The target intention can be determined according to a preset rule, and the preset rule for intention recognition can be preconfigured, for example, a regular matching rule. Some alternative sentence patterns and the intentions corresponding to the alternative sentence patterns can be preconfigured, for example, if the sentence pattern A appears, it is determined that the scene a is in, if the sentence pattern B appears, it is determined that the scene b is in, and then the input information is matched with the alternative sentence patterns according to the regular matching rule, the alternative sentence pattern corresponding to the input information is determined, and the intention corresponding to the alternative sentence pattern is determined as the target intention corresponding to the input information. Optionally, in order to improve the accuracy of intention recognition, the target intention of the input information can also be determined based on the preset rule in combination with the context information, wherein the context information can include but is not limited to the material uploaded by the user, the input information of the user last time, the historical editing task (which can include input information and / or target editing instruction) of the last round (or multiple rounds), the output result of the video editing model, the user operation, and the like.
[0045] The target intention of the input information is recognized by the intention analysis model, and the intention analysis model can be pre-trained, wherein the intention analysis model can be any known model, for example, a neural network model, a large language model, and the like, and then the input information is input into the intention analysis model, and the target intention of the input information is output. Optionally, in order to improve the accuracy of intention recognition, the context information can also be input into the intention analysis model, so that the intention analysis model determines the target intention of the input information according to the input information and the context information.
[0046] Optionally, the intention analysis model can be pre-trained according to the training data, wherein the training data can be historical data, including historical input information and the intention (which can include editing scene and editing function) corresponding to the historical input information, and further, the intention analysis model can be trained according to the training data so that the intention analysis model can implement the classification task of the single-label multi-classification problem of the intention.
[0047] Optionally, considering that the intention recognition by the intention classification model has higher cost and delay compared with the intention recognition according to the preset rule, the intention recognition according to the preset rule can be used as a pre-step of the intention recognition by the intention classification model, as shown in FIG. 3, the target intention is determined according to the preset rule based on the input information; if the target intention cannot be confirmed according to the preset rule, the intention analysis model is called with the input information as input, and the target intention is output, which can reduce the number of calls to the intention classification model, and reduce the cost and delay.
[0048] In addition, based on any of the above embodiments, if only the editing function corresponding to the input information can be determined, and the editing scene corresponding to the input information cannot be determined, since there can be the same or similar editing functions in different editing scenes, it is further needed to determine the editing scene corresponding to the input information. At this time, the historical editing scene corresponding to the historical editing instruction of the last round can be determined as the editing scene corresponding to the input information, so that the complete target intention corresponding to the input information can be obtained.
[0049] Further, after the target intention of the input information is determined, the recommendation information can be recommended according to the target intention. Specifically, a preset recommendation information library can be preconfigured, which can include some preset recommendation information, and then one or more recommendation information can be determined from the preset recommendation information library according to the target intention. Wherein, the recommendation information for the current video editing task can be determined according to the target intention by using any recommendation algorithm, such as similarity-based recommendation and the like.
[0050] Optionally, the recommendation information included in the preset recommendation information library can be classified according to intention to obtain a recommendation information set corresponding to different intentions, and then when the recommendation information for the current video editing task is determined according to the target intention, one or more recommendation information can be obtained from the target recommendation information set corresponding to the target intention in the preset recommendation information library.
[0051] S203, generating an editing result of the current video editing task according to the recommendation information and the input information.
[0052] In this embodiment, the video editing can be automatically performed based on the recommendation information and the input information to generate the editing result of the current video editing task.
[0053] Optionally, in this embodiment, the target video of the current video editing task is generated according to the recommendation information and the input information; or the target video editing draft of the current video editing task is generated according to the recommendation information and the input information; or the initial video editing draft of the current video editing task is updated to obtain the target video editing draft of the current video editing task according to the recommendation information and the input information.
[0054] That is, in this embodiment, the target video can be automatically generated directly, the target video editing draft can be generated directly, and the initial video editing draft can be further edited.
[0055] Optionally, the target editing instruction for the current video draft can be generated according to the recommendation information and the input information, and further, the editing result of the current video editing task is generated according to the target editing instruction.
[0056] In this embodiment, after determining the recommendation information, a complete target editing instruction for the current video editing task can be generated according to the recommendation information and the input information. The generation manner can be supplementing on the basis of the input information, generating a completely new editing instruction to replace the input information as the target editing instruction, or other manners, such as using a large language model to generate the target editing instruction.
[0057] Specifically, in the supplement mode, a supplement instruction for the input information can be generated according to the recommendation information and the input information, and the supplement instruction and the input information can be spliced. The supplement instruction can be spliced after or before the input information or at any other position to obtain the target editing instruction.
[0058] In the replacement mode, a complete target editing instruction can be generated according to the recommendation information and the input information, and the target editing instruction can replace the input information.
[0059] The selection of the supplement mode or the replacement mode can be selected according to the condition of the input information. For example, if the input information is relatively complete, smooth, and logical, the supplement mode can be used. If the input information is incomplete, not smooth, and not logical, or even only a few words, the replacement mode can be used.
[0060] Optionally, when generating the target editing instruction, a plurality of candidate target editing instructions can be generated according to the recommendation information and the input information, and displayed for the user to select. After receiving a selection operation instruction of the user on the plurality of candidate target editing instructions, the selected candidate target editing instruction can be determined as the final target editing instruction, so that the target editing instruction can meet the editing demand of the user as much as possible. Of course, if the user is not satisfied with the candidate target editing instruction, the user can trigger the generation of the target editing instruction again, and then the recommendation information can be obtained again, and the candidate target editing instruction can be generated again for the user to select.
[0061] After obtaining the target editing instruction, the target editing instruction can be executed to perform automatic editing and generate an editing result of the current video editing task.
[0062] Optionally, a video editing model can be called, and the target editing instruction can be input into the video editing model, so that the video editing model can edit according to the target editing instruction. The editing capabilities of adding a sticker, adding background music, material cutting and splicing, adding a transition effect, adding a subtitle, a title, or a label can be automatically implemented, the user operation can be simplified, the video editing difficulty can be reduced, the video editing efficiency can be improved, and the video editing experience can be improved.
[0063] Of course, other possible ways can also be used to automatically execute the target editing instruction and automatically edit, such as using specific program codes to run the target editing instruction, and the like, which are not listed here. In addition, the target editing instruction can not be generated, and the editing result of the current video editing task can be generated according to the recommended information and the input information by other possible ways
[0064] The video editing method provided by the embodiment can obtain input information of a user on a current video editing task and / or a historical video editing task, determine recommended information on the current video editing task according to the input information, the recommended information being used to prompt a target intention of an editing effect of the current video editing task, and generate an editing result of the current video editing task according to the recommended information and the input information. The embodiment can recommend the recommended information based on the input information of the user, and then automatically edit the video based on the input information and the recommended information, thereby simplifying the operation of the user, reducing the difficulty of video editing, improving the efficiency of video editing, and improving the experience of video editing.
[0065] On the basis of any of the above embodiments, when the preset recommended information library is preconfigured, the recommended information included in the preset recommended information library can be classified according to intentions, and each set of recommended information corresponding to an intention can be further classified according to a theme category, that is, each set of recommended information corresponding to an intention can further include a set of recommended information corresponding to different theme categories, and the theme categories can include travel, food, sports, games, and the like.
[0066] Further, when one or more recommended information is obtained from a set of target recommended information corresponding to a target intention in the preset recommended information library, the target theme category of the input information can be determined first, and then one or more recommended information is obtained from a set of target recommended information corresponding to the target theme category in the set of target recommended information, so as to improve the recommendation accuracy.
[0067] The target theme category of the input information can be determined in any way, for example, the similarity between the input information and a candidate theme category can be obtained, and the candidate theme category with the highest similarity can be taken as the target theme category of the input information. The similarity between the input information and the candidate theme category can be obtained in the form of vector similarity, that is, the input information and the candidate theme category are respectively vectorized, and the vector similarity is obtained.
[0068] On the basis of any of the above embodiments, when the preset recommended information library is preconfigured, the recommended information can be extracted or summarized based on historical editing instructions (including input information and / or target editing instructions), and the extracted recommended information can be added to the preset recommended information library.
[0069] For example, if the historical editing instruction is "The world of adults, someone helps you is love, and no help is duty, we should not resent others", the theme or keywords can be extracted to obtain recommended information such as "the value of friendship", "gratitude", and "mature interpersonal relationship"; for example, if the historical editing instruction is "animal a in region A", the theme or keywords can be extracted to obtain recommended information such as "region A animal a base visit strategy" and "exploring the living habits of animal a in region A"; for example, if the historical editing instruction is "add sticker y to video clip x", the theme or keywords can be extracted to obtain recommended information such as "add stickers to video clips". When adding the preset recommended information library, the recommended information can be converted into a structured text in a preset format, and the intent, theme category, etc. of the recommended information can be indicated.
[0070] Optionally, the extraction or summary of the recommended information based on the historical editing instruction can be implemented in any manner, for example, a large language model is used to implement the extraction or summary of the recommended information, the historical editing instruction is input, the large language model is called, and the recommended information is output. Optionally, the historical editing instruction can also be rewritten, wherein the rewriting includes but is not limited to error correction, expansion, and other expression methods of the historical editing instruction, and the recommended information is extracted or summarized based on the rewritten historical editing instruction, so as to expand the recommended information, improve the quality of the recommended information, and avoid possible errors, omissions, and misspellings. The rewriting process can also be implemented in any manner, for example, a large language model is used to implement the rewriting process. In addition, the recommended information can also be cleaned to further improve the quality.
[0071] Corresponding to the video editing method of the above embodiment, FIG. 4 is a structural block diagram of a video editing device provided by an embodiment of the present disclosure. For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to FIG. 4, the video editing device 400 includes an input unit 401, a recommendation unit 402, and an editing unit 403.
[0072] The input unit 401 is configured to obtain input information of a user on a current video editing task and / or a historical video editing task.
[0073] The recommendation unit 402 is configured to determine recommended information for the current video editing task according to the input information, and the recommended information is used to prompt a target intent of an editing effect of the current video editing task.
[0074] The editing unit 403 is configured to generate an editing result of the current video editing task according to the recommended information and the input information.
[0075] In one or more embodiments of the present disclosure, when the editing unit 403 generates the editing result of the current video editing task according to the recommended information and the input information, the editing unit 403 is configured to:
[0076] generate a target video of the current video editing task according to the recommendation information and the input information; or
[0077] generate a target video editing draft of the current video editing task according to the recommendation information and the input information; or
[0078] update an initial video editing draft of the current video editing task according to the recommendation information and the input information, to obtain a target video editing draft of the current video editing task.
[0079] In one or more embodiments of the present disclosure, the recommendation unit 402, when determining the recommendation information for the current video editing task according to the input information, is configured to:
[0080] determine the target intention according to the input information, and determine the recommendation information according to the target intention.
[0081] In one or more embodiments of the present disclosure, the recommendation unit 402, when determining the target intention according to the input information, is configured to:
[0082] determine the target intention according to a preset rule based on the input information; or
[0083] call an intention analysis model with the input information as input, and output the target intention.
[0084] In one or more embodiments of the present disclosure, the recommendation unit 402, when determining the target intention according to the input information, is configured to:
[0085] determine the target intention according to a preset rule based on the input information;
[0086] if the target intention cannot be confirmed, call an intention analysis model with the input information as input, and output the target intention.
[0087] In one or more embodiments of the present disclosure, the recommendation unit 402, when determining the target intention according to the input information, is configured to:
[0088] determine the target intention according to the input information and context information of the input information.
[0089] In one or more embodiments of the present disclosure, the recommendation unit 402, when determining the target intention according to the input information, is configured to:
[0090] determine an editing scene and an editing function associated with an editing effect of the current video editing task according to the input information, and determine the editing scene and the editing function as the target intention.
[0091] In one or more embodiments of the present disclosure, the recommendation unit 402 is configured to, when determining an editing scene and an editing function associated with an editing effect of the current video editing task according to the input information, and determining the editing scene and the editing function as the target intention:
[0092] If only the editing function can be determined, the historical editing scene is determined as the editing scene, and the editing scene and the editing function are determined as the target intention.
[0093] In one or more embodiments of the present disclosure, the recommendation unit 402 is configured to, when determining the recommendation information according to the target intention:
[0094] obtain one or more recommendation information from a target recommendation information set corresponding to the target intention in the preset recommendation information library, wherein the preset recommendation information library includes recommendation information sets corresponding to different intentions.
[0095] In one or more embodiments of the present disclosure, the recommendation unit 402 is configured to, when obtaining one or more recommendation information from a target recommendation information set corresponding to the target intention in the preset recommendation information library:
[0096] determine a target theme category of the input information;
[0097] obtain one or more recommendation information from a target recommendation information sub-set corresponding to the target theme category in the target recommendation information set, wherein the target recommendation information set includes recommendation information sub-sets corresponding to different theme categories.
[0098] In one or more embodiments of the present disclosure, the editing unit 403 is configured to, when generating an editing result of the current video editing task according to the recommendation information and the input information:
[0099] generate a target editing instruction for the current video editing task according to the recommendation information and the input information, and generate an editing result of the current video editing task according to the target editing instruction.
[0100] In one or more embodiments of the present disclosure, the editing unit 403 is configured to, when generating a target editing instruction for the current video editing task according to the recommendation information and the input information:
[0101] According to the recommendation information and the input information, a supplement instruction for the input information is generated, and the supplement instruction is spliced with the input information to obtain the target editing instruction; or
[0102] According to the recommendation information and the input information, a complete target editing instruction is generated, and the target editing instruction is used to replace the input information.
[0103] In one or more embodiments of the present disclosure, when the editing unit 403 generates a target editing instruction for the current video editing task according to the recommendation information and the input information, the editing unit 403 is configured to:
[0104] According to the recommendation information and the input information, a plurality of alternative target editing instructions are generated and displayed.
[0105] In response to a selection operation instruction of a user on the plurality of alternative target editing instructions, the selected alternative target editing instruction is determined as the final target editing instruction.
[0106] In one or more embodiments of the present disclosure, when the editing unit 403 generates an editing result of the current video editing task according to the target editing instruction, the editing unit 403 is configured to:
[0107] The video editing model is called with the target editing instruction as input, so that the video editing model generates an editing result of the current video editing task according to the target editing instruction.
[0108] The device provided in the embodiment can be used to execute the technical solutions of the above-mentioned method embodiments, and has similar implementation principles and technical effects. Details are not described herein.
[0109] Referring to FIG. 5, a structural schematic diagram of an electronic device 500 suitable for implementing embodiments of the present disclosure is shown. The electronic device 500 can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (PAD), a portable multimedia player (PMP), a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 5 is only an example, and should not limit the functions and use range of the embodiments of the present disclosure.
[0110] As shown in FIG. 5, the electronic device 500 can include a processing device (e.g., a central processor, a graphics processor, etc.) 501 that can perform various suitable actions and processes according to programs stored in a Read Only Memory (ROM) 502 or loaded into a Random Access Memory (RAM) 503 from a storage device 508. Various programs and data required by the electronic device 500 for its operation are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.
[0111] Generally, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data. Although FIG. 5 shows the electronic device 500 with various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.
[0112] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0113] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0114] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.
[0115] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0116] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0117] The flow diagrams and the block diagrams in the drawings are meant as possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0118] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0119] The functions described above in the specification of the present disclosure can be performed by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0121] In a first aspect, according to one or more embodiments of the present disclosure, a video editing method is provided, comprising:
[0122] obtaining input information of a user on a current video editing task and / or a historical video editing task;
[0123] determining, according to the input information, recommendation information on the current video editing task, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task;
[0124] generating, according to the recommendation information and the input information, an editing result of the current video editing task.
[0125] According to one or more embodiments of the present disclosure, the generating, according to the recommendation information and the input information, an editing result of the current video editing task comprises: generating, according to the recommendation information and the input information, a target video of the current video editing task; or
[0126] generating, according to the recommendation information and the input information, a target video editing draft of the current video editing task; or
[0127] updating, according to the recommendation information and the input information, an initial video editing draft of the current video editing task to obtain a target video editing draft of the current video editing task.
[0128] According to one or more embodiments of the present disclosure, the determining, according to the input information, recommendation information on the current video editing task, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task comprises:
[0129] The target intention is determined according to the input information, and the recommendation information is determined according to the target intention.
[0130] According to one or more embodiments of the present disclosure, the target intention is determined according to the input information, including:
[0131] The target intention is determined according to a preset rule based on the input information; or
[0132] An intention analysis model is called with the input information as input, and the target intention is output.
[0133] According to one or more embodiments of the present disclosure, the target intention is determined according to the input information, including:
[0134] The target intention is determined according to a preset rule based on the input information;
[0135] If the target intention cannot be confirmed, an intention analysis model is called with the input information as input, and the target intention is output.
[0136] According to one or more embodiments of the present disclosure, the target intention is determined according to the input information, including:
[0137] The target intention is determined according to the input information and context information of the input information.
[0138] According to one or more embodiments of the present disclosure, the target intention is determined according to the input information, including:
[0139] The editing scene and editing function associated with the editing effect of the current video editing task are determined according to the input information, and the editing scene and editing function are determined as the target intention.
[0140] According to one or more embodiments of the present disclosure, the editing scene and editing function associated with the editing effect of the current video editing task are determined according to the input information, and the editing scene and editing function are determined as the target intention, including:
[0141] If only the editing function can be determined, the historical editing scene is determined as the editing scene, and the editing scene and editing function are determined as the target intention.
[0142] According to one or more embodiments of the present disclosure, the recommendation information is determined according to the target intention, including:
[0143] One or more recommendation information are obtained from a target recommendation information set corresponding to the target intention in the preset recommendation information library, wherein the preset recommendation information library includes recommendation information sets corresponding to different intentions.
[0144] According to one or more embodiments of the present disclosure, the one or more recommendation information is obtained from a target recommendation information set corresponding to the target intention in the preset recommendation information library, including:
[0145] Determining a target subject category of the input information;
[0146] Obtaining one or more recommendation information from a target recommendation information sub-set corresponding to the target subject category in the target recommendation information set, wherein the target recommendation information set includes recommendation information sub-sets corresponding to different subject categories.
[0147] According to one or more embodiments of the present disclosure, the editing result of the current video editing task is generated according to the recommendation information and the input information, including:
[0148] Generating a target editing instruction for the current video editing task according to the recommendation information and the input information, and generating an editing result of the current video editing task according to the target editing instruction.
[0149] According to one or more embodiments of the present disclosure, the target editing instruction for the current video editing task is generated according to the recommendation information and the input information, including:
[0150] Generating a supplement instruction for the input information according to the recommendation information and the input information, and splicing the supplement instruction with the input information to obtain the target editing instruction; or
[0151] Generating a complete target editing instruction according to the recommendation information and the input information, and replacing the input information with the target editing instruction.
[0152] According to one or more embodiments of the present disclosure, the target editing instruction for the current video editing task is generated according to the recommendation information and the input information, including:
[0153] Generating a plurality of alternative target editing instructions according to the recommendation information and the input information, and displaying the plurality of alternative target editing instructions;
[0154] In response to a selection operation instruction of a user on the plurality of alternative target editing instructions, determining the selected alternative target editing instruction as the final target editing instruction.
[0155] According to one or more embodiments of the present disclosure, the editing result of the current video editing task is generated according to the target editing instruction, including:
[0156] inputting the target editing instruction, calling a video editing model to generate an editing result of the current video editing task according to the target editing instruction.
[0157] In a second aspect, according to one or more embodiments of the present disclosure, a video editing device is provided, comprising:
[0158] an input unit configured to obtain input information of a user on a current video editing task and / or a historical video editing task;
[0159] a recommendation unit configured to determine, according to the input information, recommendation information on the current video editing task, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task;
[0160] an editing unit configured to generate an editing result of the current video editing task according to the recommendation information and the input information.
[0161] According to one or more embodiments of the present disclosure, when the editing unit generates the editing result of the current video editing task according to the recommendation information and the input information, the editing unit is configured to:
[0162] generate a target video of the current video editing task according to the recommendation information and the input information; or
[0163] generate a target video editing draft of the current video editing task according to the recommendation information and the input information; or
[0164] update an initial video editing draft of the current video editing task according to the recommendation information and the input information to obtain a target video editing draft of the current video editing task.
[0165] According to one or more embodiments of the present disclosure, when the recommendation unit determines the recommendation information on the current video editing task according to the input information, the recommendation unit is configured to:
[0166] determine the target intention according to the input information, and determine the recommendation information according to the target intention.
[0167] According to one or more embodiments of the present disclosure, when the recommendation unit determines the target intention according to the input information, the recommendation unit is configured to:
[0168] determine the target intention according to a preset rule based on the input information; or
[0169] call an intention analysis model with the input information as input, and output the target intention.
[0170] According to one or more embodiments of the present disclosure, the recommendation unit, when determining the target intention according to the input information, is configured to:
[0171] determine the target intention according to a preset rule based on the input information;
[0172] if the target intention cannot be confirmed, call an intention analysis model with the input information as input, and output the target intention.
[0173] According to one or more embodiments of the present disclosure, the recommendation unit, when determining the target intention according to the input information, is configured to:
[0174] determine the target intention according to the input information and context information of the input information.
[0175] According to one or more embodiments of the present disclosure, the recommendation unit, when determining the target intention according to the input information, is configured to:
[0176] determine an editing scene and an editing function associated with an editing effect of the current video editing task according to the input information, and determine the editing scene and the editing function as the target intention.
[0177] According to one or more embodiments of the present disclosure, the recommendation unit, when determining an editing scene and an editing function associated with an editing effect of the current video editing task according to the input information and determining the editing scene and the editing function as the target intention, is configured to:
[0178] if only the editing function can be determined, determine a historical editing scene as the editing scene, and determine the editing scene and the editing function as the target intention.
[0179] According to one or more embodiments of the present disclosure, the recommendation unit, when determining the recommendation information according to the target intention, is configured to:
[0180] obtain one or more recommendation information from a target recommendation information set corresponding to the target intention in the preset recommendation information library, wherein the preset recommendation information library includes recommendation information sets corresponding to different intentions.
[0181] According to one or more embodiments of the present disclosure, the recommendation unit, when obtaining one or more recommendation information from a target recommendation information set corresponding to the target intention in the preset recommendation information library, is configured to:
[0182] determine a target subject category of the input information;
[0183] obtain one or more recommendation information from a target recommendation information sub-set corresponding to the target theme category in the target recommendation information set, wherein the target recommendation information set comprises recommendation information sub-sets corresponding to different theme categories.
[0184] According to one or more embodiments of the present disclosure, the editing unit, when generating the editing result of the current video editing task according to the recommendation information and the input information, is configured to:
[0185] generate a target editing instruction for the current video editing task according to the recommendation information and the input information, and generate the editing result of the current video editing task according to the target editing instruction.
[0186] According to one or more embodiments of the present disclosure, the editing unit, when generating a target editing instruction for the current video editing task according to the recommendation information and the input information, is configured to:
[0187] generate a supplement instruction for the input information according to the recommendation information and the input information, and splice the supplement instruction with the input information to obtain the target editing instruction; or
[0188] generate a complete target editing instruction according to the recommendation information and the input information, and replace the input information with the target editing instruction.
[0189] According to one or more embodiments of the present disclosure, the editing unit, when generating a target editing instruction for the current video editing task according to the recommendation information and the input information, is configured to:
[0190] generate a plurality of alternative target editing instructions according to the recommendation information and the input information, and display the plurality of alternative target editing instructions;
[0191] In response to a selection operation instruction of a user on the plurality of alternative target editing instructions, determine the selected alternative target editing instruction as the final target editing instruction.
[0192] According to one or more embodiments of the present disclosure, the editing unit, when generating the editing result of the current video editing task according to the target editing instruction, is configured to:
[0193] call a video editing model with the target editing instruction as input, so that the video editing model generates the editing result of the current video editing task according to the target editing instruction.
[0194] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising at least one processor and a memory;
[0195] The memory stores computer-executable instructions;
[0196] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video editing method as described in the first aspect and various possible designs of the first aspect.
[0197] In a fourth aspect, a computer-readable storage medium is provided according to one or more embodiments of the present disclosure, and the computer-readable storage medium stores computer-executable instructions, when a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible designs of the first aspect is implemented.
[0198] In a fifth aspect, a computer program product is provided according to one or more embodiments of the present disclosure, and the computer program product includes computer-executable instructions, when a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible designs of the first aspect is implemented.
[0199] The above description is merely exemplary of the present disclosure and the application of the principles thereof and the scope of the disclosure is not limited to the specific embodiments described herein, but only by the claims that follow. It will be readily apparent to those skilled in the art that varying substitutions and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. For example, features described herein can be combined, swapped, or eliminated in any combination.
[0200] In addition, while operations are depicted in a particular, chronological sequence in this disclosure, this should not be understood as requiring such order unless specifically specified. Certain activities can be performed in different order or concurrently with each other. Additionally, certain activities can be performed simultaneously. Similarly, while several specific implementations have been contended, these need not be the only implementations. Other implementations can be configured in accordance with the same or similar principles as discussed above. In other words, unless otherwise specified, components, qualities, and / or operations can be substituted with like components, qualities, and / or operations in order to achieve the same or similar result.
[0201] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A video editing method, comprising: obtaining input information of a user on a current video editing task and / or a historical video editing task; determining recommendation information on the current video editing task according to the input information, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task; and generating an editing result of the current video editing task according to the recommendation information and the input information.
2. The method of claim 1, wherein the generating the editing result of the current video editing task according to the recommendation information and the input information comprises: generating a target video of the current video editing task according to the recommendation information and the input information; or generating a target video editing draft of the current video editing task according to the recommendation information and the input information; or updating an initial video editing draft of the current video editing task to obtain a target video editing draft of the current video editing task according to the recommendation information and the input information.
3. The method of claim 1, wherein the determining the recommendation information on the current video editing task according to the input information comprises: determining the target intention according to the input information, and determining the recommendation information according to the target intention.
4. The method of claim 3, wherein the determining the target intention according to the input information comprises: determining the target intention according to a preset rule based on the input information; or inputting the input information into an intention analysis model to output the target intention.
5. The method of claim 3, wherein the determining the target intention according to the input information comprises: determining the target intention according to a preset rule based on the input information; and if the target intention cannot be determined, inputting the input information into an intention analysis model to output the target intention.
6. The method of any one of claims 3-5, wherein the determining the target intention according to the input information comprises: determining the target intention according to the input information and context information of the input information.
7. The method of any one of claims 3-5, wherein the determining the target intention according to the input information comprises: determining an editing scene and an editing function associated with an editing effect of the current video editing task according to the input information, and determining the editing scene and the editing function as the target intention.
8. The method of claim 7, wherein the determining the editing scene and the editing function associated with the editing effect of the current video editing task according to the input information, and determining the editing scene and the editing function as the target intention comprises: if only the editing function can be determined, determining a historical editing scene as the editing scene, and determining the editing scene and the editing function as the target intention.
9. The method of any one of claims 3-5, wherein the determining the recommendation information according to the target intention comprises: obtain one or more recommendation information from a target recommendation information set corresponding to the target intention in a preset recommendation information library, wherein the preset recommendation information library comprises recommendation information sets corresponding to different intentions.
10. The method of claim 9, wherein the obtaining one or more recommendation information from a target recommendation information set corresponding to the target intention in a preset recommendation information library comprises: determining a target subject category of the input information; and obtaining one or more recommendation information from a target recommendation information sub-set corresponding to the target subject category in the target recommendation information set, wherein the target recommendation information set comprises recommendation information sub-sets corresponding to different subject categories.
11. The method of any one of claims 1-5, wherein the generating an editing result of the current video editing task according to the recommendation information and the input information comprises: generating a target editing instruction for the current video editing task according to the recommendation information and the input information, and generating the editing result of the current video editing task according to the target editing instruction.
12. The method of claim 11, wherein the generating a target editing instruction for the current video editing task according to the recommendation information and the input information comprises: generating a supplement instruction for the input information according to the recommendation information and the input information, and splicing the supplement instruction with the input information to obtain the target editing instruction; or generating a complete target editing instruction according to the recommendation information and the input information, and replacing the input information with the target editing instruction.
13. The method of claim 11, wherein the generating a target editing instruction for the current video editing task according to the recommendation information and the input information comprises: generating a plurality of alternative target editing instructions according to the recommendation information and the input information, and displaying the plurality of alternative target editing instructions; and determining a selected alternative target editing instruction as a final target editing instruction in response to a user selection operation instruction on the plurality of alternative target editing instructions.
14. The method of claim 11, wherein the generating an editing result of the current video editing task according to the target editing instruction comprises: inputting the target editing instruction into a video editing model to enable the video editing model to generate the editing result of the current video editing task according to the target editing instruction.
15. A video editing device, comprising: an input unit configured to obtain input information of a user on a current video editing task and / or a historical video editing task; a recommendation unit configured to determine recommendation information for the current video editing task according to the input information, the recommendation information being used to prompt a target intention of an editing effect of the current video editing task; and an editing unit configured to generate an editing result of the current video editing task according to the recommendation information and the input information. at least one processor and a memory; the memory stores computer execution instructions; 16. An electronic device comprising: The at least one processor executes the computer-executable instructions stored in the memory such that the at least one processor performs the method of any one of claims 1-14.
17. A computer-readable storage medium having computer-executable instructions stored therein such that, when executed by a processor, the computer-executable instructions implement the method of any one of claims 1-14.
18. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1-14.
Citation Information
Patent Citations
Information recommendation method and device, equipment, storage medium and program product
CN117009494A
Conversation-based video editing method and device, electronic equipment and storage medium
CN117714784A
Audio and video editing method, system and device and storage medium
CN117831528A
Video editing material generation method, client, server and system
CN117835008A
Organic compound and organic electroluminescent device using the same
KR102826881B1