A media editing method, device, electronic device and storage medium based on AI big model

By applying the media editing method based on AI model in video editing, the constraints of obtaining video copy are solved, and the problems of low efficiency and poor quality caused by relying on manual input in the prior art are achieved, and efficient and accurate video copy generation are achieved.

CN118354120BActive Publication Date: 2025-05-09BEIJING QICHUANG TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410450705.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2025-05-09
Estimated Expiration
2044-04-15

AI Technical Summary

Technical Problem

Existing video editing technologies rely on user external manual input constraints when automatically generating video copy, and the generation efficiency is low, which easily leads to poor copy quality due to input errors.

Method used

Using a media editing method based on AI big model, we use the AI ​​dialogue interface to receive user input constraint content and sample videos, and automatically obtain custom and format constraints, and combine the implicit constraints of the video content itself to generate efficient and accurate video description text.

Benefits of technology

It realizes efficient generation of short video copy without relying on a large amount of external manual input, improving copy quality and generation efficiency, and reducing the impact of manual errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118354120B_ABST
    Figure CN118354120B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a media editing method based on an AI big model, the method comprising, in response to obtaining a generation instruction for generating text content for a target media segment, displaying an AI dialogue interface corresponding to the generation instruction, the AI ​​dialogue interface comprising at least an input control, the input control being used to receive user input constraint content and / or for receiving a user-specified sample video to obtain a first constraint condition for the target media segment; in response to receiving user input constraint content and / or receiving a user-specified sample video through the input control, displaying a target text content item, the target text content item being at least one text content item in a set of candidate text content items, the candidate text content item being a text content item without a display mark among the M text content items in the Nth text content set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a media editing method, device, equipment and storage medium based on an AI big model. Background Art

[0002] Currently in the field of video editing, the method for automatically generating text based on videos is relatively rough, requiring users to provide more external manual input constraints, and is highly dependent on external manual input constraints. Once the external manual constraints are input inaccurately, it is difficult to obtain good video text, and the generation efficiency is low. Summary of the invention

[0003] An embodiment of the present invention provides a media editing method based on an AI big model, which is used to use the AI ​​big model to quickly generate short video copy content from video materials shot by users.

[0004] In a first aspect, an embodiment of the present invention provides a media editing method based on an AI big model, the method comprising:

[0005] In response to obtaining a generation instruction for generating text content for a target media segment,

[0006] Displaying an AI dialogue interface corresponding to the generation instruction, the AI ​​dialogue interface at least comprising an input control, the input control being used to receive user input constraint content and / or being used to receive a sample video specified by a user, so as to obtain a first constraint condition of the target media segment, wherein the first constraint condition comprises a custom constraint condition and / or a formatting constraint condition, the custom constraint condition is extracted from the constraint content by a first large language model, and the sample video extraction formatting constraint condition is extracted from the sample video by a third large language model;

[0007] In response to receiving user input constraint content through the input control and / or receiving a user-specified sample video,

[0008] Displaying a target text content item, wherein the target text content item is at least one text content item in the set of candidate text content items, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set;

[0009] Among them, the Nth text content set is generated by the Nth input based on the first constraint and the second constraint, and is obtained by the input target large language model. The second constraint is generated based on the target media segment, and the second constraint includes at least a content constraint and a text length constraint.

[0010] In a second aspect, an embodiment of the present invention provides a media editing device, the device comprising: an acquisition module, an extraction module, and an identification module;

[0011] An input module, which is used to display an AI dialogue interface corresponding to the generation instruction in response to obtaining a generation instruction for generating text content for a target media segment, wherein the AI ​​dialogue interface at least includes an input control; the input control is used to receive user input constraint content and / or to receive a sample video specified by a user to obtain a first constraint condition for the target media segment, wherein the first constraint condition includes a custom constraint condition and / or a formatting constraint condition, the custom constraint condition is extracted from the constraint content by a first language model, and the sample video extraction formatting constraint condition is extracted from the sample video by a third language model;

[0012] A display module, which is used to display a target text content item in response to receiving user input constraint content and / or receiving a sample video specified by a user through the input control, wherein the target text content item is at least one text content item in the candidate text content item set, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set;

[0013] Among them, the Nth text content set is generated by the Nth input based on the first constraint and the second constraint, and is obtained by the input target large language model. The second constraint is generated based on the target media segment, and the second constraint includes at least a content constraint and a text length constraint.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the media editing method based on the AI ​​big model as described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor can at least implement the media editing method based on the AI ​​big model as described in the first aspect.

[0016] In the solution provided by an embodiment of the present invention, in response to obtaining the first constraint condition for the target media segment, the first constraint condition at least includes a custom constraint condition and / or a formatting constraint condition, a second constraint condition is generated according to the target media segment, the second constraint condition at least includes a content constraint condition and a text length constraint condition, and an Nth input is generated according to the first constraint condition and the second constraint condition, and input into a target large language model to obtain an Nth text content set corresponding to the target media segment.

[0017] Based on the solution provided by the embodiment of the present invention, it can be seen that the calculation solution of the present invention automatically obtains constraints from multiple dimensions, not only supports obtaining explicit first constraints, but also obtains implicit second constraints from the video content itself, and combines the first constraints and the second constraints as the input of the large model, so that the video description copy can be obtained more efficiently and accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A first exemplary diagram of a media editing method based on an AI big model provided by an editing terminal according to an embodiment of the present invention;

[0020] Figure 2 A second example diagram of a media editing method based on an AI big model provided by an editing terminal according to an embodiment of the present invention;

[0021] Figure 3 A third exemplary diagram of a media editing method based on an AI big model provided by an editing terminal according to an embodiment of the present invention;

[0022] Figure 4 A flowchart of an execution of a media editing method based on an AI big model provided by an embodiment of the present invention;

[0023] Figure 5 A schematic diagram of the structure of a media editing device provided by an embodiment of the present invention;

[0024] Figure 6 For Figure 5 A schematic structural diagram of an electronic device corresponding to the media editing device provided in the illustrated embodiment. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two, but does not exclude the inclusion of at least one.

[0027] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0028] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0029] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a product or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such a product or system. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the product or system including the elements.

[0030] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0031] The media editing method based on the AI ​​big model provided in the embodiment of the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, or a server. The server can be a physical server including an independent host, or a virtual server, or a cloud server or a server cluster.

[0032] Figure 1 The first example diagram of the media editing method based on the AI ​​big model of the editing terminal provided in the embodiment of the present invention is as follows Figure 1 The content displayed in the image user interface shown in the middle interface M at least includes an object display area and an editing area, the object display area displays the target media segment content 101, and the editing area is loaded with the media frame 102 of the target media segment, the object display area is used to display the content of each frame of each media segment to be edited, the editing area is used to edit the media frame of each media segment to be edited, and the editing area at least includes an editing line for loading the media frame of each media segment to be edited. The target media segment in this application is a media content composed of at least one frame of image.

[0033] Embodiment 1

[0034] The execution process of a media editing method based on an AI large model provided in the first embodiment of the present invention is as shown in the attached figure. Figure 4 As shown, specifically including:

[0035] Step S1, in response to obtaining a generation instruction for generating text content for a target media segment, displaying an AI dialogue interface corresponding to the generation instruction, wherein the AI ​​dialogue interface includes at least an input control; in this embodiment, the AI ​​dialogue interface includes at least a guide control, and the guide control is used to guide a user to select an input control;

[0036] For example, the AI ​​dialogue interface of the first embodiment is as shown in the attached Figure 1 As shown in the middle interface A2, the guide controls displayed respectively include a guide control 103 "Have ideas, AI assists in creation" and a guide control 104 "Find inspiration, AI randomly generates".

[0037] Step S2, in response to receiving a user input constraint content through an input control, obtaining a first constraint condition of the target media segment, wherein the first constraint condition at least includes a custom constraint condition, the custom constraint condition is a custom condition extracted according to the constraint content, and the input control includes a first input control;

[0038] Exemplarily, in this embodiment, in response to the user attaching Figure 1 In the middle interface A2, select the guide control 103 "Have ideas, AI assists creation", and the following is displayed. Figure 1 Middle interface A3, displaying a first input control 107 (exemplary input control) in the AI ​​dialogue interface, so that the user can input constraint content through the first input control 107, so that the custom constraint condition is obtained through the first language model;

[0039] Preferably, the above constraint content can be sent to the server, and the above constraint content can be processed by the first largest language model deployed on the server to obtain the first constraint condition; if the local computing power is sufficient, the first largest language model can also be deployed locally on the terminal, so that after obtaining the above constraint content, the above constraint content can be directly processed locally to obtain the first constraint condition.

[0040] The first language model is obtained after a first training set consisting of specific content input, and can be used to extract a multi-dimensional description tag of a custom description content (i.e., a content constraint prompt word generated by the first language model) as a first constraint condition. Specifically, the large model for generating a prompt according to a text description belongs to the category of open source, and will not be described in detail in this embodiment;

[0041] Preferably, it is also possible to Figure 1 A corresponding guidance message is outputted in the middle interface A3, wherein the guidance message is to guide the user to input constraint content through the first input control 107, so as to obtain the custom constraint condition through the first language model;

[0042] Furthermore, the first constraint at this time includes a custom constraint. After receiving the explicit first constraint, it is necessary to obtain an implicit second constraint according to the target video segment. Specifically, the second constraint is generated according to the target media segment. The second constraint includes at least a content constraint and a text length constraint. The content constraint is a constraint on text content extracted according to the content of each frame of the target media segment, and the length constraint is a constraint on text length extracted according to the duration of the target media segment.

[0043] Preferably, the target video segment can be sent to the server, and the target video segment can be processed by the second largest language model deployed on the server to obtain the second constraint. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above-mentioned target video segment can be directly processed locally to obtain the second constraint. Specifically, the second largest language model is a large language model that has the ability to understand the video content and is obtained after training with a second training set composed of a specific video set, and is used to identify the specific content of the video. The second largest language model can be used to extract the second constraint for the target video segment. Specifically, the second largest language model can be used to extract the multi-dimensional content label of each frame (that is, the content label prompt word generated by the second largest language model) as the content constraint, and then the length of the text (that is, the text length prompt word prompt) is determined according to the length of the video segment as the text length constraint;

[0044] It should be noted that the length constraint of the text determined according to the length of the video clip can be identified by the second largest language model, or can be directly obtained according to the video length. The target video clip can be sent to the server, and the target video clip can be processed by the second largest language model deployed on the server to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above target video clip can be directly processed locally to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length, which will not be described in detail here.

[0045] For example, in this embodiment, as shown in the attached Figure 1 The duration of the media frame 102 of the target media segment shown in the middle interface M is T, and the text length constraint condition is obtained according to the duration T; Figure 1 The content 101 of the target media segment shown in the middle interface M is identified by using the second largest language model to obtain the content label of each frame of the target media segment for generating content constraints.

[0046] Further, the Nth input is generated according to the first constraint and the second constraint, and the target large language model is input to obtain the Nth text content set corresponding to the target media segment, and the Nth text content set includes M text content items. It should be noted that the target large language model involved in this embodiment is obtained after training with a specific data set based on an open source model. In order to facilitate user selection, a text content set can be obtained after one input, and a text content set can include multiple text content items.

[0047] Specifically, the first constraint condition obtained through the first largest language model (i.e., the multi-dimensional description label of the custom description content (which is the content constraint prompt word generated by the first largest language model)) and the second constraint condition obtained through the second largest language model (the multi-dimensional content label of each frame of the target video clip (which is the content label prompt word generated by the second largest language model), the length of the text (which is the prompt word prompt for the text length)) are combined as the prompt of the target large language model to obtain a text content set adapted to the target video clip.

[0048] The calculation scheme of the present invention automatically obtains constraints from multiple dimensions, and not only supports obtaining explicit first constraints, but also obtains implicit second constraints from the video content itself. The first constraints and the second constraints are combined as the input of the large model, so that the video description text can be obtained more efficiently and accurately.

[0049] Step S3, displaying a target text content item, wherein the target text content item is at least one text content item in a set of candidate text content items, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set.

[0050] Furthermore, a display mark is added to the displayed text content items among the M text content items, and in response to receiving a switch instruction for the target text content item, at least one text content item is selected from the candidate text content item set as the target text content item.

[0051] For example, as shown in the attached Figure 1 As shown in the middle interface N, in response to receiving a switch instruction for the target text content item sent by the user through the switch control 109 "Change", the displayed text content item is marked as displayed. At this time, at least one text content item is selected from the candidate text content item set as the target text content item for display.

[0052] Further, in response to the number of candidate text content items in the set of candidate text content items reaching a preset number threshold;

[0053] Exemplarily, when only 2 of the M content items are left without display marks, it can be considered that the number of candidate text content items is 2, reaching the preset number threshold, and it is necessary to obtain more text content items again, and obtain the N+1th text content set again based on the Nth input, wherein the N+1th text content set includes M text content items; the M text content items in the N+1th text content set are used and added to the candidate text content item set for displaying the content items after the switching instruction is obtained later. In this way, by performing the Nth input once, multiple results can be obtained at one time, avoiding frequent calls to the model, simplifying the content generation process, and improving the content generation efficiency.

[0054] Embodiment 2

[0055] The second embodiment of the present invention provides an execution process of a media editing method based on an AI large model, which specifically includes:

[0056] Step K1, in response to obtaining a generation instruction for generating text content for a target media segment, displaying an AI dialogue interface corresponding to the generation instruction, wherein the AI ​​dialogue interface includes at least an input control; in this embodiment, the AI ​​dialogue interface includes at least a guide control, and the guide control is used to guide a user to select an input control;

[0057] For example, the AI ​​dialogue interface of this embodiment is as shown in the attached Figure 2 As shown in the middle interface B2, the guide controls displayed respectively include guide control 103 "Have ideas, AI assists in creation" and guide control 104 "Find inspiration, AI randomly generates".

[0058] Step K2, in response to receiving a sample video specified by a user through an input control, a first constraint condition for acquiring the target media segment, wherein the first constraint condition at least includes a formatting constraint condition, wherein the formatting constraint condition is a formatting condition extracted by parsing the sample video selected by the user; wherein the input control includes a selection control or a second input control; wherein the selection control is used to select a sample video in a video category, and the second input control is used to input a sample video address to acquire the sample video;

[0059] Exemplarily, in this embodiment, in response to the user attaching Figure 2 In the middle interface B2, the guide control 104 "Find inspiration, AI randomly generates" can be selected to display the video stream of the classified video in the AI ​​dialogue interface, so that the user can select the sample video in the video category through the selection control (exemplary input control); the second input control 108 (exemplary input control) can also be displayed in the AI ​​dialogue interface, so that the user can input the sample video address through the second input control to obtain the sample video; as shown in the attached figure Figure 2 As shown in the middle interface B3.

[0060] For example, after obtaining the sample video, the sample video can be identified and processed, and the text style type of the sample video can be extracted as the formatting constraint condition. Optionally, the sample video set can be classified and pre-processed in advance, and the sample video can be identified using a traditional artificial intelligence model to obtain a multi-dimensional style label of the sample video. This is a common prior art and will not be described in detail here.

[0061] Preferably, the above-mentioned sample video can be sent to the server, and the sample video can be processed by the third language model deployed on the server to obtain the formatting constraints. If the local computing power is sufficient, the first language model can also be deployed locally on the terminal. In this way, after obtaining the above-mentioned sample video, the above-mentioned sample video can be directly processed locally to obtain the formatting constraints.

[0062] Specifically, the third largest language model is a large language model that is obtained after training with a third training set consisting of a large number of video sample sets and has the ability to understand video content. Unlike the second largest language model mentioned above, the third largest language model is mainly used to perceive and identify the type of sample videos, such as identifying the emotional type of the video, whether it is a funny type or a literary type, etc., identifying the attribute style classification of the video, whether it is a daily life record or oral knowledge, etc. The third largest language model can be used to extract multi-dimensional style labels for sample videos (i.e., the prompt word of the style type generated by the third largest language model) as formatting constraints.

[0063] Furthermore, the first constraint at this time includes a formatting constraint. After receiving the explicit first constraint, it is necessary to obtain an implicit second constraint according to the target video segment. Specifically, the second constraint is generated according to the target media segment. The second constraint includes at least a content constraint and a text length constraint. The content constraint is a constraint on the text content extracted according to the content of each frame of the target media segment, and the length constraint is a constraint on the text length extracted according to the duration of the target media segment.

[0064] Preferably, the target video segment can be sent to the server, and the target video segment can be processed by the second largest language model deployed on the server to obtain the second constraint. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above-mentioned target video segment can be directly processed locally to obtain the second constraint. Specifically, the second largest language model is a large language model that has the ability to understand the video content and is obtained after training with a second training set composed of a specific video set, and is used to identify the specific content of the video. The second largest language model can be used to extract the second constraint for the target video segment. Specifically, the second largest language model can be used to extract the multi-dimensional content label of each frame (that is, the content label prompt word generated by the second largest language model) as the content constraint, and then the length of the text (that is, the text length prompt word prompt) is determined according to the length of the video segment as the text length constraint;

[0065] It should be noted that the length constraint of the text determined according to the length of the video clip can be identified by the second largest language model, or can be directly obtained according to the video length. The target video clip can be sent to the server, and the target video clip can be processed by the second largest language model deployed on the server to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above target video clip can be directly processed locally to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length, which will not be described in detail here.

[0066] For example, in this embodiment, as shown in the attached Figure 1 The duration of the media frame 102 of the target media segment shown in the middle interface M is T, and the second largest language model can be used to obtain the text length constraint condition according to the duration T; Figure 1 The content 101 of the target media segment shown in the middle interface M can be identified by using the second largest language model to obtain the content label of each frame of the target media segment, so as to generate content constraints.

[0067] Further, the Nth input is generated according to the first constraint and the second constraint, and the target large language model is input to obtain the Nth text content set corresponding to the target media segment, and the Nth text content set includes M text content items. It should be noted that the target large language model involved in this embodiment is obtained after training with a specific data set based on an open source model. In order to facilitate user selection, a text content set can be obtained after one input, and a text content set can include multiple text content items.

[0068] Specifically, the first constraint condition obtained through the third largest language model (i.e., the multi-dimensional style label extracted from the sample video (which is the prompt word prompt of the style type generated by the third largest language model)) and the second constraint condition obtained through the second largest language model (the multi-dimensional content label of each frame of the target video clip (which is the content label prompt word generated by the second largest language model), the length of the text (which is the prompt word prompt of the text length)) are combined as the prompt of the target large language model to obtain a text content set adapted to the target video clip.

[0069] The calculation scheme of the present invention automatically obtains constraints from multiple dimensions, and not only supports obtaining explicit first constraints, but also obtains implicit second constraints from the video content itself. The first constraints and the second constraints are combined as the input of the large model, so that the video description text can be obtained more efficiently and accurately.

[0070] Step K3, displaying a target text content item, wherein the target text content item is at least one text content item in the candidate text content item set, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set.

[0071] Furthermore, a display mark is added to the displayed text content items among the M text content items, and in response to receiving a switch instruction for the target text content item, at least one text content item is selected from the candidate text content item set as the target text content item.

[0072] For example, as shown in the attached Figure 1 As shown in the middle interface N, in response to receiving a switch instruction for the target text content item sent by the user through the switch control 109 "Change", the displayed text content item is marked as displayed. At this time, at least one text content item is selected from the candidate text content item set as the target text content item for display.

[0073] Further, in response to the number of candidate text content items in the set of candidate text content items reaching a preset number threshold;

[0074] Exemplarily, when only 2 of the M content items are left without display marks, it can be considered that the number of candidate text content items is 2, reaching the preset number threshold, and it is necessary to obtain more text content items again, and obtain the N+1th text content set again based on the Nth input, wherein the N+1th text content set includes M text content items; the M text content items in the N+1th text content set are used and added to the candidate text content item set for displaying the content items after the switching instruction is obtained later. In this way, by performing the Nth input once, multiple results can be obtained at one time, avoiding frequent calls to the model, simplifying the content generation process, and improving the content generation efficiency.

[0075] Embodiment 3

[0076] The execution process of a media editing method based on an AI large model provided in the third embodiment of the present invention specifically includes:

[0077] Step P1, in response to obtaining a generation instruction for generating text content for a target media segment, displaying an AI dialogue interface corresponding to the generation instruction, wherein the AI ​​dialogue interface includes at least an input control; in this embodiment, the AI ​​dialogue interface includes at least a guide control, wherein the guide control is used to guide a user to select an input control, wherein the input control includes a first input control and a second input control; or the input control includes a first input control and a selection control;

[0078] For example, the AI ​​dialogue interface of this embodiment is as shown in the attached Figure 3 As shown in the middle interface C2, the guide controls displayed respectively include a guide control 106 “My Ideas” and a guide control 105 “Find Inspiration”.

[0079] Step P2, in response to receiving a user input constraint content through a first input control to obtain a custom constraint condition of the target media segment, and receiving a user-specified sample video through a second input control to obtain a formatting constraint condition of the target media segment, wherein the custom constraint condition is a custom condition extracted based on the constraint content, and the formatting constraint condition is a formatting condition extracted by parsing the sample video selected by the user;

[0080] Exemplarily, in this embodiment, in response to the user attaching Figure 3In the middle interface C2, the guide control 105 "Find Inspiration" can be selected, and the video stream of the classified video can be displayed on the AI ​​dialogue interface so that the user can select the target video in the sample video category; the second input control 108 can also be displayed on the AI ​​dialogue interface so that the user can enter the sample video address through the second input control to obtain the sample video; the video stream of the classified video and the second input control 108 can also be displayed on the AI ​​dialogue interface, as shown in the attached figure. Figure 3 As shown in the middle interface C3.

[0081] Exemplarily, after obtaining the sample video, the sample video may be subjected to recognition processing, and the text style type of the sample video may be extracted as the formatting constraint condition.

[0082] Optionally, the sample video set may be classified and preprocessed in advance, and the sample videos may be identified using a traditional artificial intelligence model to obtain multi-dimensional style labels for the sample videos. This is a relatively common prior art and will not be elaborated here.

[0083] Preferably, the above-mentioned sample video can be sent to the server, and the sample video can be processed by the third language model deployed on the server to obtain the formatting constraints. If the local computing power is sufficient, the first language model can also be deployed locally on the terminal. In this way, after obtaining the above-mentioned sample video, the above-mentioned sample video can be directly processed locally to obtain the formatting constraints.

[0084] Specifically, the third largest language model is a large language model that is obtained after training with a third training set consisting of a large number of video sample sets and has the ability to understand video content. Unlike the second largest language model mentioned above, the third largest language model is mainly used to perceive and identify the type of sample videos, such as identifying the emotional type of the video, whether it is a funny type or a literary type, etc., identifying the attribute style classification of the video, whether it is a daily life record or oral knowledge, etc. The third largest language model can be used to extract multi-dimensional style labels for sample videos (i.e., the prompt word of the style type generated by the third largest language model) as formatting constraints.

[0085] In addition, further, in this embodiment, the user continues to input the custom constraint condition. For example, in this embodiment, in response to the user further inputting the custom constraint condition through the attachment, Figure 3 In the middle interface C3, select the guide control 106 "My Idea", then the following is displayed Figure 3 Middle interface C4, displaying the first input control 107 in the AI ​​dialogue interface, so that the user can input constraint content through the first input control 107 to generate a custom constraint condition;

[0086] Preferably, the above constraint content can be sent to the server, and the above constraint content can be processed by the first largest language model deployed on the server to obtain the first constraint condition; if the local computing power is sufficient, the first largest language model can also be deployed locally on the terminal, so that after obtaining the above constraint content, the above constraint content can be directly processed locally to obtain the first constraint condition.

[0087] The first language model is obtained after training with a first training set consisting of specific content input, which can be used to extract

[0088] A multi-dimensional description tag for a custom description content (i.e., a content constraint prompt word generated by the first language model) is used as the first constraint condition. Specifically, the large model for generating prompts based on text descriptions belongs to the category of open source, and will not be described in detail in this embodiment.

[0089] Preferably, it is also possible to Figure 1 A corresponding guidance message is outputted in the middle interface A3, wherein the guidance message is for guiding the user to input content through the first input control 107, so as to generate the custom constraint condition through the first large language model;

[0090] Furthermore, the first constraint at this time includes a custom constraint and a formatted constraint. After receiving the explicit first constraint, it is necessary to obtain an implicit second constraint according to the target video segment. Specifically, the second constraint is generated according to the target media segment. The second constraint includes at least a content constraint and a text length constraint. The content constraint is a constraint on the text content extracted according to the content of each frame of the target media segment, and the length constraint is a constraint on the text length extracted according to the duration of the target media segment.

[0091] Preferably, the target video segment can be sent to the server, and the target video segment can be processed by the second largest language model deployed on the server to obtain the second constraint. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above-mentioned target video segment can be directly processed locally to obtain the second constraint. Specifically, the second largest language model is a large language model that has the ability to understand the video content and is obtained after training with a second training set composed of a specific video set, and is used to identify the specific content of the video. The second largest language model can be used to extract the second constraint for the target video segment. Specifically, the second largest language model can be used to extract the multi-dimensional content label of each frame (that is, the content label prompt word generated by the second largest language model) as the content constraint, and then the length of the text (that is, the text length prompt word prompt) is determined according to the length of the video segment as the text length constraint;

[0092] It should be noted that the length constraint of the text determined according to the length of the video clip can be identified by the second largest language model, or can be directly obtained according to the video length. The target video clip can be sent to the server, and the target video clip can be processed by the second largest language model deployed on the server to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length. If the local computing power is sufficient, the second largest language model can also be deployed locally on the terminal, so that the above target video clip can be directly processed locally to obtain the content constraint condition. In addition, the length constraint condition can be directly obtained according to the video length, which will not be described in detail here.

[0093] For example, in this embodiment, as shown in the attached Figure 1 The duration of the media frame 102 of the target media segment shown in the middle interface M is T, and the text length constraint condition is obtained according to the duration T; Figure 1 The content 101 of the target media segment shown in the middle interface M is identified by using the second largest language model to obtain the content label of each frame of the target media segment for generating content constraints.

[0094] Further, the Nth input is generated according to the first constraint and the second constraint, and the target large language model is input to obtain the Nth text content set corresponding to the target media segment, and the Nth text content set includes M text content items. It should be noted that the target large language model involved in this embodiment is obtained after training with a specific data set based on an open source model. In order to facilitate user selection, a text content set can be obtained after one input, and a text content set can include multiple text content items.

[0095] Specifically, the formatting constraints in the first constraints obtained through the third largest language model (i.e., extracting multi-dimensional style labels for the sample video (i.e., the prompt word prompt of the style type generated by the third largest language model)), the custom constraints in the first constraints obtained through the first largest language model (i.e., the multi-dimensional description labels for custom description of content (which are the content constraint prompt word prompted generated by the first largest language model)), and the second constraints obtained through the second largest language model (the multi-dimensional content labels of each frame of the target video clip (which are the content label prompt word prompted generated by the second largest language model), the length of the text (which are the prompt word prompt of the text length)) are combined as the prompt of the target large language model to obtain a text content set adapted to the target video clip.

[0096] The calculation scheme of the present invention automatically obtains constraints from multiple dimensions, and not only supports obtaining explicit first constraints, but also obtains implicit second constraints from the video content itself. The first constraints and the second constraints are combined as the input of the large model, so that the video description text can be obtained more efficiently and accurately.

[0097] Step P3, displaying a target text content item, wherein the target text content item is at least one text content item in a set of candidate text content items, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set.

[0098] Furthermore, a display mark is added to the displayed text content items among the M text content items, and in response to receiving a switch instruction for the target text content item, at least one text content item is selected from the candidate text content item set as the target text content item.

[0099] For example, as shown in the attached Figure 1 As shown in the middle interface N, in response to receiving a switch instruction for the target text content item sent by the user through the switch control 109 "Change", the displayed text content item is marked as displayed. At this time, at least one text content item is selected from the candidate text content item set as the target text content item for display.

[0100] Further, in response to the number of candidate text content items in the set of candidate text content items reaching a preset number threshold;

[0101] Exemplarily, when only 2 of the M content items are left without display marks, it can be considered that the number of candidate text content items is 2, reaching the preset number threshold, and it is necessary to obtain more text content items again, and obtain the N+1th text content set again based on the Nth input, wherein the N+1th text content set includes M text content items; the M text content items in the N+1th text content set are used and added to the candidate text content item set for displaying the content items after the switching instruction is obtained later. In this way, by performing the Nth input once, multiple results can be obtained at one time, avoiding frequent calls to the model, simplifying the content generation process, and improving the content generation efficiency.

[0102] The following will describe in detail the media editing device of one or more embodiments of the present invention. Those skilled in the art will appreciate that these devices can be configured using commercially available hardware components through the steps taught in this solution.

[0103] Figure 5 A structural diagram of a media editing device provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the device comprises: an input module 11, a display module 12;

[0104] An input module 11, which is used to display an AI dialogue interface corresponding to the generation instruction in response to obtaining a generation instruction for generating text content for a target media segment, wherein the AI ​​dialogue interface at least includes an input control; the input control is used to receive user input constraint content and / or to receive a sample video specified by a user to obtain a first constraint condition for the target media segment, wherein the first constraint condition includes a custom constraint condition and / or a formatting constraint condition, the custom constraint condition is extracted from the constraint content by a first language model, and the sample video extraction formatting constraint condition is extracted from the sample video by a third language model;

[0105] A display module 12, which is used to display a target text content item in response to receiving user input constraint content and / or receiving a sample video specified by a user through the input control, wherein the target text content item is at least one text content item in the candidate text content item set, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set;

[0106] Among them, the Nth text content set is generated by the Nth input based on the first constraint and the second constraint, and is obtained by the input target large language model. The second constraint is generated based on the target media segment, and the second constraint includes at least a content constraint and a text length constraint.

[0107] The input module 11 is further specifically configured to display a first input control on the AI ​​dialogue interface, so that the user can input content through the first input control to form the custom constraint condition;

[0108] The input module 11 is further specifically used to output a corresponding guidance message in the AI ​​dialogue interface, wherein the guidance message is to guide the user to input content through the first input control to generate the custom constraint condition.

[0109] The input module 11 is also specifically used for

[0110] Displaying a video stream of the classified video and / or a second input control on the AI ​​dialogue interface,

[0111] So that the user can select the sample video in the sample video category of the video stream through the selection control, or input the sample video address through the second input control to obtain the sample video, so that the formatting constraint condition is generated according to the sample video.

[0112] The display module 12 is further specifically configured to add a display mark to the displayed text content strips among the M text content strips,

[0113] In response to receiving a switch instruction for the target text content item, at least one text content item is selected from the set of candidate text content items as the target text content item.

[0114] The display module 12 is further specifically configured to respond to the number of candidate text content items in the candidate text content item set reaching a preset number threshold;

[0115] Acquire an N+1th text content set again according to the Nth input, wherein the N+1th text content set includes M text content items;

[0116] M text content items in the N+1 text content sets are adopted and added to the candidate text content item set.

[0117] Figure 5 The device shown can execute the steps introduced in the aforementioned embodiments. For detailed execution process and technical effects, please refer to the description in the aforementioned embodiments, which will not be repeated here.

[0118] In one possible design, the above Figure 6 The structure of the media editing device shown can be implemented as an electronic device, such as Figure 6As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores executable code, and when the executable code is executed by the processor 22, the processor 22 can at least implement the media editing method based on the AI ​​large model provided in the above-mentioned embodiment.

[0119] In addition, an embodiment of the present invention provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the media editing method based on the AI ​​big model provided in the aforementioned embodiment.

[0120] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art may understand and implement the present invention without creative effort.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a computer product, and the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A media editing method based on an AI large model, which is applied to an editing terminal, characterized in that: The method comprises: In response to obtaining a generation instruction for generating text content for a target media segment, Displaying an AI dialogue interface corresponding to the generation instruction, the AI ​​dialogue interface at least comprising an input control, the input control being used to receive a user input constraint content and / or being used to receive a user specified sample video, so as to obtain a first constraint condition of the target media segment, wherein the first constraint condition comprises a custom constraint condition and / or a formatting constraint condition, the custom constraint condition is extracted from the constraint content by a first language model, and the formatting constraint condition is extracted from the sample video by a third language model; In response to receiving user input constraint content through the input control and / or receiving a user-specified sample video, Displaying a target text content item, wherein the target text content item is at least one text content item in the candidate text content item set, and the candidate text content item is a text content item without a display mark among the M text content items in the Nth text content set; In response to the number of candidate text content items in the set of candidate text content items reaching a preset number threshold; Acquire the N+1th text content set again according to the Nth input, wherein the N+1th text content set includes M text content items; Adopting M text content items from the N+1 text content sets, and adding them to the candidate text content item set; Among them, the Nth text content set is generated by the Nth input based on the first constraint and the second constraint, and is obtained by the input target large language model. The second constraint is generated based on the target media segment, and the second constraint includes at least a content constraint and a text length constraint.

2. The media editing method based on AI big model according to claim 1, characterized in that: The content constraint condition is a constraint condition on text content extracted from each frame content of the target media segment through the second largest language model, and the length constraint condition is a constraint condition on text length extracted from the duration of the target media segment.

3. The media editing method based on AI big model according to claim 1, characterized in that: The input control comprises a first input control, A first input control is displayed on the AI ​​dialogue interface so that the user can input content through the first input control to generate the custom constraint condition, or a corresponding guidance message is output on the AI ​​dialogue interface, wherein the guidance message guides the user to input content through the first input control to generate the custom constraint condition.

4. The media editing method based on AI big model according to claim 1, characterized in that: The input control includes a selection control, and the selection control is used to select the sample video in the sample video category of the video stream to generate the formatting constraint condition according to the sample video.

5. The media editing method based on AI big model according to claim 1, characterized in that: The input control includes a second input control, and the second input control is used to input a sample video address to obtain a sample video, so as to generate the formatting constraint condition according to the sample video.

6. The media editing method based on AI big model according to claim 1, characterized in that: Adding a display mark to the displayed text content strips among the M text content strips, In response to receiving a switch instruction for the target text content item, at least one text content item is selected from the set of candidate text content items as the target text content item.

7. A media editing device based on an AI large model, characterized in that: The device comprises: an acquisition module, an extraction module, and an identification module; wherein, An input module, which is used to display an AI dialogue interface corresponding to the generation instruction in response to obtaining a generation instruction for generating text content for a target media segment, wherein the AI ​​dialogue interface at least includes an input control; the input control is used to receive user input constraint content and / or to receive a sample video specified by a user to obtain a first constraint condition for the target media segment, wherein the first constraint condition includes a custom constraint condition and / or a formatting constraint condition, the custom constraint condition is extracted from the constraint content by a first language model, and the sample video extraction formatting constraint condition is extracted from the sample video by a third language model; A display module, which is used to display a target text content item in response to receiving user input constraint content and / or receiving a sample video specified by a user through the input control, wherein the target text content item is at least one text content item in a set of candidate text content items, and the candidate text content item is a text content item without a display mark among M text content items in an Nth text content set; It is also used to respond to the number of candidate text content items in the set of candidate text content items reaching a preset number threshold; Acquire the N+1th text content set again according to the Nth input, wherein the N+1th text content set includes M text content items; Adopting M text content items from the N+1 text content sets, and adding them to the candidate text content item set; Among them, the Nth text content set is generated by the Nth input based on the first constraint and the second constraint, and is obtained by the input target large language model. The second constraint is generated based on the target media segment, and the second constraint includes at least a content constraint and a text length constraint.

8. An electronic device, characterized in that: include: A memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the media editing method based on the AI ​​large model as described in any one of claims 1 to 6.

9. A non-transitory machine-readable storage medium, characterized in that: The non-temporary machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor executes the media editing method based on the AI ​​big model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image sharing method and device, terminal and storage medium

    CN114117270A

  • Video dubbing method and related device, electronic equipment and storage medium

    CN117177024A

  • Big language model-based copywriting generation method, apparatus and device, and storage medium

    CN117611254A

  • Method and device for converting text style, equipment and medium

    CN117829101A