Video processing method and device, electronic equipment and storage medium
By obtaining semantic description information of video and user video editing problems, and using a large language model to generate data to answer video editing problems, the problem of high operating threshold for existing video editing applications is solved and the user's video editing efficiency is improved.
Patent Information
- Application Number
- CN202311460063.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
The operating threshold of existing video editing applications is high, resulting in inefficient video editing for users and it takes a long time to edit satisfactory video works.
By obtaining semantic description information of the video and the user's video editing problems, a large language model is used to generate data to answer the video editing problems, thereby improving the user's video editing efficiency.
Through the semantic understanding ability of the large language model, accurate video editing data is generated, users' video editing process is simplified, and users' editing efficiency is improved.
Smart Images

Figure CN119946366A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video processing method, device, electronic device and storage medium. Background Art
[0002] Currently, users can record videos on their mobile phones, edit them as needed through video editing applications, and publish the edited videos on social platforms to meet users' social needs. However, the video editing functions of existing video editing applications are complex, and the user's video editing operation threshold is high, resulting in low video editing efficiency for users, and it takes a long time to edit a satisfactory video work. Summary of the invention
[0003] The embodiments of the present disclosure provide a video processing method, device, electronic device and storage medium, which can improve the editing efficiency of users in video editing.
[0004] In a first aspect, an embodiment of the present disclosure provides a video processing method, including:
[0005] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0006] Based on a pre-created video clip knowledge base, acquiring video clip knowledge information associated with the first data;
[0007] Through a large language model, based on the first data and the video clip knowledge information, second data corresponding to the video is generated, and the second data is configured to answer the video clip question.
[0008] In a second aspect, an embodiment of the present disclosure provides a video processing method, including:
[0009] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0010] Sending the first data to a server;
[0011] Receive second data corresponding to the video returned by the server according to the first data, where the second data is configured to answer the video editing question.
[0012] In a third aspect, an embodiment of the present disclosure provides a video processing device, including:
[0013] A first acquisition unit is used to acquire first data corresponding to a video; the first data includes semantic description information corresponding to the video and a video editing question of a user regarding the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to a historical editing operation of the video;
[0014] A second acquisition unit, configured to acquire video clip knowledge information associated with the first data based on a pre-created video clip knowledge base;
[0015] A data generating unit is used to generate second data corresponding to the video based on the first data and the video clip knowledge information through a large language model, wherein the second data is configured to answer the video clip question.
[0016] In a fourth aspect, an embodiment of the present disclosure provides a video processing device, including:
[0017] A third acquisition unit is used to acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0018] A sending unit, configured to send the first data to a server;
[0019] A receiving unit is used to receive second data corresponding to the video returned by the server based on the first data, where the second data is configured to answer the video editing problem.
[0020] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, enable the processor to implement the steps of the method described in the first aspect or the second aspect above.
[0021] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0022] In one or more embodiments of the present disclosure, first, the first data corresponding to the video is obtained, the first data corresponding to the video includes semantic description information corresponding to the video and video editing questions of the user for the video, the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video, then, based on a pre-created video editing knowledge base, video editing knowledge information associated with the first data corresponding to the video is obtained, and finally, through a large language model, based on the first data and the obtained video editing knowledge information, second data corresponding to the video is generated, and the second data is configured to answer the video editing question. It can be seen that through this embodiment, the second data required by the user can be generated by the large language model according to the semantic description information corresponding to the video and the video editing questions of the user for the video, and the second data can answer the video editing question. Based on the advantages of the large language model's strong semantic understanding ability and high accuracy of the output data, the accuracy of the second data is effectively improved, thereby facilitating the user to perform video editing according to the second data and improving the user's editing efficiency in video editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate one or more embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor.
[0024] Figure 1 A schematic diagram of a flow chart of a video processing method provided by an embodiment of the present disclosure;
[0025] Figure 2 A schematic diagram of a video processing method according to another embodiment of the present disclosure;
[0026] Figure 3 A schematic diagram of an application scenario of a video processing method provided by an embodiment of the present disclosure;
[0027] Figure 4 A schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure;
[0028] Figure 5 A schematic diagram of the structure of a video processing device provided by another embodiment of the present disclosure;
[0029] Figure 6 A schematic diagram of the structure of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the protection scope of the present disclosure.
[0031] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0032] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0033] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0035] The disclosed embodiment provides a video processing method, which can improve the editing efficiency of a user in video editing. Figure 1 A schematic diagram of a video processing method according to an embodiment of the present disclosure is shown in FIG. Figure 1 As shown, the method includes:
[0036] Step S102, obtaining first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0037] Step S104, acquiring video clip knowledge information associated with the first data based on a pre-created video clip knowledge base;
[0038] Step S106, using a large language model, based on the first data and video editing knowledge information, generates second data corresponding to the video, where the second data is configured to answer the video editing question.
[0039] In this embodiment, first, the first data corresponding to the video is obtained, and the first data corresponding to the video includes the semantic description information corresponding to the video and the video editing questions of the user for the video. The semantic description information corresponding to the video includes the first description information corresponding to the video content of the video and / or the second description information corresponding to the historical editing operation of the video. Then, based on the pre-created video editing knowledge base, the video editing knowledge information associated with the first data corresponding to the video is obtained. Finally, through the large language model, based on the first data and the obtained video editing knowledge information, the second data corresponding to the video is generated, and the second data is configured to answer the video editing questions. It can be seen that through this embodiment, the second data required by the user can be generated by the large language model according to the semantic description information corresponding to the video and the video editing questions of the user for the video. The second data can answer the video editing questions. Based on the advantages of the large language model's strong semantic understanding ability and high accuracy of the output data, the accuracy of the second data is effectively improved, thereby facilitating the user to perform video editing according to the second data and improving the user's editing efficiency in video editing.
[0040] The video processing method in this embodiment can be applied to a server and executed by the server. The specific process of the video processing method in this embodiment is described in detail below.
[0041] In the above step S102, the server obtains first data corresponding to the video, the first data including semantic description information corresponding to the video and video editing questions of the user for the video. The first data can be understood as problem description data corresponding to the video, including semantic description information corresponding to the video and video editing questions of the user for the video.
[0042] The semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video. In one embodiment, the semantic description information corresponding to the video includes first description information corresponding to the video content of the video. In another embodiment, the semantic description information corresponding to the video includes second description information corresponding to the historical editing operation of the video. In yet another embodiment, the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and second description information corresponding to the historical editing operation of the video.
[0043] The first description information corresponding to the video content of the video may be in text form, used to indicate the category of the video content. The second description information corresponding to the historical editing operation of the video may be in text form, used to indicate at least one of the operation name and operation category of the historical editing operation that the video has undergone. In one embodiment, after obtaining the first description information corresponding to the video content of the video, the first description information may be used as the semantic description information corresponding to the video. In another embodiment, after obtaining the second description information corresponding to the historical editing operation of the video, the second description information may be used as the semantic description information corresponding to the video. In yet another embodiment, after obtaining the first description information corresponding to the video content of the video and the second description information corresponding to the historical editing operation of the video, the first description information and the second description information are combined into the semantic description information corresponding to the video in a certain text format.
[0044] In one embodiment, the semantic description information corresponding to the video, in addition to the first description information and / or the second description information, may also include third description information corresponding to the basic information of the video. The third description information may be in text form and is used to indicate at least one of the number of videos and the length of the video.
[0045] In a specific example, the semantic description information corresponding to the video is: there are two videos on the main track, the first is a video of shooting food, with a length of 10 seconds, and the second is a video of shooting animals, with a length of 5 seconds. Background music has been added to the video, and the name of the background music is XXX, and the music category is XXX. Among them, "the first video is a video of shooting food, and the second video is a video of shooting animals" is the first description information, indicating that the content categories of the video content are food and animals, respectively, "the video has been added with background music, the name of the background music is XXX, and the music category is XXX" is the second description information, indicating that the historical editing operation of the video includes the operation of adding music, and "there are two videos on the main track, with a length of 10 seconds and a length of 5 seconds" is the third description information, indicating that there are two videos, with a length of 10 seconds and 5 seconds, respectively.
[0046] The video editing problem of the user for the video is used to indicate the problem encountered by the user during the video editing process. For example, the video editing problem may be "how to move the camera", "how to add cartoon face effects", etc. The video editing problem of the user for the video may be the question text.
[0047] In one embodiment, a user inputs a video editing question for a video in a terminal device, and the video editing question is in text form. The user's terminal device also obtains semantic description information corresponding to the video, combines the semantic description information corresponding to the video and the user's video editing question for the video into first data corresponding to the video in a certain text format, and sends the first data corresponding to the video to a server. The semantic description information corresponding to the video is also in text form, and the first data corresponding to the video is also in text form.
[0048] In another embodiment, the user inputs a video editing question for the video in a terminal device, and the video editing question is in text form. The user's terminal device sends the video editing question for the video input by the user to the server, and the server also obtains the semantic description information corresponding to the video, and combines the semantic description information corresponding to the video and the video editing question of the user for the video in a certain text format into the first data corresponding to the video. Among them, the semantic description information corresponding to the video is also in text form, and the first data corresponding to the video is also in text form.
[0049] In the above step S104, video clipping knowledge information associated with the first data is obtained based on a pre-created video clipping knowledge base. In this embodiment, a video clipping knowledge base is pre-created, and the video clipping knowledge base stores relevant knowledge information about video clipping. Therefore, in this step, video clipping knowledge information associated with the first data can be obtained based on the pre-created video clipping knowledge base, and the video clipping knowledge information associated with the first data includes at least one of relevant information about the video clipping process, relevant information about the video clipping effect, and relevant information about the video clipping application.
[0050] Among them, the relevant information of the video editing process includes but is not limited to at least one of the following information: process description information of the video editing process, applicable video description information of the video editing process, effect description information of the video editing process, and editing skills information of the video editing process.
[0051] The information related to the video editing effect includes but is not limited to at least one of the following information: information on the implementation process of the video editing effect, information on applicable video instructions for the video editing effect, and information on editing techniques for the video editing effect.
[0052] Among them, the relevant information of the video editing application includes but is not limited to at least one of the following information: operation instruction information of the video editing application, instruction information of the terminal device applicable to the video editing application, and instruction information of the operating system applicable to the video editing application.
[0053] In the above step S106, the second data corresponding to the video is generated based on the first data and the acquired video editing knowledge information through the large language model, and the second data is configured to answer the video editing problem. For example, the first data and the acquired video editing knowledge information are input into the large language model, and the large language model generates and outputs the second data corresponding to the video based on the first data and the acquired video editing knowledge information, thereby acquiring the second data. The second data can be a natural language in text form, which is used to answer the problems encountered by the user during the video editing process, and is mainly used to answer the user's video editing problems. For example, the video editing problem in the first data is: "How to add special effects to my video", and the second data can be related instructions on how to add special effects. In one embodiment, the large language model (LLM, Large Language Model) can be ChatGPT (Chat Generative Pre-trained Transformer).
[0054] It can be seen that through this embodiment, the large language model can be used to generate the second data required by the user according to the semantic description information corresponding to the video and the video editing questions of the user for the video. The second data can answer the video editing questions. Based on the advantages of the large language model's strong semantic understanding ability and high output data accuracy, the accuracy of the second data is effectively improved, thereby facilitating the user to perform video editing according to the second data and improving the user's video editing efficiency.
[0055] In one embodiment, the method further includes:
[0056] Determine a content category to which video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category;
[0057] and / or,
[0058] A historical editing operation of the video is determined, the historical editing operation of the video is semanticized, and second description information corresponding to the historical editing operation of the video is generated according to the semanticization result.
[0059] In this embodiment, in one case, the content category to which the video content of the video belongs can be determined, and first description information corresponding to the video content of the video can be generated based on the content category. In another case, the historical editing operations of the video can be determined, the historical editing operations of the video can be semanticized, and second description information corresponding to the historical editing operations of the video can be obtained. In yet another case, the content category to which the video content of the video belongs can be determined, and first description information corresponding to the video content of the video can be generated based on the content category, and the historical editing operations of the video can be determined, the historical editing operations of the video can be semanticized, and second description information corresponding to the historical editing operations of the video can be obtained.
[0060] In this embodiment, the content category to which the video content of the video belongs can be identified by a classification algorithm such as the Bach Scene Recognition algorithm, and the content category includes but is not limited to multiple preset categories such as animals, people, scenery, food, buildings, parks, forests, etc. After determining the content category to which the video content belongs, the first description information corresponding to the video content of the video is generated according to the content category of the video content. For example, a corresponding description text is generated according to the content category of the video content, and the description text represents the content category of the video content in the form of text. The description text can be exemplified as "the content category is animals", and the description text is determined as the first description information corresponding to the video content, or, according to a certain text format, the first description information corresponding to the video content is generated based on the description text, and the first description information is also in the form of text.
[0061] In this embodiment, the historical editing operations of the video can be determined, for example, the operation name and at least one of the operation categories of the historical editing operations of the video can be determined. The historical editing operations of the video are semanticized, and the operation name and at least one of the operation categories of the historical editing operations of the video are expressed in a semanticized form to obtain a semantic result, which can be a text. The text can be exemplified as "The historical editing operations of the video include XXX, which belongs to the XXX category of operations." The second description information is generated according to the semanticization result, for example, the semanticization result is determined as the second description information corresponding to the historical editing operations of the video, or the second description information is generated based on the semanticization result according to a certain text format, and the second description information is also in text form.
[0062] It can be seen that through this embodiment, it is possible to determine the content category to which the video content of a video belongs, and based on the content category, accurately generate first description information corresponding to the video content of the video, and / or determine the historical editing operations of the video, perform semanticization on the historical editing operations of the video, and accurately generate second description information corresponding to the historical editing operations of the video based on the semanticization results, so as to generate semantic description information corresponding to the video based on the first description information and / or the second description information, thereby improving the accuracy of generating semantic description information corresponding to the video.
[0063] In this embodiment, after determining the first description information and / or the second description information, the first description information can be used as the semantic description information corresponding to the video, or the semantic description information corresponding to the video can be generated based on the first description information in a certain text format, as described above, or the second description information can be used as the semantic description information corresponding to the video, or the semantic description information corresponding to the video can be generated based on the second description information in a certain text format, or the first description information and the second description information can be combined into the semantic description information corresponding to the video in a certain text format.
[0064] In one embodiment, the above method flow further includes the following steps, and the video clip knowledge base is generated by the following steps:
[0065] Create a video clip knowledge base based on at least one of the following information:
[0066] Information related to at least one video editing process, information related to at least one video editing effect, and information related to video editing applications.
[0067] In this embodiment, at least one of the following information is obtained: information related to at least one video editing process, information related to at least one video editing effect, and information related to a video editing application.
[0068] The at least one video editing process includes but is not limited to: a video editing process for at least one video content category, for example, an editing process for animal videos, an editing process for person videos. The relevant information of the at least one video editing process includes but is not limited to at least one of the following information: process description information of at least one video editing process, applicable video description information of at least one video editing process, effect description information of at least one video editing process, and editing technique information of at least one video editing process.
[0069] The at least one video editing effect may be, for example, a camera effect, a filter effect, a sound effect, etc. The relevant information of the at least one video editing effect includes, but is not limited to, at least one of the following information: implementation process information of the at least one video editing effect, applicable video description information of the at least one video editing effect, and editing technique information of the at least one video editing effect.
[0070] Among them, the relevant information of the video editing application includes but is not limited to at least one of the following information: operation instruction information of the video editing application, instruction information of the terminal device applicable to the video editing application, and instruction information of the operating system applicable to the video editing application.
[0071] After obtaining at least one of relevant information of at least one video editing process, relevant information of at least one video editing effect, and relevant information of video editing applications, a video editing knowledge base can be created based on the acquired information, and the created video editing knowledge base can store the acquired information.
[0072] In one embodiment, a video editing knowledge base is created based on a LangChain database, specifically including: loading, segmenting, and vectorizing the acquired information in sequence to obtain a vectorized expression of the information, creating a LangChain database, and storing the vectorized expression of the information in the LangChain database. The stored LangChain database is the video editing knowledge base.
[0073] After creating a video editing knowledge base based on the LangChain database, the first data corresponding to the video can be directly input into the LangChain database. The LangChain database can automatically output the video editing knowledge information associated with the first data, and the output video editing knowledge information is in text form, which has the advantage of efficient and fast acquisition of video editing knowledge information.
[0074] It can be seen that through this embodiment, a video clip knowledge base that stores rich and diverse information can be created, so that when subsequently obtaining video clip knowledge information associated with the first data based on the video clip knowledge base, as much video clip knowledge information as possible can be obtained, thereby improving the accuracy of the second data corresponding to the generated video.
[0075] After the video clip knowledge base is created, in one embodiment, based on the pre-created video clip knowledge base, obtaining video clip knowledge information associated with the first data includes:
[0076] Extracting a first keyword of semantic description information corresponding to the video in the first data and a second keyword of the video editing problem in the first data;
[0077] Retrieving first knowledge information associated with the first keyword in the video clip knowledge base, and retrieving second knowledge information associated with the second keyword in the video clip knowledge base;
[0078] Based on the first knowledge information and the second knowledge information, video clip knowledge information associated with the first data is generated.
[0079] As mentioned above, the first data corresponding to the video includes the semantic description information corresponding to the video and the video editing questions of the user regarding the video. In this embodiment, the first keyword of the semantic description information corresponding to the video in the first data and the second keyword of the video editing questions in the first data are first extracted. When extracting the first keyword of the semantic description information, a general keyword extraction method can be used for extraction, which is not limited in this embodiment. When extracting the second keyword of the video editing questions, a general keyword extraction method can be used for extraction, which is not limited in this embodiment.
[0080] Next, the first knowledge information associated with the first keyword is retrieved from the video clip knowledge base, and the second knowledge information associated with the second keyword is retrieved from the video clip knowledge base. In this embodiment, the video clip knowledge base may be retrieved for information containing at least one of the first keyword, its synonyms, and near synonyms, and the information may be used as the first knowledge information, and the video clip knowledge base may be retrieved for information containing at least one of the second keyword, its synonyms, and near synonyms, and the information may be used as the second knowledge information.
[0081] Finally, based on the first knowledge information and the second knowledge information, video clip knowledge information associated with the first data is generated. For example, the first knowledge information and the second knowledge information are combined in a certain text format to obtain video clip knowledge information associated with the first data. The first knowledge information, the second knowledge information, and the video clip knowledge information are all in text form.
[0082] It can be seen that through this embodiment, it is possible to obtain as much video clip knowledge information associated with the first data as possible based on the video clip knowledge base through keyword extraction and information retrieval, thereby improving the accuracy of the second data subsequently generated based on the video clip knowledge information.
[0083] After obtaining the video editing knowledge information associated with the first data, in one embodiment, the video editing problem includes behavioral intention information for the video. Accordingly, second data corresponding to the video is generated based on the first data and the video editing knowledge information through a large language model, including:
[0084] extracting third knowledge information for realizing the behavior intention corresponding to the above-mentioned behavior intention information from the video clip knowledge information through the large language model;
[0085] The second data is generated through a large language model according to the semantic description information corresponding to the video and the third knowledge information.
[0086] In this embodiment, the video editing problem includes behavioral intention information for the video, and the behavioral intention information corresponds to behavioral intentions, which can be exemplified as adding filters, realizing camera movements, etc. Based on this, the behavioral intention represented by the behavioral intention information in the video editing problem is first determined through a large language model, and the third knowledge information used to realize the behavioral intention is extracted from the video editing knowledge information. Based on the previous example, the third knowledge information can be information describing how to add filters, or information describing how to realize camera movements. Then, the second data is generated through the large language model according to the semantic description information corresponding to the video and the third knowledge information.
[0087] It can be seen that through this embodiment, when the video editing problem includes behavioral intention information for the video, the third knowledge information for realizing the behavioral intention corresponding to the above-mentioned behavioral intention information can be extracted from the video editing knowledge information through the large language model, and the second data can be generated according to the semantic description information corresponding to the video and the third knowledge information, so as to accurately determine the user's behavioral intention for the video through the large language model and give the second data in combination with the semantic description information corresponding to the video, so that the user can perform video editing according to the second data, thereby improving the convenience and efficiency of video editing for the user.
[0088] In one embodiment, the second data is generated by using a large language model according to the semantic description information corresponding to the video and the third knowledge information, including:
[0089] Through the large language model, according to the semantic description information corresponding to the video, the information related to the video is filtered out from the third-party knowledge information;
[0090] The second data is generated based on the filtered information through the large language model.
[0091] In this embodiment, first, through the large language model, according to the semantic description information corresponding to the video, the information related to the video is filtered in the third knowledge information. The information related to the video may include keywords of the content category of the video content or keywords of the historical editing operations of the video. Then, through the large language model, the second data is generated based on the filtered information. For example, the filtered information is in text form, and the filtered information is processed into second data according to a certain text format. In one example, the filtered information can be processed according to the tone and expression of the video editing question sent by the user to generate the second data. The second data is also in text form, so that the second data is easier to be accepted by the user, improving the user's experience of browsing the question and receiving the data.
[0092] It can be seen that through this embodiment, it is possible to use a large language model to filter out information related to the video in the third knowledge information according to the semantic description information corresponding to the video, and generate second data based on the filtered information, so as to determine the second data based on the information related to the video, so that the second data conforms to the actual situation of the video, thereby improving the convenience and efficiency of video editing for users.
[0093] In one embodiment, by using a large language model, according to the semantic description information corresponding to the video, information related to the video is filtered out from the third knowledge information, including:
[0094] By using a large language model, according to the first description information in the semantic description information, information related to the video content of the video is filtered out in the third knowledge information;
[0095] and / or,
[0096] Through the large language model, information related to the historical editing operation of the video is filtered out in the third knowledge information according to the second description information in the semantic description information.
[0097] In this embodiment, through the large language model, according to the first description information, information related to the video content represented by the semantic description information corresponding to the video is filtered in the third knowledge information, and the filtered related information may include keywords in the first description information or keywords of the content category of the video content, or, through the large language model, according to the second description information, information related to the historical editing operation represented by the semantic description information corresponding to the video is filtered in the third knowledge information, and the filtered related information may include keywords in the second description information or keywords of the historical editing operation. Of course, in this embodiment, through the large language model, according to the first description information in the semantic description information, information related to the video content of the video can be filtered in the third knowledge information, and, through the large language model, according to the second description information in the semantic description information, information related to the historical editing operation of the video can be filtered in the third knowledge information.
[0098] For example, the third knowledge information indicates how to add filters, and indicates how to add filters to animal videos and person videos respectively, and also indicates how to add filters to videos with special effects added and how to add filters to videos with text added. The semantic description information corresponding to the video includes first description information and second description information. The first description information is used to indicate that the video content of the video is a person video, and the second description information is used to indicate that the historical editing operations of the video include special effects adding operations. Based on this, the information on how to add filters to person videos and how to add filters to videos with special effects added are filtered out in the third knowledge information, and the filtered information is information related to the video.
[0099] It can be seen that through this embodiment, it is possible to accurately identify information related to the video content and / or information related to the historical editing operations of the video in the third knowledge information through a large language model, thereby preparing for the generation of accurate second data and improving the accuracy of the second data.
[0100] In one embodiment, the video editing question includes query information for the next step of processing suggestions for the video. Accordingly, second data corresponding to the video is generated based on the first data and the video editing knowledge information through a large language model, including:
[0101] Through the large language model, according to the semantic description information corresponding to the video, the fourth knowledge information for optimizing the video is extracted from the video clip knowledge information;
[0102] The second data is generated based on the fourth knowledge information through the large language model.
[0103] In this embodiment, the video editing question includes a query information for the next processing suggestion for the video. For example, the video editing question may be "What is the best way to edit the video next?" Based on this, the fourth knowledge information for optimizing the video can be extracted from the video editing knowledge information according to the semantic description information corresponding to the video through a large language model. Then, the second data is generated based on the fourth knowledge information through the large language model. For example, the fourth knowledge information is in text form, and the fourth knowledge information is processed into second data according to a certain text format. In one example, the fourth knowledge information can be processed according to the tone and expression of the video editing question sent by the user to generate the second data. The second data is also in text form, which makes the second data easier for users to accept and improves the user's experience of browsing questions and receiving data.
[0104] In one embodiment, the large language model can be combined with currently popular video editing techniques and, based on the semantic description information corresponding to the video, extract fourth knowledge information for optimizing the video from the video editing knowledge information. The fourth knowledge information is used to explain how to implement currently popular and video-suitable editing techniques in the video.
[0105] It can be seen that through this embodiment, when the video editing problem includes inquiry information on the next step of processing suggestions for the video, the large language model can be used to extract the fourth knowledge information for optimizing the video from the video editing knowledge information according to the semantic description information corresponding to the video, and the second data can be generated based on the fourth knowledge information, so that the second data can help the user to further optimize the video and achieve the effect of providing the user with specific video editing suggestions.
[0106] In one embodiment, fourth knowledge information for optimizing the video is extracted from the video clip knowledge information through a large language model according to the semantic description information corresponding to the video, including:
[0107] Extracting fourth knowledge information for optimizing video content of the video from the video clip knowledge information based on the first description information in the semantic description information through the large language model;
[0108] and / or,
[0109] Through the large language model, fourth knowledge information for optimizing the editing results of historical editing operations of the video is extracted from the video editing knowledge information according to the second description information in the semantic description information.
[0110] In one case, if the semantic description information corresponding to the video includes first description information, the fourth knowledge information for optimizing the video content of the video is extracted from the video clipping knowledge information according to the first description information in the semantic description information through a large language model. In one case, if the semantic description information corresponding to the video includes second description information, the fourth knowledge information for optimizing the editing results of historical editing operations of the video is extracted from the video clipping knowledge information according to the second description information in the semantic description information through a large language model. In one case, if the semantic description information corresponding to the video includes first description information and second description information, the fourth knowledge information for optimizing the video content of the video is extracted from the video clipping knowledge information according to the first description information in the semantic description information through a large language model, and the fourth knowledge information for optimizing the editing results of historical editing operations of the video is extracted from the video clipping knowledge information according to the second description information in the semantic description information through a large language model.
[0111] In one embodiment, keywords of the first description information can be obtained, and the keywords of the first description information can be used as keywords of the content category of the video content. Based on the keywords of the content category of the video content, fourth knowledge information of the video content for optimizing the video can be extracted from the video editing knowledge information. The large language model has good semantic understanding and semantic analysis capabilities, and can accurately identify the fourth knowledge information of the video content for optimizing the video based on the keywords of the content category of the video content. For example, if the content category of the video content is food, then information on editing techniques to increase the atmosphere of food videos can be extracted from the video editing knowledge information as the fourth knowledge information.
[0112] In one embodiment, keywords of the second description information can be obtained, and the keywords of the second description information can be used as keywords of the historical editing operation. Based on the keywords of the historical editing operation, fourth knowledge information for optimizing the editing results of the historical editing operation is extracted from the video editing knowledge information. After editing the video based on the fourth knowledge information, the editing results of the historical editing operation can be further improved, presenting a better editing effect. The large language model has good semantic understanding and semantic analysis capabilities, and can accurately identify the fourth knowledge information for optimizing the editing results of the historical editing operation of the video based on the keywords of the historical editing operation of the video. For example, if the historical editing operation is an operation of adding background music, the fourth knowledge information can be information related to camera movement. By coordinating camera movement with background music, the expression effect of background music can be improved.
[0113] It can be seen that through this embodiment, the fourth knowledge information for optimizing the video content of the video can be extracted from the video editing knowledge information, and / or the fourth knowledge information for optimizing the editing results of the historical editing operations of the video can be extracted from the video editing knowledge information, and second data can be generated based on the fourth knowledge information, so that the user can further optimize the video according to the second data and improve the video editing efficiency.
[0114] In this embodiment, after the second data corresponding to the video is generated, the second data may also be pushed to the user's terminal device. In one embodiment, the above method flow further includes:
[0115] Determine a processing control for the video associated with the second data, and push the second data and the processing control to the user;
[0116] and / or,
[0117] Determine the editing process for the video represented by the second data, execute the editing process, and push the second data and the execution result of the editing process to the user.
[0118] In one case, when the second data is associated with a processing control for a video, for example, the second data is used to explain how to move the camera, then the second data is associated with an entry control for the camera movement function. Based on this, the processing control for the video associated with the second data is determined, and the second data and the processing control are pushed to the user, so that the user can directly perform video editing through the processing control, thereby improving video editing efficiency.
[0119] In another case, the second data represents a video editing process, for example, the second data represents a specific process of adjusting the volume of a video, then the editing process can be directly executed, and the second data and the execution result of the editing process can be pushed to the user, eliminating the need for the user to manually refer to the second data to execute the editing process. The execution result includes execution success and execution failure.
[0120] Of course, in this embodiment, it is also possible to determine the processing controls for the video associated with the second data, push the second data and the processing controls to the user, and determine the editing process for the video represented by the second data, execute the editing process, and push the second data and the execution results of the editing process to the user.
[0121] It can be seen that through this embodiment, the second data can be desemanticized and the processing controls associated with the second data can be displayed to the user to facilitate the user to perform video editing, or the editing process represented by the second data can be directly and automatically executed to improve the efficiency of video editing.
[0122] In summary, the above video processing method has at least the following technical effects:
[0123] 1. Users can raise video editing questions, and the large language model will provide users with specific editing processes or processing controls. It can also automatically execute the editing process on behalf of users, improving video editing efficiency.
[0124] 2. Users can not only ask video editing questions with specific behavioral intentions, but also ask for suggestions on the next step of video processing. The large language model helps users analyze the next specific editing operation and improve the quality of users' video editing;
[0125] 3. The video editing questions raised by users can also be questions about how to use specific functions, so that the large language model can be used to answer various questions encountered by users in the video editing process, reducing the user's video editing learning cost.
[0126] The above describes the specific process of the video processing method executed by the server. The above description of pushing certain information to the user refers to pushing the information to the user's terminal device and displaying it to the user through the terminal device. In one embodiment, the video processing method can also be implemented by the server and the user's terminal device. The user's terminal device includes but is not limited to a mobile phone, a desktop computer, a PAD, a tablet computer, a laptop computer, a car computer, a wearable device, etc. Figure 2 A flowchart of a video processing method provided by another embodiment of the present disclosure is shown as follows: Figure 2 As shown, the method includes:
[0127] Step S202, obtaining first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0128] Step S204, sending the first data to the server;
[0129] Step S206: receiving second data corresponding to the video returned by the server based on the first data, where the second data is configured to answer the video editing question.
[0130] In this embodiment, the first data corresponding to the video is obtained; the first data includes the semantic description information corresponding to the video and the video editing questions of the user for the video; the semantic description information corresponding to the video includes the first description information corresponding to the video content of the video and / or the second description information corresponding to the historical editing operations of the video, then, the first data is sent to the server, and finally, the second data corresponding to the video returned by the server according to the first data is received, and the second data is configured to answer the video editing questions. It can be seen that through this embodiment, the second data required by the user can be generated according to the semantic description information corresponding to the video and the video editing questions of the user for the video, and the second data can answer the video editing questions, thereby facilitating the user to perform video editing according to the second data and improving the editing efficiency of the user in performing video editing.
[0131] In the above step S202, first data corresponding to the video is obtained, the first data including semantic description information corresponding to the video and video editing questions of the user for the video, the semantic description information corresponding to the video including first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video. In the above step S204, the first data is sent to the server.
[0132] The detailed process of step S202 and step S204 can refer to the previous introduction. Different from the previous embodiment, in this embodiment, the user inputs a video editing question for the video in the terminal device, and the video editing question is in text form. The user's terminal device also obtains the semantic description information corresponding to the video, combines the semantic description information corresponding to the video and the user's video editing question for the video in a certain text format into the first data corresponding to the video, and sends the first data corresponding to the video to the server. Among them, the semantic description information corresponding to the video is also in text form, and the first data corresponding to the video is also in text form.
[0133] The server can generate the second data according to the method flow described above, and return the second data to the terminal device. In the above step S206, the second data corresponding to the video returned by the server according to the first data is received, and the second data is displayed to the user.
[0134] In one embodiment, the above method flow further includes:
[0135] Determine a content category to which video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category;
[0136] and / or,
[0137] A historical editing operation of the video is determined, the historical editing operation of the video is semanticized, and second description information corresponding to the historical editing operation of the video is generated according to the semanticization result.
[0138] In this embodiment, in one case, the content category to which the video content of the video belongs can be determined, and first description information corresponding to the video content of the video can be generated based on the content category. In another case, the historical editing operations of the video can be determined, the historical editing operations of the video can be semanticized, and second description information corresponding to the historical editing operations of the video can be obtained. In yet another case, the content category to which the video content of the video belongs can be determined, and first description information corresponding to the video content of the video can be generated based on the content category, and the historical editing operations of the video can be determined, the historical editing operations of the video can be semanticized, and second description information corresponding to the historical editing operations of the video can be obtained.
[0139] The specific process of generating the first description information and the second description information in this embodiment can refer to the previous introduction on the server side, which will not be repeated here.
[0140] It can be seen that through this embodiment, it is possible to determine the content category to which the video content of a video belongs, and based on the content category, accurately generate first description information corresponding to the video content of the video, and / or determine the historical editing operations of the video, perform semanticization on the historical editing operations of the video, and accurately generate second description information corresponding to the historical editing operations of the video based on the semanticization results, so as to generate semantic description information corresponding to the video based on the first description information and / or the second description information, thereby improving the accuracy of generating semantic description information corresponding to the video.
[0141] In one embodiment, the above method flow further includes:
[0142] determining a processing control for the video associated with the second data, and displaying the second data and the processing control;
[0143] and / or,
[0144] Determine the editing process for the video represented by the second data, execute the editing process, and display the second data and the execution result of the editing process.
[0145] In one case, when the second data is associated with a processing control for a video, for example, the second data is used to explain how to add a filter, the second data is associated with an entry control for the filter function. Based on this, the processing control for the video associated with the second data is determined, and the second data and the processing control are displayed, so that the user can directly edit the video through the processing control, thereby improving the efficiency of video editing.
[0146] In another case, the second data represents a video editing process, for example, the second data represents a specific process of adjusting the video playback speed, then the editing process can be directly executed, and the second data and the execution result of the editing process are displayed, eliminating the need for the user to manually refer to the second data to execute the editing process. The execution result includes execution success and execution failure.
[0147] Of course, in this embodiment, it is also possible to determine the processing controls for the video associated with the second data, display the second data and the processing controls, and determine the editing process for the video represented by the second data, execute the editing process, and display the execution results of the second data and the editing process.
[0148] It can be seen that through this embodiment, the second data can be desemanticized and the processing controls associated with the second data can be displayed to the user to facilitate the user to perform video editing, or the editing process represented by the second data can be directly and automatically executed to improve the efficiency of video editing.
[0149] Figure 3 A schematic diagram of an application scenario of a video processing method provided by an embodiment of the present disclosure, such as Figure 3 As shown, the scenario includes a user's terminal device and a server. The user's terminal device can generate first data corresponding to a video and send the first data corresponding to the video to the server. The server obtains video clip knowledge information associated with the first data based on a pre-created video clip knowledge base, and generates second data corresponding to the video based on the first data and the video clip knowledge information through a large language model, and returns the second data to the terminal device. The terminal device determines a processing control for the video associated with the second data, displays the second data and the processing control, and determines a clipping process for the video represented by the second data, executes the clipping process, and displays the second data and the execution result of the clipping process.
[0150] exist Figure 3 In the illustrated embodiment, when the server generates second data corresponding to a video based on the first data and video clip knowledge information through a large language model, the server inputs the first data, the video clip knowledge information and the constraint information of the large language model into the large language model. The constraint information of the large language model is used to indicate information such as the types of questions that the large language model can answer, the language of the answers, etc., so that the large language model generates the second data through natural language based on the first data and video clip knowledge information within the constraint range of the constraint information.
[0151] pass Figure 3The embodiments in the embodiment can generate the second data required by the user by using the large language model according to the semantic description information corresponding to the video and the video editing questions of the user for the video through the mutual cooperation between the terminal device and the server. Based on the advantages of the large language model with strong semantic understanding ability and high accuracy of output data, the accuracy of the second data is effectively improved, thereby facilitating the user to edit the video according to the second data and improving the editing efficiency of the user in video editing.
[0152] Figure 4 A schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure is shown in FIG. Figure 4 As shown, the device comprises:
[0153] A first acquisition unit 41 is used to acquire first data corresponding to a video; the first data includes semantic description information corresponding to the video and a video editing question of a user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to a historical editing operation of the video;
[0154] A second acquisition unit 42 is used to acquire video clip knowledge information associated with the first data based on a pre-created video clip knowledge base;
[0155] The data generating unit 43 is used to generate second data corresponding to the video based on the first data and the video clip knowledge information through a large language model, wherein the second data is configured to answer the video clip question.
[0156] Optionally, the device further includes a first information generating unit, configured to:
[0157] Determine a content category to which the video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category;
[0158] and / or,
[0159] Determine the historical editing operation of the video, perform semanticization on the historical editing operation of the video, and generate second description information corresponding to the historical editing operation of the video according to the semanticization result.
[0160] Optionally, the device further comprises a library creation unit, configured to:
[0161] The video clip knowledge base is created based on at least one of the following information:
[0162] Information related to at least one video editing process, information related to at least one video editing effect, and information related to video editing applications.
[0163] Optionally, the second acquiring unit 42 is specifically configured to:
[0164] Extracting a first keyword of the semantic description information corresponding to the video in the first data and a second keyword of the video editing problem in the first data;
[0165] Retrieving first knowledge information associated with the first keyword in the video clip knowledge base, and retrieving second knowledge information associated with the second keyword in the video clip knowledge base;
[0166] Video clip knowledge information associated with the first data is generated based on the first knowledge information and the second knowledge information.
[0167] Optionally, the video editing question includes behavioral intention information for the video; the data generating unit 43 is specifically used to:
[0168] extracting, from the video clip knowledge information, third knowledge information for realizing the behavior intention corresponding to the behavior intention information by using the large language model;
[0169] The second data is generated through the large language model according to the semantic description information corresponding to the video and the third knowledge information.
[0170] Optionally, the data generating unit 43 is further specifically configured to:
[0171] By using the large language model, according to the semantic description information corresponding to the video, filtering information related to the video in the third knowledge information;
[0172] The second data is generated based on the filtered information through the large language model.
[0173] Optionally, the data generating unit 43 is further specifically configured to:
[0174] Using the large language model, according to the first description information in the semantic description information, filtering information related to the video content of the video in the third knowledge information;
[0175] and / or,
[0176] By using the large language model, information related to the historical editing operation of the video is filtered out from the third knowledge information according to the second description information in the semantic description information.
[0177] Optionally, the video editing question includes query information of a next processing suggestion for the video; the data generating unit 43 is specifically used for:
[0178] Extracting fourth knowledge information for optimizing the video from the video clip knowledge information through the large language model according to the semantic description information corresponding to the video;
[0179] The second data is generated based on the fourth knowledge information through the large language model.
[0180] Optionally, the data generating unit 43 is further specifically configured to:
[0181] extracting, from the video clip knowledge information, fourth knowledge information for optimizing the video content of the video according to the first description information in the semantic description information by using the large language model;
[0182] and / or,
[0183] By using the large language model, fourth knowledge information for optimizing the editing results of the historical editing operations of the video is extracted from the video editing knowledge information according to the second description information in the semantic description information.
[0184] Optionally, the device further comprises a pushing unit, configured to:
[0185] Determine a processing control for the video associated with the second data, and push the second data and the processing control to a user;
[0186] and / or,
[0187] Determine the editing process for the video represented by the second data, execute the editing process, and push the second data and the execution result of the editing process to the user.
[0188] The video processing device in the embodiment of the present disclosure can implement each process of the above-mentioned video processing method embodiment applied to the server and achieve the same effects and functions, which will not be repeated here.
[0189] Figure 5 A schematic diagram of the structure of a video processing device provided by another embodiment of the present disclosure is shown in FIG. Figure 5 As shown, the device comprises:
[0190] The third acquisition unit 51 is used to acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and a video editing question of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0191] A sending unit 52, configured to send the first data to a server;
[0192] The receiving unit 53 is used to receive second data corresponding to the video returned by the server according to the first data, where the second data is configured to answer the video editing problem.
[0193] Optionally, the device further includes a second information generating unit, configured to:
[0194] Determine a content category to which the video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category;
[0195] and / or,
[0196] Determine the historical editing operation of the video, perform semanticization on the historical editing operation of the video, and generate second description information corresponding to the historical editing operation of the video according to the semanticization result.
[0197] Optionally, the device further comprises a display unit, configured to:
[0198] determining a processing control for the video associated with the second data, and displaying the second data and the processing control;
[0199] and / or,
[0200] Determine the editing process for the video represented by the second data, execute the editing process, and display the second data and the execution result of the editing process.
[0201] The video processing device in the embodiment of the present disclosure can implement each process of the above-mentioned video processing method embodiment applied to the terminal device and achieve the same effects and functions, which will not be repeated here.
[0202] An embodiment of the present disclosure further provides an electronic device, Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure, such as Figure 6As shown, the electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors 601 and memory 602, and one or more applications or data may be stored in the memory 602. Among them, the memory 602 may be a short-term storage or a persistent storage. The application stored in the memory 602 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the electronic device. Furthermore, the processor 601 may be configured to communicate with the memory 602 to execute a series of computer executable instructions in the memory 602 on the electronic device. The electronic device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input or output interfaces 605, one or more keyboards 606, etc.
[0203] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer executable instructions, wherein when the computer executable instructions are executed, the processor implements the following process:
[0204] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0205] Based on a pre-created video clip knowledge base, acquiring video clip knowledge information associated with the first data;
[0206] Through a large language model, based on the first data and the video clip knowledge information, second data corresponding to the video is generated, and the second data is configured to answer the video clip question.
[0207] The electronic device in the embodiment of the present disclosure can implement each process of the above-mentioned video processing method embodiment applied to the server and achieve the same effects and functions, which will not be repeated here.
[0208] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer executable instructions, wherein when the computer executable instructions are executed, the processor implements the following process:
[0209] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0210] Sending the first data to a server;
[0211] Receive second data corresponding to the video returned by the server according to the first data, where the second data is configured to answer the video editing question.
[0212] The electronic device in the embodiment of the present disclosure can implement the various processes of the above-mentioned video processing method embodiment applied to the terminal device and achieve the same effects and functions, which will not be repeated here.
[0213] Another embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following process is implemented:
[0214] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0215] Based on a pre-created video clip knowledge base, acquiring video clip knowledge information associated with the first data;
[0216] Through a large language model, based on the first data and the video clip knowledge information, second data corresponding to the video is generated, and the second data is configured to answer the video clip question.
[0217] The storage medium in the embodiment of the present disclosure can implement each process of the above-mentioned video processing method embodiment applied to the server and achieve the same effect and function, which will not be repeated here.
[0218] Another embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following process is implemented:
[0219] Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video;
[0220] Sending the first data to a server;
[0221] Receive second data corresponding to the video returned by the server according to the first data, where the second data is configured to answer the video editing question.
[0222] The storage medium in the embodiment of the present disclosure can implement each process of the above-mentioned video processing method embodiment applied to the terminal device and achieve the same effect and function, which will not be repeated here.
[0223] In various embodiments of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0224] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0225] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0226] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0227] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0228] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0229] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0230] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0231] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0232] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0233] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0234] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0235] The above description is only an embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the scope of the claims of the present disclosure.
Claims
1. A video processing method, characterized in that: include: Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user regarding the video; The semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video; Based on a pre-created video clip knowledge base, acquiring video clip knowledge information associated with the first data; Through a large language model, based on the first data and the video clip knowledge information, second data corresponding to the video is generated, and the second data is configured to answer the video clip question.
2. The method according to claim 1, characterized in that The method further comprises: Determine a content category to which the video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category; and / or, Determine the historical editing operation of the video, perform semanticization on the historical editing operation of the video, and generate second description information corresponding to the historical editing operation of the video according to the semanticization result.
3. The method according to claim 1, characterized in that The method further comprises: The video clip knowledge base is created based on at least one of the following information: Information related to at least one video editing process, information related to at least one video editing effect, and information related to video editing applications.
4. The method according to claim 1, characterized in that: The acquiring the video clip knowledge information associated with the first data based on the pre-created video clip knowledge base includes: Extracting a first keyword of the semantic description information corresponding to the video in the first data and a second keyword of the video editing problem in the first data; Retrieving first knowledge information associated with the first keyword in the video clip knowledge base, and retrieving second knowledge information associated with the second keyword in the video clip knowledge base; Video clip knowledge information associated with the first data is generated based on the first knowledge information and the second knowledge information.
5. The method according to claim 1, characterized in that The video editing problem includes behavioral intention information for the video; and the generating second data corresponding to the video based on the first data and the video editing knowledge information through the large language model includes: extracting, from the video clip knowledge information, third knowledge information for realizing the behavior intention corresponding to the behavior intention information by using the large language model; The second data is generated through the large language model according to the semantic description information corresponding to the video and the third knowledge information.
6. The method according to claim 5, characterized in that The generating the second data by using the large language model according to the semantic description information corresponding to the video and the third knowledge information includes: By using the large language model, according to the semantic description information corresponding to the video, filtering information related to the video in the third knowledge information; The second data is generated based on the filtered information through the large language model.
7. The method according to claim 6, characterized in that The filtering of information related to the video in the third knowledge information by using the large language model according to the semantic description information corresponding to the video includes: Using the large language model, according to the first description information in the semantic description information, filtering information related to the video content of the video in the third knowledge information; and / or, By using the large language model, information related to the historical editing operation of the video is filtered out from the third knowledge information according to the second description information in the semantic description information.
8. The method according to claim 1, characterized in that The video editing question includes query information of a next processing suggestion for the video; the generating of second data corresponding to the video based on the first data and the video editing knowledge information by using a large language model includes: Extracting fourth knowledge information for optimizing the video from the video clip knowledge information through the large language model according to the semantic description information corresponding to the video; The second data is generated based on the fourth knowledge information through the large language model.
9. The method according to claim 8, characterized in that The extracting, from the video clip knowledge information, fourth knowledge information for optimizing the video according to the semantic description information corresponding to the video by using the large language model, includes: extracting, from the video clip knowledge information, fourth knowledge information for optimizing the video content of the video according to the first description information in the semantic description information by using the large language model; and / or, By using the large language model, fourth knowledge information for optimizing the editing results of the historical editing operations of the video is extracted from the video editing knowledge information according to the second description information in the semantic description information.
10. The method according to claim 1, characterized in that The method further comprises: Determine a processing control for the video associated with the second data, and push the second data and the processing control to a user; and / or, Determine the editing process for the video represented by the second data, execute the editing process, and push the second data and the execution result of the editing process to the user.
11. A video processing method, characterized in that: include: Acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user regarding the video; The semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video; Sending the first data to a server; Receive second data corresponding to the video returned by the server according to the first data, where the second data is configured to answer the video editing question.
12. The method according to claim 11, characterized in that The method further comprises: Determine a content category to which the video content of the video belongs, and generate first description information corresponding to the video content of the video according to the content category; and / or, Determine the historical editing operation of the video, perform semanticization on the historical editing operation of the video, and generate second description information corresponding to the historical editing operation of the video according to the semanticization result.
13. The method according to claim 11, characterized in that The method further comprises: determining a processing control for the video associated with the second data, and displaying the second data and the processing control; and / or, Determine the editing process for the video represented by the second data, execute the editing process, and display the second data and the execution result of the editing process.
14. A video processing device, characterized in that: include: A first acquisition unit is used to acquire first data corresponding to a video; the first data includes semantic description information corresponding to the video and a video editing question of a user regarding the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to a historical editing operation of the video; A second acquisition unit, configured to acquire video clip knowledge information associated with the first data based on a pre-created video clip knowledge base; A data generating unit is used to generate second data corresponding to the video based on the first data and the video clip knowledge information through a large language model, wherein the second data is configured to answer the video clip question.
15. A video processing device, characterized in that: include: A third acquisition unit is used to acquire first data corresponding to the video; the first data includes semantic description information corresponding to the video and video editing questions of the user for the video; the semantic description information corresponding to the video includes first description information corresponding to the video content of the video and / or second description information corresponding to the historical editing operation of the video; A sending unit, configured to send the first data to a server; A receiving unit is used to receive second data corresponding to the video returned by the server based on the first data, where the second data is configured to answer the video editing problem.
16. An electronic device, characterized in that: include: processor; as well as, A memory configured to store computer executable instructions, wherein the computer executable instructions, when executed, cause the processor to implement the steps of the method described in any one of claims 1 to 10, or implement the steps of the method described in any one of claims 11 to 13.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the method described in any one of claims 1 to 10 above are implemented, or the steps of the method described in any one of claims 11 to 13 above are implemented.