Video processing method and device, computer equipment and storage medium
Through the large language model framework proxy, different tools are called to generate diverse video processing results, solving the problem of single application functions of the small assistant with graphic and text summary, and improving the user interaction experience.
Patent Information
- Application Number
- CN202311868090.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
The existing graphic summary assistant has a single application function, insufficient diversity, and lacks interaction with users.
The corresponding tools are called through the large language model framework agent, and a variety of processing results such as mind maps, video summary texts, question answers, video questions and answers, recipe templates, video covers, music links and high-energy barrage points are generated based on user instructions.
It realizes diversification of video processing, meets users' personalized needs, and increases the interaction between users and videos.
Smart Images

Figure CN120277236A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technologies, and particularly to a video processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of Large Language Model (LLM) technology, the graphic summary assistant application developed based on LLM has become a popular application. Through this graphic summary assistant application, video content with a long duration, boring content, or strong professionalism can be summarized into text and output. However, the inventor found that the functions of this graphic summary assistant application are relatively single, lacking diversity, and lacking other interactions with users. Summary of the Invention
[0003] In view of this, a video processing method, apparatus, computer device, and computer-readable storage medium are provided to solve the above problems.
[0004] This application provides a video processing method, and the method includes:
[0005] Obtain a user instruction sent by a user based on a target video;
[0006] Create a large language model framework proxy according to the user instruction;
[0007] Input the user instruction and the description information of all tools that the large language model framework proxy can call into the large language model through the large language model framework proxy, so as to output a target tool for responding to the user instruction through the large language model;
[0008] Call the target tool through the large language model framework proxy to execute a task associated with the target tool, and obtain a task execution result, where the task execution result is one of a mind map, a video summary text, an answer to a question, at least one video question and the corresponding answer to each video question, a recipe template, a video cover, a music link, and high-energy bullet points;
[0009] Return the task execution result to the user.
[0010] Optionally, outputting a target tool for responding to the user instruction through the large language model includes:
[0011] Analyze the user instruction through the large language model to obtain a user intention, where the user intention is the output result required by the user;
[0012] Based on the user intention and the description information of all the tools, the large language model determines the target tool for responding to the user instruction.
[0013] Optionally, when the target tool is a mind map generation tool, the large language model framework proxy invokes the target tool to execute the task associated with the target tool, and the obtained task execution result includes:
[0014] The large language model framework proxy invokes the mind map generation tool to execute the task associated with the mind map generation tool, and obtains the task execution result;
[0015] Among them, the mind map generation tool obtains the task execution result by performing the following steps:
[0016] Obtain the video ID of the target video;
[0017] Obtain the audio data of the target video according to the video ID;
[0018] Call the speech recognition service to convert the audio data into text information;
[0019] Segment the text information into multiple text blocks;
[0020] Input multiple text blocks and the content summary prompt words in a preset format into the large language model;
[0021] The large language model outputs the video summary content in the preset format;
[0022] Convert the video summary content into a mind map;
[0023] Use the mind map as the task execution result.
[0024] Optionally, when the target tool is a video content summary tool, the large language model framework proxy invokes the target tool to execute the task associated with the target tool, and the obtained task execution result includes:
[0025] The large language model framework proxy invokes the video content summary tool to execute the task associated with the video content summary tool, and obtains the task execution result;
[0026] Among them, the video content summary tool obtains the task execution result by performing the following steps:
[0027] Obtain the video ID of the target video;
[0028] Obtain the audio data of the target video according to the video ID;
[0029] Call the speech recognition service to convert the audio data into text information;
[0030] Segment the text information into multiple text blocks;
[0031] Convert multiple text blocks into corresponding vector data, and save multiple text blocks and corresponding vector data in a preset vector library in an associated manner;
[0032] Input multiple text blocks into the large language model;
[0033] Output video summary text through the large language model;
[0034] Use the video summary text as the task execution result.
[0035] Optionally, when the target tool is a video content question tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes:
[0036] Proxy call the video content question tool through the large language model framework to execute a task associated with the video content question tool, and obtain a task execution result;
[0037] Among them, the video content question tool obtains the task execution result by performing the following steps:
[0038] Query whether there is vector data corresponding to the target video in a preset vector library;
[0039] If there is vector data corresponding to the target video in the vector library, convert the user instruction into corresponding vector data;
[0040] Select text blocks that meet preset conditions from multiple text blocks based on the vector data corresponding to the user instruction and the vector data corresponding to the target video;
[0041] Input the selected text blocks and the user instruction into the large language model;
[0042] Output the question answer corresponding to the user instruction through the large language model;
[0043] Use the question answer as the task execution result.
[0044] Optionally, when the target tool is a video self-question-and-answer tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes:
[0045] The large language model framework is used to proxy - call the video Q&A tool to execute a task associated with the video Q&A tool, and a task execution result is obtained;
[0046] Among them, the video Q&A tool obtains the task execution result by performing the following steps:
[0047] Query whether the video summary text exists in a preset video summary text database;
[0048] If the video summary text exists in the video summary text database, input the video summary text into the large language model;
[0049] Summarize at least one video question through the large language model;
[0050] Call the video content question - answering tool to obtain the question answers corresponding to each of the video questions;
[0051] Use each video question and the corresponding question answer as the task execution result output by the video Q&A tool.
[0052] Optionally, when the target tool is a recipe generation tool, the step of using the large language model framework to proxy - call the target tool to execute a task associated with the target tool and obtaining a task execution result includes:
[0053] The large language model framework is used to proxy - call a recipe generation tool to execute a task associated with the recipe generation tool, and a task execution result is obtained;
[0054] Among them, the recipe generation tool obtains the task execution result by performing the following steps:
[0055] Obtain the video ID of the target video;
[0056] Obtain the audio data of the target video according to the video ID;
[0057] Call a speech recognition service to convert the audio data into text information;
[0058] Segment the text information into multiple text blocks;
[0059] Convert the multiple text blocks into corresponding vector data, and associate and save the multiple text blocks and the corresponding vector data in a preset vector library;
[0060] Input the multiple text blocks and prompt words describing the recipe making process into the large language model;
[0061] Output the summary content of the recipe making process through the large language model;
[0062] Fill the summary content of the recipe making process into a preset template to obtain a recipe template;
[0063] Use the recipe template as the result of the task execution.
[0064] Optionally, when the target tool is a video cover extraction tool, the proxy call of the target tool through the large language model framework to execute the task associated with the target tool to obtain the task execution result includes:
[0065] Proxy call the video cover extraction tool through the large language model framework to execute the task associated with the video cover extraction tool to obtain the task execution result;
[0066] Among them, the video cover extraction tool obtains the task execution result by performing the following steps:
[0067] Obtain the video ID of the target video;
[0068] Obtain the video cover of the target video according to the video ID;
[0069] Use the video cover as the task execution result.
[0070] Optionally, when the target tool is a background music extraction tool, the proxy call of the target tool through the large language model framework to execute the task associated with the target tool to obtain the task execution result includes:
[0071] Proxy call the background music extraction tool through the large language model framework to execute the task associated with the background music extraction tool to obtain the task execution result;
[0072] Among them, the background music extraction tool obtains the task execution result by performing the following steps:
[0073] Obtain the video ID of the target video;
[0074] Obtain the music name of the background music of the target video according to the video ID;
[0075] Obtain the music link of the background music according to the music name;
[0076] Use the music link as the task execution result.
[0077] Optionally, when the target tool is a high-energy barrage point extraction tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes:
[0078] Proxy calling the high-energy barrage point extraction tool through the large language model framework to execute a task associated with the high-energy barrage point extraction tool, and obtaining a task execution result;
[0079] Among them, the high-energy barrage point extraction tool obtains the task execution result by performing the following steps:
[0080] Obtain the video ID of the target video;
[0081] Obtain the barrage of the target video according to the video ID;
[0082] Analyze the barrage to obtain high-energy barrage points;
[0083] Use the high-energy barrage points as the task execution result.
[0084] This application also provides a video processing method, and the method includes:
[0085] Receive a user instruction input by a user based on a target video;
[0086] Input the user instruction into a large language model, so that the large language model outputs a processing result based on the user instruction and the video content of the target video, and the processing result is one of a mind map, a video summary text, a question answer, at least one video question and the question answers corresponding to each video question, and a recipe template;
[0087] Display the processing result.
[0088] This application also provides a video processing device, and the video processing device includes:
[0089] An acquisition module, configured to acquire a user instruction sent by a user based on a target video;
[0090] A creation module, configured to create a large language model framework proxy according to the user instruction;
[0091] An input module, configured to input the user instruction and the description information of all tools included in the large language model framework proxy into the large language model through the large language model framework proxy, so as to output a target tool for responding to the user instruction through the large language model;
[0092] An execution module, configured to proxy - call the target tool through the large - language model framework to execute a task associated with the target tool, and obtain a task execution result, where the task execution result is one of a mind map, a video summary text, an answer to a question, at least one video question and the corresponding answer to each video question, a recipe template, a video cover, a music link, and high - energy barrage points;
[0093] A return module, configured to return the task execution result to the user.
[0094] This application also provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above - mentioned method are implemented.
[0095] This application also provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above - mentioned method are implemented.
[0096] In the video processing method of the embodiment of this application, by obtaining a user instruction sent by a user based on a target video; creating a large - language model framework proxy according to the user instruction; analyzing the user instruction by calling a large - language model through the large - language model framework proxy to obtain a user intention; calling a tool matching the user intention through the large - language model to execute a task corresponding to the user intention to obtain a task execution result; and returning the task execution result to the user. By adopting the above video processing method, different tools can be called according to different user intentions to execute tasks corresponding to the user intentions, so as to realize diversified processing of videos, meet the personalized needs of users, and increase the interaction methods between users and videos. Description of the Drawings
[0097] Figure 1 It is a schematic diagram of the application environment of an embodiment of the video processing method of the embodiment of this application;
[0098] Figure 2 It is a flowchart of an embodiment of the video processing method described in this application;
[0099] Figure 3 It is a schematic diagram of the step refinement of outputting a target tool for responding to the user instruction through the large - language model in an embodiment of this application;
[0100] Figure 4 It is a schematic diagram of the step refinement of the mind - map generation tool obtaining the task execution result in an embodiment of this application;
[0101] Figure 5Schematic diagram for the refinement of steps for a video content question tool to obtain the task execution result in an embodiment of this application;
[0102] Figure 6 Schematic diagram for the refinement of steps for a video self-questioning and self-answering tool to obtain the task execution result in an embodiment of this application;
[0103] Figure 7 Schematic diagram for the refinement of steps for a recipe generation tool to obtain the task execution result in an embodiment of this application;
[0104] Figure 8 Program module diagram of an embodiment of the video processing device described in this application;
[0105] Figure 9 Schematic diagram of the hardware structure of a computer device for executing the video processing method provided in an embodiment of this application. Detailed implementation manners
[0106] The advantages of this application are further elaborated below in conjunction with the accompanying drawings and specific embodiments.
[0107] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0108] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit this disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0109] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0110] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step. Therefore, it should not be construed as a limitation to the present application.
[0111] The following are the term explanations of the present application:
[0112] LLM (Large Language Model): A large language model is a natural language processing model constructed using deep learning technology. It can process a large amount of text data and has powerful language understanding and generation capabilities. It is trained with billions of parameters and can understand semantics, context, and generate natural language responses. Large language models have achieved remarkable results in multiple fields, such as natural language understanding, dialogue systems, text summarization, etc.
[0113] ASR (Auto Speech Recognition): Its goal is to convert the lexical content in human speech into computer-readable input.
[0114] prompt: A prompt word given to the LLM.
[0115] Figure 1 A schematic diagram showing an application scenario provided by an embodiment of the present application is shown.
[0116] In an exemplary embodiment, the system of this application environment may include a terminal device 10 and a server 20. Among them, the terminal device 10 is connected to the server 20 through a wireless or wired network. The terminal device 10 includes but is not limited to smartphones, tablets, laptop computers, desktop computers, etc. The server 20 can be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. The network may include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, and / or proxy devices, etc. The network may also include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and their combinations and / or analogs.
[0117] Next, several embodiments will be provided in the above exemplary application environment to illustrate the video processing solution in the present application. Refer to Figure 2 , which is a flowchart of the video processing method according to an embodiment of the present application. It should be noted that the flowchart in the embodiment of this method is not used to limit the order of execution of the steps. As can be seen from the figure, the video processing method provided in this embodiment includes:
[0118] Step S20: Obtain a user instruction sent by the user based on the target video.
[0119] Specifically, the user instruction may be an instruction for extracting video-related content of the target video. The video-related content of the target video includes, but is not limited to, the cover of the target video, the background music of the target video, and the high-energy bullet screen points of the target video.
[0120] Among them, the high-energy bullet screen point refers to the bullet screen playback time point when the number of bullet screens meets a preset condition. The preset condition is a condition for determining whether the current bullet screen playback time point is a high-energy bullet screen point. The preset condition can also be set in advance according to the actual situation.
[0121] In one embodiment, the preset condition may be that the number of bullet screens is greater than a preset number. Among them, the preset number can also be set and modified according to the actual situation.
[0122] In another embodiment, the preset condition may also be that the number of bullet screens is among the top N.
[0123] The user instruction may also be an instruction for text summarization or picture summarization of the video content of the target video. Among them, the picture summarization instruction includes, but is not limited to, the instruction for summarization in the form of a mind map, the instruction for summarization in the form of a food map, and the instruction for summarization in the form of an evaluation report.
[0124] The user instruction may also be an instruction for answering questions raised by the user based on the video content.
[0125] The user instruction may also be an instruction for self-questioning and self-answering based on the video content.
[0126] It should be noted that the user instruction contains a user intention, and the user intention is the output result required by the user.
[0127] In a specific scenario, when the user is watching the target video, the user can send the user instruction in the comment area of the target video.
[0128] In another scenario, when the user is watching the target video, the user can send the user instruction in the bullet screen sending area of the target video.
[0129] Step S21, create a large language model framework proxy according to the user instruction.
[0130] Specifically, after obtaining the user instruction, a large language model framework proxy will be created according to the user instruction, so that the large language model framework proxy can call corresponding tools to respond to the user instruction.
[0131] Among them, the large language model framework agent is an entity that promotes decision-making in the large language model framework. The large language model framework can be LangChain or other LLM development frameworks such as AutoChain, etc.
[0132] Step S22: Input the user instruction and the description information of all tools that can be called by the large language model framework agent into the large language model, so as to output a target tool for responding to the user instruction through the large language model.
[0133] Specifically, the large language model framework agent will input the user instruction and the description information of all tools that it can call into the large language model, so that the large language model can decide what action to take, that is, which tool to use to respond to the user instruction.
[0134] Among them, the description information describes the tasks that the current tool can handle and the input parameters of the tool, etc.
[0135] In a specific scenario, the content of the description information of a certain tool is as follows: "When you need to obtain a mind map, use this tool. The input of this tool is a BVID number, for example: BVxxxxx".
[0136] In an exemplary embodiment, in order to enable the large language model to extract the video ID, the large language model framework agent can also input the video link into the large language model.
[0137] In an exemplary embodiment, refer to Figure 3 , the target tool for responding to the user instruction output by the large language model includes:
[0138] Step S30: Analyze the user instruction through the large language model to obtain the user intention, and the user intention is the output result required by the user.
[0139] Specifically, after receiving the user instruction, the large language model will perform semantic analysis on the user instruction to obtain the user intention, that is, obtain the output result required by the user.
[0140] Step S31: Determine a target tool for responding to the user instruction based on the user intention and the description information of all tools through the large language model.
[0141] Specifically, since the description information of each tool describes the tasks that the tool can perform, when the large language model obtains the user's intention, it can match the user's intention with the description information of each tool, and thus find the target tool that matches the user's intention. The target tool can be used to respond to the user's instruction.
[0142] In this embodiment, the large language model first analyzes the user's instruction to obtain the user's intention, and then determines the target tool according to the user's intention, so that instructions with different text contents but the same intention can all determine the same target tool.
[0143] In this embodiment, the target tools corresponding to the user's instructions with different user intentions are different.
[0144] When the user's intention is to generate a mind map, the corresponding target tool is a mind map generation tool; when the user's intention is to summarize video content, the corresponding target tool is a video content summary tool; when the user's intention is video content Q&A, the corresponding target tool is a video content question tool; when the user's intention is to ask and answer questions about video content by oneself, the corresponding target tool is a video self-question and self-answer tool; when the user's intention is to generate a recipe, the corresponding target tool is a recipe generation tool; when the user's intention is to extract a video cover, the corresponding target tool is a video cover extraction tool; when the user's intention is to extract background music, the corresponding target tool is a background music extraction tool; when the user's intention is to provide high-energy barrage points, the corresponding target tool is a high-energy barrage point extraction tool.
[0145] Step S23, the large language model framework proxy calls the target tool to execute the task associated with the target tool, and obtains the task execution result, where the task execution result is one of a mind map, a video summary text, a question answer, at least one video question and the question answers corresponding to each of the video questions, a recipe template, a video cover, a music link, and high-energy barrage points.
[0146] Specifically, each target tool is associated with a task, and the tasks associated with different target tools are different.
[0147] In this embodiment, after the large language model determines the target tool for responding to the user's instruction, the large language model framework proxy will call the target tool to execute the corresponding task to implement the response to the user's instruction.
[0148] Among them, the meanings of the mind map, the video summary text, the question answer, at least one video question and the question answers corresponding to each of the video questions, the recipe template, the video cover, the music link, and the high-energy barrage points will be described in detail in the following embodiments.
[0149] In an exemplary embodiment, when the target tool is a mind map generation tool, the task execution result obtained by the large language model framework proxy calling the target tool to execute a task associated with the target tool includes: the large language model framework proxy calling a mind map generation tool to execute a task associated with the mind map generation tool to obtain a task execution result.
[0150] Specifically, when it is determined that the target tool for responding to the user instruction is a mind map generation tool, the large language model framework proxy calls the mind map generation tool to execute a mind map generation task, thereby obtaining a task execution result.
[0151] In a specific embodiment, refer to Figure 4 , the mind map generation tool obtains the task execution result by performing the following steps:
[0152] Step S40, obtain the video ID of the target video.
[0153] Specifically, the video ID is identification information for distinguishing different videos.
[0154] In this embodiment, the video ID can be extracted from the video link of the target video. Among them, the video ID can be extracted by the large language model from the video link, or can be extracted by the mind map generation tool from the video link.
[0155] It can be understood that when the video ID is extracted by the large language model, when the large language model framework proxy calls the mind map generation tool to execute the mind map generation task, the video ID extracted by the large language model will be transmitted to the mind map generation tool as an input parameter of the mind map generation tool, so that the mind map generation tool obtains the video ID.
[0156] Step S41, obtain the audio data of the target video according to the video ID.
[0157] Specifically, after obtaining the video ID, the video file of the target video can be obtained according to the video ID, and then the audio data therein can be obtained from the video file.
[0158] Step S42, call a speech recognition service to convert the audio data into text information.
[0159] Specifically, the speech recognition service is a service that can convert audio data into text information.
[0160] Step S43, segment the text information into multiple text blocks.
[0161] Specifically, a text segmentation tool may be called to segment text information into multiple text blocks, wherein the text segmentation tool may be any tool in the prior art that can segment text, and the specific selected text segmentation tool is not limited in this embodiment.
[0162] Step S44, converting the plurality of text blocks into corresponding vector data, and associating the plurality of text blocks and the corresponding vector data and saving them in a preset vector library.
[0163] Specifically, a text vector model may be called to perform vectorization processing on the text block, thereby converting it into corresponding vector data.
[0164] The text vector model may be a word embedding model, a bag-of-words model (BOW), a word frequency-inverse document frequency model N-Gram language model, or the like.
[0165] Step S45: input the plurality of text blocks and the content summary prompt words in a preset format into the large language model.
[0166] Specifically, the preset format is a format determined in advance according to actual conditions.
[0167] As an example, the preset format may be a lightweight markup format, such as a markdown format.
[0168] Markdown is a lightweight markup language.
[0169] Step S46: output the video summary content in the preset format through the large language model.
[0170] Specifically, by inputting the plurality of text blocks and content summary prompts in a preset format into the large language model, the large language model can summarize the plurality of text blocks and generate the video summary content in the preset format.
[0171] In an exemplary embodiment, when the preset format is the markdown format, LLM can generate an elegant markdown format video content context based on the text information.
[0172] It should be noted that the large language model cannot process too long text at one time, so when summarizing multiple text blocks, the large language model is called multiple times to iteratively summarize multiple text blocks to generate the video summary content.
[0173] Step S47: Convert the video summary content into a mind map.
[0174] Specifically, the markup library can be called to first convert the video summary content into a mind map in HTML format to achieve visual processing of the text, and then the playwright library is called to render and capture a screenshot of the mind map in HTML format, so as to obtain a mind map in picture format.
[0175] Among them, the markup library is a tool library in Python.
[0176] Among them, Playwright is a powerful Python library. Through the Python library, the mind map in HTML format can be rendered into an html web page. After the rendering of the mind map is completed, the playwright library is also called to capture a screenshot of the rendered mind map web page, so as to obtain a mind map in picture form.
[0177] It should be noted that the above-mentioned markup library and playwright library are only exemplary, and the video summary content can also be converted into a mind map through other libraries.
[0178] Step S48: Use the mind map as the task execution result.
[0179] In this embodiment, by summarizing the video content of the target video in the form of a mind map, it is convenient for users to understand the video content.
[0180] In an exemplary embodiment, when the target tool is a video content summarization tool, the task execution result obtained by the large language model framework proxy calling the target tool to execute a task associated with the target tool includes: the large language model framework proxy calling the video content summarization tool to execute a task associated with the video content summarization tool to obtain a task execution result.
[0181] Specifically, when it is determined that the target tool for responding to the user instruction is a video content summarization tool, the large language model framework proxy calls the video content summarization tool to execute the video content summarization task, so as to obtain a task execution result.
[0182] In a specific embodiment, the video content summarization tool obtains the task execution result by performing the following steps: obtaining the video ID of the target video; obtaining the audio data of the target video according to the video ID; calling a speech recognition service to convert the audio data into text information; splitting the text information into multiple text blocks; converting the multiple text blocks into corresponding vector data, and associatively storing the multiple text blocks and the corresponding vector data in a preset vector library; inputting the multiple text blocks into the large language model; outputting a video summary text through the large language model; and using the video summary text as the task execution result.
[0183] Specifically, the above steps are similar to steps S40 - S46 in the above embodiment and will not be elaborated in this embodiment.
[0184] In an exemplary embodiment, when the target tool is a video content question tool, the task execution result obtained by proxy - calling the target tool to execute a task associated with the target tool through the large language model framework includes: proxy - calling the video content question tool through the large language model framework to execute a task associated with the video content question tool to obtain the task execution result.
[0185] Specifically, when it is determined that the target tool for responding to the user instruction is a video content question tool, the large language model framework proxy - calls the video content question tool to execute a video content question task, thereby obtaining the task execution result.
[0186] In a specific embodiment, referring to Figure 5 , the video content question tool obtains the task execution result by performing the following steps:
[0187] Step S50, query whether there is vector data corresponding to the target video in a preset vector library.
[0188] Specifically, since other tools may have already converted multiple text blocks into vector data. Therefore, to avoid repeatedly performing vectorization processing on multiple text blocks, it is first queried whether there is vector data corresponding to the target video in the vector library.
[0189] In an exemplary embodiment, it is possible to query whether there is vector data corresponding to the video ID in the vector library according to the video ID of the target video.
[0190] It can be understood that in order to query whether there is vector data corresponding to the video ID in the vector library through the video ID, when associatively storing multiple text blocks and the corresponding vector data, it will also be associated with the video ID.
[0191] In step S51, if there is vector data corresponding to the target video in the vector library, the user instruction is converted into corresponding vector data.
[0192] Specifically, a text vector model can be called to convert the user instruction into corresponding vector data.
[0193] In step S52, text blocks meeting preset conditions are selected from multiple text blocks based on the vector data corresponding to the user instruction and the vector data corresponding to the target video.
[0194] Specifically, text blocks meeting preset conditions can be selected from multiple text blocks based on the similarity between the vector data corresponding to the user instruction and the vector data corresponding to the target video.
[0195] Among them, the preset condition can be that the similarity value is greater than a preset value, or that the similarity value is among the top N.
[0196] In step S53, the selected text blocks and the user instruction are input into the large language model.
[0197] In step S54, the question answer corresponding to the user instruction is output through the large language model.
[0198] Specifically, the large language model will search for the question answer corresponding to the user instruction from the selected text blocks.
[0199] In step S55, the question answer is used as the task execution result.
[0200] Specifically, when answering questions, the video content question tool in this embodiment does not process all text blocks, but only searches for answers from the selected text blocks, thereby improving the question answering speed.
[0201] It should be noted that in this embodiment, the user instruction serves as a video question.
[0202] In an exemplary implementation, when the target tool is a video self-questioning and self-answering tool, the task execution result obtained by proxy calling the target tool to execute a task associated with the target tool through the large language model framework includes:
[0203] The video self-questioning and self-answering tool is proxy called through the large language model framework to execute a task associated with the video self-questioning and self-answering tool, and the task execution result is obtained.
[0204] Specifically, when the target tool for responding to the user instruction is determined to be the video Q&A tool, the large language model framework agent invokes the video Q&A tool to execute the video Q&A task, thereby obtaining the task execution result.
[0205] In a specific embodiment, refer to Figure 6 , the video Q&A tool obtains the task execution result by performing the following steps:
[0206] Step S60, query whether the video summary text exists in the preset video summary text database;
[0207] Step S61, if the video summary text exists in the video summary text database, input the video summary text into the large language model;
[0208] Step S62, summarize at least one video question through the large language model;
[0209] Step S63, invoke the video content question tool to obtain the question answer corresponding to each video question;
[0210] Step S64, use each video question and the question answer corresponding to each video question as the task execution result output by the video Q&A tool.
[0211] Specifically, if the video summary text does not exist in the video summary text database, the above-mentioned video content summarization tool will be invoked to execute the video summarization task, thereby obtaining the video summary text. If the video summary text database exists the video summary text, the video summary text will be directly used, and at least one video question will be further summarized from the video summary text through the large language model. After obtaining the video question, the above-mentioned video content question tool will be invoked to obtain the question answer corresponding to the video question.
[0212] In this embodiment, the video question and the corresponding question answer are output through the video Q&A tool, so that the user can quickly understand the main content contained in the target video.
[0213] In an exemplary embodiment, when the target tool is a recipe generation tool, the large language model framework agent invokes the target tool to execute the task associated with the target tool, and the obtained task execution result includes: the large language model framework agent invokes the recipe generation tool to execute the task associated with the recipe generation tool, and obtains the task execution result.
[0214] Specifically, when the target tool for responding to the user instruction is determined to be a recipe generation tool, the large language model framework agent invokes the recipe generation tool to execute the recipe generation task, thereby obtaining the task execution result.
[0215] In a specific embodiment, referring to Figure 7 , the recipe generation tool obtains the task execution result by performing the following steps:
[0216] Step S70, obtain the video ID of the target video;
[0217] Step S71, obtain the audio data of the target video according to the video ID;
[0218] Step S72, call the speech recognition service to convert the audio data into text information;
[0219] Step S73, segment the text information into multiple text blocks;
[0220] Step S74, convert the multiple text blocks into corresponding vector data, and associate and save the multiple text blocks and the corresponding vector data to a preset vector library.
[0221] Specifically, the above steps S70 - S74 are the same as the above steps S40 - S44, and will not be elaborated in this embodiment.
[0222] Step S75, input the multiple text blocks and the prompt words describing the recipe making process into the large language model.
[0223] Step S76, output the summary content of the recipe making process through the large language model.
[0224] Specifically, by inputting the multiple text blocks and the prompt words (prompt) describing the recipe making process into the large language model, the large language model can summarize the multiple text blocks and generate the summary content of the recipe making process.
[0225] Among them, the prompt words describing the recipe making process focus on prompting the step - by - step summary of the food making process, so that the LLM can understand the core food making methods and processes introduced by the blogger and make a point - by - point summary.
[0226] Step S77, fill the summary content of the recipe making process into a preset template to obtain a recipe template.
[0227] Specifically, the template is preset according to the recipe making process.
[0228] In an embodiment, for the convenience of subsequent rendering, the template can be an HTML template.
[0229] Step S78, use the recipe template as the task execution result.
[0230] In one embodiment, for the convenience of user understanding, after obtaining the recipe template, the playwright library can be called to render the recipe template in HTML format and take a screenshot of the rendering result, so as to obtain a recipe template in image format.
[0231] In this embodiment, by summarizing the video content of the target video in the form of a recipe, it is convenient for users to understand the recipe making process.
[0232] In an exemplary embodiment, when the target tool is a video cover extraction tool, the task execution result obtained by proxy calling the target tool to execute the task associated with the target tool through the large language model framework includes: proxy calling the video cover extraction tool through the large language model framework to execute the task associated with the video cover extraction tool, and obtaining the task execution result.
[0233] Among them, the video cover extraction tool obtains the task execution result by performing the following steps: obtaining the video ID of the target video; obtaining the video cover of the target video according to the video ID; using the video cover as the task execution result.
[0234] Specifically, after obtaining the video ID, all content related to the target video can be found according to the video ID, and then the video cover can be extracted from these contents.
[0235] In an exemplary embodiment, when the target tool is a background music extraction tool, the task execution result obtained by proxy calling the target tool to execute the task associated with the target tool through the large language model framework includes: proxy calling the background music extraction tool through the large language model framework to execute the task associated with the background music extraction tool, and obtaining the task execution result.
[0236] Among them, the background music extraction tool obtains the task execution result by performing the following steps: obtaining the video ID of the target video; obtaining the music name of the background music of the target video according to the video ID; obtaining the music link of the background music according to the music name; using the music link as the task execution result.
[0237] Specifically, after obtaining the video ID, all content related to the target video can be found according to the video ID, and then the music name of the background music can be obtained from these contents. Then, the music link of the background music can be obtained from at least one music platform according to the music name.
[0238] In an exemplary embodiment, when the target tool is a high-energy barrage point extraction tool, the task execution result obtained by proxy-calling the target tool through the large language model framework to execute the task associated with the target tool includes: proxy-calling the high-energy barrage point extraction tool through the large language model framework to execute the task associated with the high-energy barrage point extraction tool, and obtaining the task execution result.
[0239] Among them, the high-energy barrage point extraction tool obtains the task execution result by performing the following steps: obtaining the video ID of the target video; obtaining the barrage of the target video according to the video ID; analyzing the barrage to obtain high-energy barrage points; and using the high-energy barrage points as the task execution result.
[0240] Specifically, after obtaining the video ID, all content related to the target video can be found according to the video ID. Then, the barrage of the target video can be obtained from this content. Next, all the barrages can be analyzed to obtain high-energy barrage points.
[0241] Step S24: Return the task execution result to the user.
[0242] Specifically, the task execution result can be returned to the user by means of a reply or a note.
[0243] In a specific scenario, the task execution result can be returned to the user in the form of a reply according to the user ID and the comment area ID.
[0244] The video processing method of the embodiment of the present application includes: obtaining a user instruction sent by a user based on a target video; creating a large language model framework proxy according to the user instruction; analyzing the user instruction through the large language model framework proxy to obtain a user intention; calling a tool matching the user intention through the large language model to execute a task corresponding to the user intention to obtain a task execution result; and returning the task execution result to the user. By adopting the above video processing method, different tools can be called according to different user intentions to execute tasks corresponding to the user intentions, so as to realize diversified processing of videos, meet the personalized needs of users, and increase the interaction methods between users and videos.
[0245] An embodiment of the present application also provides a video processing method, which is applied to a client. The method includes: receiving a user instruction input by a user based on a target video; inputting the user instruction into a large language model, so that the large language model outputs a processing result based on the user instruction and the video content of the target video, and the processing result is one of a mind map, a video summary text, an answer to a question, at least one video question and the corresponding answer to each video question, and a recipe template; and displaying the processing result.
[0246] In this embodiment, after the user inputs a user instruction, the user instruction will be input into the large language model, so that the large language model can output a processing result matching the user instruction based on the user instruction and the video content of the target video, achieving the effect of outputting a corresponding processing result according to the user's intention.
[0247] It should be noted that the definitions of the user instruction, the large language model, the video content, the mind map, the video summary text, the answer to a question, at least one video question and the corresponding answer to each video question, and the recipe template in this embodiment are the same as those of the corresponding terms in the above embodiment, and will not be elaborated in this embodiment.
[0248] Refer to Figure 8 As shown, it is a program module diagram of an embodiment of a video processing device 80 of the present application.
[0249] In this embodiment, the video processing device 80 includes a series of computer program instructions stored in a memory. When the computer program instructions are executed by a processor, the video processing functions of the embodiments of the present application can be implemented. In some embodiments, based on the specific operations implemented by each part of the computer program instructions, the video processing device 80 can be divided into one or more modules. The specific modules that can be divided are as follows:
[0250] An acquisition module 81, configured to acquire a user instruction sent by a user based on a target video;
[0251] A creation module 82, configured to create a large language model framework proxy according to the user instruction;
[0252] An input module 83, configured to input the user instruction and the description information of all tools included in the large language model framework proxy into the large language model through the large language model framework proxy, so as to output a target tool for responding to the user instruction through the large language model;
[0253] An execution module 84, configured to proxy-invoke the target tool through the large language model framework to execute a task associated with the target tool, and obtain a task execution result, where the task execution result is one of a mind map, a video summary text, a question answer, at least one video question and the corresponding question answer for each video question, a recipe template, a video cover, a music link, and high-energy barrage points;
[0254] A return module 85, configured to return the task execution result to the user.
[0255] In an exemplary embodiment, the target tool output for responding to the user instruction through the large language model includes: analyzing the user instruction through the large language model to obtain a user intention, where the user intention is the output result required by the user; and determining, based on the user intention and the description information of all the tools through the large language model, the target tool for responding to the user instruction.
[0256] In an exemplary embodiment, when the target tool is a mind map generation tool, the proxy-invoking the target tool through the large language model framework to execute a task associated with the target tool and obtaining a task execution result includes: proxy-invoking a mind map generation tool through the large language model framework to execute a task associated with the mind map generation tool, and obtaining a task execution result;
[0257] Wherein, the mind map generation tool obtains the task execution result by performing the following steps:
[0258] Obtain the video ID of the target video;
[0259] Obtain the audio data of the target video according to the video ID;
[0260] Call a speech recognition service to convert the audio data into text information;
[0261] Segment the text information into multiple text blocks;
[0262] Input multiple text blocks and a content summary prompt word in a preset format into the large language model;
[0263] Output the video summary content in the preset format through the large language model;
[0264] Convert the video summary content into a mind map;
[0265] Use the mind map as the task execution result.
[0266] In an exemplary embodiment, when the target tool is a video content summarization tool, the task execution result obtained by proxy-calling the target tool through the large language model framework to execute a task associated with the target tool includes: proxy-calling the video content summarization tool through the large language model framework to execute a task associated with the video content summarization tool, and obtaining a task execution result;
[0267] Wherein, the video content summarization tool obtains the task execution result by performing the following steps:
[0268] Obtain the video ID of the target video;
[0269] Obtain the audio data of the target video according to the video ID;
[0270] Call a speech recognition service to convert the audio data into text information;
[0271] Segment the text information into multiple text blocks;
[0272] Convert multiple text blocks into corresponding vector data, and associate and save multiple text blocks and corresponding vector data to a preset vector library;
[0273] Input multiple text blocks into the large language model;
[0274] Output video summary text through the large language model;
[0275] Use the video summary text as the task execution result.
[0276] In an exemplary embodiment, when the target tool is a video content question tool, the task execution result obtained by proxy-calling the target tool through the large language model framework to execute a task associated with the target tool includes: proxy-calling the video content question tool through the large language model framework to execute a task associated with the video content question tool, and obtaining a task execution result;
[0277] Wherein, the video content question tool obtains the task execution result by performing the following steps:
[0278] Query whether there is vector data corresponding to the target video in a preset vector library;
[0279] If there is vector data corresponding to the target video in the vector library, convert the user instruction into corresponding vector data;
[0280] Select text blocks that meet the preset conditions from multiple text blocks based on the vector data corresponding to the user instruction and the vector data corresponding to the target video;
[0281] Input the selected text blocks and the user instruction into the large language model;
[0282] Output the question answer corresponding to the user instruction through the large language model;
[0283] Use the question answer as the task execution result.
[0284] In an exemplary embodiment, when the target tool is a video Q&A tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute the task associated with the target tool includes: proxy calling the video Q&A tool through the large language model framework to execute the task associated with the video Q&A tool, and obtaining the task execution result;
[0285] Among them, the video Q&A tool obtains the task execution result by performing the following steps:
[0286] Query whether the video summary text exists in the preset video summary text database;
[0287] If the video summary text exists in the video summary text database, input the video summary text into the large language model;
[0288] Summarize at least one video question through the large language model;
[0289] Call the video content question tool to obtain the question answer corresponding to each video question;
[0290] Use each video question and the question answer corresponding to each video question as the task execution result output by the video Q&A tool.
[0291] In an exemplary embodiment, when the target tool is a recipe generation tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute the task associated with the target tool includes: proxy calling the recipe generation tool through the large language model framework to execute the task associated with the recipe generation tool, and obtaining the task execution result;
[0292] Among them, the recipe generation tool obtains the task execution result by performing the following steps:
[0293] Obtain the video ID of the target video;
[0294] Obtain the audio data of the target video according to the video ID;
[0295] Call the speech recognition service to convert the audio data into text information;
[0296] Segment the text information into multiple text blocks;
[0297] Convert the multiple text blocks into corresponding vector data, and associate and save the multiple text blocks and the corresponding vector data to a preset vector library;
[0298] Input the multiple text blocks and the prompt words describing the recipe making process into the large language model;
[0299] Output the summary content of the recipe making process through the large language model;
[0300] Fill the summary content of the recipe making process into a preset template to obtain a recipe template;
[0301] Use the recipe template as the task execution result.
[0302] In an exemplary embodiment, when the target tool is a video cover extraction tool, the proxy call of the target tool through the large language model framework to execute the task associated with the target tool to obtain the task execution result includes: proxy call of the video cover extraction tool through the large language model framework to execute the task associated with the video cover extraction tool to obtain the task execution result;
[0303] Among them, the video cover extraction tool obtains the task execution result by performing the following steps:
[0304] Obtain the video ID of the target video;
[0305] Obtain the video cover of the target video according to the video ID;
[0306] Use the video cover as the task execution result.
[0307] In an exemplary embodiment, when the target tool is a background music extraction tool, the proxy call of the target tool through the large language model framework to execute the task associated with the target tool to obtain the task execution result includes: proxy call of the background music extraction tool through the large language model framework to execute the task associated with the background music extraction tool to obtain the task execution result;
[0308] Among them, the background music extraction tool obtains the task execution result by performing the following steps:
[0309] Obtain the video ID of the target video;
[0310] Obtain the music name of the background music of the target video according to the video ID;
[0311] Obtain the music link of the background music according to the music name;
[0312] Use the music link as the task execution result.
[0313] In an exemplary embodiment, when the target tool is a high-energy barrage point extraction tool, the task execution result obtained by proxy calling the target tool to execute a task associated with the target tool through the large language model framework includes: proxy calling the high-energy barrage point extraction tool through the large language model framework to execute a task associated with the high-energy barrage point extraction tool, and obtaining a task execution result;
[0314] Among them, the high-energy barrage point extraction tool obtains the task execution result by performing the following steps:
[0315] Obtain the video ID of the target video;
[0316] Obtain the barrage of the target video according to the video ID;
[0317] Analyze the barrage to obtain high-energy barrage points;
[0318] Use the high-energy barrage points as the task execution result.
[0319] Figure 9 Schematically shows a hardware architecture diagram of a computer device 9 suitable for implementing the video processing method according to an embodiment of the present application. In this embodiment, the computer device 9 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Such as Figure 9 shown, the computer device 9 at least includes but is not limited to: a memory 120, a processor 121, and a network interface 122 that can communicate with each other through a system bus. Among them:
[0320] The memory 120 includes at least one type of computer-readable storage medium, which can be volatile or non-volatile. Specifically, the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 120 can be an internal storage module of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 120 can also be an external storage device of the computer device 9, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the computer device 9. Of course, the memory 120 can also include both the internal storage module and the external storage device of the computer device 9. In this embodiment, the memory 120 is generally used to store the operating system and various application software installed on the computer device 9, such as the program code of the video processing method. In addition, the memory 120 can also be used to temporarily store various data that have been output or will be output.
[0321] In some embodiments, the processor 121 can be a central processing unit (CPU), controller, microcontroller, microprocessor, or other video processing chip. The processor 121 is generally used to control the overall operation of the computer device 9, such as performing control and processing related to data interaction or communication with the computer device 9. In this embodiment, the processor 121 is used to run the program code stored in the memory 120 or process data.
[0322] The network interface 122 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 9 and other computer devices. For example, the network interface 122 is used to connect the computer device 9 to an external terminal via a network, and establish a data transmission channel and a communication link between the computer device 9 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet, the Internet, Global System of Mobile communication (abbreviated as GSM), Wideband Code Division Multiple Access (abbreviated as WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0323] It should be noted that Figure 9 only the computer device having components 120 to 122 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.
[0324] In this embodiment, the video processing method stored in the memory 120 may be divided into one or more program modules and executed by one or more processors (processor 121 in this embodiment) to complete the present application.
[0325] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the video processing method in the embodiment are implemented.
[0326] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the video processing method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or are to be output.
[0327] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to at least two network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application. A person of ordinary skill in the art can understand and implement it without creative effort.
[0328] Through the description of the above embodiments, a person of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. A person of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.
[0329] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A video processing method, characterized in that, The method includes: Obtaining a user instruction sent by the user based on the target video; Creating a large language model framework agent according to the user instruction; Inputting the user instruction and the description information of all tools that the large language model framework agent can call into the large language model through the large language model framework agent, so as to output a target tool for responding to the user instruction through the large language model; Calling the target tool through the large language model framework agent to execute a task associated with the target tool, and obtaining a task execution result, where the task execution result is one of a mind map, a video summary text, a question answer, at least one video question and the corresponding question answers for each video question, a recipe template, a video cover, a music link, and high-energy barrage points; Returning the task execution result to the user.
2. The video processing method according to claim 1, wherein Outputting a target tool for responding to the user instruction through the large language model includes: Analyzing the user instruction through the large language model to obtain a user intention, where the user intention is the output result required by the user; Determining a target tool for responding to the user instruction based on the user intention and the description information of all tools through the large language model.
3. The video processing method according to claim 1, wherein When the target tool is a mind map generation tool, the calling the target tool through the large language model framework agent to execute a task associated with the target tool and obtaining a task execution result includes: Calling a mind map generation tool through the large language model framework agent to execute a task associated with the mind map generation tool, and obtaining a task execution result; Wherein, the mind map generation tool obtains the task execution result by performing the following steps: Obtaining the video ID of the target video; Obtaining the audio data of the target video according to the video ID; Calling a speech recognition service to convert the audio data into text information; Segmenting the text information into multiple text blocks; Converting multiple text blocks into corresponding vector data, and associating and storing multiple text blocks and corresponding vector data in a preset vector library; Inputting multiple text blocks and a content summary prompt in a preset format into the large language model; Outputting the video summary content in the preset format through the large language model; Converting the video summary content into a mind map; Using the mind map as the task execution result.
4. The video processing method according to claim 1, wherein When the target tool is a video content summary tool, the calling the target tool through the large language model framework agent to execute a task associated with the target tool and obtaining a task execution result includes: Calling the video content summary tool through the large language model framework agent to execute a task associated with the video content summary tool, and obtaining a task execution result; Wherein, the video content summary tool obtains the task execution result by performing the following steps: Obtaining the video ID of the target video; Obtaining the audio data of the target video according to the video ID; Calling a speech recognition service to convert the audio data into text information; Segmenting the text information into multiple text blocks; Convert multiple of the text blocks into corresponding vector data, and associate and save the multiple text blocks and the corresponding vector data to a preset vector library; Input multiple of the text blocks into the large language model; Output video summary text through the large language model; Use the video summary text as the task execution result.
5. The video processing method according to claim 1, characterized in that When the target tool is a video content question tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes: Proxy call the video content question tool through the large language model framework to execute a task associated with the video content question tool, and obtain a task execution result; Among them, the video content question tool obtains the task execution result by performing the following steps: Query whether there is vector data corresponding to the target video in a preset vector library; If there is vector data corresponding to the target video in the vector library, convert the user instruction into corresponding vector data; Select text blocks that meet preset conditions from multiple of the text blocks based on the vector data corresponding to the user instruction and the vector data corresponding to the target video; Input the selected text blocks and the user instruction into the large language model; Output the question answer corresponding to the user instruction through the large language model; Use the question answer as the task execution result.
6. The video processing method according to claim 5, wherein When the target tool is a video self-question-and-answer tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes: Proxy call the video self-question-and-answer tool through the large language model framework to execute a task associated with the video self-question-and-answer tool, and obtain a task execution result; Among them, the video self-question-and-answer tool obtains the task execution result by performing the following steps: Query whether the video summary text exists in a preset video summary text database; If the video summary text exists in the video summary text database, input the video summary text into the large language model; Summarize at least one video question through the large language model; Call the video content question tool to obtain the question answer corresponding to each video question; Use each video question and the question answer corresponding to each video question as the task execution result output by the video self-question-and-answer tool.
7. The video processing method according to claim 1, wherein When the target tool is a recipe generation tool, the task execution result obtained by proxy calling the target tool through the large language model framework to execute a task associated with the target tool includes: Proxy call the recipe generation tool through the large language model framework to execute a task associated with the recipe generation tool, and obtain a task execution result; Among them, the recipe generation tool obtains the task execution result by performing the following steps: Obtain the video ID of the target video; Obtain the audio data of the target video according to the video ID; Call a speech recognition service to convert the audio data into text information; Segment the text information into multiple text blocks; Convert the multiple text blocks into corresponding vector data, and associate and save the multiple text blocks and the corresponding vector data to a preset vector library; Input the multiple text blocks and a prompt word describing the recipe making process into the large language model; Output a summary of the recipe making process through the large language model; Fill the summary of the recipe making process into a preset template to obtain a recipe template; Use the recipe template as the task execution result.
8. The video processing method according to claim 1, characterized in that When the target tool is a video cover extraction tool, the task execution result obtained by proxy calling the target tool to execute a task associated with the target tool through the large language model framework includes: Proxy call a video cover extraction tool through the large language model framework to execute a task associated with the video cover extraction tool to obtain a task execution result; Among them, the video cover extraction tool obtains the task execution result by performing the following steps: Obtain the video ID of the target video; Obtain the video cover of the target video according to the video ID; Use the video cover as the task execution result.
9. The video processing method according to claim 1, wherein When the target tool is a background music extraction tool, the task execution result obtained by proxy calling the target tool to execute a task associated with the target tool through the large language model framework includes: Proxy call a background music extraction tool through the large language model framework to execute a task associated with the background music extraction tool to obtain a task execution result; Among them, the background music extraction tool obtains the task execution result by performing the following steps: Obtain the video ID of the target video; Obtain the music name of the background music of the target video according to the video ID; Obtain the music link of the background music according to the music name; Use the music link as the task execution result.
10. The video processing method according to claim 1, wherein When the target tool is a high-energy barrage point extraction tool, the task execution result obtained by proxy calling the target tool to execute a task associated with the target tool through the large language model framework includes: Proxy call a high-energy barrage point extraction tool through the large language model framework to execute a task associated with the high-energy barrage point extraction tool to obtain a task execution result; Among them, the high-energy barrage point extraction tool obtains the task execution result by performing the following steps: Obtain the video ID of the target video; Obtain the barrage of the target video according to the video ID; Analyze the barrage to obtain high-energy barrage points; Use the high-energy barrage points as the task execution result.
11. A video processing method, characterized in that, The method includes: Receive a user instruction input by the user based on the target video; Input the user instruction into the large language model for the large language model to output a processing result based on the user instruction and the video content of the target video, and the processing result is one of a mind map, a video summary text, a question answer, at least one video question and the question answer corresponding to each video question, and a recipe template; Display the processing result.
12. A video processing device, characterized in that, The video processing device includes: An acquisition module, configured to acquire a user instruction sent by a user based on a target video; A creation module, configured to create a large language model framework proxy according to the user instruction; An input module, configured to input the user instruction and description information of all tools included in the large language model framework proxy into a large language model through the large language model framework proxy, so as to output a target tool for responding to the user instruction through the large language model; An execution module, configured to call the target tool through the large language model framework proxy to execute a task associated with the target tool, and obtain a task execution result, where the task execution result is one of a mind map, a video summary text, a question answer, at least one video question and a question answer corresponding to each video question, a recipe template, a video cover, a music link, and a high-energy barrage point; A return module, configured to return the task execution result to the user.
13. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.