Video processing method and system based on large language model
Through a large language model combined with vectorized private domain knowledge data, a video processing task execution plan is generated, which solves the problems of poor intelligent collaboration and information security of traditional tools, and realizes efficient and secure closed-loop optimization of video processing.
Patent Information
- Application Number
- CN202510904501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Traditional video processing tools have poor intelligent synergy, which is difficult to meet the needs of complex scenarios, have low utilization of private domain data, lack feedback mechanisms, and large language models have information security problems in the application of professional fields.
Using a large language model for video processing, vectorize private domain knowledge information data, combine with a video processing tool set, a task execution plan is generated, and a closed-loop optimization is formed after execution.
Improve the accuracy and efficiency of video processing, realize multi-tool management and secure utilization of private domain data, forming a complete closed loop from demand to feedback.
Smart Images

Figure CN120407848A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing, and in particular to a video processing method and system based on a large language model. Background Art
[0002] Amid the explosive growth in video content creation and distribution, video processing needs are becoming increasingly diverse and complex, encompassing, but not limited to, video transcoding, video enhancement, video restoration, and video super-resolution. Traditional video processing methods face numerous technical bottlenecks: 1. The diverse nature of video processing requirements requires manual understanding and time-consuming adjustments to the various parameters of video processing tools to meet these diverse needs. 2. A wide variety of video processing tools with diverse functions are available to meet these diverse needs. Traditional video processing technologies are highly targeted and excel at addressing single video issues. However, in reality, improving video quality, for example, may involve multiple processing steps, including video enhancement, denoising, and sharpening. Some of these issues may even require conflicting processing principles, necessitating a balanced approach across multiple processing methods to ensure optimal results. However, the lack of intelligent collaboration between individual video processing tools makes it difficult to meet the demands of complex scenarios. 3. Enterprises accumulate large amounts of private data related to video processing (such as historical transcoding cases, software parameter manuals, and industry standard documents). Due to a lack of effective data processing and management methods, it is difficult to quickly extract key information to guide new video processing tasks, resulting in low utilization of private data. 4. Traditional video processing tools lack effective feedback mechanisms, making them unable to optimize based on actual video processing results and user feedback, making them difficult to adapt to technological developments and changing user needs.
[0003] On the other hand, with the advancement of AI technology, large language models (LLMs), as large-scale neural network models, have demonstrated impressive performance in processing human language. However, their application in real-world tasks still faces challenges. Furthermore, while large language models utilize extensive training data, they are also highly dependent on that data. When applied to specialized fields, these models are limited by specialized knowledge and lack in-depth domain knowledge. This results in limited usability in these fields. In order to utilize existing large language models in specialized fields, specialized and complex fine-tuning training is required. In the field of video processing, relevant companies and institutions have accumulated vast amounts of specialized, highly confidential, and private data related to video processing through their production, operations, and research. However, model training presents a challenge: while training data is typically de-identified, the model may retain certain information, which can then be leaked during text generation. Consequently, the application of large language models in specialized fields presents an information security dilemma. Summary of the Invention
[0004] Due to the above problems existing in the existing methods, embodiments of the present invention propose a video processing method and system based on a large language model.
[0005] Specifically, embodiments of the present invention provide the following technical solutions:
[0006] In a first aspect, embodiments of the present invention provide a video processing method based on a large language model, including:
[0007] Using natural language to input video processing task instructions.
[0008] Performing semantic parsing on the task instructions and vectorizing them into query vectors.
[0009] Performing similarity search in a vector database according to the query vectors.
[0010] The vector database stores privately-owned knowledge information data related to the field of video processing that has been pre-vectorized, including: video processing operation guides, video processing tool parameter descriptions, video pre-processing tool descriptions, historical test data, industry technical specifications, and video processing experience summaries.
[0011] Semantically associating the query vectors with the similarity search results, integrating them in combination with video processing tool set information according to a preset prompt word template, and outputting a prompt word context.
[0012] The video processing tool set information includes, but is not limited to, the functional attributes of video processing tools, input / output interface specifications, call functions, call protocols, and performance indicators.
[0013] Using a large language model to perform semantic parsing on the prompt word context, generating and executing a video processing task execution plan, and outputting a processed video file.
[0014] In another aspect, embodiments of the present invention provide a video processing system based on a large language model, including:
[0015] A user interaction module for using natural language to input video processing task instructions.
[0016] A vectorization module for performing semantic parsing on the task instructions and vectorizing them into query vectors.
[0017] A data retrieval module for performing similarity search in a vector database according to the query vectors.
[0018] The vector database stores privately-owned knowledge information data related to the field of video processing that has been vectorized.
[0019] A prompt word generation module, which is used to semantically associate the query vector with the similarity search results, integrate the information of the video processing tool set, and output the prompt word context according to a preset prompt word template.
[0020] A video processing tool management module, which is used to register and uniformly manage video processing tools, and define in detail the information related to video processing tools.
[0021] An intelligent processing module, which is used to semantically parse the prompt word context by using a large language model, generate and execute a video processing task execution plan, and output a processed video file.
[0022] As can be seen from the above technical solutions, the embodiments of the present invention have the following beneficial effects: on the one hand, using the artificial intelligence technology of the large language model for video processing greatly reduces the time cost of manual parameter adjustment and improves the accuracy of video processing; the technical solution provided by the present invention also realizes the unified management and intelligent scheduling of multiple video processing tools, automatically selects the most suitable one or more tools and the most suitable parameter configuration for video processing tasks; on the other hand, using the vectorization model to vectorize the professional knowledge in this field, combined with the vector database, while avoiding the need for professional training of the existing large language model, it also avoids the information security problems commonly existing in current large models, and improves the utilization of enterprise private domain professional data; finally, using the video processing results to perform self-learning and optimization on the entire video processing system, realizing a complete closed loop from demand processing, task execution to data feedback optimization of the video processing system. Description of the Drawings
[0023] Figure 1 It is a flowchart of a video processing method based on a large language model provided by an embodiment of the present invention.
[0024] Figure 2 It is a flowchart of a video processing method based on a large language model provided by another embodiment of the present invention.
[0025] Figure 3 It is a structural diagram of a video processing system based on a large language model provided by an embodiment of the present invention.
[0026] Figure 4 It is a structural diagram of a video processing system based on a large language model provided by another embodiment of the present invention. Detailed Embodiments
[0027] The following combines the drawings to further describe the specific embodiments of the present invention. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0028] Figure 1The flowchart of the video processing method based on the large language model provided by an embodiment of the present invention is shown. As Figure 1 shown, the video processing method based on the large speech model provided by an embodiment of the present invention specifically includes the following content:
[0029] Step 1, input a video processing task instruction using natural language.
[0030] It should be noted that the embodiments of the present invention do not limit the specific task instructions for video processing. According to actual needs, the video processing tasks can be any related tasks in the field of video processing, including but not limited to: video encoding tasks, video denoising tasks, image enhancement tasks, image super-resolution tasks, video frame interpolation tasks.
[0031] Exemplarily, the task instruction using natural language: Transcode the video file to reduce the volume by 30% while ensuring that the subjective quality index remains unchanged.
[0032] Exemplarily, the task instruction using natural language: Remove the mosaic noise in the picture so that the VMAF score of the denoised picture reaches more than 90 points.
[0033] Exemplarily, the task instruction using natural language: Perform super-resolution processing on a video file with a resolution of 720P and output a video file with a resolution of 1080P.
[0034] Exemplarily, the task instruction using natural language: Encode the video into a format suitable for publishing on social media platforms and highlight the color expressiveness of the video.
[0035] Step 2, perform semantic parsing on the task instruction and vectorize it into a query vector.
[0036] Step 3, perform a similarity search in the vector database according to the query vector.
[0037] Among them, the vector database stores private domain knowledge information data related to the field of video processing that has been pre-vectorized, including but not limited to video processing operation guides, video processing tool parameter descriptions, video pre-processing tool descriptions, historical test data, industry technical specifications, and video processing experience summaries.
[0038] It should be noted that the embodiments of the present invention do not limit the vectorization algorithm for private domain knowledge information data, and vectorize it according to actual needs.
[0039] For example, for data in PDF format, a PDF parsing library can be used to extract text content, and semantic analysis can be performed through a pre-trained model based on BERT to convert content such as operation steps and parameter descriptions into semantic vectors.
[0040] For example, for data in XLS format, you can use the Pandas library to read the data and convert each row of records into a feature vector through a custom data mapping function.
[0041] Furthermore, in order to improve search efficiency and accuracy, hierarchical storage can be performed according to data type and importance, and a multi-level index can be established.
[0042] Step 4: semantically associate the query vector with the similarity search results, combine them with the video processing tool set information, integrate them according to the preset prompt word template, and output the prompt word context.
[0043] The video processing tool set is composed of various video processing tools and is managed in a unified manner. The video processing tool set information includes but is not limited to: functional attributes, input and output interface specifications, calling functions, calling protocols, and performance indicators of each video processing tool.
[0044] It is understandable that the embodiment of the present invention does not limit which video processing tools are specifically included in the video processing tool set, and corresponding video processing tools can be integrated and registered according to actual needs.
[0045] It should also be noted that the embodiment of the present invention does not limit the specific form of the prompt word template, and it can be set according to actual usage scenarios.
[0046] Exemplarily, the prompt word template may be:
[0047] Video processing requirements input by the user: {video processing task instructions described by the user in natural language}.
[0048] Related data retrieved from the vector database: 1. {search result 1}; 2. {search result 2}; ...; n. {search result n}.
[0049] For example, the video processing requirements input by the user are: {encode the video into a format suitable for publishing on social media platforms and highlight the color expression of the video}.
[0050] Relevant data retrieved from the vector database: 1. {Related knowledge: Social media platforms (such as TikTok and WeChat Video) recommend the MP4 video format, 1080×1920 resolution, and 30fps frame rate; video color expression can be improved by adjusting color correction parameters (such as saturation, contrast, and hue); 2. {Video processing tools: video encoding tools, video color processing tools}.
[0051] Furthermore, in order to more intuitively display the correlation between the retrieval results and the video processing task instructions, the retrieved data may be arranged in descending order of relevance.
[0052] Step 5: Use a large language model to perform semantic parsing on the context of the prompt words, generate a video processing task execution plan and execute it, and output the processed video file.
[0053] It should be noted that the embodiments of the present invention do not limit the specific large language model, and any large language model can be used, including but not limited to Transformer, MoE, hybrid architectures, etc. The large language model has the ability to obtain knowledge such as relevant basic concepts and terms in the field of video processing.
[0054] Specifically, the large language model accepts the integrated context of the prompt words, performs in-depth semantic analysis and logical reasoning based on its powerful natural language understanding and generation capabilities. According to the video processing task instructions, appropriate tools are selected from the video processing tool set, and a detailed task execution plan including the call order, call method, parameter configuration, data flow, etc. of the video processing tools is generated.
[0055] Exemplarily, the task execution plan includes but is not limited to the following content:
[0056] Task steps: List the specific step operations to complete the video processing requirements.
[0057] Use of video processing tools: Specify the video tools used in each step and their call methods.
[0058] Parameter settings: Parameter settings of the called video tools and relevant parameter settings of each model.
[0059] Data dependencies: Explain the data sources and processing methods required for each step.
[0060] Expected output: Describe the results obtained after each step is completed.
[0061] Among them, the task execution plan can be one or more feasible solutions, listed in descending order of priority.
[0062] Exemplarily, the video processing tool call method is expressed as follows:
[0063] {Tool 1 name}: {Tool 1 function description}, call method: {Tool 1 call parameter or instruction}; {Tool 2 name}: {Tool 2 function description}, call method: {Tool 2 call parameter or instruction};...; {Tool n name}: {Tool n function description}, call method: {Tool n call parameter or instruction}.
[0064] For example, for the task instruction "Encode the video into a format suitable for publishing on a social media platform and highlight the color expressiveness of the video", the video processing tools to be called and the calling methods are as follows: {Video Encoding Tool}: {Supports MP4 format conversion, can adjust resolution, frame rate, and color correction parameters}, and the calling method is {VideoTranscoder, with parameters passed in JSON format}; {Video Color Processing Tool}: {Specifically used for video color optimization}, and the calling method is {VideoProcessor, with parameters passed in JSON format}.
[0065] Furthermore, to improve security, the task execution plan can be presented to the user before executing the task execution plan, and the user is requested to authorize the execution.
[0066] Furthermore, to ensure smooth coordination between video processing tools, during the task execution process, the tool call status and data transmission situation can be monitored. If an abnormality occurs, the tool can be reselected or the parameters can be adjusted.
[0067] Furthermore, as Figure 2 shown, in another embodiment, after the processed video file is output, the processed video file can be tested for video quality, a test report can be formed, and the test report can be vectorized and then updated back to the vector database.
[0068] The test report includes but is not limited to: task instructions, parameter settings of video tools (such as resolution, frame rate, bit rate), video quality scores (such as PSNR, VMAF), video processing performance metrics (such as processing time, resource occupancy rate), user evaluations, and other information.
[0069] The process of vectorizing the test report and then updating it back to the vector database includes:
[0070] Classify the test report according to the video processing task type and assign corresponding task type labels. <00001
[0071] It should be noted that the specific classification of the video processing task type in the embodiments of the present invention is not limited, and can be classified according to actual situations. For example: video encoding task, video denoising task, image enhancement task, image super-resolution task, video frame interpolation task.
[0072] Map the video quality score to the relevance between the task execution plan of the current video processing task and the current task category and assign a relevance label.<000
[0073] Among them, the higher the video quality score, the higher the relevance, and the lower the video quality score, the lower the relevance; the video quality score is the comprehensive score and / or single-dimensional score of the processed video file.
[0074] Vectorize the content of the test report and the corresponding tags and update them to the vector database.
[0075] Use the vectorized content of the test report as new private domain knowledge information data, and compare and analyze it with the historical data in the vector database to improve the ability of video processing tasks, forming a complete closed loop from requirement processing, task execution to data feedback optimization.
[0076] For example, if it is found that the current encoding task has a long encoding time while improving the color expressiveness, through an algorithm based on reinforcement learning, it is analyzed that the complexity of the color processing algorithm can be appropriately reduced to shorten the encoding time.
[0077] Figure 3 FIG. shows the architecture diagram of a video processing system based on a large language model provided by an embodiment of the present invention. As Figure 3 shown, a video processing system based on a large speech model provided by an embodiment of the present invention specifically includes:
[0078] A user interaction module for using natural language to input video processing task instructions.
[0079] It should be noted that the embodiments of the present invention do not limit the specific task instructions for video processing. According to actual needs, the video processing tasks can be any related tasks in the field of video processing, including but not limited to: video encoding tasks, video denoising tasks, image enhancement tasks, image super-resolution tasks, video frame interpolation tasks.
[0080] Exemplarily, a task instruction using natural language: Transcode a video file to reduce its volume by 30% while ensuring that the subjective quality index remains unchanged.
[0081] Exemplarily, a task instruction using natural language: Remove the mosaic noise in a picture so that the VMAF score of the denoised picture reaches more than 90 points.
[0082] Exemplarily, a task instruction using natural language: Perform super-resolution processing on a video file with a resolution of 720P and output a video file with a resolution of 1080P.
[0083] Exemplarily, a task instruction using natural language: Encode a video into a format suitable for publishing on a social media platform and highlight the color expressiveness of the video.
[0084] A vectorization module for semantically parsing the task instruction and vectorizing it into a query vector.
[0085] A data retrieval module for performing similarity search in the vector database according to the query vector.
[0086] Among them, the vector database stores private domain knowledge information data related to the video processing field that has been vectorized in advance, including but not limited to video processing operation guides, video processing tool parameter descriptions, video pre-processing tool descriptions, historical test data, industry technical specifications, and video processing experience summaries.
[0087] It should be noted that the embodiment of the present invention does not limit the vectorization algorithm of private domain knowledge information data, and vectorization is performed according to actual needs.
[0088] For example, for data in PDF format, the PDF parsing library can be used to extract text content, and semantic analysis can be performed through a BERT-based pre-trained model to convert operation steps, parameter descriptions, and other content into semantic vectors.
[0089] For example, for data in XLS format, you can use the Pandas library to read the data and convert each row of records into a feature vector through a custom data mapping function.
[0090] Furthermore, in order to improve search efficiency and accuracy, hierarchical storage can be performed according to data type and importance, and a multi-level index can be established.
[0091] The prompt word generation module is used to semantically associate the query vector with the similarity search results, combine the video processing tool set information, integrate them according to the preset prompt word template, and output the prompt word context.
[0092] It should be noted that the embodiment of the present invention does not limit the specific form of the prompt word template, and it can be set according to actual usage scenarios.
[0093] Exemplarily, the prompt word template may be:
[0094] Video processing requirements input by the user: {video processing task instructions described by the user in natural language}.
[0095] Related data retrieved from the vector database: 1. {search result 1}; 2. {search result 2}; ...; n. {search result n}.
[0096] For example, the video processing requirements input by the user are: {encode the video into a format suitable for publishing on social media platforms and highlight the color expression of the video}.
[0097] Relevant data retrieved by the vector database: 1. {Relevant knowledge: The recommended video format for social media platforms (such as Douyin and WeChat Video Channels) is MP4, with a recommended resolution of 1080×1920 and a frame rate of 30fps; the color expressiveness of the video can be enhanced by adjusting color correction parameters (such as saturation, contrast, and hue)}; 2. {Video processing tools: Video encoding tools, video color processing tools}.
[0098] Furthermore, to more intuitively display the correlation between the retrieved results and the video processing task instructions, it can be set to sort the retrieved data in descending order of relevance.
[0099] Video processing tool management module, used to register and uniformly manage video processing tools, and define in detail the relevant information of the video processing tool set.
[0100] The relevant information of the video processing tool set includes but is not limited to: the functional attributes of each video processing tool, the input / output interface specifications, call functions, call protocols, and performance metrics.
[0101] It can be understood that the specific video processing tools included in the video processing tool set are not limited in this embodiment of the present invention, and corresponding video processing tools can be integrated and registered according to actual needs.
[0102] Intelligent processing module, used to semantically parse the context of the prompt words using a large language model, generate a video processing task execution plan and execute it, and output the processed video file.
[0103] It should be noted that the specific large language model is not limited in this embodiment of the present invention, and it can be any large language model, including but not limited to Transformer, MoE, hybrid architectures, etc. The large language model has the ability to obtain knowledge such as basic concepts and terms in the field of video processing.
[0104] Specifically, the large language model accepts the integrated context of the prompt words, and performs in-depth semantic analysis and logical reasoning based on its powerful natural language understanding and generation capabilities. According to the video processing task instructions, it screens suitable tools from the video processing tool set and generates a detailed task execution plan including the call order, call method, parameter configuration, data flow, etc. of the video processing tools.
[0105] Exemplarily, the task execution plan includes but is not limited to the following content:
[0106] Task steps: List the specific step operations to complete the video processing requirements.
[0107] Use of video processing tools: Specify the video tools used in each step and their call methods.
[0108] Parameter settings: Parameter settings of the called video tool and related parameter settings of each model.
[0109] Data dependencies: Describe the data sources required for each step and the processing methods.
[0110] Expected outputs: Describe the results obtained after each step is completed.
[0111] Among them, the task execution plan can be one or more feasible solutions, listed in descending order of priority.
[0112] Exemplarily, the video processing tool calling method is shown as follows: {Tool 1 name}: {Tool 1 function description}, the calling method is {Tool 1 calling parameter or instruction}; {Tool 2 name}: {Tool 2 function description}, the calling method is {Tool 2 calling parameter or instruction};...; {Tool n name}: {Tool n function description}, the calling method is {Tool n calling parameter or instruction}.
[0113] For example, for the task instruction "Encode the video into a format suitable for publishing on a social media platform and highlight the color expressiveness of the video", the video processing tools to be called and the calling methods are: {Video encoding tool}: {Supports MP4 format conversion, can adjust resolution, frame rate, color correction parameters}, the calling method is {VideoTranscoder, parameters are passed in JSON format}; {Video color processing tool}: {Specifically for video color optimization}, the calling method is {VideoProcessor, parameters are passed in JSON format}.
[0114] Further, to improve security, the task execution plan can be presented to the user before executing the task execution plan, and the user is requested to authorize the execution.
[0115] Further, to ensure smooth coordination between video processing tools, during the task execution process, the tool call status and data transmission situation can be monitored. If an abnormality occurs, tools can be reselected or parameters can be adjusted.
[0116] Further, as Figure 4 shown, in another embodiment, after the intelligent processing module, there is also a data update module, which is used to perform video quality testing on the processed video file, form a test report, and re-update the test report to the vector database after quantization.
[0117] The test report includes but is not limited to: task instructions, parameter settings of video tools (such as resolution, frame rate, bit rate), video quality scores (such as PSNR, VMAF), video processing performance metrics (such as processing time, resource occupancy rate), user evaluations and other information.
[0118] The vectorization of the test report and re-updating the vector database includes: The test reports are classified according to the video processing task types and labeled with corresponding task type labels.
[0119] The video quality score is mapped to the correlation between the task execution plan of the current video processing task and the current task category and a correlation label is added.
[0120] The higher the video quality score, the higher the correlation, and the lower the video quality score, the lower the correlation; the video quality score is a comprehensive score of the processed video file and / or a single dimension score.
[0121] The content of the test report and the corresponding labels are vectorized and updated in the vector database.
[0122] The vectorized test report content is used as new private domain knowledge information data and compared with the historical data in the vector database to improve the ability of video processing tasks and form a complete closed loop from demand processing, task execution to data feedback optimization.
[0123] For example, if it is found that the current encoding task improves color expression while the encoding time is too long, through the reinforcement learning-based algorithm, it can be analyzed that the complexity of the color processing algorithm can be appropriately reduced to shorten the encoding time.
[0124] The term "and / or" in the embodiment of the present invention describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A video processing method based on a large language model, characterized in that, Including: Using natural language to input video processing task instructions; Semantically parsing the task instructions and vectorizing them into query vectors; Performing similarity search in the vector database according to the query vectors; Semantically associating the query vectors with the similarity search results, integrating them in combination with video processing toolset information according to a preset prompt template, and outputting prompt context; The video processing toolset is composed of various video processing tools and is uniformly managed, and details of the information related to the video processing tools are defined; Using a large language model to semantically parse the prompt context, generate and execute a video processing task execution plan, and output a processed video file.
2. The video processing method based on a large language model according to claim 1, wherein The vector database stores pre-vectorized private domain knowledge information data related to the field of video processing, including: video processing operation guides, video processing tool parameter descriptions, video pre-processing tool descriptions, historical test data, industry technical specifications, video processing experience summaries.
3. The video processing method based on a large language model according to claim 1, wherein The video processing toolset information includes: the functional attributes of each video processing tool, input / output interface specifications, call functions, call protocols, and performance indicators.
4. The video processing method based on a large language model according to claim 1, wherein After outputting the processed video file, it also includes: performing video quality testing on the processed video file, forming a test report, and re-updating the test report to the vector database after vectorization.
5. The video processing method based on a large language model according to claim 4, wherein The re-updating the test report to the vector database after vectorization includes: Classifying the test report according to the video processing task type and attaching corresponding task type tags; Mapping the video quality score to the relevance between the task execution plan of the current video processing task and the current task category to which it belongs and attaching a relevance tag; Among them, the higher the video quality score, the higher the relevance, and the lower the video quality score, the lower the relevance; the video quality score is the comprehensive score and / or single-dimensional score of the processed video file; Vectorizing the content and corresponding tags of the test report and updating them to the vector database.
6. A video processing system based on a large language model, characterized in that, Including: A user interaction module for using natural language to input video processing task instructions; A vectorization module for semantically parsing the task instructions and vectorizing them into query vectors; A data retrieval module for performing similarity search in the vector database according to the query vectors; A prompt generation module for semantically associating the query vectors with the similarity search results, integrating them in combination with video processing toolset information according to a preset prompt template, and outputting prompt context; A video processing tool management module for registering and uniformly managing video processing tools, and defining details of the information related to the video processing tools; An intelligent processing module for using a large language model to semantically parse the prompt context, generate and execute a video processing task execution plan, and output a processed video file.
7. The video processing system based on a large language model according to claim 6, wherein The vector database stores pre-vectorized private domain knowledge information data related to the field of video processing, including: video processing operation guides, video processing tool parameter descriptions, video pre-processing tool descriptions, historical test data, industry technical specifications, video processing experience summaries.
8. The video processing system based on a large language model according to claim 6, wherein The information of the video processing tool set includes: the functional attributes of each video processing tool, the input / output interface specifications, the calling functions, the calling protocols, and the performance metrics.
9. The video processing system based on a large language model according to claim 6, wherein After the intelligent processing module, there is also a data update module, which is used to perform video quality tests on the processed video files, generate test reports, and re-update the vectorized test reports into the vector database.
10. The video processing system based on a large language model according to claim 9, wherein, The re-updating the vectorized test reports into the vector database includes: Classifying the test reports according to the video processing task types and attaching corresponding task type labels; Mapping the video quality scores to the relevance between the task execution plan of the current video processing task and the current task category to which it belongs and attaching relevance labels; Among them, a high video quality score indicates a high relevance, and a low video quality score indicates a low relevance; the video quality score is the comprehensive score and / or the single-dimensional score of the processed video file; Vectorizing the content of the test report and the corresponding labels and updating them into the vector database.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN118314229A
Creation method and system based on multi-modal data processing and dynamic task planning
CN119515303A
Langchain-based intelligent interactive Agent system and implementation method
CN119646123A
Short video automatic generation method and device based on new media content creation model
CN119865672A
Data processing method based on big language model dialogue robot and related equipment
CN120104748A
Cited By
Training and pushing integrated video analysis method and system based on multi-modal large model
CN121074764A
Multimodal large model-based training and inference integrated video analysis method and system
CN121074764B