Video generation method based on multi-Agent collaborative technology
Through the video generation method based on multi-Agent collaborative technology, the problem of low efficiency of traditional video generation methods is solved, and efficient video generation is achieved to meet the creation needs of complex scenarios.
Patent Information
- Application Number
- CN202510175163.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional video generation methods mainly rely on manual editing and special effects production, resulting in low video generation efficiency and difficulty in meeting the video creation needs of large-scale and complex scenes.
The video generation method based on multi-agent collaboration technology is adopted. By receiving video generation requests, the operation status information of the agent is obtained, the multi-agent collaboration index is calculated, the video collection, timing action detection and rendering agent is activated, the video generation task is decomposed and the video generation task is completed in a coordinated manner.
It improves video generation efficiency, can meet the video creation needs of large-scale and complex scenarios, and achieves the consistency and consistency of tasks.
Smart Images

Figure CN120034689A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video generation, and in particular to a video generation method based on multi-agent collaboration technology. Background Art
[0002] With the rapid development of artificial intelligence technology, the field of video generation is undergoing unprecedented changes. Traditional video generation methods mainly rely on manual editing and special effects production, which is not only time-consuming and labor-intensive, but also difficult to cope with the needs of large-scale and complex scene video creation. In recent years, video generation methods based on multi-agent collaborative technology have emerged, bringing new opportunities and challenges to the field of video generation.
[0003] Multi-agent collaboration technology is a distributed artificial intelligence technology that completes complex tasks through the collaboration of multiple agents. Each agent has a certain degree of autonomy and intelligence, can make decisions based on its own state and environmental information, and achieve information sharing and collaborative work through interaction with other agents. Multi-agent collaboration technology is highly flexible and scalable, and can cope with complex and changing environments and task requirements.
[0004] Therefore, traditional video generation methods mainly rely on manual editing and special effects production, which is not only time-consuming and labor-intensive, but also difficult to cope with the difficulties of video creation needs of large-scale and complex scenes. There is an urgent need for a video generation method based on multi-agent collaborative technology to solve the above problems. Summary of the invention
[0005] The purpose of the present invention is to provide a video generation method based on multi-agent collaborative technology: to solve the technical problem that the traditional video generation method mainly relies on manual editing and special effects production, resulting in low video generation efficiency and difficulty in meeting the video creation needs of large-scale and complex scenes.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A video generation method based on multi-agent collaborative technology, the method comprising:
[0008] Receive a request for generating a video, obtain the current running status information of the Agent based on the request for generating the video, and calculate the multi-agent coordination index XTZ based on the current running status information of the Agent;
[0009] Based on the multi-agent synergy index XTZ, determine whether the conditions for activating the Agent are met. If so, activate the Agent, where the Agent includes the video collection Agent, the video temporal action detection Agent, and the video rendering Agent;
[0010] The video collection agent obtains the video data to be processed based on the request to generate the video;
[0011] The video temporal action detection agent extracts the feature sequence at the video clip level based on the video data to be processed, inputs the obtained feature sequence into the score generation module to obtain the action score at the video clip level, and generates the final candidate proposal set based on the action score. Based on the final candidate proposal set, the preprocessed video containing the specified action is obtained;
[0012] The video rendering agent renders the pre-processed video and outputs a video file.
[0013] Furthermore, the calculation of the multi-agent synergy index XTZ based on the current running status information of the Agent includes the following process:
[0014] Based on the current running status information of the Agent, the CPU occupancy rate of each Agent is obtained, and the number S of Agents whose CPU occupancy rate exceeds the preset occupancy rate is calculated, where the CPU occupancy rate of the Agent is the occupancy rate of the CPU resources during the running process of the Agent;
[0015] Based on the current running status information of the Agent, the current task completion rate W of the Agent is obtained. The current task completion rate of the Agent is the sum of the progress ratios of each Agent currently completing the specified task;
[0016] Based on the current running status information of the Agent, obtain the average response time T between each Agent;
[0017] Based on the current running status information of the Agent, the current throughput L of the Agent is obtained. The current throughput L of the Agent is the sum of the number of tasks processed by each Agent per unit time.
[0018] Substitute S, W, T and L into the multi-agent synergy index calculation formula to calculate the multi-agent synergy index XTZ. The calculation formula is as follows:
[0019]
[0020] Among them, the value of e is 2.72.
[0021] Furthermore, judging whether the conditions for activating an Agent are met based on the multi-agent synergy index XTZ specifically includes the following process:
[0022] Get the multi-agent collaboration index threshold XTZ min XTZ min is the preset parameter value;
[0023] Determine whether the multi-agent synergy index XTZ exceeds the multi-agent synergy index threshold XTZ min If so, the Agent activation requirements are met and the Agent is activated; if not, the Agent activation requirements are not met and the Agent is not activated.
[0024] Furthermore, extracting a feature sequence at the video segment level based on the video data to be processed specifically includes the following process:
[0025] The convolutional neural network BN-Inception is selected as the network for extracting feature sequences at the video clip level;
[0026] The sliding window method is used to sample the processed video data at a certain interval to obtain a series of video clips. Where K is the total number of video clips, the value of k is 1, 2...K, and each clip contains several adjacent video frames;
[0027] The static RGB image and optical flow stack are used as the input of the convolutional neural network BN-Inception to extract the appearance and action content of the video clip, where the RGB image is randomly sampled from the video clip;
[0028] The video features are forward calculated for the appearance and action content of the video clip to obtain the spatial and temporal feature sequence representation of the video clip.
[0029] Furthermore, the obtained feature sequence is input into the score generation module to obtain the action score at the video clip level, which specifically includes the following process:
[0030] The obtained feature sequence is input into the deconvolution layer for feature enhancement, where the step size of the deconvolution layer is 1;
[0031] Based on the attention mechanism, we focus on the key information of the video and obtain the action score prediction results at the video clip level based on the key information of the video:
[0032] The feature sequence after feature enhancement is recorded as a series of key-value pairs Key, Value. For each feature in the feature sequence after feature enhancement, the correlation between it and each Key value is calculated to obtain the weight coefficient of the corresponding Value value, and the weighted sum of each Value value is performed to obtain the final attention value;
[0033] The key information of the video corresponding to the final attention value is output to the fully connected layer for feature integration to obtain the action score prediction result at the video clip level.
[0034] Furthermore, the key information of the video corresponding to the final attention value is output to the fully connected layer for feature integration, and the action score prediction result at the video clip level is obtained. The specific process includes the following:
[0035] The fully connected layer contains neurons related to the prediction task. In the video clip level action score prediction, the number of neurons in the output layer corresponds to the number of action categories or a single score value. The output layer will generate the action score prediction result for each video clip based on the feature representation of the previous layer.
[0036] Furthermore, the final candidate proposal set is generated based on the action score, and the preprocessed video containing the specified action is obtained based on the final candidate proposal set, which specifically includes the following processes:
[0037] The video clips with action scores higher than the preset action scores are grouped into subsequences and stored in a set. An action clip in the set is selected as the start, and it is recursively extended by adding the continuous clips after the action clip. When the ratio of the continuous part of the low action probability clip to the length of the candidate region proposal formed at a certain moment exceeds a given tolerance threshold, the extension is stopped to obtain the final candidate proposal set.
[0038] The SoftNMS algorithm with Gaussian weighting function is used to remove the redundant proposal frames of the final candidate proposal set to obtain the time series proposals in the final candidate proposal set, and the preprocessed video containing the specified action is obtained based on the time series proposals:
[0039] Among them, Y i Represents the final score of the proposal box, y i represents the basic score of the pre-set proposal box, i represents the proposal box number in the final candidate proposal set, σ represents the Gaussian attenuation function parameter, the value of e is 2.72, -iou(i) 2 Represents the confidence score obtained by the IoU function;
[0040] The proposal frames whose final scores are lower than the preset final scores are deleted and formed into new time sequence proposals. The frame sequence is reassembled according to the time sequence to obtain the preprocessed video containing the specified action.
[0041] Furthermore, the video rendering agent renders the pre-processed video, and the output video file specifically includes the following process:
[0042] The video rendering agent includes a rendering module and an image redundancy removal module;
[0043] The rendering module extracts the depth information of the rendering result from the preprocessed video, and extracts the color information of the rendering result by using the frame buffer off-screen rendering method:
[0044] The rendering module renders a frame of the scene, and the result is drawn to the video rendering agent frame buffer;
[0045] Through the pixel reading function, the depth information is obtained and passed to the image redundancy removal module;
[0046] Create a new framebuffer object, bind it as the default framebuffer, create a new texture attachment, and bind the new texture attachment to the new framebuffer;
[0047] Create a new texture object, set the width and height parameters, bind the texture object to the new frame buffer, and keep its data in an uninitialized state;
[0048] Create a new stencil buffer object and bind it to the new framebuffer as its depth attachment and stencil attachment.
[0049] Draw a quad to the new framebuffer, with the size equal to the window size, so that it tiles to the entire window;
[0050] Copy the color buffer in the video rendering agent buffer to the texture object of the new frame buffer, and then pass the new frame buffer to the image de-redundancy module for image de-redundancy processing;
[0051] Bind the video rendering Agent frame buffer to the default frame buffer;
[0052] The image de-redundancy module performs image de-redundancy processing and outputs a video file.
[0053] Compared with the existing solutions, the present invention achieves the following beneficial effects:
[0054] The present invention designs different types of intelligent agents according to the needs of video generation, decomposes the video generation task into multiple subtasks according to the complexity of the task and the ability of the intelligent agent, and assigns them to corresponding intelligent agents, thereby further improving the efficiency of video generation.
[0055] The various intelligent agents of the present invention share information and collaborate with each other through a communication protocol to ensure the continuity and consistency of task execution and further improve the efficiency of video generation.
[0056] The present invention can meet the needs of video creation for large-scale and complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0058] Figure 1 It is a workflow diagram of a video generation method based on multi-agent collaborative technology according to an embodiment of the present invention;
[0059] Figure 2 is a workflow diagram of another video generation method based on multi-agent collaborative technology according to an embodiment of the present invention;
[0060] Figure 3 This is a workflow diagram of another video generation method based on multi-agent collaborative technology according to an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] In addition, the described features, structures or characteristics may be combined in one or more example embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the example embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, steps, etc. may be adopted. In other cases, well-known structures, methods, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0063] Multi-agent collaboration technology is a technology that uses collaboration and communication between multiple agents in a distributed system to jointly complete complex tasks. In the field of video generation, multi-agent collaboration technology can be applied to multiple links such as video content creation, editing, and rendering to improve the efficiency and quality of video generation.
[0064] This embodiment provides a video generation method based on multi-agent collaboration technology. Figure 1 is a workflow diagram of a video generation method based on multi-agent collaborative technology according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:
[0065] Step S101: receiving a request for generating a video, and obtaining the current running status information of the Agent based on the request for generating the video;
[0066] Step S102: Calculate the multi-agent coordination index XTZ based on the current running status information of the Agent;
[0067] Step S103: Based on the multi-agent synergy index XTZ, determine whether the conditions for activating the Agent are met. If yes, proceed to step S104; if no, proceed to step S108;
[0068] Step S104: Activate Agent;
[0069] Step S105: the video collection agent obtains the video data to be processed based on the request for generating the video;
[0070] Step S106: The video temporal action detection agent extracts a feature sequence at the video segment level based on the video data to be processed, inputs the obtained feature sequence into the score generation module to obtain the action score at the video segment level, and generates a final candidate proposal set based on the action score, and obtains a preprocessed video containing the specified action based on the final candidate proposal set;
[0071] Step S107: the video rendering agent renders the pre-processed video and outputs a video file;
[0072] Step S108: Refuse to activate the Agent.
[0073] In summary, the present invention receives a request for generating a video, obtains the current running status information of the Agent based on the request for generating a video, and calculates the multi-agent synergy index XTZ based on the current running status information of the Agent; determines whether the conditions required for activating the Agent are met based on the multi-agent synergy index XTZ, and if so, activates the Agent, and the video collection Agent obtains the video data to be processed based on the request for generating a video; the video time-series action detection Agent extracts the feature sequence at the video clip level based on the video data to be processed, inputs the obtained feature sequence into the score generation module to obtain the action score at the video clip level, and generates the final candidate proposal set based on the action score, and obtains the pre-processed video containing the specified action based on the final candidate proposal set; the video rendering Agent renders the pre-processed video and outputs the video file. The present invention can improve the efficiency of video generation to meet the needs of video creation for large-scale and complex scenes.
[0074] In some embodiments, in step S102, Figure 2 FIG. 1 is another workflow diagram of a video generation method based on multi-agent collaborative technology according to an embodiment of the present invention. Figure 2 As shown, the calculation of the multi-agent synergy index XTZ based on the current running status information of the Agent includes the following steps:
[0075] Step S201: based on the current running status information of the Agent, the CPU occupancy rate of each Agent is obtained, and the number S of Agents whose CPU occupancy rate exceeds the preset occupancy rate is calculated, wherein the CPU occupancy rate of the Agent is the occupancy rate of the CPU resources during the running process of the Agent;
[0076] Step S202: obtaining the current task completion rate W of the Agent based on the current running status information of the Agent, where the current task completion rate of the Agent is the sum of the progress ratios of each Agent currently completing the specified task;
[0077] Step S203: obtaining the average response time T between each Agent based on the current running status information of the Agent;
[0078] Step S204: obtaining the current throughput L of the Agent based on the current running status information of the Agent, where the current throughput L of the Agent is the sum of the number of tasks processed by each Agent per unit time;
[0079] Step S205: Substitute S, W, T and L into the multi-agent synergy index calculation formula to calculate the multi-agent synergy index XTZ. The calculation formula is as follows:
[0080]
[0081] Among them, the value of e is 2.72.
[0082] In some embodiments, judging whether the conditions for activating an Agent are met based on the multi-agent synergy index XTZ specifically includes the following process:
[0083] Get the multi-agent collaboration index threshold XTZ min XTZ min is the preset parameter value;
[0084] Determine whether the multi-agent synergy index XTZ exceeds the multi-agent synergy index threshold XTZ min If so, the Agent activation requirements are met and the Agent is activated; if not, the Agent activation requirements are not met and the Agent is not activated.
[0085] In some embodiments, extracting a feature sequence at the video segment level based on the video data to be processed specifically includes the following process:
[0086] The convolutional neural network BN-Inception is selected as the network for extracting feature sequences at the video clip level;
[0087] The sliding window method is used to sample the processed video data at a certain interval to obtain a series of video clips. Where K is the total number of video clips, the value of k is 1, 2...K, and each clip contains several adjacent video frames;
[0088] The static RGB image and optical flow stack are used as the input of the convolutional neural network BN-Inception to extract the appearance and action content of the video clip, where the RGB image is randomly sampled from the video clip;
[0089] The video features are forward calculated for the appearance and action content of the video clip to obtain the spatial and temporal feature sequence representation of the video clip.
[0090] In some embodiments, Figure 3 FIG. 1 is another workflow diagram of a video generation method based on multi-agent collaborative technology according to an embodiment of the present invention. Figure 3 As shown, inputting the obtained feature sequence into the score generation module to obtain the action score at the video clip level specifically includes the following steps:
[0091] Step S301: input the obtained feature sequence into the deconvolution layer for feature enhancement, wherein the step length of the deconvolution layer is 1;
[0092] Step S302: focusing on the key information of the video based on the attention mechanism, and obtaining the action score prediction result at the video clip level based on the key information of the video:
[0093] Step S303: The feature sequence after feature enhancement is recorded as a series of key-value pairs Key, Value. For each feature in the feature sequence after feature enhancement, the correlation between it and each Key value is calculated to obtain the weight coefficient of the corresponding Value value, and the weighted sum of each Value value is performed to obtain the final attention value;
[0094] Step S304: Output the key information of the video corresponding to the final attention value to the fully connected layer for feature integration to obtain the action score prediction result at the video clip level.
[0095] In some embodiments, the key information of the video corresponding to the final attention value is output to the fully connected layer for feature integration, and the action score prediction result at the video clip level is obtained, which specifically includes the following process:
[0096] The fully connected layer contains neurons related to the prediction task. In the video clip level action score prediction, the number of neurons in the output layer corresponds to the number of action categories or a single score value. The output layer will generate the action score prediction result for each video clip based on the feature representation of the previous layer.
[0097] Furthermore, the final candidate proposal set is generated based on the action score, and the preprocessed video containing the specified action is obtained based on the final candidate proposal set, which specifically includes the following processes:
[0098] The video clips with action scores higher than the preset action scores are grouped into subsequences and stored in a set. An action clip in the set is selected as the start, and it is recursively extended by adding the continuous clips after the action clip. When the ratio of the continuous part of the low action probability clip to the length of the candidate region proposal formed at a certain moment exceeds a given tolerance threshold, the extension is stopped to obtain the final candidate proposal set.
[0099] The SoftNMS algorithm with Gaussian weighting function is used to remove the redundant proposal frames of the final candidate proposal set to obtain the time series proposals in the final candidate proposal set, and the preprocessed video containing the specified action is obtained based on the time series proposals:
[0100] Among them, Y i Represents the final score of the proposal box, y i represents the basic score of the pre-set proposal box, i represents the proposal box sequence number in the final candidate proposal set, σ represents the Gaussian attenuation function parameter, the value of e is 2.72, -iou(i) 2 Represents the confidence score obtained by the IoU function;
[0101] The proposal frames whose final scores are lower than the preset final scores are deleted and formed into new time sequence proposals. The frame sequence is reassembled according to the time sequence to obtain the preprocessed video containing the specified action.
[0102] In some embodiments, the video rendering agent renders the pre-processed video, and the output video file specifically includes the following process:
[0103] The video rendering agent includes a rendering module and an image redundancy removal module;
[0104] The rendering module extracts the depth information of the rendering result from the preprocessed video, and extracts the color information of the rendering result by using the frame buffer off-screen rendering method:
[0105] The rendering module renders a frame of the scene, and the result is drawn to the video rendering agent frame buffer;
[0106] Through the pixel reading function, the depth information is obtained and passed to the image redundancy removal module;
[0107] Create a new framebuffer object, bind it as the default framebuffer, create a new texture attachment, and bind the new texture attachment to the new framebuffer;
[0108] Create a new texture object, set the width and height parameters, bind the texture object to the new frame buffer, and keep its data in an uninitialized state;
[0109] Create a new stencil buffer object and bind it to the new framebuffer as its depth attachment and stencil attachment.
[0110] Draw a quad to the new framebuffer, with the size equal to the window size, so that it tiles to the entire window;
[0111] Copy the color buffer in the video rendering agent buffer to the texture object of the new frame buffer, and then pass the new frame buffer to the image de-redundancy module for image de-redundancy processing;
[0112] Bind the video rendering Agent frame buffer to the default frame buffer;
[0113] The image de-redundancy module performs image de-redundancy processing and outputs a video file.
[0114] It is worth mentioning that image de-redundancy processing is an important part of image processing, which aims to reduce the redundant information in image data, thereby reducing storage and transmission costs while maintaining the image quality as much as possible. Image de-redundancy processing is mainly based on the various types of redundancy existing in the image, including coding redundancy, inter-pixel redundancy and psychological visual redundancy.
[0115] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0116] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0117] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0118] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only some logical function divisions. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0119] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0120] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A video generation method based on multi-agent collaborative technology, characterized in that: Methods include: Receive a request for generating a video, obtain the current running status information of the Agent based on the request for generating the video, and calculate the multi-agent coordination index XTZ based on the current running status information of the Agent; Based on the multi-agent synergy index XTZ, determine whether the conditions for activating the Agent are met. If so, activate the Agent, where the Agent includes the video collection Agent, the video temporal action detection Agent, and the video rendering Agent; The video collection agent obtains the video data to be processed based on the request to generate the video; The video temporal action detection agent extracts the feature sequence at the video clip level based on the video data to be processed, inputs the obtained feature sequence into the score generation module to obtain the action score at the video clip level, and generates the final candidate proposal set based on the action score. Based on the final candidate proposal set, the preprocessed video containing the specified action is obtained; The video rendering agent renders the pre-processed video and outputs a video file.
2. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: The calculation of the multi-agent synergy index XTZ based on the current running status information of the Agent includes the following process: Based on the current running status information of the Agent, the CPU occupancy rate of each Agent is obtained, and the number S of Agents whose CPU occupancy rate exceeds the preset occupancy rate is calculated, where the CPU occupancy rate of the Agent is the occupancy rate of the CPU resources during the running process of the Agent; Based on the current running status information of the Agent, the current task completion rate W of the Agent is obtained. The current task completion rate of the Agent is the sum of the progress ratios of each Agent currently completing the specified task; Based on the current running status information of the Agent, obtain the average response time T between each Agent; Based on the current running status information of the Agent, the current throughput L of the Agent is obtained. The current throughput L of the Agent is the sum of the number of tasks processed by each Agent per unit time. Substitute S, W, T and L into the multi-agent synergy index calculation formula to calculate the multi-agent synergy index XTZ. The calculation formula is as follows: Among them, the value of e is 2.
72.
3. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: Based on the multi-agent synergy index XTZ, it is determined whether the conditions for activating the agent are met. The process includes: Get the multi-agent collaboration index threshold XTZ min XTZ min is the preset parameter value; Determine whether the multi-agent synergy index XTZ exceeds the multi-agent synergy index threshold XTZ min If so, the Agent activation requirements are met and the Agent is activated; if not, the Agent activation requirements are not met and the Agent is not activated.
4. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: Extracting the feature sequence at the video segment level based on the video data to be processed specifically includes the following processes: The convolutional neural network BN-Inception is selected as the network for extracting feature sequences at the video clip level; The sliding window method is used to sample the processed video data at a certain interval to obtain a series of video clips. Where K is the total number of video clips, the value of k is 1, 2...K, and each clip contains several adjacent video frames; The static RGB image and optical flow stack are used as the input of the convolutional neural network BN-Inception to extract the appearance and action content of the video clip, respectively, where the RGB image is randomly sampled from the video clip; The video features are forward calculated for the appearance and action content of the video clip to obtain the spatial and temporal feature sequence representation of the video clip.
5. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: The obtained feature sequence is input into the score generation module to obtain the action score at the video clip level, which specifically includes the following process: The obtained feature sequence is input into the deconvolution layer for feature enhancement, where the step size of the deconvolution layer is 1; Based on the attention mechanism, we focus on the key information of the video and obtain the action score prediction results at the video clip level based on the key information of the video: The feature sequence after feature enhancement is recorded as a series of key-value pairs Key, Value. For each feature in the feature sequence after feature enhancement, the correlation between it and each Key value is calculated to obtain the weight coefficient of the corresponding Value value, and the weighted sum of each Value value is performed to obtain the final attention value; The key information of the video corresponding to the final attention value is output to the fully connected layer for feature integration to obtain the action score prediction result at the video clip level.
6. The video generation method based on multi-agent collaborative technology according to claim 5 is characterized in that: The key information of the video corresponding to the final attention value is output to the fully connected layer for feature integration, and the action score prediction result at the video clip level is obtained, which specifically includes the following process: The fully connected layer contains neurons related to the prediction task. In the video clip level action score prediction, the number of neurons in the output layer corresponds to the number of action categories or a single score value. The output layer will generate the action score prediction result for each video clip based on the feature representation of the previous layer.
7. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: The final candidate proposal set is generated based on the action score, and the specific preprocessed video containing the specified action is obtained based on the final candidate proposal set. The process includes: The video clips with action scores higher than the preset action scores are grouped into subsequences and stored in a set. An action clip in the set is selected as the start, and it is recursively extended by adding the continuous clips after the action clip. When the ratio of the continuous part of the low action probability clip to the length of the candidate region proposal formed at a certain moment exceeds a given tolerance threshold, the extension is stopped to obtain the final candidate proposal set. The SoftNMS algorithm with Gaussian weighting function is used to remove the redundant proposal frames of the final candidate proposal set to obtain the time series proposals in the final candidate proposal set, and the preprocessed video containing the specified action is obtained based on the time series proposals: Among them, Y i Represents the final score of the proposal box, y i represents the basic score of the pre-set proposal box, i represents the proposal box number in the final candidate proposal set, σ represents the Gaussian attenuation function parameter, the value of e is 2.72, -iou(i) 2 Represents the confidence score obtained by the IoU function; The proposal frames whose final scores are lower than the preset final scores are deleted and formed into new time sequence proposals. The frame sequence is reassembled according to the time sequence to obtain the preprocessed video containing the specified action.
8. The video generation method based on multi-agent collaborative technology according to claim 1 is characterized in that: The video rendering agent renders the pre-processed video, and the output video file specifically includes the following process: The video rendering agent includes a rendering module and an image redundancy removal module; The rendering module extracts the depth information of the rendering result from the preprocessed video, and extracts the color information of the rendering result by using the frame buffer off-screen rendering method: The rendering module renders a frame of the scene, and the result is drawn to the video rendering agent frame buffer; Through the pixel reading function, the depth information is obtained and passed to the image redundancy removal module; Create a new framebuffer object, bind it as the default framebuffer, create a new texture attachment, and bind the new texture attachment to the new framebuffer; Create a new texture object, set the width and height parameters, bind the texture object to the new frame buffer, and keep its data in an uninitialized state; Create a new stencil buffer object and bind it to the new framebuffer as its depth attachment and stencil attachment. Draw a quad to the new framebuffer, with the size equal to the window size, so that it tiles to the entire window; Copy the color buffer in the video rendering agent buffer to the texture object of the new frame buffer, and then pass the new frame buffer to the image de-redundancy module for image de-redundancy processing; Bind the video rendering Agent frame buffer to the default frame buffer; The image de-redundancy module performs image de-redundancy processing and outputs a video file.