Task deployment method and system based on artificial intelligence large model and medium
By deploying a task deployment method based on artificial intelligence large-scale models on edge devices, combining multi-threaded management and sliding window voting mechanisms, the problems of real-time and inefficiency on edge devices are solved, and efficient and accurate video analysis is achieved.
Patent Information
- Application Number
- CN202510363479.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When deploying AI big models on resource-constrained edge devices, the prior art has problems of real-time and inefficiency, limited feature extraction capabilities, and traditional algorithms are prone to gradient disappearance or explosion when dealing with dependencies over a long span, resulting in poor recognition of certain behaviors or events that last longer.
By deploying task deployment methods based on artificial intelligence large models on edge devices, including video frame preprocessing, similar frame algorithm screening, multimodal model inference and sliding window statistics, combining multithreading and process management, optimizing resource utilization, and introducing a sliding window voting mechanism to integrate multi-frame inference results.
It improves the real-time and accuracy of video analysis, can efficiently process complex video content on edge devices with resource-constrained resources, reduces misjudgment of single-frame inference results, and improves the ability to identify scenes with inconspicuous features or complex backgrounds.
Smart Images

Figure CN120279461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a task deployment method, system and medium based on an artificial intelligence large model. Background Art
[0002] With the rapid development of artificial intelligence technology, AI video analysis has been widely applied in many fields, such as intelligent security, behavior monitoring, abnormal event warning, etc. The current technology mainly uses deep learning algorithms, combined with architectures such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to achieve spatio-temporal feature extraction and analysis of video data. These technologies can identify targets, behaviors and events in videos, providing analysis and recognition functions for complex scenarios.
[0003] However, the feature extraction ability of the existing technology is limited. For example, when a convolutional neural network (CNN) processes image data, it mainly relies on the extraction of local features, and has a weak ability to capture global structural information. In addition, traditional RNNs are prone to problems such as gradient vanishing or gradient explosion when dealing with long-term dependencies, which may lead to poor recognition effects for some behaviors or events with long durations. When deploying an artificial intelligence large model on resource-constrained edge devices, there will also be a problem of low inference efficiency, which not only increases the demand for hardware resources, but also deteriorates the real-time performance. Summary of the Invention
[0004] The present invention aims to solve the problems of low real-time performance and efficiency of task deployment in the above-mentioned existing technology, and limited feature extraction ability, and provides a task deployment method, system and medium based on an artificial intelligence large model.
[0005] The present invention provides a task deployment method based on an artificial intelligence large model, including the following steps:
[0006] Obtain video data through a preset video source; wherein, the video data includes a plurality of video frames;
[0007] Preprocess the video frames and store them in a preset first queue;
[0008] Read the video frames in the first queue, and screen the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames;
[0009] Store the target video frames in a preset second queue;
[0010] Call a preset multimodal model to infer the target video frames and obtain an inference result;
[0011] Store the inference result in a preset sliding window; wherein, the size of the sliding window corresponds to the number of inference times of the inference result;
[0012] Count the inference result with the highest proportion to obtain a real-time classification result and output it.
[0013] Further, in the step of preprocessing the video frame and storing it in a preset first queue, it includes:
[0014] Perform image correction and formatting on the video frame.
[0015] Further, in the step of reading the video frame from the first queue and screening the video frames that meet the preset conditions through a preset similar frame algorithm to obtain the target video frame, it includes:
[0016] Calculate the similarity between the video frames;
[0017] Screen out the video frames with a similarity lower than the preset similarity threshold.
[0018] Further, after the step of storing the target video frame in a preset second queue, it includes:
[0019] Judge whether the target video frame exists in the second queue;
[0020] If so, judge whether the second queue is full;
[0021] If not, increase the similarity threshold.
[0022] Further, in the step of if so, judge whether the second queue is full, it includes:
[0023] If so, lower the similarity threshold and replace the video frame with the highest similarity;
[0024] If not, store the target video frame in the second queue.
[0025] Further, in the step of calculating the similarity between the video frames, it includes:
[0026] Extract the feature information of the video frame and generate a corresponding sequence;
[0027] Compare the sequences between multiple video frames to obtain the similarity.
[0028] Further, after the step of screening out the video frames with a similarity lower than the preset similarity threshold, it further includes:
[0029] Analyze the historical state of the video frame and the similarity threshold to obtain an analysis result;
[0030] Adjust the similarity threshold according to the analysis result.
[0031] Further, in the step of calling a preset multimodal model to infer the target video frame and obtaining an inference result, it includes:
[0032] Extract the local detail information and global context information of the target video frame.
[0033] Further, before the step of storing the target video frame into a preset second queue, it includes:
[0034] Perform abnormal frame detection on the target video frame and filter out the abnormal frames.
[0035] The present invention also provides a task deployment system based on an artificial intelligence large model, including an edge device, a first thread, a second thread, and an inference process; the first thread and the second thread are respectively arranged on the edge device; the first thread includes a data acquisition module, a data storage module, and a preprocessing module, and the data acquisition module is used to obtain video data through a preset video source; the data storage module is used to store video frames into a preset first queue; the preprocessing module is used to preprocess the video frames; the second thread includes a retrieval module and a similarity detection module; the retrieval module is used to read the video frames in the first queue; the similarity detection module is used to screen video frames that meet preset conditions through a preset similar frame algorithm to obtain target video frames; the inference process includes an inference module, a sliding window storage module, and a statistics module, the inference module is used to call a preset multimodal model to infer the target video frame and obtain an inference result; the sliding window storage module is used to store the inference result into a preset sliding window; the statistics module is used to count the inference result with the highest proportion to obtain a classification result and output it.
[0036] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor executes the computer program to implement the steps in any one of the above methods.
[0037] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in any one of the above methods are implemented.
[0038] The present invention provides a task deployment method, system, and medium based on an artificial intelligence large model, having the following beneficial effects:
[0039] Calculating the read video frames to filter out the target video frames that meet the conditions can improve the efficiency of subsequent processing. By using a multimodal large model to infer the target video frames, it is possible to simultaneously extract local details and global context information, enhance the recognition ability for scenes with unclear features or complex backgrounds, be more efficient and accurate when processing complex video content, and better handle diverse behavior analysis and event detection tasks.
[0040] This application deploys data collection, preprocessing, and similarity detection tasks to edge devices close to the data source, which can reduce data transmission latency, improve real-time performance and response speed. By reasonably allocating tasks to different threads and processes, the computing resources of edge devices can be fully utilized, the resource utilization efficiency can be optimized, and a voting mechanism based on a sliding window is introduced to integrate the inference results of multiple frames and obtain the final real-time classification result. It can not only reduce misjudgments caused by single-frame inference results but also achieve high detection accuracy with low computational overhead, enabling efficient and real-time video analysis on resource-constrained edge devices. Brief Description of the Drawings
[0041] Figure 1 It is a schematic diagram of the method steps of a task deployment method based on an artificial intelligence large model in the present invention;
[0042] Figure 2 It is a block diagram of the structure of a task deployment system based on an artificial intelligence large model in the present invention;
[0043] Figure 3 It is a block diagram of the structure of a computer device of the present invention;
[0044] Figure 4 It is a schematic diagram of the working process of an embodiment of a task deployment system based on an artificial intelligence large model of the present invention.
[0045] Marking Explanation: Edge device 1, First thread 2, Second thread 3, Inference process 4, Data collection module 201, Data storage module 202, Preprocessing module 203, Retrieval module 301, Similarity detection module 302, Inference module 401, Sliding window storage module 402, Statistical module 403. Detailed Description of the Preferred Embodiments
[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0048] Refer to the attached Figure 1 , which is a task deployment method based on an artificial intelligence large model in an embodiment of the present invention, including:
[0049] S1. Obtain video data through a preset video source; wherein, the video data includes multiple video frames;
[0050] S2. Preprocess the video frames and store them in a preset first queue;
[0051] S3. Read the video frames in the first queue and screen the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames;
[0052] S4. Store the target video frames in a preset second queue;
[0053] S5. Invoke a preset multi-modal model to infer the target video frames and obtain an inference result;
[0054] S6. Store the inference result in a preset sliding window; wherein, the size of the sliding window corresponds to the number of inference times of the inference result;
[0055] S7. Statistically analyze the inference result with the highest proportion to obtain a real-time classification result and output it.
[0056] In the above steps, video data is first obtained through a preset video source. Here, the video source is a camera. In a specific embodiment, a video is captured by the camera to obtain video data. Further, the video data includes multiple video frames. Then, the video frames are preprocessed and stored in a preset first queue; wherein, the preprocessing includes image correction and formatting. Storing the preprocessed video frames in the first queue facilitates subsequent processing. Then, the video frames in the first queue are read and screened through a similar frame algorithm to screen out the video frames that meet the preset conditions, obtaining target video frames, and the target video frames are stored in a preset second queue. Screening the video frames can remove duplicate or highly similar frames, reduce redundant data, and improve the efficiency of subsequent processing. When there are target video frames stored in the second queue, a preset multimodal large model is called to perform inference on the video frames; wherein, the modality is a type, such as video, text, image, audio; the multimodal large model can be an existing artificial intelligence model. By integrating data of multiple modalities, it can simultaneously extract local details and global information of the target video frames, obtaining an inference result, and can understand and analyze complex scenarios more comprehensively. Among them, the inference result can be the category to which the video frame belongs, such as "competition", "news", or generate subtitles. Then, the inference result is saved in a sliding window, and the result with the highest proportion is counted in the sliding window and output as the final real-time classification result, which can solve the problem that the inference result of a single-frame picture is prone to misjudgment. By reasonably allocating tasks to different threads and processes, the computing resources of the edge device 1 can be fully utilized, optimizing the resource utilization efficiency, and a voting mechanism based on a sliding window is introduced to synthesize the inference results of multiple frames and obtain the final detection conclusion. It can achieve efficient and real-time video analysis on the resource-constrained edge device 1.
[0057] In one embodiment, in the step of preprocessing the video frames and storing them in a preset first queue, it includes:
[0058] Perform image correction and formatting on the video frames.
[0059] In this embodiment, the preprocessing includes image correction and formatting. Preprocessing the video frames can repair the deformation problems in the images. By analyzing the parameters and distortion conditions of the camera, the bending or distortion of the picture caused by the lens is adjusted back, which can make the image more natural. Among them, the correction methods include rotation correction, flipping correction, scaling correction, and translation correction. In a specific embodiment, rotation correction can rotate the image and correct the tilt of the image. In another specific embodiment, flipping correction can flip the image and correct the mirror image of the image. Repairing the distortion or distortion in the video frames to restore the image to a state closer to the real scene significantly improves the quality of the video frames and makes subsequent analysis and inference more accurate.
[0060] In one embodiment, in the step of reading the video frames of the first queue and screening the video frames that meet the preset conditions through a preset similar frame algorithm to obtain the target video frames, it includes:
[0061] Calculate the similarity between video frames;
[0062] Screen out the video frames whose similarity is lower than the preset similarity threshold.
[0063] In this embodiment, read the video frames from the first queue, calculate the similarity between the video frames through the similar frame algorithm, screen out the video frames whose similarity is lower than the preset similarity threshold, and store these video frames in queue B. It can remove duplicate or highly similar video frames, reduce redundant data, and improve the efficiency of subsequent processing.
[0064] Specifically, the similar frame algorithm is the hash algorithm, which can extract the feature information of the video frames and generate the corresponding hash value sequence. The hash value sequence can represent the content features of each frame. Then, by comparing the hash value sequences, the similarity between two video frames can be quickly judged, so as to identify duplicate or highly similar video frames, which can significantly improve the screening efficiency and accuracy of video frames.
[0065] In one embodiment, after the step of storing the target video frames in the preset second queue, it includes:
[0066] Judge whether there are target video frames in the second queue;
[0067] If so, judge whether the second queue is full;
[0068] If not, increase the similarity threshold.
[0069] In this embodiment, in order to control the inflow speed of video frames and optimize queue management, the present application dynamically adjusts the similarity threshold to optimize the screening and storage efficiency of video frames.
[0070] More specifically, after storing the target video frames in the second queue, judge whether there are target video frames in the second queue. If there are target video frames in the second queue, then judge whether the second queue is full. If there are no target video frames in the second queue, increase the similarity threshold to allow more video frames to enter the queue. By real-time monitoring the storage state of video frames, such as the queue is empty, the queue is full or the queue is not full, and then dynamically adjusting the threshold, the inflow speed of video frames can be accurately controlled.
[0071] In one embodiment, if so, in the step of judging whether the second queue is full, it includes:
[0072] If so, lower the similarity threshold and replace the video frame with the highest similarity;
[0073] Otherwise, store the target video frame in the second queue.
[0074] In this embodiment, if the second queue is full, then reduce the similarity threshold and replace the video frame with the highest similarity, that is, replace the most similar video frame to reduce the number of video frames in the queue. If the second queue is not full, that is, in a state of being neither empty nor full, directly store the video frame in the second queue and keep the similarity threshold unchanged. By monitoring the storage state of video frames in real time and dynamically adjusting the similarity threshold, the inflow speed of video frames can be accurately controlled.
[0075] In one embodiment, in the step of calculating the similarity between video frames, it includes:
[0076] Extract the feature information of the video frame and generate a corresponding sequence;
[0077] Compare the sequences between multiple video frames to obtain the similarity.
[0078] In this embodiment, preferably, calculating the similarity between video frames through the hash algorithm can solve the efficiency and accuracy problems of key frame extraction and duplicate frame detection in video processing. Specifically, extract the feature information of the video frame and generate a corresponding hash value, and then compare the hash values between multiple video frames to obtain the similarity. Among them, the comparison between video frames can be two or multiple. By calculating and comparing the hash values, the similarity between video frames can be quickly judged, thereby improving the screening efficiency and accuracy of video frames.
[0079] It is worth mentioning that in addition to the hash algorithm mentioned in the embodiment, the similarity between video frames can also be calculated through other algorithms, such as SIFT (Scale-Invariant Feature Transform).
[0080] In one embodiment, after the step of screening out video frames with similarity lower than the preset similarity threshold, it further includes:
[0081] Analyze the historical state of the video frame and the similarity threshold to obtain an analysis result;
[0082] Adjust the similarity threshold according to the analysis result.
[0083] In this embodiment, in addition to monitoring the storage state of video frames and dynamically adjusting the similarity threshold, it also analyzes the historical state of video frame storage and the change of the similarity threshold to obtain an analysis result, and optimizes the adjustment of the similarity threshold according to the analysis result to further improve the accuracy of dynamic adjustment. In a specific embodiment, if the historical states stored in the second queue are all not full, then reduce the similarity threshold to store more video frames.
[0084] In one embodiment, in the step of calling a preset multimodal model to infer a target video frame and obtaining an inference result, it includes:
[0085] Extract the local detail information and global context information of the target video frame.
[0086] In this embodiment, during the process of inferring the target video frame, the local details and global context information of the target video frame are extracted to overcome the limitations of strong dependence on local features and difficulty in capturing global information. Specifically, the local detail information can be feature events. In a specific embodiment, key events in the video are detected, such as emergencies, specific actions, or scene changes. In another specific embodiment, the behavior patterns of people or objects in the video are analyzed to identify behavior sequences or abnormal behaviors.
[0087] More specifically, the multimodal model can be an existing large multimodal model, which can simultaneously extract the local details and global context information of the target video frame, improving the recognition ability for scenes with unclear features or complex backgrounds.
[0088] Specifically, the large multimodal model can be an input classification head, such as an MLP (Multilayer Perceptron) or a Transformer (deep learning model) for classification inference; when it is single-label classification, such as action recognition or sentiment analysis, it uses Softmax, the normalized exponential function, and cross-entropy loss. When it is multi-label classification, such as object detection or event classification, it uses Sigmoid (S-shaped function) and BCE Loss (binary cross-entropy loss function).
[0089] In one embodiment, in the step of storing the inference result in a preset sliding window, it includes:
[0090] The size of the sliding window corresponds to the number of inference times of the inference result.
[0091] In this embodiment, specifically, the sliding window is "Sliding window", and the size of the sliding window corresponds to the number of receivable data. The sender determines the number of bytes of data to be sent according to the size of the sliding window. When the second queue is not empty, that is, when the target video frame is stored, the multi-modal large model is called to perform inference on the video frame, and the inference result is stored in the sliding window. The size of the sliding window corresponds to the number of inference times of the inference result. In a specific embodiment, the size of the sliding window is X, and the sliding window is used to store the inference results of the most recent X inferences. In another specific embodiment, the sliding window saves the inference results within a period of time, and the inference result with the largest proportion is selected as the real-time classification result through voting. It can reduce the misjudgment of single-frame inference results, and at the same time achieve high detection accuracy with low computational overhead, and can be applied to resource-constrained edge devices 1.
[0092] In one embodiment, before the step of storing the target video frame into the preset second queue, it includes:
[0093] Perform abnormal frame detection on the target video frame and screen out the abnormal frames.
[0094] In this embodiment, performing abnormal frame detection on the target video frame can identify damaged or misclassified video frames and improve the accuracy of subsequent inferences. In a specific embodiment, the abnormal frame detection includes brightness and color abnormal detection, such as the video frame being too bright or too dark, black screen or white screen caused by the failure of the acquisition device; at this time, by calculating the brightness mean value and color histogram of the target video frame, it can be judged whether it deviates from the normal range. More specifically, the calculation can be performed through opencv (cross-platform computer vision library). In another specific embodiment, the abnormal frame detection includes motion blur detection, such as device vibration or too fast movement of the shooting target. Specifically, the Laplace transform of the target video frame is calculated to obtain the image edge sharpness. When the calculated value is lower than the preset threshold, this target video frame is determined to be a blurred frame and screened out to improve the accuracy of the inference.
[0095] In summary, in specific implementation, video data is obtained through a preset video source; wherein, the video data includes multiple video frames; then the video frames are preprocessed and stored in a preset first queue; wherein, the preprocessing includes image correction and formatting of the video frames and storing them in the first queue to ensure the accuracy and consistency of the data. Then the video frames in the first queue are read, and the video frames meeting the preset conditions are screened through a preset similar frame algorithm to obtain target video frames; then the similarity between the video frames is calculated, and the video frames with a similarity lower than the preset similarity threshold are screened out. Among them, the calculation process extracts the feature information of the video frames and generates corresponding hash values. Then the hash values between multiple video frames are compared to obtain the similarity. Then the target video frames are stored in a preset second queue; then it is determined whether there are target video frames in the second queue; if so, it is determined whether the second queue is full; if not, the similarity threshold is increased. When determining whether the second queue is full, if it is full, the similarity threshold is decreased. If it is not full, the target video frames are stored in the second queue. After the step of screening out the video frames with a similarity lower than the preset similarity threshold, the historical state of the video frames and the similarity threshold are analyzed to obtain an analysis result. Then the similarity threshold is adaptively adjusted according to the analysis result. Then a preset multi-modal model is called to infer the target video frames and an inference result is obtained; during this period, the local detail information and global context information of the target video frames are extracted. Then the inference result is stored in a preset sliding window; wherein, the size of the sliding window corresponds to the number of inference times of the inference result. Finally, the inference result with the highest proportion is counted to obtain a real-time classification result and output.
[0096] Refer to the appendix Figure 2 , a task deployment system based on an artificial intelligence large model, includes an edge device 1, a first thread 2, a second thread 3 and an inference process 4; the first thread 2 and the second thread 3 are respectively arranged on the edge device 1; the first thread 2 includes a data acquisition module 201, a data storage module 202 and a preprocessing module 203, and the data acquisition module 201 is used to obtain video data through a preset video source; the data storage module 202 is used to store the video frames in a preset first queue; the preprocessing module 203 is used to preprocess the video frames; the second thread 3 includes a retrieval module 301 and a similarity detection module 302; the retrieval module 301 is used to read the video frames in the first queue; the similarity detection module 302 is used to screen the video frames meeting the preset conditions through a preset similar frame algorithm to obtain target video frames; the inference process 4 includes an inference module 401, a sliding window storage module 402 and a statistics module 403, the inference module 401 is used to call a preset multi-modal model to infer the target video frames and obtain an inference result; the sliding window storage module 402 is used to store the inference result in a preset sliding window; the statistics module 403 is used to count the inference result with the highest proportion to obtain a classification result and output.
[0097] In this embodiment, the present application deploys the first thread 2 and the second thread 3 on the edge device 1, which can reduce the dependence on the server, reduce data transmission latency, improve the real-time performance and response speed of the system, and at the same time enhance the privacy and security of data, and can be applied to video analysis scenarios with high real-time requirements.
[0098] Specifically, as Figure 4 shown, the system includes two threads and one process, namely thread A, thread B, and process A. The threads and the process work together to achieve the acquisition, processing, and inference of video data. Thread A is mainly responsible for the acquisition of video data and saves the acquired video data to queue A. Further, thread A obtains video data in real time through a video source, extracts each frame of video data, and stores the acquired video frames in queue A for subsequent processing. When thread A stores video frames in queue A, a synchronization mechanism can be set, that is, multiple execution units coordinate and cooperate to complete a task together to ensure that the data is correctly stored and avoid data conflicts with other threads.
[0099] Thread B is responsible for taking out video frames from queue A and comparing their similarity with the last video frame in queue B. When queue B is empty, the current video frame is directly stored in queue B, and the similarity threshold is increased to allow more video frames to enter the queue. When queue B is full: replace the most similar video frame in queue B and decrease the similarity threshold to reduce the number of video frames in queue B and ensure that the video frames stored in the queue are non-repetitive. When queue B is neither empty nor full, the current video frame is directly stored in queue B and the similarity threshold remains unchanged.
[0100] The main function of process A is to take out video frames from queue B and perform inference through a multimodal large model. The inference results are stored in a sliding window, and finally the classification result of the entire video is determined by counting the inference result with the highest proportion. Specifically, take out the video frame from queue B as the input data for inference; use the multimodal large model to perform inference on the video frame to generate a classification result. Store the inference result in the sliding window for subsequent voting processing. Voting processing means counting the inference result with the highest proportion from the sliding window and using it as the final classification result of the entire video. It can effectively reduce the misjudgment of single-frame inference results and improve the overall classification accuracy. Through the multi-threaded and multi-process method, efficient data processing and inference can be achieved.
[0101] Further, the threads and processes of the entire program are controlled to end through software interrupts. Among them, the software interrupt method is an interrupt method triggered by software rather than hardware. When the system receives an end signal, Thread A and Thread B will stop collecting and processing video frames, and Process A will also stop reasoning to ensure that the system can end its operation safely and smoothly.
[0102] Refer to the appendix Figure 3 , in an embodiment of the present application, a computer device is further provided. The computer device may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as templates, tables, and preset fields. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a task deployment method based on an artificial intelligence large model, including the following steps:
[0103] Obtain video data through a preset video source; among them, the video data includes multiple video frames;
[0104] Preprocess the video frames and store them in a preset first queue;
[0105] Read the video frames in the first queue and screen the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames;
[0106] Store the target video frames in a preset second queue;
[0107] Call a preset multi-modal model to reason about the target video frames and obtain a reasoning result;
[0108] Store the reasoning result in a preset sliding window; among them, the size of the sliding window corresponds to the number of reasoning times of the reasoning result;
[0109] Statistically analyze the reasoning result with the highest proportion to obtain a real-time classification result and output it.
[0110] In one embodiment, in the step of preprocessing the video frames and storing them in a preset first queue, it includes:
[0111] Perform image correction and formatting on the video frames.
[0112] In one embodiment, in the step of reading video frames of the first queue and screening out video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames, it includes:
[0113] Calculate the similarity between video frames;
[0114] Screen out video frames whose similarity is lower than the preset similarity threshold.
[0115] In one embodiment, after the step of storing the target video frames into a preset second queue, it includes:
[0116] Determine whether there are target video frames in the second queue;
[0117] If so, determine whether the second queue is full;
[0118] If not, increase the similarity threshold.
[0119] In one embodiment, in the step of determining whether the second queue is full if so, it includes:
[0120] If so, lower the similarity threshold and replace the video frame with the highest similarity;
[0121] If not, store the target video frames into the second queue.
[0122] In one embodiment, in the step of calculating the similarity between video frames, it includes:
[0123] Extract the feature information of the video frames and generate corresponding sequences;
[0124] Compare the sequences between multiple video frames to obtain the similarity.
[0125] In one embodiment, after the step of screening out video frames whose similarity is lower than the preset similarity threshold, it further includes:
[0126] Analyze the historical state of the video frames and the similarity threshold to obtain an analysis result;
[0127] Adjust the similarity threshold according to the analysis result.
[0128] In one embodiment, in the step of calling a preset multi-modal model to infer the target video frames and obtaining an inference result, it includes:
[0129] Extract the local detail information and global context information of the target video frames.
[0130] In one embodiment, in the step of storing the inference result into a preset sliding window, it includes:
[0131] The size of the sliding window corresponds to the number of inference times of the inference result.
[0132] Those skilled in the art can understand that Figure 3 the structure shown in
[0133] An embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a task deployment method based on an artificial intelligence large model, including the following steps:
[0134] Obtain video data through a preset video source; wherein, the video data includes multiple video frames;
[0135] Preprocess the video frames and store them in a preset first queue;
[0136] Read the video frames in the first queue, and filter the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames;
[0137] Store the target video frames in a preset second queue;
[0138] Call a preset multi-modal model to infer the target video frames and obtain an inference result;
[0139] Store the inference result in a preset sliding window; wherein, the size of the sliding window corresponds to the number of inference times of the inference result;
[0140] Count the inference result with the highest proportion to obtain a real-time classification result and output it.
[0141] In one embodiment, in the step of preprocessing the video frames and storing them in a preset first queue, it includes:
[0142] Perform image correction and formatting on the video frames.
[0143] In one embodiment, in the step of reading the video frames in the first queue and filtering the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames, it includes:
[0144] Calculate the similarity between video frames;
[0145] Filter out the video frames with a similarity lower than a preset similarity threshold.
[0146] In one embodiment, after the step of storing the target video frames in a preset second queue, it includes:
[0147] Judge whether there are target video frames in the second queue;
[0148] If so, judge whether the second queue is full;
[0149] If not, increase the similarity threshold.
[0150] In one embodiment, if so, in the step of determining whether the second queue is full, it includes:
[0151] If so, decrease the similarity threshold and replace the video frame with the highest similarity.
[0152] If not, store the target video frame in the second queue.
[0153] In one embodiment, in the step of calculating the similarity between video frames, it includes:
[0154] Extract the feature information of the video frame and generate a corresponding sequence.
[0155] Compare the sequences between multiple video frames to obtain the similarity.
[0156] In one embodiment, after the step of screening out the video frames with similarity lower than the preset similarity threshold, it further includes:
[0157] Analyze the historical state of the video frame and the similarity threshold to obtain an analysis result.
[0158] Adjust the similarity threshold according to the analysis result.
[0159] In one embodiment, in the step of invoking a preset multi-modal model to infer the target video frame and obtaining an inference result, it includes:
[0160] Extract the local detail information and global context information of the target video frame.
[0161] In one embodiment, in the step of storing the inference result in a preset sliding window, it includes:
[0162] The size of the sliding window corresponds to the number of inference times of the inference result.
[0163] In summary, a task deployment method, system, and medium based on an artificial intelligence large model provided in the embodiments of the present application calculate the read video frames to screen out target video frames that meet the conditions, which can improve the efficiency of subsequent processing. By using a multi-modal large model to infer the target video frames, local details and global context information can be extracted simultaneously, enhancing the recognition ability for scenes with unclear features or complex backgrounds, being more efficient and accurate in processing complex video content, and being able to better handle diverse behavior analysis and event detection tasks. The present application deploys data collection, preprocessing, and similarity detection tasks to edge device 1 close to the data source, which can reduce data transmission latency, improve real-time performance and response speed. By reasonably allocating tasks to different threads and processes, the computing resources of edge device 1 can be fully utilized, optimizing the resource utilization efficiency, and a voting mechanism based on a sliding window is introduced to synthesize the inference results of multiple frames and obtain the final real-time classification result. It can not only reduce misjudgments caused by single-frame inference results, but also achieve high detection accuracy with low computing overhead, and can realize efficient and real-time video analysis on resource-constrained edge device 1.
[0164] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided in the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0165] It should be noted that in this text, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, device, article or method comprising such element.
[0166] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A task deployment method based on a large artificial intelligence model, characterized in that Including the following steps: Obtain video data through a preset video source; wherein, the video data includes multiple video frames; Preprocess the video frames and store them in a preset first queue; Read the video frames in the first queue, and screen the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames; Store the target video frames in a preset second queue; Call a preset multi-modal model to infer the target video frames and obtain an inference result; Store the inference result in a preset sliding window; wherein, the size of the sliding window corresponds to the number of inference times of the inference result; Count the inference result with the highest proportion to obtain a real-time classification result and output it.
2. The task deployment method based on an artificial intelligence large model according to claim 1, wherein In the step of preprocessing the video frames and storing them in a preset first queue, it includes: Perform image correction and formatting on the video frames.
3. The task deployment method based on the artificial intelligence large model according to claim 1, wherein In the step of reading the video frames in the first queue and screening the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames, it includes: Calculate the similarity between the video frames; Screen out the video frames with a similarity lower than a preset similarity threshold.
4. The task deployment method based on the large artificial intelligence model according to claim 3, wherein After the step of storing the target video frames in a preset second queue, it includes: Judge whether the target video frames exist in the second queue; If so, judge whether the second queue is full; If not, increase the similarity threshold.
5. The task deployment method based on the large artificial intelligence model according to claim 4, wherein In the step of if so, judge whether the second queue is full, it includes: If so, lower the similarity threshold and replace the video frame with the highest similarity; If not, store the target video frames in the second queue.
6. The task deployment method based on the large artificial intelligence model according to claim 3, wherein In the step of calculating the similarity between the video frames, it includes: Extract the feature information of the video frames and generate corresponding sequences; Compare the sequences between multiple video frames to obtain the similarity.
7. The task deployment method based on the large artificial intelligence model according to claim 3, wherein After the step of screening out the video frames with a similarity lower than a preset similarity threshold, it further includes: Analyze the historical state of the video frames and the similarity threshold to obtain an analysis result; Adjust the similarity threshold according to the analysis result.
8. The task deployment method based on the artificial intelligence large model according to claim 1, characterized in that In the step of calling a preset multi-modal model to infer the target video frames and obtain an inference result, it includes: Extract the local detail information and global context information of the target video frames.
9. The task deployment method based on the artificial intelligence large model according to claim 1, wherein, Before the step of storing the target video frames in a preset second queue, it includes: Perform abnormal frame detection on the target video frames and screen out the abnormal frames.
10. A system for the task deployment method based on the artificial intelligence large model according to any one of claims 1-9, characterized in that, Including an edge device, a first thread, a second thread, and an inference process; the first thread and the second thread are respectively arranged on the edge device; The first thread includes a data acquisition module, a data storage module, and a preprocessing module. The data acquisition module is used to obtain video data through a preset video source; the data storage module is used to store video frames in a preset first queue; The preprocessing module is used to preprocess the video frames; The second thread includes a retrieval module and a similarity detection module; the retrieval module is used to read the video frames in the first queue; the similarity detection module is used to screen the video frames that meet the preset conditions through a preset similar frame algorithm to obtain target video frames; The inference process includes an inference module, a sliding window storage module, and a statistics module. The inference module is used to call a preset multimodal model to infer the target video frames and obtain an inference result; the sliding window storage module is used to store the inference result in a preset sliding window; The statistics module is used to count the inference result with the highest proportion to obtain a classification result and output it.
11. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps in the task deployment method based on the artificial intelligence large model according to any one of claims 1 to 9.
12. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the task deployment method based on the artificial intelligence large model according to any one of claims 1 to 9.