Intelligent video analysis method based on large model scheduling and storage medium
Through intelligent video analysis methods based on large-model scheduling, dynamically select and schedule large models, combined with distributed computing and model collaborative optimization, the problems of poor model universality, low analysis efficiency and insufficient adaptability to complex scenarios in the existing technology are solved, and efficient and accurate video analysis and rapid adaptability are achieved.
Patent Information
- Application Number
- CN202510360895.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-25
AI Technical Summary
There are problems in existing intelligent video analysis technologies that have poor model universality, low analysis efficiency and insufficient adaptability to complex scenarios.
Using an intelligent video analysis method based on large-model scheduling, we collect and preprocess video data, extract images and spatiotemporal features, dynamically select the optimal model from the large-model library, and conduct in-depth analysis and results fusion through distributed computing and model collaborative optimization, and finally generate a structured analysis report, and support online incremental learning and model updates.
Significantly improve the accuracy and efficiency of video analysis, enhance the flexibility and adaptability of the system, effectively improve the utilization rate of computing resources, and can quickly adapt to various complex scenarios and task needs.
Smart Images

Figure CN120220031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and particularly to an intelligent video analysis method and a storage medium based on large model scheduling. Background Art
[0002] With the wide application of video surveillance technology in many fields, such as security, transportation, industrial production, etc., the demand for intelligent video analysis is increasing day by day. Traditional intelligent video analysis methods usually rely on small models designed for specific tasks and scenarios. These models often show limitations when dealing with complex and variable actual video data. On the one hand, a single model is difficult to meet the requirements of multiple analysis tasks simultaneously, resulting in the need to deploy multiple models, increasing the complexity of the system and resource consumption; on the other hand, when the video scene or task changes, it is relatively difficult to adjust and update traditional models and they cannot quickly adapt to new situations. In addition, in the face of large-scale video data, traditional computing architectures and model collaboration methods cannot fully utilize the efficiency of computing resources, further restricting the efficiency and accuracy of video analysis.
[0003] Therefore, the applicant proposes an intelligent video analysis method based on large model scheduling, aiming to solve the problems of poor model generality, low analysis efficiency, and insufficient adaptability to complex scenarios existing in the existing intelligent video analysis technology. Summary of the Invention
[0004] The present invention proposes an intelligent video analysis method based on large model scheduling, which solves the problems of poor model generality, low analysis efficiency, and insufficient adaptability to complex scenarios existing in the existing intelligent video analysis technology. The technical solution of the present invention is realized as follows:
[0005] An intelligent video analysis method based on large model scheduling includes: collecting original video data and performing preprocessing to split a continuous video stream into a series of video segments with independent analysis value; for each video segment, parallelly extracting image features and spatio-temporal features, quantifying and normalizing the features, screening highly correlated features, and reducing redundancy; based on the video content description report, combining the task priority and model performance indicators, selecting the optimal model from the large model library, and dynamically allocating the model to a suitable node according to the resource status of the computing node; calling the selected large model for in-depth analysis, outputting the target position, category, confidence level, and behavior description results, and generating a structured analysis report, including event time, location, abnormal behavior label, and confidence information; integrating the output results of different models through a fusion model method to eliminate conflicts and noises; collecting samples that are not accurately recognized, inputting them into the incremental learning module after manual annotation, and adopting a knowledge distillation or adaptive parameter update strategy to update the model parameters and structure while retaining the original performance, and improving the system adaptability.
[0006] As a preferred technical solution, the preprocessing includes: using Gaussian filtering to remove noise interference in the video image, and applying histogram equalization technology to enhance the contrast and brightness of the image, making the target object more clearly distinguishable and laying a good foundation for subsequent feature extraction and analysis; it also includes splitting the continuous video stream into a series of video segments with independent analysis value according to the timestamp of the video data and the key event detection algorithm.
[0007] As a preferred technical solution, in the traffic monitoring scenario, the video segments are divided according to the time points when vehicles pass through specific intersections or sections; in the industrial production line, the segmentation is carried out according to the production cycle of the product or the completion nodes of key processes.
[0008] As a preferred technical solution, the video content feature extraction is achieved in the following way: the feature extraction module based on the deep convolutional neural network CNN uses convolutional kernels of different sizes and a multi-layer convolutional pooling structure to perform multi-scale feature extraction on the video image, obtaining rich low-level and high-level features; the spatio-temporal feature extraction module based on the recurrent neural network RNN and its variants models the time dimension information in the video sequence; then, the features output by different feature extraction modules are quantized and normalized to make their numerical ranges unified and comparable, and then a feature subset with relatively high relevance to the current video analysis task is selected through information gain, chi-square test or Relief algorithm.
[0009] As a preferred technical solution, the specific method for generating the video content description report is as follows:
[0010] Step S1: Target recognition and localization. The system automatically detects the main targets in the video, marks their positions and classifies them, and generates a confidence score for each recognition result to reflect the accuracy of the recognition.
[0011] Step S2: Scene classification. According to the overall characteristics of the video frame, the scene is classified into a specific type to assist in understanding the background environment of the target.
[0012] Step S3: Behavior analysis. Analyze the behavior patterns of the targets, record the duration and dynamic trends of the behaviors, and mark abnormal behaviors and their probabilities, providing a preliminary behavior explanation.
[0013] Step S4: Additional information extraction. Supplement the shooting time, location, weather conditions and context information of the video to enhance the integrity of the report.
[0014] Step S5: Structured integration. Organize the above information into a table or list form, clearly presenting data such as target categories, positions, behavior descriptions, confidence levels, and scene types, providing a unified basis for subsequent analysis.
[0015] As a preferred technical solution, the optimal model should be selected from the large model library according to the following method: for tasks that require high-precision object detection and classification, schedule a large object detection model based on the Transformer architecture with high-resolution feature extraction capabilities and a large amount of training data for a large number of object categories; for complex behavior analysis tasks, select a large behavior recognition model based on 3D convolutional neural networks that has been widely trained on various behavior patterns and has good generalization capabilities.
[0016] As a preferred technical solution, the specific methods for behavior description and probability estimation are as follows:
[0017] Step a) Behavior feature extraction and encoding: Based on a deep learning-based feature extraction model, multi-level features of the target behavior in the video segment are extracted, and the extracted behavior features are quantified and encoded into a numerical vector form that can be processed by a computer for subsequent model analysis and calculation.
[0018] Step b) Behavior classification and recognition: Using a pre-trained behavior classification model, the encoded behavior feature vector is input into the model. Based on the various behavior patterns and feature distributions it has learned, the model classifies and recognizes the target behavior. For each possible behavior category, the model calculates its corresponding probability score.
[0019] Step c) Generation of behavior description: According to the behavior classification results and probability scores, combined with a predefined behavior description template, a detailed behavior description statement is generated.
[0020] Step d) Uncertainty processing and supplementary information: When the behavior probability distribution is relatively dispersed, that is, no behavior category has a significantly high probability, the model will reflect this uncertainty in the behavior description; at the same time, the model will also combine other information in the video to supplement and improve the behavior description to provide more behavior analysis results.
[0021] As a preferred technical solution, the specific implementation steps for result fusion processing are as follows:
[0022] Step a) Data preparation and standardization: Collect the analysis results generated by different large models for the same video segment, including the location information of the target, the target category label, the behavior recognition result, and the relevant confidence scores; perform standardization processing on the output results of different models to ensure the consistency and comparability of the data formats.
[0023] Step b) Fusion based on probability statistics, including target location fusion, target category fusion, and behavior recognition result fusion.
[0024] Step c) Deep learning model assisted fusion: construct a deep learning model dedicated to result fusion, use the output results of different models as the input features of this fusion model, and train the fusion model on a large number of labeled video data to enable it to learn the optimal fusion method between the results of different models;
[0025] Step d) Conflict resolution and result optimization: for conflicts, use rule-based methods or further data analysis to resolve them;
[0026] Step e) Fusion result output: after the processing of the above steps, obtain the final fusion result, including information such as the fused target position, category, behavior, etc., and organize it into a unified format for output for subsequent application processing.
[0027] As a preferred technical solution, during the video analysis process, the online monitoring module continuously evaluates and validates the analysis results of the large model. The collected incremental data will enter the online incremental learning module after manual annotation and preprocessing. After online incremental learning, the updated large model will be reinvested in the video analysis task, continuously improving the intelligence level and adaptability of the system to ensure that it can cope with the ever-changing video data and analysis task requirements.
[0028] A non-transitory storage medium is used to store a program for executing the above-mentioned intelligent video analysis method based on large model scheduling.
[0029] Compared with the prior art, the present solution has the following beneficial effects:
[0030] (1) Significantly improve the accuracy of video analysis: The dynamic large model scheduling strategy can accurately select the most suitable model according to the real-time characteristics of the video content, ensuring that the most appropriate analysis means can be used in various complex scenarios, thus greatly reducing the probability of misjudgment and missed judgment. The multi-level feature fusion and adaptive adjustment mechanism fully exploits the rich information in the video data. By reasonably assigning weights to different-level features, the model can understand the video content more comprehensively and accurately, thereby improving the accuracy of tasks such as object detection and behavior recognition. For example, in industrial production quality inspection, it can more accurately identify the subtle defects and assembly problems of products and reduce the outflow of defective products. Distributed computing and model collaborative optimization enable each model to complement and cooperate with each other during the analysis process. Through information interaction and collaborative work, the overall analysis accuracy is further improved. For example, in traffic scenarios, multiple models can cooperate to more accurately judge the driving intentions and traffic conditions of vehicles and reduce the misjudgment of traffic accidents.
[0031] (2) Greatly enhance the flexibility and adaptability of the system: According to different application scenarios and diverse task requirements, flexibly schedule appropriate models from a rich large model library to easily handle various complex and changing situations. Whether it is security, transportation, industrial production or other fields, it can quickly adapt to new tasks and scenario changes, meeting the personalized needs of different users. Support online incremental learning and model update functions, which can collect newly emerging video data and samples that are not accurately recognized in real time, and update the parameters and structure of the model in a timely manner, enabling it to quickly adapt to new targets, behaviors and scenarios, maintaining the efficient analysis ability for continuously changing video data, and always providing accurate and up-to-date analysis results.
[0032] (3) Effectively improve the utilization rate of computing resources and analysis efficiency: The distributed computing architecture divides and processes video data in parallel, avoiding the resource bottleneck of traditional centralized computing, giving full play to the advantages of cluster computing resources, and greatly improving the data processing speed. For example, in a large-scale video surveillance system, it can process video streams from multiple cameras simultaneously to ensure real-time requirements. The model collaborative optimization algorithm reduces duplicate calculations and resource waste through efficient information sharing and collaborative work among models, further improving the utilization rate of computing resources and the overall analysis efficiency. For example, in an intelligent transportation system, multiple related models work together to quickly process a large amount of traffic video data, timely feedback traffic conditions, and provide strong support for traffic management. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0034] Figure 1 It is a method flow chart of an intelligent video analysis method based on large model scheduling of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] Refer to Figure 1, the present invention proposes an intelligent video analysis method based on large model scheduling, including steps such as video data collection and preprocessing, video content feature extraction and analysis, large model dynamic scheduling and task allocation, video analysis and result generation based on large models, result fusion and post-processing, and online incremental learning and model update, etc., which solves the problems existing in the intelligent video analysis technology in the prior art, such as poor model generality, low analysis efficiency, and insufficient adaptability to complex scenarios.
[0037] The specific method steps are as follows:
[0038] 1. Video data collection and preprocessing:
[0039] 1) Use cameras distributed at different locations with different perspectives and parameter settings to collect video data, ensuring comprehensive coverage of the monitoring area and diverse information collection. Through an optimized data transmission protocol, the collected video data is transmitted to the data processing center in real time to ensure the integrity and timeliness of the data.
[0040] 2) In the data processing center, perform preliminary preprocessing operations on the received video data, including but not limited to denoising, filtering, image enhancement, and color correction, etc. For example, use Gaussian filtering to remove noise interference in the video image, and apply histogram equalization technology to enhance the contrast and brightness of the image, making the target objects more clearly distinguishable, laying a good foundation for subsequent feature extraction and analysis.
[0041] 3) According to the time stamps of the video data and the key event detection algorithm, segment the continuous video stream into a series of video segments with independent analysis value. For example, in the traffic monitoring scenario, video segments can be divided according to the time points when vehicles pass through specific intersections or sections; in the industrial production line, segmentation can be carried out according to the production cycle of products or the completion nodes of key processes.
[0042] 2. Video content feature extraction and analysis:
[0043] 1) For each video segment, start multiple parallel feature extraction modules simultaneously. The feature extraction module based on the deep convolutional neural network (CNN) uses convolutional kernels of different sizes and a multi-layer convolutional pooling structure to perform multi-scale feature extraction on the video image, obtaining rich low-level and high-level features. For example, use 3x3, 5x5, and 7x7 convolutional kernels to extract local detail features, medium-scale features, and global semantic features of the image respectively, and reduce the resolution of the feature map through a series of pooling operations while retaining key information.
[0044] 2) A spatio-temporal feature extraction module based on recurrent neural networks (RNN) and its variants (such as long short-term memory network LSTM, gated recurrent unit GRU) models the temporal dimension information in the video sequence. By processing the video frames sequentially, spatio-temporal features such as the movement trajectory of the target, speed changes, temporal sequence patterns of behaviors, and interaction relationships between targets are captured. For example, in pedestrian behavior analysis, the LSTM network can remember the walking path and speed changes of pedestrians to judge their behavior intentions, such as whether they are wandering, running, or walking normally.
[0045] 3) Quantify and normalize the features output by different feature extraction modules to make their numerical ranges unified and comparable. Then, through feature selection methods such as information gain, chi-square test, or Relief algorithm, a feature subset with high relevance to the current video analysis task is selected to reduce the subsequent computational amount and interference of redundant information, and improve the analysis efficiency and accuracy of the model.
[0046] 3. Dynamic Scheduling and Task Allocation of Large Models:
[0047] 1) After the video content feature extraction and analysis are completed, the video content real-time analysis system quickly conducts a comprehensive and in-depth interpretation of the feature information of the video clip. Determine the main target types and quantities in the video through object detection algorithms, use scene classification algorithms to judge the environmental scene categories where the video is located, combine behavior analysis algorithms to initially identify the behavior patterns and trends of the targets, and generate a detailed video content description report based on this information.
[0048] Among them, the specific method for generating the video content description report is as follows:
[0049] First, for the object detection part, the system will accurately determine the position information of each object appearing in the video, and detail the specific position of each object in the video frame through the form of bounding box coordinates, and at the same time clarify its category, such as identifying the object as a pedestrian, vehicle (further subdivided into sedan, truck, motorcycle, etc.), animal, or specific industrial product, etc. And for each object, a confidence score will also be generated, which reflects the certainty degree of the model's object recognition result, presented in the form of a percentage, to help judge the reliability of the object recognition result in subsequent analysis. For example, an object identified as a sedan has a confidence score of 90%, indicating that the model has a high degree of confidence in this recognition result.
[0050] Secondly, in terms of scene classification, the system will accurately classify the environmental scene where the video is located into specific categories according to the pre-trained scene classification model, such as indoor office scenes, outdoor park scenes, urban street traffic scenes, industrial production workshop scenes, etc. This classification process is not only based on the overall features of the image, but also takes into account various factors such as lighting conditions, background elements, and target distribution. For example, for a video clip with a large number of traffic signs, road markings, and frequent vehicle driving, the system will accurately classify it as an urban street traffic scene.
[0051] Furthermore, the behavior analysis module will deeply analyze the behavior patterns and trends of the targets in the video. For pedestrians, it will judge whether they are in a walking, running, standing, sitting, wandering state, and whether there are specific actions such as waving, jumping, fighting, etc.; for vehicles, it will analyze their driving directions (due east, due south, due west, due north, etc.), speed changes (acceleration, deceleration, constant speed), whether they comply with traffic rules (such as running red lights, illegal lane changes, etc.), and the interaction behaviors with other targets (such as overtaking and avoiding between vehicles, crossing between pedestrians and vehicles, etc.). At the same time, corresponding behavior labels and behavior durations will be assigned to each behavior, such as "pedestrian wandering at the intersection, duration 10 seconds", "vehicle accelerating on the main road, duration 5 seconds", etc., in order to more clearly describe the behavior dynamics of the targets.
[0052] In addition, the system will also analyze and record other important elements in the video, such as the shooting time, location of the video (by associating with the geographical location information of the camera), weather conditions (if inferable from the video frame, such as rainy days, sunny days, snowy days, etc.), and visual features such as light intensity and color distribution in the frame. These additional information will further enrich the video content description report and provide a more comprehensive context basis for subsequent large model scheduling and video analysis.
[0053] Finally, all the above information will be integrated to generate a structured video content description report. The report is presented in a clear and easy-to-understand format, such as in tabular form, with each row recording information such as the target number, category, location, confidence level, behavior label, behavior duration, as well as scene category, shooting time, location, weather conditions, light intensity, etc., so as to provide accurate and detailed basis for subsequent large model selection and task allocation, ensuring that the entire intelligent video analysis system can make optimal decisions according to the actual situation of the video content, and improving the accuracy and efficiency of video analysis.
[0054] Such a video content description report can comprehensively reflect the key information of the video clip, making the system more accurate and efficient in large model scheduling and task allocation, giving full play to the advantages of each large model, and enhancing the performance and effect of the entire intelligent video analysis method.
[0055] 2) According to the video content description report, combined with the preset task priority list and model performance evaluation metrics, the intelligent scheduling algorithm quickly searches and matches in the large model library to select the large model most suitable for the current video analysis task. Each large model in the model library is equipped with metadata information such as detailed performance metric descriptions, applicable scenario ranges, and input / output requirements, so that the scheduling algorithm can accurately select the model. For example, for tasks that require high-precision object detection and classification, a large object detection model based on the Transformer architecture with high-resolution feature extraction capabilities and a large amount of training data for target categories is scheduled; for complex behavior analysis tasks, a large behavior recognition model based on 3D convolutional neural networks that has been widely trained on multiple behavior patterns and has good generalization capabilities is selected.
[0056] 3) In the task allocation stage, considering the load balancing and real-time requirements of computing resources, an algorithm based on dynamic resource allocation and task queue scheduling is adopted to allocate the selected large model to appropriate computing nodes. The resource status of the computing nodes (including CPU, GPU utilization, memory capacity, network bandwidth, etc.) is fed back to the scheduling algorithm through a real-time monitoring system to ensure that each large model can run in a computing environment with sufficient and stable resources, giving full play to its performance advantages and improving the overall response speed and analysis efficiency of the system.
[0057] 4. Video analysis and result generation based on large models:
[0058] 1) Input the preprocessed and feature-extracted video data into the selected large model. The large model deeply analyzes and infers the video features based on its pre-trained parameters and complex network architecture. For example, the object detection large model accurately determines the location, category, and confidence score of the objects in the video through multi-layer convolution and fully connected operations, and at the same time outputs the feature vectors of the objects for subsequent model interaction and analysis; the behavior recognition large model uses the learned behavior patterns and semantic information to classify and label the behaviors of the objects, determines whether they belong to normal or abnormal behaviors, and gives detailed behavior descriptions and probability estimates.
[0059] The specific methods for behavior description and probability estimation are as follows:
[0060] a) Behavior feature extraction and encoding:
[0061] A deep learning-based feature extraction model (such as a 3D convolutional neural network or a convolutional network combined with a spatio-temporal attention mechanism) performs multi-level feature extraction on the target behavior in the video clip. These features not only cover the appearance changes of the target (such as human body postures, limb movements, steering angles of vehicles, etc.), but also include the movement trajectories, speed changes of the target in the time series, and interaction features with the surrounding environment and other targets (such as distance changes between people, relative position relationships between vehicles and traffic facilities, etc.).
[0062] Quantify and encode the extracted behavior features and convert them into a numerical vector form that can be processed by a computer for subsequent model analysis and calculation. For example, for human behavior, information such as the position changes of joint points, movement directions and speeds of limbs may be encoded into a vector of fixed length, and each dimension represents a specific behavior feature parameter.
[0063] b) Behavior classification and recognition:
[0064] Using a pre-trained behavior classification model, input the encoded behavior feature vector into the model. Based on the various behavior patterns and feature distributions it has learned, the model classifies and recognizes the target behavior. For example, in a security monitoring scenario, the model can distinguish between normal behaviors such as walking, standing, and talking and abnormal behaviors such as fighting, running, and breaking in.
[0065] For each possible behavior category, the model calculates its corresponding probability score. This process is usually implemented by using the softmax function in the output layer of the model. The softmax function converts the prediction scores of the model for each behavior category into a probability distribution, making the sum of the probabilities of all behavior categories equal to 1. For example, if the model identifies that the human behavior in the video may belong to three categories: "normal walking", "running", and "wandering", after softmax calculation, the probability of "normal walking" is 0.6, the probability of "running" is 0.2, and the probability of "wandering" is 0.2.
[0066] c) Behavior description generation:
[0067] Generate a detailed behavior description statement based on the behavior classification results and probability scores, combined with a predefined behavior description template. For example, if the model determines that the target behavior belongs to the "running" category and has a relatively high probability (e.g., above 0.8), the generated behavior description may be "The target runs at a relatively fast speed in the video frame. The probability estimate of the running behavior is 0.85, and its behavior shows obvious signs of panic, indicating that there may be an emergency." For complex behavior scenarios, such as interactions between multiple people, the model will further analyze the behaviors of each target and their relationships to generate a more detailed and accurate description, such as "Three people have a fierce physical conflict in the video area. One person shows aggressive behavior, and the other two are in a defensive state. The probability estimate of the conflict behavior is 0.9."
[0068] d) Uncertainty handling and supplementary information:
[0069] When the behavior probability distribution is relatively dispersed, that is, no single behavior category has a significantly high probability, the model will reflect this uncertainty in the behavior description. For example, if the probabilities of the three behavior categories of "walking normally", "moving slowly", and "staying briefly" are 0.35, 0.3, and 0.35 respectively, the behavior description may be "The behavior of the target in the video frame has a certain degree of uncertainty. The possibilities of walking normally, moving slowly, and staying briefly are relatively close, approximately 0.35, 0.3, and 0.35 respectively. Further observation is needed to clarify its behavior intention."
[0070] At the same time, the model will also combine other information in the video, such as the appearance characteristics of the target (wearing a specific uniform may imply its professional identity, which in turn affects the interpretation of the behavior), the scene environment (behavior in a dangerous area may have a higher risk meaning), etc., to supplement and improve the behavior description, so as to provide a richer, more accurate and practically meaningful behavior analysis result.
[0071] Through the above steps, the system can give a detailed behavior description and probability estimate for the target behavior in the video, providing key decision-making basis and information support for video analysis applications in the fields of security monitoring, intelligent transportation, industrial production, etc., helping users to more accurately understand the events and behaviors occurring in the video, timely discover potential problems or abnormal situations, and take corresponding measures.
[0072] 2) During the analysis process, the large model will also generate intermediate results and auxiliary information, which can be utilized by subsequent processing modules or other relevant models. For example, when the semantic segmentation large model segments the video scene, it will simultaneously output the semantic category to which each pixel belongs and the boundary information of the region. These information are of great importance for target localization, scene understanding, and subsequent behavior analysis.
[0073] 3) Finally, the large model generates a detailed video analysis report based on the analysis results, including key information such as the location, category, behavior labels of the targets, the time and location of the event occurrence, and the confidence level of the analysis results, for subsequent result fusion and application processing.
[0074] 5. Result Fusion and Post-Processing:
[0075] 1) Since different large models may generate multiple analysis results for the same video clip, result fusion processing is required, and the specific implementation steps are as follows:
[0076] a) Data Preparation and Standardization:
[0077] Collect the analysis results generated by different large models for the same video clip. These results may include the location information of the targets (represented in coordinate form, such as the upper-left and lower-right coordinates of the bounding box), target category labels (such as "pedestrian", "vehicle", "animal", etc.), behavior recognition results (such as "walking", "running", "staying", etc.), and related confidence scores (usually a value between 0 and 1, indicating the degree of certainty of the model for this result).
[0078] Perform standardization processing on the output results of different models to ensure the consistency and comparability of the data formats. For example, if the target location coordinates output by some models are relative coordinates with the image center as the origin, while other models output absolute coordinates, they need to be unified into the same coordinate system; for the confidence scores, if there are different calculation methods or value ranges, normalization processing is also required to make them all within the same 0 to 1 interval.
[0079] b) Probability-Statistics-Based Fusion:
[0080] Target Location Fusion: For the location information of the same target detected by multiple models, a weighted average method is used for fusion. Assign corresponding weights according to the performance evaluation indicators of each model (such as the average accuracy rate, recall rate, etc. in the target localization task). For example, if Model A has a higher accuracy rate of 90% in previous target localization tests, a higher weight, such as 0.6, is assigned to it; if the accuracy rate of Model B is 80%, a weight of 0.4 is assigned. When calculating the fused target location coordinates, multiply the target location coordinates output by each model by its corresponding weight, and then add them up to obtain the final fused location coordinates.
[0081] Target category fusion: Statistically analyze the prediction results of each model for the target category, and make a fusion decision by combining their confidence scores. A common method is the majority voting method. If the majority of models predict the target as a certain category, then that category is taken as the fused target category. However, considering the confidence factor at the same time, if a model has a very high confidence in predicting a certain category (e.g., higher than 0.9), even if that category does not have a majority in the voting, it will be given a large weight for comprehensive judgment. For example, three models respectively predict the target as "pedestrian", "vehicle", "pedestrian", but the confidence of the model predicting "vehicle" is only 0.6, while the confidence of the other two models predicting "pedestrian" is 0.8 and 0.7 respectively. Then, the comprehensively judged fused target category is "pedestrian".
[0082] Behavior recognition result fusion: Similar to target category fusion, statistically analyze the recognition results of each model for the target behavior and their confidence scores. A fusion method based on probability distribution can be adopted. For example, the probability distributions of the behavior categories output by each model are weighted and summed to obtain the fused behavior probability distribution, and then the behavior category with the highest probability is selected as the final fused behavior result. At the same time, considering the temporal continuity and logical consistency of the behavior, if a certain behavior is recognized as "walking" by the majority of models at the previous moment in a video segment, and although some models predict other behaviors at the current moment, these behaviors have a certain logical continuity with "walking" (such as "walking faster", "turning while walking", etc.), then the fusion will be more inclined to the category related to the behavior at the previous moment.
[0083] c) Deep learning model-assisted fusion:
[0084] Build a deep learning model dedicated to result fusion, such as a multi-layer perceptron (MLP) or a convolutional neural network (CNN). Take the output results of different models (including information such as target location, category, behavior, etc. and their corresponding confidences) as the input features of this fusion model.
[0085] Train the fusion model on a large amount of labeled video data so that it learns the optimal fusion method between the results of different models. For example, through training, the model can learn which model results are more reliable in certain specific scenarios, and how to adjust the fusion weights and decision-making strategies according to the output features of different models to obtain more accurate fusion results.
[0086] The trained fusion model can further optimize and integrate the analysis results of different large models in practical applications, improving the accuracy and robustness of result fusion, especially applicable to situations where there are large differences or high uncertainties in the results of multiple models in complex scenarios.
[0087] d) Conflict resolution and result optimization:
[0088] During the fusion process, some conflict situations may occur. For example, the predictions of the target position by two models may differ significantly, or the judgments of the target category are completely different and the confidence levels are similar. For these conflicts, rule-based methods or further data analysis are used to resolve them. For example, if the difference in the target position between the two models exceeds a certain threshold (determined according to the video resolution and target size), then recheck the feature extraction process and analysis results of the two models. It may be found that there is a problem of false detection or inaccurate feature extraction in one of the models, thus excluding its results; for target category conflicts, if it cannot be clearly judged through confidence levels and majority voting, then combine the context information in the video (such as the environment around the target, the categories of other relevant targets, etc.) for auxiliary judgment.
[0089] Optimize and post-process the fused results to remove possible noise and redundant information. For example, if there are outliers in the fused target position coordinates (such as exceeding the video frame range), then correct or remove them; for some target detection results or behavior recognition results with low confidence levels, if their impact on the overall analysis result is small and there is a certain degree of uncertainty, their weights can also be appropriately reduced or directly ignored to improve the accuracy and reliability of the final fused result.
[0090] e) Output of the fused result:
[0091] After the processing of the above steps, the final fused result is obtained, including information such as the fused target position, category, behavior, etc., and it is organized into a unified format for output, so as to facilitate subsequent application processing. For example, output a structured data list, where each element contains information such as the unique identifier of the target, the fused position coordinates, the finally determined target category, behavior description, and the corresponding confidence score, providing an accurate, reliable, and consistent decision-making basis for subsequent video analysis applications (such as security event alarms, traffic flow statistics, industrial production quality inspection, etc.).
[0092] 3) According to specific application requirements, further post-process the finally fused output result. In the field of security monitoring, if the analysis result shows the existence of abnormal behaviors or security events, the system will automatically generate alarm information and send the relevant video clips and analysis reports to the monitoring personnel's terminal devices via text messages, emails, or dedicated monitoring software, etc., and at the same time trigger other security systems (such as access control systems, alarm systems, etc.) for emergency response. In industrial production quality inspection, compare the analysis result with the preset quality standards to generate a product quality inspection report, including information such as the type, location, and severity of product defects, and feedback it to the production control system for adjusting and optimizing the production process.
[0093] 6. Online incremental learning and model update:
[0094] 1) During the video analysis process, the online monitoring module continuously evaluates and validates the analysis results of the large model. By comparing with the manually annotated standard answers or accurate results in historical data, it identifies the targets, behaviors, or video segments with new features that have not been accurately recognized. For example, in an intelligent transportation system, if a new model of vehicle is misidentified as another type of vehicle, or a new type of traffic violation is not detected by the existing model, these video samples will be automatically collected.
[0095] The collected incremental data will be manually annotated and preprocessed before entering the online incremental learning module. This module adopts advanced incremental learning algorithms, such as knowledge distillation-based methods or adaptive parameter update strategies, to transfer the knowledge in the new data to the existing large model while avoiding negative impacts on the performance of the original model. During the training process, according to the characteristics of the new data and the current state of the model, the parameters and structure of the model are adaptively adjusted. For example, for newly emerging target categories.
[0096] 2) Add corresponding neurons to the last layer of the model, and use the backpropagation algorithm to train and update the new parameters so that the model can accurately recognize these new targets; for new behavior patterns, adjust the parameters of the feature extraction and classifier parts in the middle layer of the model to enhance the model's understanding and recognition ability of new behaviors.
[0097] 3) After online incremental learning, the updated large model will be reinvested in the video analysis task, continuously improving the intelligence level and adaptability of the system to ensure that it can handle the ever-changing video data and analysis task requirements.
[0098] Beneficial effects:
[0099] 1) Significantly improve video analysis accuracy:
[0100] The dynamic large model scheduling strategy can accurately select the most suitable model based on the real-time characteristics of the video content, ensuring that the most appropriate analysis means can be used in various complex scenarios, thus greatly reducing the probability of misjudgment and missed judgment. For example, in security monitoring, the analysis of crowd behavior can more accurately identify abnormal behaviors, avoiding false alarms or missed dangerous situations caused by model mismatches.
[0101] The multi-level feature fusion and adaptive adjustment mechanism fully exploits the rich information in the video data. By reasonably allocating weights to different-level features, the model can understand the video content more comprehensively and accurately, thereby improving the accuracy of tasks such as target detection and behavior recognition. For example, in industrial production quality inspection, it can more accurately identify the subtle defects and assembly problems of products, reducing the outflow of defective products.
[0102] Distributed computing and model collaborative optimization enable each model to complement and collaborate with each other during the analysis process. Through information interaction and collaborative work, the accuracy of the overall analysis is further improved. For example, in a traffic scenario, multiple models collaborating can more accurately judge the driving intentions of vehicles and traffic conditions, reducing misjudgments of traffic accidents.
[0103] 2) Greatly enhance the flexibility and adaptability of the system:
[0104] This solution can flexibly schedule appropriate models from a rich large model library according to different application scenarios and diverse task requirements, easily coping with various complex and changing situations. Whether it is in security, traffic, industrial production or other fields, it can quickly adapt to new tasks and scenario changes, meeting the personalized needs of different users.
[0105] Support online incremental learning and model update functions, which can collect newly emerging video data and samples that have not been accurately identified in real time, and update the parameters and structure of the model in a timely manner, enabling it to quickly adapt to new targets, behaviors and scenarios, maintaining the efficient analysis ability for continuously changing video data, and always providing accurate and up-to-date analysis results.
[0106] 3) Effectively improve the utilization rate of computing resources and analysis efficiency:
[0107] The distributed computing architecture divides video data for parallel processing, avoiding the resource bottleneck of traditional centralized computing, giving full play to the advantages of cluster computing resources, and greatly improving the data processing speed. For example, in a large-scale video surveillance system, it can process video streams from multiple cameras simultaneously to ensure real-time requirements.
[0108] The model collaborative optimization algorithm reduces duplicate calculations and resource waste through efficient information sharing and collaborative work among models, further improving the utilization rate of computing resources and the overall analysis efficiency. For example, in an intelligent transportation system, multiple related models collaborate to quickly process a large amount of traffic video data and timely feedback traffic conditions, providing strong support for traffic management.
[0109] The following is illustrated with a specific embodiment:
[0110] Example: Urban security monitoring system
[0111] In the security monitoring network of a modern city, a large number of high-definition cameras are deployed, covering all key areas of the city, such as commercial centers, residential communities, transportation hubs, etc. The intelligent video analysis method of the present invention is applied to this security monitoring system, and the specific implementation process is as follows:
[0112] 1. Video data collection and preprocessing:
[0113] The camera captures video data at a high frame rate and transmits the data to the server cluster in the monitoring center in real time through a wired network. The server performs real-time Gaussian filtering and median filtering on the received video data to remove noise interference caused by factors such as light flickering and wind blowing the grass. At the same time, the adaptive histogram equalization algorithm is used to enhance the image, improving the contrast and clarity of the target object for subsequent feature extraction and analysis.
[0114] According to the urban traffic flow and the patterns of people's activities, the video data is segmented into clips of 30 seconds each, ensuring that each clip can reflect relatively complete scene information and the dynamic behavior of people.
[0115] 2. Feature Extraction and Analysis of Video Content:
[0116] For each video clip, a CNN-based feature extraction module is launched to perform multi-scale feature extraction on the video image using different-sized convolutional kernels (such as 3x3, 5x5, 7x7) to obtain the texture, color, and shape features of the image. At the same time, an LSTM-based spatio-temporal feature extraction module models the movement trajectories, speed changes, and behavior patterns of people and vehicles in the video sequence to capture their time-series features.
[0117] The extracted features are normalized so that their numerical range is between [0,1], and a feature selection algorithm based on information gain is used to screen out a feature subset closely related to the security monitoring task, such as the facial features, clothing features, and behavioral action features of people, as well as the model type, color, and driving direction features of vehicles.
[0118] 3. Dynamic Scheduling and Task Allocation of Large Models:
[0119] When a video clip enters the analysis stage, the task perception module quickly analyzes the video content. If it detects signs of abnormal behavior (such as arguing and shoving) during a crowd gathering in the video scene, the system immediately schedules a large model with high-precision behavior recognition capabilities from the model library and allocates it to a computing node with sufficient GPU resources for task processing. At the same time, for video clips in other normal scenarios, such as monitoring the driving conditions of vehicles at a traffic intersection, a lightweight object detection and traffic flow analysis large model is scheduled to be processed on the CPU resources to achieve reasonable utilization of computing resources and efficient execution of tasks.
[0120] 4. Video Analysis and Result Generation Based on Large Models:
[0121] The large behavior recognition model deeply analyzes the input video features. Leveraging its pre-training advantages on crowd behavior data in a large number of different scenarios, it accurately determines whether there are abnormal behaviors in the crowd and outputs a detailed behavior description and the confidence score of the abnormal behavior. For example, for a conflict behavior involving multiple people, the model may output "A fierce physical conflict among a group of people was detected at [specific location]. The duration of the behavior was approximately [X] seconds. The confidence in this abnormal behavior is 0.92. Approximately [specific number] people were involved, and the actions of some of them had aggressive characteristics, which may cause personal injuries and public order chaos." The large object detection model quickly and accurately detects and identifies objects such as people and vehicles in the video, determines their positions, categories, and related attribute information, such as the license plate number of the vehicle, the gender and age range of the people, etc., and generates a corresponding confidence score for each detection result. For example, "A black car with the license plate number [specific license plate number] was detected at [specific coordinate location]. The confidence in identifying the vehicle type is 0.95. The driving direction is due east, and the speed is approximately [specific speed] km / h."
[0122] Generate a corresponding security monitoring report based on the analysis results, including the occurrence time, location, behavior type of the abnormal behavior, and information about the people and vehicles involved. At the same time, give a confidence estimate for the analysis results of each object and behavior for the monitoring personnel to conduct further verification and processing.
[0123] 5. Result Fusion and Post-processing:
[0124] Fuse the analysis results output by different large models. Through a fusion algorithm based on probability statistics, eliminate possible duplicate object information and inconsistent behavior judgment results to obtain an accurate and complete security monitoring report. For example, for the detection results of the same vehicle by multiple large object detection models, determine the final vehicle position and category information by calculating the weighted average of the position and category confidence; for the behavior recognition results, use a deep learning model to fuse the behavior labels and probabilities of different models to obtain the most reliable behavior judgment conclusion.
[0125] According to the severity of the analysis results, the system automatically generates corresponding alarm information and sends it to the mobile terminal devices of the monitoring personnel. At the same time, store the video analysis results in the database for subsequent query and statistical analysis. In addition, the system can also be linked with other security systems. For example, when an abnormal behavior is detected, it automatically triggers the alarm devices in the vicinity and notifies the surrounding security personnel to conduct on-site disposal.
[0126] 6. Online Incremental Learning and Model Update:
[0127] During the video analysis process, the system continuously collects video samples of abnormal behaviors that have not been accurately identified and video data of newly emerging target types (such as new models of drones, electric scooters, etc.). After these samples are manually annotated, they enter the online incremental learning module.
[0128] The online incremental learning module adopts an incremental learning algorithm based on knowledge distillation to transfer the knowledge in the new data to the existing large model. By adding a small number of neurons to the original model and adjusting some connection weights, the model can quickly learn new behavior patterns and target features without significantly affecting the performance of the original model on the learned tasks. For example, when a new model of drone flies in a no-fly zone, through online incremental learning, the model can quickly identify this new target and behavior pattern, include it in the scope of abnormal behavior monitoring, and update the relevant behavior recognition and warning mechanisms. After incremental learning, the updated large model will be reinvested in the security monitoring task to continuously improve the system's recognition and processing capabilities for new situations.
[0129] Specifically, it has the following advantages:
[0130] 1) Efficient and automated analysis;
[0131] 2) Quantify behavior reliability;
[0132] 3) Flexibly adapt to scenarios;
[0133] 4) Support complex decision-making;
[0134] 5) Multi-dimensional information fusion;
[0135] The practical application value part includes security, transportation, and industry.
[0136] It can be seen from the above embodiments that the intelligent video analysis method based on large model scheduling of the present invention can effectively improve the accuracy, efficiency, and flexibility of video analysis in the urban security monitoring system, provide strong technical support for the security guarantee of the city, and has significant practical application value and social benefits.
[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent video analysis method based on large model scheduling, characterized in that: include: Collect and pre-process the original video data to segment the continuous video stream into a series of video segments with independent analysis value; For each video clip, image features and spatiotemporal features are extracted in parallel and quantified and normalized to screen highly relevant features and reduce redundancy. Based on the video content description report combined with task priority and model performance indicators, the optimal model is selected from the large model library, and the model is dynamically allocated to the appropriate node according to the computing node resource status. The selected large model is called for in-depth analysis, and the target location, category, confidence, and behavior description results are output to generate a structured analysis report, including event time, location, abnormal behavior label, and confidence information. Integrate the output results of different models through the fusion model method to eliminate conflicts and noise; Collect samples that are not accurately identified, input them into the incremental learning module after manual annotation, and adopt knowledge distillation or adaptive parameter update strategy to update the model parameters and structure while retaining the original performance, thereby improving the system's adaptability.
2. The intelligent video analysis method based on large model scheduling as claimed in claim 1, characterized in that: The preprocessing includes: using Gaussian filtering to remove noise interference in video images, using histogram equalization technology to enhance the contrast and brightness of the image, making the target object clearer and more identifiable, and laying a good foundation for subsequent feature extraction and analysis; and also includes dividing the continuous video stream into a series of video clips with independent analysis value based on the timestamp and key event detection algorithm of the video data.
3. The intelligent video analysis method based on large model scheduling as claimed in claim 2, characterized in that: In traffic monitoring scenarios, video clips are divided according to the time points when vehicles pass through specific intersections or sections; in industrial production lines, they are segmented according to the product's production cycle or the completion nodes of key processes.
4. The intelligent video analysis method based on large model scheduling as claimed in claim 1, characterized in that: Video content feature extraction is achieved in the following ways: the feature extraction module based on the deep convolutional neural network (CNN) uses convolution kernels of different sizes and multi-layer convolution pooling structures to perform multi-scale feature extraction on video images and obtain rich low-level and high-level features; the spatiotemporal feature extraction module based on the recurrent neural network (RNN) and its variants models the time dimension information in the video sequence; then, the features output by different feature extraction modules are quantized and normalized to make their numerical ranges uniform and comparable, and then the feature subsets with high relevance to the current video analysis task are screened out based on information gain, chi-square test or relief algorithm.
5. The intelligent video analysis method based on large model scheduling as claimed in claim 1, characterized in that: The specific method for generating the video content description report is as follows: Step S1: Target recognition and positioning, the system automatically detects the main targets in the video, marks their locations and classifies them, and generates a credibility score for each recognition result to reflect the accuracy of the recognition; Step S2: scene classification, classifying the scene into a specific type based on the overall characteristics of the video image to assist in understanding the background environment of the target; Step S3: Behavior analysis, analyzing the target's behavior pattern, recording the duration and dynamic trend of the behavior, marking abnormal behaviors and their probabilities, and providing preliminary behavioral explanations; Step S4: extracting additional information to supplement the video’s shooting time, location, weather conditions, and context information to enhance the completeness of the report; Step S5: Structured integration, organize the above information into a table or list form, clearly present the target category, location, behavior description, confidence, and scene type data, and provide a unified basis for subsequent analysis.
6. The intelligent video analysis method based on large model scheduling as claimed in claim 1, characterized in that: The optimal model should be selected from the large model library in the following way: For tasks that require high-precision target detection and classification, a large target detection model based on the Transformer architecture with high-resolution feature extraction capabilities and a large amount of target category training data should be scheduled; for complex behavior analysis tasks, a large behavior recognition model based on a 3D convolutional neural network that has been extensively trained on multiple behavior patterns and has good generalization capabilities should be selected.
7. The method for digital management of ultrasound images according to claim 1, characterized in that: The specific methods of behavior description and probability estimation are as follows: Step a) Behavior feature extraction and encoding: The feature extraction model based on deep learning performs multi-level feature extraction on the target behavior in the video clip, quantifies and encodes the extracted behavior features, and converts them into numerical vector form that can be processed by the computer for subsequent model analysis and calculation. Step b) Behavior classification and identification: Using the pre-trained behavior classification model, the encoded behavior feature vector is input into the model. The model classifies and identifies the target behavior based on the various behavior patterns and feature distributions it has learned. For each possible behavior category, the model calculates its corresponding probability score; Step c) generating a behavior description, generating a detailed behavior description statement based on the behavior classification result and probability score combined with a predefined behavior description template; Step d) Uncertainty processing and supplementary information: When the behavior probability distribution is relatively dispersed, that is, no behavior category has a significantly high probability, the model will reflect this uncertainty in the behavior description; at the same time, the model will also combine other information in the video to supplement and improve the behavior description to provide more behavior analysis results.
8. The intelligent video analysis method based on large model scheduling as described in claim 1 is characterized in that: The specific implementation steps of result fusion processing are as follows: Step a) Data preparation and standardization: Collect the analysis results of different large models on the same video clip, including the location information of the target, the target category label, the behavior recognition result and the related confidence score; Standardize the output results of different models to ensure consistency and comparability of data formats; Step b) Fusion based on probability statistics, including target position fusion, target category fusion, and behavior recognition result fusion; Step c) Deep learning model assisted fusion: a deep learning model specifically used for result fusion is constructed, the output results of different models are used as input features of the fusion model, and the fusion model is trained on a large amount of annotated video data to learn the optimal fusion method between the results of different models; Step d) conflict resolution and result optimization, using rule-based methods or further data analysis to resolve conflicts; Step e) fusion result output: after the above steps are processed, the final fusion result is obtained, including the fused target position, category, behavior and other information, and it is organized into a unified format for output for subsequent application processing.
9. The intelligent video analysis method based on large model scheduling as described in claim 1 is characterized in that: During the video analysis process, the online monitoring module continuously evaluates and verifies the analysis results of the large model. The collected incremental data will be manually labeled and preprocessed before entering the online incremental learning module. After online incremental learning, the updated large model will be re-invested in the video analysis task, continuously improving the intelligence level and adaptability of the system, ensuring that it can cope with the ever-changing video data and analysis task requirements.
10. A non-temporary storage medium, characterized in that: It is used to store a program, which is used to: execute an intelligent video analysis method based on large model scheduling as described in any one of claims 1 to 9 above.
Citation Information
Patent Citations
Video classification method and system for intelligent elevator passenger intention analysis
CN118155119A
Video vector fusion analysis method and system based on deep learning
CN118366076A
Emergency scene intelligent analysis and decision support method based on AI large model
CN118378912A
Incremental learning method and device based on target detection and computer equipment
CN118941881A
Safety monitoring method based on mixing of large model and neural network algorithm
CN119380166A
Cited By
Intelligent marketing management system and method for text travel activities based on multi-dimensional distribution
CN120856900A