Video feature analysis method based on multi-modal large model

Through the video feature analysis method of multimodal large model, the problems of incomplete information capture, poor adaptability and insufficient security in the traditional video feature analysis method are solved, and efficient, accurate and safe video feature analysis is achieved.

CN120339902AInactive Publication Date: 2025-07-18SHANGHAI QUANZONG SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510351400.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional video feature analysis methods are based on a single mode, and are difficult to fully capture the rich information in the video, have poor adaptability and generalization capabilities, low computing efficiency, and insufficient data security and management.

Method used

The video feature analysis method based on multimodal large models is adopted, including video data acquisition, multimodal data preprocessing, multimodal large model construction and training, video feature extraction and fusion, feature analysis and classification, result visualization, model evaluation and optimization, data storage and management, user interaction and system security protection, etc., and feature analysis and security management are used to fuse images, audio and text information.

Benefits of technology

It improves the accuracy and adaptability of video feature analysis, reduces the rate of misjudgment and misjudgment, realizes efficient feature extraction and analysis, provides intuitive results display, and ensures data security and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339902A_ABST
    Figure CN120339902A_ABST
Patent Text Reader

Abstract

The invention provides a video feature analysis method based on a multi-modal large model, and relates to the technical field of video feature analysis methods. Comprising a video data acquisition module, a multi-modal data preprocessing module, a multi-modal large model construction and training module, a video feature extraction and fusion module, a feature analysis and classification module, a result visualization module, a model evaluation and optimization module, a data storage and management module, a user interaction module and a system security protection module. The video data acquisition module can improve the accuracy of feature analysis from various video sources such as a monitoring camera, a network video platform and a local video file library, features and semantics in videos can be captured more comprehensively by fusing multi-modal information and a powerful multi-modal large model, and compared with a single-modal analysis method, the video data acquisition module has the advantages that the accuracy of feature analysis is improved. The recognition accuracy of objects, scenes and behaviors in the video is remarkably improved, misjudgment and missed judgment conditions are reduced, and a more reliable basis is provided for subsequent decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video feature analysis methods. More specifically, it particularly relates to a video feature analysis method based on a multimodal large model. Background Art

[0002] With the rapid development of video technology, video data has been widely used in various fields such as security monitoring, intelligent transportation, video entertainment, and education. However, traditional video feature analysis methods are mainly based on a single modality, such as only using image information for analysis, and have many limitations. On the one hand, a single modality is difficult to comprehensively capture the rich information in the video. For example, it is unable to effectively combine audio and text information to understand the complete semantics of the video, resulting in inaccurate and incomplete analysis results. On the other hand, for complex scenarios and diverse video content, the adaptability and generalization ability of traditional methods are poor, and it is difficult to meet the high-precision and high-efficiency requirements for video feature analysis in practical applications. In addition, when dealing with large-scale video data, the computational efficiency of traditional methods is low, and fast feature extraction and analysis cannot be achieved. At the same time, the security and management of video data are also important issues. Existing methods have certain security risks during data storage and transmission, and lack an effective data management mechanism. Therefore, an innovative video feature analysis method based on a multimodal large model is needed to overcome these problems and improve the performance and application value of video feature analysis. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a video feature analysis method based on a multimodal large model to solve the above problems.

[0004] A video feature analysis method based on a multimodal large model includes a video data acquisition module, a multimodal data preprocessing module, a multimodal large model construction and training module, a video feature extraction and fusion module, a feature analysis and classification module, a result visualization module, a model evaluation and optimization module, a data storage and management module, a user interaction module, and a system security protection module. Each module works together to achieve multimodal feature analysis of the video.

[0005] Preferably, the video data acquisition module can obtain video data from various video sources, such as surveillance cameras, online video platforms, local video file libraries, etc., using distributed acquisition technology and video stream processing algorithms. The acquisition process follows specific data acquisition specifications and frame rate requirements. The formula is: , where represents the video source list, represents the acquisition algorithm, Represents the frame rate and has functions of data integrity verification and video format conversion. The multi-modal data preprocessing module preprocesses the collected video data using algorithms such as image denoising, video stabilization, and audio denoising. Its image denoising algorithm is based on a deep learning convolutional neural network, which removes noise by learning the noise pattern. The formula is: , where represents the original video data, represents the image denoising algorithm, represents the video stabilization algorithm, represents the audio denoising operation. At the same time, the text information in the video is extracted and standardized to improve the quality and usability of the data. The multi-modal large model construction and training module is based on a deep learning architecture, integrating model structures such as convolutional neural network (CNN), recurrent neural network (RNN) and its variants (such as LSTM, GRU), and uses a large-scale multi-modal video dataset for training. During the training process, an adaptive learning rate adjustment strategy and regularization technology are adopted. The formula is: , where represents the model architecture, represents the multi-modal dataset, represents the learning rate, represents the regularization technology, enabling the model to learn the multi-modal feature representation and semantic information of the video.

[0006] Preferably, the video feature extraction and fusion module uses the trained multi-modal large model to extract features from the preprocessed video data, including image features, audio features, text features, etc., and uses a feature fusion algorithm to fuse features of different modalities. The fusion algorithm is based on an attention mechanism and a multi-modal information weighting strategy. The formula is: , where represents the preprocessed video, represents the trained model, represents the attention mechanism, represents the weighting strategy, obtaining a comprehensive video feature vector. The feature analysis and classification module uses machine learning classification algorithms such as support vector machine (SVM), decision tree, random forest, etc. to analyze and classify the fused video feature vector. During the classification process, cross-validation and hyperparameter optimization techniques are adopted. The formula is: , where represents the fused feature vector, represents the classification algorithm, represents cross-validation, It represents hyperparameter optimization, identifies category information such as objects, scenes, and behaviors in the video, and the result visualization module converts the analyzed and classified results into intuitive visual forms such as charts, graphs, or video annotations, supporting a variety of visualization tools and technologies, such as matplotlib, seaborn, OpenCV, etc. The visualization conversion formula is: , where represents the classification result, represents the visualization tool, facilitating users to intuitively understand the results of video feature analysis.

[0007] Preferably, the model evaluation and optimization module uses evaluation metrics such as accuracy, recall rate, and F1 value to evaluate the performance of the trained multi-modal large model, and adjusts and optimizes the model parameters according to the evaluation results using the gradient descent method or other optimization algorithms. The formula is: , where represents the trained model, represents the evaluation metric, represents the optimization algorithm, improving the accuracy and generalization ability of the model. The data storage and management module adopts a distributed storage architecture and a database management system to store and manage the collected video data, preprocessed data, model parameters, analysis results, etc. During the storage process, data encryption and redundant backup technologies are used. Its data encryption adopts the Advanced Encryption Standard algorithm, and the redundant backup formula is: , where represents the stored data, represents the redundancy factor, ensuring the security and reliability of the data. The user interaction module provides a friendly user interface, supporting operations such as user input for video queries, setting analysis parameters, viewing analysis results, and providing feedback. The interface design follows the principles of human-computer interaction, with good usability and operability. Its interaction response formula is: , where represents the user request, represents the system function, ensuring that users can use the system efficiently.

[0008] Compared with the prior art, the present invention has the following beneficial effects:

[0009] Improve the accuracy of feature analysis: By integrating multi-modal information and a powerful multi-modal large model, it can capture features and semantics in the video more comprehensively. Compared with single-modal analysis methods, the recognition accuracy of objects, scenes, and behaviors in the video is significantly improved, reducing misjudgment and missed judgment situations, and providing a more reliable basis for subsequent decision-making.

[0010] Enhance the adaptability and generalization ability of analysis: The multi-modal large model trained with a large-scale dataset can learn rich feature patterns and semantic relationships, be able to adapt to different types of video data and complex scene changes, maintain good performance in various actual application scenarios, and reduce the dependence on specific scenarios.

[0011] Achieve efficient video processing: The distributed acquisition technology and optimized algorithm process enable the rapid acquisition and preprocessing of video data. The efficient feature extraction and fusion of the multi-modal large model and the application of classification algorithms greatly shorten the time of video feature analysis, improve the processing efficiency, and meet the application requirements with high real-time requirements.

[0012] Provide intuitive result display: The result visualization module converts complex analysis results into intuitive and easy-to-understand forms such as charts, graphs, or video annotations, facilitating users with different professional backgrounds to quickly understand the results of video feature analysis, promoting the effective communication and utilization of information, and improving user satisfaction with the system.

[0013] Ensure data security and reliability: The encryption and redundant backup technologies of the data storage and management module ensure the security and integrity of video data and analysis results, prevent data leakage and loss, maintain the stable operation of the system, enhance users' trust in the system, and are applicable to application scenarios with high requirements for data security. Brief Description of the Drawings

[0014] Figure 1 is the overall architecture flowchart of the present invention;

[0015] Figure 2 is the flowchart of the video data acquisition module of the present invention;

[0016] Figure 3 is the flowchart of the multi-modal data preprocessing module of the present invention;

[0017] Figure 4 is the flowchart of the multi-modal large model construction and training module of the present invention;

[0018] Figure 5 is the flowchart of the video feature extraction and fusion module of the present invention;

[0019] Figure 6 is the flowchart of the feature analysis and classification module of the present invention;

[0020] Figure 7 is the flowchart of the result visualization module of the present invention;

[0021] Figure 8 is the flowchart of the model evaluation and optimization module of the present invention;

[0022] Figure 9It is the flowchart of the data storage and management module of the present invention;

[0023] Figure 10 It is the flowchart of the user interaction module of the present invention. Specific embodiments

[0024] The following further describes the embodiments of the present invention in detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0025] Please refer to Figures 1-10 , the present invention provides a video feature analysis method based on a multimodal large model, including a video data acquisition module, a multimodal data preprocessing module, a multimodal large model construction and training module, a video feature extraction and fusion module, a feature analysis and classification module, a result visualization module, a model evaluation and optimization module, a data storage and management module, a user interaction module, and a system security protection module. Each module works together to achieve multimodal feature analysis of videos.

[0026] The video data acquisition module can obtain video data from various video sources, such as surveillance cameras, network video platforms, local video file libraries, etc., using distributed acquisition technology and video stream processing algorithms. The acquisition process follows specific data acquisition specifications and frame rate requirements. The formula is: , where represents the video source list, represents the acquisition algorithm, represents the frame rate, and has data integrity verification and video format conversion functions. The multimodal data preprocessing module uses algorithms such as image denoising, video stabilization, and audio denoising to preprocess the acquired video data. Its image denoising algorithm is based on a deep learning convolutional neural network and denoises by learning noise patterns. The formula is: , where represents the original video data, represents the image denoising algorithm, represents the video stabilization algorithm, represents the audio denoising operation, and at the same time extracts and standardizes the text information in the video to improve the quality and usability of the data. The multimodal large model construction and training module is based on a deep learning architecture, integrates model structures such as convolutional neural networks (CNN), recurrent neural networks (RNN) and their variants (such as LSTM, GRU), and uses a large-scale multimodal video dataset for training. During the training process, an adaptive learning rate adjustment strategy and regularization techniques are adopted. The formula is: , where represents the model architecture, represents the multimodal dataset, represents the learning rate, Represents a regularization technique that enables the model to learn the multi-modal feature representation and semantic information of the video.

[0027] The video feature extraction and fusion module uses a trained multi-modal large model to extract features from the preprocessed video data, including image features, audio features, text features, etc., and uses a feature fusion algorithm to fuse features of different modalities. The fusion algorithm is based on an attention mechanism and a multi-modal information weighting strategy, and the formula is: , where represents the preprocessed video, represents the trained model, represents the attention mechanism, represents the weighting strategy, obtaining a comprehensive video feature vector. The feature analysis and classification module uses machine learning classification algorithms, such as support vector machine (SVM), decision tree, random forest, etc., to analyze and classify the fused video feature vector. Cross-validation and hyperparameter optimization techniques are used during the classification process, and the formula is: , where represents the fused feature vector, represents the classification algorithm, represents cross-validation, represents hyperparameter optimization, identifying category information such as objects, scenes, and behaviors in the video. The result visualization module converts the analysis and classification results into intuitive visual forms such as charts, graphs, or video annotations, supporting a variety of visualization tools and technologies, such as matplotlib, seaborn, OpenCV, etc. The visualization conversion formula is: , where represents the classification result, represents the visualization tool, facilitating users to intuitively understand the results of video feature analysis.

[0028] The model evaluation and optimization module uses evaluation metrics such as accuracy, recall rate, F1 value, etc. to evaluate the performance of the trained multi-modal large model, and adjusts and optimizes the model parameters according to the evaluation results using gradient descent method or other optimization algorithms. The formula is: , where represents the trained model, represents the evaluation metric, represents the optimization algorithm, improving the accuracy and generalization ability of the model. The data storage and management module uses a distributed storage architecture and a database management system to store and manage the collected video data, preprocessed data, model parameters, analysis results, etc. Data encryption and redundant backup technologies are used during the storage process. Its data encryption uses the Advanced Encryption Standard algorithm, and the redundant backup formula is: , where represents the stored data, Denotes the redundancy factor, ensuring data security and reliability. The user interaction module provides a friendly user interface, supporting operations such as user input for video queries, setting analysis parameters, viewing analysis results, and providing feedback. The interface design follows the principles of human-computer interaction, featuring good usability and operability. Its interaction response formula is: , where Denotes the user request, Denotes the system function, ensuring that users can use the system efficiently.

[0029] Example:

[0030] Example 1:

[0031] · Hardware configuration:

[0032] · Server: Select a Dell PowerEdge R740 server, equipped with an Intel Xeon E5 - 2620 v4 processor, 64GB of memory, and a 2TB hard drive, providing basic computing power and storage support for system operation.

[0033] · GPU: Adopt an NVIDIA GeForce GTX 1080 Ti to accelerate the operation of deep learning models and meet the processing requirements of a certain scale of multi-modal data.

[0034] · Network device: Select a Huawei S5720 - 56C - PWR Ethernet switch to ensure stable and efficient data transmission and avoid the impact of network congestion on video data transmission.

[0035] · Software system:

[0036] · Video data acquisition module: Develop an acquisition program based on the Python OpenCV library to sequentially read video data from the local stored monitoring video file library, set the frame rate to 20 frames per second, and verify data integrity through a simple checksum algorithm during the acquisition process. Convert the acquired videos to the.avi format for subsequent processing.

[0037] · Multi-modal data preprocessing module: Use the median filtering algorithm built into OpenCV for image denoising, use a traditional video stabilization algorithm based on feature point matching to stabilize the video frame, and use a simple high-pass filter for the audio part to remove low-frequency noise. For the text information in the video, extract it using the Tesseract OCR engine and perform preliminary standardization processing using a custom rule dictionary to correct common spelling mistakes and format problems.

[0038] · Multimodal Large Model Construction and Training Module: Build a relatively simple multimodal large model under the TensorFlow framework. For the image feature extraction part, use a 3-layer convolutional neural network (CNN); for audio feature extraction, utilize a 2-layer bidirectional GRU network; for text feature extraction, base on a 1-layer LSTM network, and perform feature fusion through a fully connected layer. Train it using a small self-made dataset containing 500 video samples, adopt a fixed learning rate of 0.001 and L1 regularization technology, and train for 8 epochs to initially enable the model to learn the basic representations of multimodal features.

[0039] · Video Feature Extraction and Fusion Module: Input the preprocessed video data into the trained model. After obtaining the feature vectors of each modality, use a simple equal-weight fusion method (the weights of image, audio, and text are all 1 / 3) to merge the features and obtain a comprehensive video feature vector. This method is computationally simple and suitable for scenarios with limited resources.

[0040] · Feature Analysis and Classification Module: Use the decision tree classification algorithm to classify the fused feature vectors. Manually set the depth of the tree to 5 to limit the model complexity, and identify the human actions in the video (such as walking, running, standing) and scene types (such as indoor rooms, outdoor streets, parks, etc.). Use 3-fold cross-validation to evaluate the performance of the model under limited data.

[0041] · Result Visualization Module: Use the matplotlib library in Python to display the classification results in the form of a simple bar chart. The abscissa is the class label, and the ordinate is the number of samples in each class. At the same time, use OpenCV to annotate the recognized results in text form on the original video frame, which is convenient for users to intuitively compare the video content with the analysis results.

[0042] · Model Evaluation and Optimization Module: Adopt accuracy and recall as the main evaluation indicators. After calculation on the test set, the accuracy reaches 65% and the recall is 60%. According to the evaluation results, use the stochastic gradient descent method to slightly adjust the parameters of the fully connected layer of the model, adjust the learning rate to 0.0008, and then train for 3 more epochs to attempt to improve the model performance.

[0043] · Data Storage and Management Module: Use the SQLite database to store the collected video data, preprocessed data, model parameters, and analysis results. Use the AES-128 encryption algorithm to encrypt sensitive data, and set the redundancy backup factor to 2 to ensure the security and reliability of local data storage and prevent data loss due to hardware failures and other reasons.

[0044] · User Interaction Module: Develop a basic web application based on HTML, CSS, and JavaScript as the user interface, providing a video upload button, a simple dropdown menu for classification categories (such as human actions, scene types) for users to select the analysis direction of interest, and a result display area. Users access the system through a browser. After submitting a video, the system gives an analysis result within 15 seconds. Although the response speed needs to be improved, it can meet simple analysis requirements.

[0045] Example 2:

[0046] · Hardware Optimization:

[0047] · Upgrade the server to a Huawei FusionServer Pro 2288H V5 server, equipped with two Intel Xeon Silver 4210R processors, expand the memory to 128GB, and increase the hard disk capacity to 4TB, significantly enhancing the overall computing performance and data storage capacity of the system to handle more complex multi-modal video analysis tasks.

[0048] · Replace the GPU with an NVIDIA A100. Its powerful computing ability can significantly accelerate the training and inference processes of deep learning models. Especially when dealing with large-scale multi-modal data, it can effectively shorten the operation time and improve the timeliness of the system.

[0049] · Upgrade the network device to a Cisco Catalyst 9300 - L switch, which supports higher network bandwidth and lower transmission latency, ensuring the efficiency and stability during the transmission of a large amount of video data and communication between multiple modules, and preventing data transmission from becoming a bottleneck in system performance.

[0050] · Software Upgrade:

[0051] · Video Data Acquisition Module: Introduce the Apache Flume distributed acquisition framework combined with the Sqoop tool to achieve video data acquisition from multiple different types of data sources (such as video clusters in a distributed file system, video metadata in a relational database, and network video streaming platforms). The acquisition frame rate can be dynamically adjusted according to the characteristics of the data source, up to 30 frames per second. Use a Bloom filter based on the hash algorithm for efficient data deduplication, and adopt the SSL / TLS encryption protocol during the acquisition process to ensure data transmission security. Convert the acquired video data into a more efficient H.265 format for storage and transmission.

[0052] · Multi-modal data preprocessing module: The DnCNN model based on deep learning is used for image deep denoising. The video stabilization algorithm based on deep learning optical flow method is utilized to improve the stability of video frames. The audio noise is removed by the Wave-U-Net audio denoising model based on deep learning, significantly enhancing the quality of each modal data. For the text information in the video, the CRNN OCR model based on deep learning is applied for extraction, and natural language processing toolkits (such as AllenNLP) are used for semantic analysis and standardization processing to further explore the value of text information.

[0053] · Multi-modal large model construction and training module: A deeply integrated multi-modal large model is constructed based on the PyTorch framework, including a CNN with multiple residual block structures for image feature extraction, multi-layer bidirectional LSTM combined with attention mechanism for audio and text sequence processing, and deep fusion of multi-modal features is achieved through an adaptive attention fusion module. It is trained using a large-scale multi-modal dataset containing 5000 video samples, adopting an adaptive learning rate adjustment strategy (such as the AdamW optimizer) and various regularization techniques (such as Dropout, Layer Normalization), and training for 30 epochs to enable the model to fully learn complex multi-modal features and semantic relationships.

[0054] · Video feature extraction and fusion module: Using the trained multi-modal large model, the adaptive attention mechanism automatically learns the weights of different modal features in different scenarios to achieve dynamic feature fusion. Compared with the fixed-weight fusion method, it can capture the key features of the video more accurately and obtain high-quality comprehensive video feature vectors.

[0055] · Feature analysis and classification module: The random forest and XGBoost classifiers in ensemble learning are used to classify the fused feature vectors. Hyperparameter optimization is carried out through 5-fold cross-validation and grid search techniques, which can accurately identify complex behaviors (such as group aggregation, abnormal movement, specific gestures, etc.) and special scenarios (such as fire scenes, traffic accidents, construction sites, etc.) in the video, improving the accuracy and practicality of the analysis.

[0056] · Result visualization module: Combining the seaborn library and the Plotly library, the classification results are presented in a variety of interactive visualization charts (such as heatmaps showing the frequency distribution of behaviors, pie charts presenting the proportion of scene categories, and line charts reflecting the trend of features over time). At the same time, OpenCV and augmented reality (AR) technologies are used to overlay 3D annotations and virtual prompt information on the video frames, providing users with a more intuitive and immersive result viewing experience and enhancing the information transmission effect.

[0057] · Model Evaluation and Optimization Module: Use a comprehensive evaluation index system, including accuracy, recall, F1 value, mean average precision (mAP), etc. to evaluate the model. On the test set, the calculated accuracy is 80%, recall is 75%, F1 value is 77%, and mAP is 73%. According to the evaluation results, use model pruning technology to remove redundant parameters in the model, and at the same time use knowledge distillation technology to transfer the knowledge of complex models to simplified models, improving the inference speed and deployment efficiency of the model without significantly reducing performance, and further optimizing the model performance.

[0058] · Data Storage and Management Module: Adopt an architecture that combines the Hadoop Distributed File System (HDFS) and the Hive database to store data. Use a distributed lock mechanism and a consistent hashing algorithm to ensure data consistency and integrity, and ensure reliable reading and writing of data in a multi-node storage environment. Use a more advanced RSA - 2048 encryption algorithm to encrypt the data, and use a data life cycle management strategy to automatically clean up expired or low-value data, improving the utilization rate of storage resources and the intelligent level of data management.

[0059] · User Interaction Module: Develop a responsive web application that adapts to multiple terminal devices (computers, tablets, mobile phones), provides an intelligent search box to support users to search for video analysis results by keywords, and at the same time provides detailed result filtering and sorting functions (such as by time, category, confidence, etc.). Support users to comment on, favorite, and share analysis results, enhancing interaction and communication among users. The system gives preliminary analysis results within 8 seconds after the user submits the video, significantly improving the interaction response speed and user experience.

[0060] Example 3:

[0061] · Hardware Innovation:

[0062] · On edge computing devices, use NVIDIA Jetson Xavier NX as an edge node, deploy it near the video data source (such as a surveillance camera). It has 8GB of memory and 16GB of eMMC storage, and communicates with the central server through a 5G network. It can realize the near-source collection and preprocessing of video data, effectively reducing data transmission latency and the computing pressure on the central server.

[0063] · The central server adopts a cloud computing architecture, rents elastic computing resources from Tencent Cloud, and dynamically adjusts the server configuration according to the load of the actual business. For example, it automatically expands CPU and memory resources during the peak period of video data processing, and releases idle resources during the business low period, reducing the hardware procurement and maintenance costs and improving the resource utilization efficiency.

[0064] · Software Expansion:

[0065] · Video data acquisition module: Integrated with Internet of Things devices (such as temperature and humidity sensors, light sensors, passive infrared sensors, etc.), it acquires environment-related data while collecting video data, and fuses these multi-source data with video data to provide richer background information for video analysis. Using edge computing technology, it performs real-time analysis and preliminary processing of the collected data locally. For example, it detects abnormal situations through simple threshold judgments, and only transmits valuable data and preliminary analysis results to the central server, greatly reducing the network transmission burden.

[0066] · Multimodal data preprocessing module: Adopting federated learning technology, it collaborates with other edge nodes or data owners to perform data preprocessing and model training without transmitting raw data, effectively protecting data privacy. Using image enhancement technology based on Generative Adversarial Networks (GANs) to enhance video images, improving the clarity and detail expressiveness of the images. At the same time, using audio super-resolution technology to improve audio quality, laying a better foundation for subsequent feature extraction and analysis.

[0067] · Multimodal large model construction and training module: Constructs a multimodal large model based on Transformer, and uses the multi-head attention mechanism to better fuse information from different modalities to achieve cross-modal semantic understanding. Adopting transfer learning technology, it fine-tunes on video datasets in specific domains (such as security monitoring, industrial inspection, etc.) using the model parameters pre-trained on large-scale general video datasets, reducing training time and resource consumption, and improving the adaptability and accuracy of the model in specific domains.

[0068] · Video feature extraction and fusion module: Introduces Dynamic Graph Convolutional Networks (DGCNNs) to extract more refined spatio-temporal features of videos, combines reinforcement learning algorithms to dynamically adjust the weights and strategies of feature fusion, and optimizes the feature fusion process in real time according to the changes in video content, further improving the effect and adaptability of feature fusion, and obtaining more representative comprehensive video feature vectors.

[0069] · Feature analysis and classification module: Applies semantic segmentation models and object detection models based on deep learning to perform more refined analysis and classification of objects and scenes in videos, capable of identifying tiny objects and complex scene structures in videos, such as identifying the component status of specific devices, subtle movements and expressions of personnel, etc., to meet the requirements of high-precision analysis.

[0070] · Result Visualization Module: Develop an immersive visualization platform using the Unity 3D engine and AR / VR technology, and display the video analysis results in the form of 3D models, virtual scenes, etc. Users can view and analyze the videos in a virtual environment through a head-mounted device or an interactive terminal, providing an unprecedented interactive experience and greatly enhancing users' understanding and application capabilities of complex video analysis results.

[0071] · Model Evaluation and Optimization Module: Adopt a continuous evaluation and adaptive optimization strategy. During the operation of the system, continuously collect new data and user feedback, use online learning technology to update model parameters in real time, and at the same time combine an automatic hyperparameter adjustment algorithm (such as a method based on Bayesian optimization) to continuously optimize the model performance, ensuring that the model always maintains high accuracy and adaptability.

[0072] · Data Storage and Management Module: Use a distributed storage architecture combined with blockchain technology to store data. Utilize the immutable and traceable characteristics of the blockchain to ensure the integrity and security of the data. At the same time, adopt a distributed caching technology to improve the data reading speed, meeting the high requirements of the system for data storage and access.

[0073] · User Interaction Module: Develop a multimodal interaction application program that supports multiple interaction methods such as voice, gesture, and touch. Users can quickly query video analysis results and set analysis parameters through voice commands, and perform operations such as zooming, rotating, and marking in the 3D visualization results through gesture operations, providing a more natural and convenient interactive experience and further improving the efficiency and satisfaction of users using the system.

[0074] Example 4:

[0075] · Hardware Configuration Adjustment:

[0076] · The server uses a Supermicro SuperServer 7049GP-TRT server, configured with an AMD EPYC 7742 processor, 512GB of memory, and 16TB of hard disk, providing the system with powerful computing capabilities and massive storage resources, suitable for processing large-scale, long-term video data and complex multimodal analysis tasks, such as urban-level security monitoring video analysis, large-scale industrial production process video monitoring, etc.

[0077] · The storage device selects a Dell EMC PowerMax storage array, which has high-speed data reading and writing capabilities and high reliability, provides data redundancy and snapshot functions, can ensure the security and recoverability of data during storage, and prevent data loss due to hardware failures or unexpected situations, meeting application scenarios with extremely high requirements for data security and integrity.

[0078] · Software System Customization:

[0079] · Video data acquisition module: Develop dedicated acquisition programs for specific professional video data sources (such as high-definition satellite videos, medical imaging videos, professional sports event videos, etc.). Use high-speed data acquisition cards and professional video acquisition protocols to ensure the complete and accurate acquisition of high-resolution and high-frame-rate video data. And use hardware encryption chips to encrypt the acquired data in real time to ensure the security and confidentiality of the data.

[0080] · Multimodal data preprocessing module: Develop specialized preprocessing processes and algorithms according to the characteristics of different professional video data. For example, for medical imaging videos, use specific image enhancement algorithms to highlight the diseased areas and use professional knowledge in the medical field for text extraction and annotation; for sports event videos, perform key frame extraction and action classification preprocessing on the videos through athlete action analysis algorithms. At the same time, use domain ontology knowledge for semantic annotation and classification of the data to improve the usability and comprehensibility of the data.

[0081] · Multimodal large model construction and training module: Build customized multimodal large models based on domain-specific knowledge and a large amount of professional video data. For example, in the medical field, combine a medical knowledge graph and a deep learning model to build a multimodal large model for disease diagnosis; in the sports field, use an athlete action library and game rules to build a multimodal large model for sports event analysis. Adopt domain adaptation training methods and few-shot learning techniques to fully utilize limited professional data to train high-performance models and improve the accuracy and reliability of the models in specific domains.

[0082] · Video feature extraction and fusion module: Use the customized multimodal large model and adopt feature extraction methods and fusion strategies based on domain knowledge. For example, in medical imaging analysis, focus on fusing image features and text description features related to diseases; in sports event analysis, highlight the fusion of athlete action features and game scene features. In this way, obtain more targeted and discriminative comprehensive video feature vectors to improve the accuracy and effectiveness of the analysis.

[0083] · Feature analysis and classification module: Analyze and classify the fused feature vectors using classification criteria and models in the professional field. For example, in the medical field, diagnose and classify the diseases in the video according to disease diagnosis criteria and classification systems; in the sports field, analyze and evaluate the events according to game rules and athlete performance evaluation criteria. Adopt a combination of expert evaluation and model evaluation to continuously optimize the classification model and improve the reliability of the analysis results.

[0084] · Result Visualization Module: Develop customized visualization tools and interfaces according to the needs of different professional fields and user habits. For example, in the medical field, use 3D modeling and visualization technologies to display the pathological conditions and analysis results of human organs; in the sports field, display the performance of athletes and key information of competitions through a combination of dynamic charts and video replays. Adopt an intuitive and easy-to-understand visualization method to facilitate professional users to quickly understand and apply the analysis results.

[0085] · Model Evaluation and Optimization Module: Establish an evaluation index system that conforms to the characteristics of the professional field. For example, in the medical field, use indicators such as sensitivity, specificity, and accuracy to evaluate the diagnostic performance of the model; in the sports field, use indicators such as the improvement degree of athlete performance and the prediction accuracy of competition results to evaluate the analysis effect of the model. According to the evaluation results, use domain expertise and optimization algorithms to conduct targeted optimization and improvement of the model. For example, in the medical field, adjust the model parameters according to the new characteristics of diseases and changes in diagnostic criteria; in the sports field, update the model structure and training data according to the new competition rules and the development trends of athletes' techniques.

[0086] · Data Storage and Management Module: Use a professional database management system (such as the PACS system in the medical field and the event database in the sports field) to store and manage data. Design a reasonable storage structure and indexing mechanism according to the characteristics and usage requirements of domain data to improve the storage efficiency and query speed of data. At the same time, use strict data access control and auditing mechanisms to ensure the security and compliance of data, and prevent data leakage and illegal use.

[0087] · User Interaction Module: Develop dedicated client software for professional users, providing a user interface and functional modules that conform to professional operation habits. For example, in the medical field, provide functions such as case management and diagnostic report generation; in the sports field, provide functions such as event analysis report generation and athlete data statistics. Support efficient interaction and data sharing between users and the system, and improve the work efficiency and analysis quality of professional users.

[0088] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A video feature analysis method based on a multimodal large model, characterized in that, It includes a video data acquisition module, a multimodal data preprocessing module, a multimodal large model construction and training module, a video feature extraction and fusion module, a feature analysis and classification module, a result visualization module, a model evaluation and optimization module, a data storage and management module, a user interaction module, and a system security protection module. Each module works collaboratively to achieve multimodal feature analysis of videos.

2. The video feature analysis method according to claim 1, wherein the video data acquisition module can obtain video data from various video sources, such as surveillance cameras, network video platforms, local video file libraries, etc., using distributed acquisition technology and video stream processing algorithms. The acquisition process follows specific data acquisition specifications and frame rate requirements, and the formula is: , where represents the video source list, represents the acquisition algorithm, represents the frame rate, and has data integrity verification and video format conversion functions.

3. The video feature analysis method according to claim 1, wherein the multimodal data preprocessing module preprocesses the collected video data by using algorithms such as image denoising, video stabilization, and audio denoising. The image denoising algorithm is based on a deep learning convolutional neural network and denoises by learning the noise pattern. The formula is: , where represents the original video data, represents the image denoising algorithm, represents the video stabilization algorithm, represents the audio denoising operation, and at the same time extracts and standardizes the text information in the video to improve the quality and usability of the data.

4. The video feature analysis method according to claim 1, wherein the multi-modal large model construction and training module is based on a deep learning architecture, integrating model structures such as convolutional neural network (CNN), recurrent neural network (RNN) and its variants (such as LSTM, GRU), and is trained using a large-scale multi-modal video dataset. During the training process, an adaptive learning rate adjustment strategy and regularization technology are adopted, and the formula is: , where represents the model architecture, represents the multi-modal dataset, represents the learning rate, represents the regularization technology, enabling the model to learn the multi-modal feature representation and semantic information of the video.

5. The video feature analysis method according to claim 1, wherein the video feature extraction and fusion module uses a trained multi-modal large model to extract features from the pre-processed video data, including image features, audio features, text features, etc., and adopts a feature fusion algorithm to fuse features of different modalities. The fusion algorithm is based on an attention mechanism and a multi-modal information weighting strategy, and the formula is: , where represents the pre-processed video, represents the trained model, represents the attention mechanism, represents the weighting strategy, and a comprehensive video feature vector is obtained.

6. The video feature analysis method according to claim 1, wherein the feature analysis and classification module uses machine learning classification algorithms, such as support vector machine (SVM), decision tree, random forest, etc., to analyze and classify the fused video feature vectors. During the classification process, cross-validation and hyperparameter optimization techniques are adopted, and the formula is: , where represents the fused feature vector, represents the classification algorithm, represents cross-validation, represents hyperparameter optimization, and identifies category information such as objects, scenes, behaviors, etc. in the video.

7. The video feature analysis method according to claim 1, wherein the result visualization module converts the analyzed and classified results into visual forms such as intuitive charts, graphs, or video annotations, supporting a variety of visualization tools and technologies, such as matplotlib, seaborn, OpenCV, etc. The visualization conversion formula is: , where represents the classification result, represents the visualization tool, facilitating users to intuitively understand the results of video feature analysis.

8. The video feature analysis method according to claim 1, wherein the model evaluation and optimization module uses evaluation metrics such as accuracy, recall rate, and F1 value to evaluate the performance of the trained multi-modal large model, and adjusts and optimizes the model parameters according to the evaluation results using the gradient descent method or other optimization algorithms. The formula is: , where represents the trained model, represents the evaluation metric, represents the optimization algorithm, which improves the accuracy and generalization ability of the model.

9. The video feature analysis method according to claim 1, wherein the data storage and management module adopts a distributed storage architecture and a database management system to store and manage the collected video data, preprocessed data, model parameters, analysis results, etc. During the storage process, data encryption and redundant backup technologies are used. The data encryption adopts the Advanced Encryption Standard algorithm, and the redundant backup formula is: , where represents the stored data, represents the redundancy factor, ensuring the security and reliability of the data.

10. The video feature analysis method according to claim 1, wherein the user interaction module provides a friendly user interface, supports operations such as user input of video queries, setting of analysis parameters, viewing of analysis results, and feedback of opinions, the interface design follows the principles of human-computer interaction, has good usability and operability, and its interaction response formula is: , where represents a user request, represents a system function, ensuring that users can use the system efficiently.

Citation Information

Patent Citations

  • Internet of Things big data intelligent video monitoring system and method

    CN117319609A

  • Data analysis system based on multi-modal fusion

    CN118981743A

  • Multi-modal file data co-processing management system

    CN119537311A

  • Big data advanced modeling method for multi-dimensional data fusion

    CN119580043A

  • Audio-visual assisted fine-grained tactile signal reconstruction method

    WO2024104376A1

Cited By

  • Call visualization model generation method and device and call visualization processing method and device

    CN121506112A

  • Video monitoring data acquisition and fusion method based on multi-source probe and large model

    CN121881346A

  • Training method, audiovisual segmentation method, electronic device and storage medium

    CN122090357A