High-definition video management system for multi-modal data acquisition and processing based on GPU (Graphics Processing Unit)
Through the GPU-based multimodal data acquisition and processing system, the information redundancy and real-time processing problems of multimodal data processing in high-definition video management systems are solved, and efficient multimodal data fusion and decision support are achieved, improving the robustness and processing speed of the system.
Patent Information
- Application Number
- CN202510209913.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-04
AI Technical Summary
The existing high-definition video management system is difficult to effectively process multimodal data, resulting in negative impact on information redundancy and classification results. In addition, traditional CPU processing methods are difficult to meet the real-time processing needs of large data volume and high-resolution video streams.
A multimodal data acquisition and processing system based on GPU is adopted, including data acquisition, GPU acceleration processing, deep learning analysis, data fusion and decision support, and other modules, provides richer input features through multimodal data fusion, and uses CUDA programming model and low-rank multimodal fusion method to build a high-precision classifier.
It realizes efficient fusion of multimodal data, reduces the risk of overfitting, improves the robustness and processing speed of the system, and provides more comprehensive scenario understanding and decision-making support.
Smart Images

Figure CN120259819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video management, and particularly to a high-definition video management system for multi-modal data acquisition and processing based on GPU. Background Art
[0002] With the development of Internet of Things and artificial intelligence technologies, the acquisition and processing of multi-modal data have become a key part of modern information systems. However, the traditional CPU processing method is difficult to meet the real-time processing requirements of large amounts of data and high-resolution video streams. GPU shows great potential in high-definition video management and multi-modal data analysis due to its powerful parallel computing ability.
[0003] Existing high-definition video management systems only perform linear or non-linear mapping on single-modal information, unable to extract the features of multi-modal data and perform fusion processing. There are also high-definition video management systems that only use single-modal data for stitching research, directly stitching multi-modal features without processing, which cannot reflect the information complementarity of mutual dependence between different modalities, resulting in a large amount of information redundancy and having a negative impact on the classification results. Therefore, there is room for improvement. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and a high-definition video management system for multi-modal data acquisition and processing based on GPU is proposed. Its advantage is that multi-modal data fusion can accelerate the training process of neural networks by providing richer input features and reduce the risk of overfitting.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A high-definition video management system for multi-modal data acquisition and processing based on GPU includes a data acquisition module, a GPU acceleration processing module, a deep learning analysis module, a data fusion and decision support module, a peripheral interface module, a display module, a storage module, a Bluetooth module, a 5G wireless transmission module, a WIFI module, and a power module;
[0007] The data acquisition module is responsible for obtaining raw data from various sensors and devices, processing multiple data sources, including high-definition video streams, audio streams, and other sensor data; the GPU acceleration processing module utilizes the highly parallel computing ability of GPU to be able to real-time decode and compress high-definition video streams, audio streams, and other sensor data; the deep learning analysis module uses a convolutional neural network for object detection and classification, and constructs a high-precision classifier with a neural network structure having fine-grained feature discovery ability through highly efficient object detection and recognition algorithms; the data fusion and decision support module includes a data fusion module and a decision support module;
[0008] The data fusion module combines information from multiple data sources and fuses data of multiple modalities using a feature fusion algorithm; the decision support module, based on data fusion, provides customized information and suggestions for decision-makers through analysis and modeling;
[0009] The steps for obtaining multi-modal fusion features are as follows:
[0010] Step 1: First, create three single-modal sub-embedding networks f m , f p and f c to extract the modal feature representations Z m , Z p and Z c of the preprocessed original MRI, PET, and CSF data;
[0011] Step 2: These three sub-embedding networks are two-layer feedforward neural networks; then, the three obtained modal feature representations are fused by a low-rank multi-modal fusion method to obtain a fused feature;
[0012] Step 3: The high-precision classifier treats each modality as a perspective for multi-perspective collaborative learning; a sub-classifier is learned for each perspective to obtain the importance weights of each perspective, and finally, a comprehensive decision is made through the sub-classifiers to obtain the classification result.
[0013] The present invention is further configured such that the steps for the data acquisition module to obtain data are as follows:
[0014] Step 1: Determine the data sources; including video sources: high-definition cameras, webcams, video files, etc.; audio sources: microphones, audio files, etc.; other sensors: temperature, humidity, and motion detectors;
[0015] Step 2: Select appropriate hardware and drivers; for each data source, select an appropriate hardware device, and the HDMS operating system comes with driver programs to control these devices;
[0016] Step 3: Design the data flow; decide on buffer strategies, data format conversion, data compression, and how the data flows from the data acquisition module to the GPU acceleration processing module.
[0017] The present invention is further configured such that the recognition algorithm makes full use of the parallel architecture of the GPU and uses the CUDA programming model.
[0018] The present invention is further configured such that the CUDA programming model mainly consists of three levels: grids, blocks, and threads; a grid is composed of several thread blocks and represents the overall scope of execution of a CUDA program; a thread block contains a group of parallel threads and is a parallel execution unit; a thread is the smallest execution unit; threads within a block perform high-speed information exchange through shared memory.
[0019] The present invention is further configured such that the data fusion is divided into the following three main levels:
[0020] Data-level fusion: directly merges and corrects the original sensor data to eliminate redundancy and inconsistencies and improve data quality;
[0021] Feature-level fusion: extracts the key features of each data source and then combines these features to identify patterns or trends;
[0022] Decision-level fusion: comprehensively evaluates the decision-making suggestions from different sources to obtain the final decision conclusion.
[0023] The present invention is further configured such that the decision support module includes the following components:
[0024] Data analysis: uses techniques such as statistical methods and machine learning algorithms to deeply analyze the fused data to discover hidden patterns and associations;
[0025] Prediction model: based on historical data and current trends, establishes a prediction model to predict possible future developments;
[0026] Optimization algorithm: for specific objectives such as cost minimization and profit maximization, uses optimization algorithms to find the best solutions;
[0027] Interactive interface: provides a user-friendly interface that enables decision-makers to intuitively view the analysis results, adjust parameters, conduct scenario simulations, and receive real-time updates.
[0028] The present invention is further configured such that the convolutional neural network is a neural network layer that extends the convolution of the frequency-domain GCN to hypergraphs and, after optimization and approximation using truncated Chebyshev polynomials, obtains the propagation method between layers of the HGCN.
[0029] The present invention is further configured such that the function of the propagation method is:
[0030] where X (l+1) ∈R N×C represents the output of the l-th layer and is also the input of the (l + 1)-th layer; X (O) = X represents the initial input data; are the parameters to be learned during the training process; σ(·) is the activation function.
[0031] The present invention is further configured such that the objective function of the multiple modality fusion features is:
[0032]
[0033] Where:
[0034]
[0035] Where, represents the consequent parameter of the output of the j-th class of the m-th modality; w m is the weight of the m-th modality; is the fixed prior knowledge of the l-th modality; represents the i-th sample of the m-th modality; if the i-th modality belongs to the j-th class, then the value of y ij is 1, otherwise it is 0; M is the total number of modalities, C is the number of sample classes, and N represents the total number of data samples; λ CO and λ w are regularization parameters, and their values are all greater than 0, and appropriate values are obtained through cross-validation; the optimal consequent parameters and the modality weight w m are solved through cross-iteration method; finally, according to the obtained and w m .
[0036] The present invention is further configured such that the global decision value is calculated from the decision values of each modality, and the calculation formula is:
[0037] The beneficial effects of the present invention are:
[0038] 1. Data of different modalities often carry different information. Fusing this data can provide a more comprehensive scene understanding and make up for the limitations of single-modal data.
[0039] 2. Multi-modal data fusion can accelerate the training process of the neural network by providing richer input features and reduce the risk of overfitting.
[0040] 3. Multi-modal fusion can make the system more robust in the face of noise, missing or outliers in single-modal data, because data of other modalities can be used as supplementary or verification information. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is the overall structural schematic diagram of the high-definition video management system for multi-modal data acquisition and processing based on GPU proposed by the present invention;
[0042] Figure 2 Schematic diagram of the organizational relationship among grids, thread blocks, and threads of the high-definition video management system for multi-modal data acquisition and processing based on GPU proposed by the present invention;
[0043] Figure 3 Schematic diagram of the overall framework of the AD classification method that takes into account both individual characteristics and fusion characteristics of the high-definition video management system for multi-modal data acquisition and processing based on GPU proposed by the present invention. Specific implementation manners
[0044] The technical solutions of this patent will be further described in detail below in conjunction with specific implementation manners.
[0045] The embodiments of this patent will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain this patent and should not be construed as a limitation of this patent.
[0046] Referring to Figure 1 , the high-definition video management system for multi-modal data acquisition and processing based on GPU includes a data acquisition module, a GPU acceleration processing module, a deep learning analysis module, a data fusion and decision support module, a peripheral interface module, a display module, a storage module, a Bluetooth module, a 5G wireless transmission module, a WIFI module, and a power supply module;
[0047] The data acquisition module is responsible for obtaining raw data from various sensors and devices and processing multiple data sources, including high-definition video streams, audio streams, and other sensor data; the GPU acceleration processing module utilizes the highly parallel computing power of the GPU to significantly improve the processing speed and can real-time decode and compress high-definition video streams, audio streams, and other sensor data to improve the data processing efficiency; the deep learning analysis module uses a convolutional neural network for object detection and classification and constructs a high-precision classifier with a neural network structure capable of fine-grained feature discovery through highly efficient object detection and recognition algorithms; the data fusion and decision support module includes a data fusion module and a decision support module;
[0048] The data fusion module combines information from multiple data sources and fuses data of multiple modalities using a feature fusion algorithm, which can avoid the problem that the internal interaction within the modality is potentially suppressed due to direct data splicing; the decision support module, based on data fusion, provides customized information and suggestions for decision-makers through analysis and modeling;
[0049] The steps for the data acquisition module to obtain data are as follows:
[0050] Step 1: Determine the data sources; including video sources: high-definition cameras, webcams, video files, etc.; audio sources: microphones, audio files, etc.; other sensors: temperature, humidity, and motion detectors;
[0051] Step 2: Select appropriate hardware and drivers; for each data source, select an appropriate hardware device, and the HDMS operating system comes with driver programs to control these devices;
[0052] Step 3: Design the data flow; decide on including buffer strategies, data format conversion, data compression, and how the data flows from the data acquisition module to the GPU acceleration processing module.
[0053] Refer to Figure 2 , an identification algorithm to make full use of the parallel architecture of the GPU, using the CUDA programming model; the CUDA programming model mainly consists of three levels: grids, blocks, and threads; a grid consists of several thread blocks and is the overall scope of a CUDA program execution; a thread block contains a group of parallel threads and is a parallel execution unit; a thread is the smallest execution unit; the threads in a block exchange high-speed information through shared memory.
[0054] Data fusion is divided into the following three main levels:
[0055] Data-level fusion: directly merge and correct the original sensor data, eliminate redundancy and inconsistencies, and improve data quality;
[0056] Feature-level fusion: extract the key features of each data source and then combine these features to identify patterns or trends;
[0057] Decision-level fusion: comprehensively evaluate the decision-making suggestions from different sources to draw a final decision conclusion.
[0058] The decision support module includes the following components:
[0059] Data analysis: use techniques such as statistical methods and machine learning algorithms to deeply analyze the fused data and discover hidden patterns and associations;
[0060] Prediction model: based on historical data and current trends, establish a prediction model to predict possible future developments;
[0061] Optimization algorithm: for specific goals, such as cost minimization and profit maximization, use optimization algorithms to find the best solutions;
[0062] Interactive interface: provide a user-friendly interface that enables decision-makers to intuitively view the analysis results, adjust parameters, conduct scenario simulations, and receive real-time updates.
[0063] The convolutional neural network is a neural network layer that extends the convolution of the frequency-domain GCN to hypergraphs. After optimizing and approximating using truncated Chebyshev polynomials, the propagation method between layers of the HGCN is obtained.
[0064] The function of the propagation method is:
[0065] where X (l+1) ∈ R N×C represents the output of the l-th layer and is also the input of the (l + 1)-th layer; X (O) = X, representing the initial input data; are the parameters to be learned during the training process; σ(·) is the activation function.
[0066] Let The hypergraph convolution formula for two layers can be written as:
[0067] Z = f(X, A H )
[0068] = soft max(ΔReLU(ΔXΘ (0) )Θ (1) )
[0069] Given a set of vector representations, representing the encoded information of M single modalities. The tensor representation is obtained by taking the outer product of the input multi-modal features. A vector 1 is added to the end of each modal feature to simulate the interaction between any subset of modalities. The input tensor Z in the form of single-modal representation can be calculated as:
[0070]
[0071] where represents the tensor outer product, z m represents the input representation concatenated with the vector 1. The input tensor generates a vector representation through a linear layer g(·):
[0072] h = g(Z; W, b) = W · Z + b
[0073] where W is the weight and b is the offset. Since Z is an M-order tensor, W is an (M + 1)-order tensor with dimensions d1 × d2 × … × d M × d h The additional dimension of the (M + 1)-th layer corresponds to the size d of the output representation h . This method requires explicitly creating a high-dimensional tensor Z, and its dimension is It will grow exponentially with the increase in the number of modalities. At this time, the number of parameters to be learned in the weight tensor W will also increase exponentially accordingly. LMF decomposes W into a set of modality-specific low-rank factors, and can directly calculate h without explicitly obtaining the high-dimensional tensor, thus reducing the computational complexity.
[0074] Regard W as a tensor of order M with dimension d h where each tensor of order M can be expressed as k = 1, …, d h There exists an exact decomposition into vectors:
[0075]
[0076] where the smallest R that makes the decomposition valid is called the rank of the tensor. After fixing R as r, use r decomposition factors to reconstruct the low-rank version of Then the low-rank weight tensor can be reconstructed by the following formula
[0077]
[0078] Based on the decomposition of W, and according to the fact that the tensor Z can also be decomposed into The original formula for calculating h can be deduced as follows:
[0079]
[0080] where represents the element-wise product of a series of tensors represents the i-th low-rank decomposition factor corresponding to modality m, and this parameter can be learned through backpropagation.
[0081] Referring to Figure 3 The steps for obtaining the multi-modal fusion feature are as follows:
[0082] Step 1: First, create three single-modal sub-embedding networks, and, to extract the modal feature representations,, of the preprocessed original MRI, PET, and CSF data;
[0083] Step 2: These three sub-embedding networks are two-layer feed-forward neural networks; then, the three obtained modal feature representations are fused through the low-rank multi-modal fusion method to obtain a fused feature;
[0084] Step 3: The high-precision classifier regards each modality as a perspective for multi-perspective collaborative learning; learns a sub-classifier for each perspective to obtain the importance weights of each perspective, and finally makes a comprehensive decision through the sub-classifiers to obtain the classification result. Its objective function is:
[0085]
[0086] Wherein:
[0087]
[0088]
[0089] wherein, represents the consequent parameter of the j-th class output of the m-th modality; w m is the weight of the m-th modality; is the fixed prior knowledge of the l-th modality; represents the i-th sample of the m-th modality; if the i-th modality belongs to the j-th class, then y ij has a value of 1, otherwise 0; M is the total number of modalities, C is the number of sample classes, and N represents the total number of data samples; λ CO and λ w are regularization parameters, both of which have values greater than 0, and appropriate values are obtained through cross-validation; the optimal consequent parameters and the modality weight w m are solved by the cross-iteration method; finally, according to the obtained and w m ; the global decision value is calculated from the decision values of each modality, and the calculation formula is:
[0090] As described above, only the preferred specific embodiments of the present invention are shown, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A high-definition video management system for multi-modal data acquisition and processing based on GPU, characterized in that, It includes a data acquisition module, a GPU acceleration processing module, a deep learning analysis module, a data fusion and decision support module, a peripheral interface module, a display module, a storage module, a Bluetooth module, a 5G wireless transmission module, a WIFI module, and a power module; The data acquisition module is responsible for obtaining raw data from various sensors and devices, and processing multiple data sources, including high-definition video streams, audio streams, and other sensor data; The GPU acceleration processing module utilizes the highly parallel computing power of the GPU to be able to decode and compress high-definition video streams, audio streams, and other sensor data in real time; The deep learning analysis module uses a convolutional neural network for object detection and classification, and constructs a high-precision classifier with a neural network structure capable of fine-grained feature discovery through highly efficient object detection and recognition algorithms; The data fusion and decision support module includes a data fusion module and a decision support module; The data fusion module combines information from multiple data sources and fuses data of multiple modalities using a feature fusion algorithm; The decision support module, based on data fusion, provides customized information and suggestions for decision-makers through analysis and modeling; The steps for obtaining multi-modal fusion features are as follows: Step 1: First, create three single-modal sub-embedding networks f m , f p and f c to extract the modality feature representations Z m , Z p and Z c of the preprocessed original MRI, PET, and CSF data; Step 2: These three sub-embedding networks are two-layer feedforward neural networks; Then, the three modality feature representations obtained are fused through a low-rank multi-modal fusion method to obtain a fused feature; Step 3: The high-precision classifier treats each modality as a perspective for multi-perspective collaborative learning; A sub-classifier is learned for each perspective to obtain the importance weights of each perspective, and finally, a comprehensive decision is made through the sub-classifiers to obtain the classification result.
2. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, characterized in that, The steps for the data acquisition module to obtain data are as follows: Step 1: Determine the data sources; including video sources: high-definition cameras, network cameras, video files, etc.; Audio sources: microphones, audio files, etc.; Other sensors: temperature, humidity, and motion detectors; Step 2: Select appropriate hardware and drivers; For each data source, select an appropriate hardware device, and the HDMS operating system comes with driver programs to control these devices; Step 3: Design the data flow; Decide on including buffer strategies, data format conversion, data compression, and how to flow from the data acquisition module to the GPU acceleration processing module.
3. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, wherein, The recognition algorithm makes full use of the parallel architecture of the GPU and uses the CUDA programming model.
4. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 3, wherein The CUDA programming model mainly consists of three levels: grids, blocks, and threads; A grid consists of several thread blocks and is the overall scope of execution of a CUDA program; A thread block contains a group of parallel threads and is a parallel execution unit; A thread is the smallest execution unit; The threads in a block exchange information at high speed through shared memory.
5. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, wherein The data fusion is divided into the following three main levels: Data-level fusion: Directly merge and correct the original sensor data to eliminate redundancy and inconsistencies and improve data quality; Feature-level fusion: Extract the key features of each data source and then combine these features to identify patterns or trends; Decision-level fusion: Comprehensively evaluate decision-making suggestions from different sources to draw a final decision conclusion.
6. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, wherein The decision support module includes the following components: Data analysis: Use techniques such as statistical methods and machine learning algorithms to deeply analyze the fused data and discover hidden patterns and associations; Prediction model: Based on historical data and current trends, establish a prediction model to predict possible future developments; Optimization algorithm: For specific goals, such as cost minimization and profit maximization, use optimization algorithms to find the best solutions; Interactive interface: Provide a user-friendly interface that enables decision-makers to intuitively view analysis results, adjust parameters, conduct scenario simulations, and receive real-time updates.
7. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, characterized in that, The convolutional neural network is a neural network layer that extends the convolution of the frequency-domain GCN to hypergraphs and uses truncated Chebyshev polynomials for optimization and approximation to obtain the propagation method between layers of the HGCN.
8. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 7, characterized in that, The function of the propagation mode is as follows: where X (l+1) ∈R N×C represents the output of the l-th layer and is also the input of the (l + 1)-th layer; X (O) = X, representing the initial input data; are the parameters to be learned during the training process; σ(·) is the activation function.
9. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 1, characterized in that, The objective function of the multiple modality fusion features is as follows: Where: Among them, represents the consequent parameter of the output of the j-th class in the m-th mode; w m is the m-th The weight of the modality; is the fixed prior knowledge of the l-th modality; represents the i-th sample of the m-th modality; if the i-th modality belongs to the j-th class, then the value of y ij is 1, otherwise it is 0; M is the total number of modalities, C is the number of sample classes, and N represents the total number of data samples; λ CO and λ w are regularization parameters, and their values are all greater than 0. Appropriate values are obtained through cross-validation; the optimal consequent parameters and the modality weight w m are solved by the cross-iteration method; finally, according to the obtained and w m .
10. The high-definition video management system for multi-modal data acquisition and processing based on GPU according to claim 9, characterized in that, The global decision value is calculated from the decision values of each modality, and the calculation formula is: